{
  "id": 570508,
  "title": "🧬 Estimating Covarying Base Pairs Based on MSA",
  "url": "/competitions/stanford-rna-3d-folding/discussion/570508",
  "author_name": "Bilzard",
  "post_date": "2025-03-28T10:39:19.243000",
  "votes": 8,
  "comment_count": 4,
  "views": 0,
  "content": "<p>I implemented the idea of <strong>covarying mutation patterns</strong> proposed by <a href=\"https://www.kaggle.com/tilii7\" target=\"_blank\">@tilii7</a>.</p>\n<p>Notebook: <a href=\"https://www.kaggle.com/code/tatamikenn/estimating-covarying-base-pairs-based-on-msa\" target=\"_blank\">https://www.kaggle.com/code/tatamikenn/estimating-covarying-base-pairs-based-on-msa</a></p>\n<p>Without going to details, I defined a <strong>pairing score</strong> across homologous RNA sequences . Using this score, we can estimate <strong>how likely each pair of bases is to form a structural base-pair</strong>. For more details see the notebook shared above.</p>\n<p>A sample pairing score matrix (target_id = R1190) is shown below:</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4910466%2Fdc0a43798f37499730c9e285f1a2b975%2FR1190.png?generation=1743220491874910&amp;alt=media\" alt=\"\"></p>\n<p>In this matrix, you can clearly see an <strong>anti-diagonal</strong> pattern, which indicates a continuous sequence of base-pairs.  <br>\nThis pattern is likely a result of covarying mutations.</p>\n<p>I also visualized the 3D structures using the competition's labeled data, and found that the <strong>predicted base-pairs are actually located close to each other</strong> in 3D space.</p>\n<p>This suggests that <a href=\"https://www.kaggle.com/tilii7\" target=\"_blank\">@tilii7</a>'s original idea really works on actual PDB data: <strong>information about covarying base-pairs can help constrain the RNA 3D structure.</strong></p>\n<hr>\n<h2>Current Limitations</h2>\n<ol>\n<li>Estimated co-varying base pairs do not always match with spatially proximal C1'-C1' pairs.  <br>\n(This might be because the actual base-pairing interaction occurs away from the C1' atoms.)</li>\n<li>In some cases, estimated co-varying pairs are even farther apart than other covarying base pairs.</li>\n</ol>\n<h3>Possible Explanation of Limitation #2</h3>\n<p>Covariation reflects evolutionary constraints, not necessarily the current physical configuration. It's possible that such base pairs were close in space <strong>at some point in evolutionary history</strong>, but later diverged due to structural rearrangements or alternative conformations.</p>\n<p>In other words, <strong>covariation signals the potential for pairing</strong>, but doesn't guarantee that those bases are currently interacting in the observed 3D structure.</p>\n<hr>\n<h2>Sample 3D Structures</h2>\n<p>Below are examples of predicted base-pairs (in cyan lines) overlaid on 3D structures.</p>\n<p><strong>1A51_A</strong>:</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4910466%2F127a74cc0a06f0c0cbd3473b090071bc%2FScreenshot%202025-03-29%20at%2012.52.35.png?generation=1743220518911275&amp;alt=media\" alt=\"\"></p>\n<p><strong>R1190</strong>:</p>\n<p><strong>Note</strong>: Most of the candidate pairs shown here are located close together in 3D space, except for the pair (39, 60), which are relatively <strong>far apart</strong> despite having high pairing scores.</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4910466%2F8b7ef4647a42af36a344f3af68cd2298%2FScreenshot%202025-03-29%20at%2012.54.10.png?generation=1743220534933611&amp;alt=media\" alt=\"\"></p>",
  "messages": [
    {
      "id": 3161738,
      "postDate": "2025-03-28T10:39:19.243Z",
      "content": "<p>I implemented the idea of <strong>covarying mutation patterns</strong> proposed by <a href=\"https://www.kaggle.com/tilii7\" target=\"_blank\">@tilii7</a>.</p>\n<p>Notebook: <a href=\"https://www.kaggle.com/code/tatamikenn/estimating-covarying-base-pairs-based-on-msa\" target=\"_blank\">https://www.kaggle.com/code/tatamikenn/estimating-covarying-base-pairs-based-on-msa</a></p>\n<p>Without going to details, I defined a <strong>pairing score</strong> across homologous RNA sequences . Using this score, we can estimate <strong>how likely each pair of bases is to form a structural base-pair</strong>. For more details see the notebook shared above.</p>\n<p>A sample pairing score matrix (target_id = R1190) is shown below:</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4910466%2Fdc0a43798f37499730c9e285f1a2b975%2FR1190.png?generation=1743220491874910&amp;alt=media\" alt=\"\"></p>\n<p>In this matrix, you can clearly see an <strong>anti-diagonal</strong> pattern, which indicates a continuous sequence of base-pairs.  <br>\nThis pattern is likely a result of covarying mutations.</p>\n<p>I also visualized the 3D structures using the competition's labeled data, and found that the <strong>predicted base-pairs are actually located close to each other</strong> in 3D space.</p>\n<p>This suggests that <a href=\"https://www.kaggle.com/tilii7\" target=\"_blank\">@tilii7</a>'s original idea really works on actual PDB data: <strong>information about covarying base-pairs can help constrain the RNA 3D structure.</strong></p>\n<hr>\n<h2>Current Limitations</h2>\n<ol>\n<li>Estimated co-varying base pairs do not always match with spatially proximal C1'-C1' pairs.  <br>\n(This might be because the actual base-pairing interaction occurs away from the C1' atoms.)</li>\n<li>In some cases, estimated co-varying pairs are even farther apart than other covarying base pairs.</li>\n</ol>\n<h3>Possible Explanation of Limitation #2</h3>\n<p>Covariation reflects evolutionary constraints, not necessarily the current physical configuration. It's possible that such base pairs were close in space <strong>at some point in evolutionary history</strong>, but later diverged due to structural rearrangements or alternative conformations.</p>\n<p>In other words, <strong>covariation signals the potential for pairing</strong>, but doesn't guarantee that those bases are currently interacting in the observed 3D structure.</p>\n<hr>\n<h2>Sample 3D Structures</h2>\n<p>Below are examples of predicted base-pairs (in cyan lines) overlaid on 3D structures.</p>\n<p><strong>1A51_A</strong>:</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4910466%2F127a74cc0a06f0c0cbd3473b090071bc%2FScreenshot%202025-03-29%20at%2012.52.35.png?generation=1743220518911275&amp;alt=media\" alt=\"\"></p>\n<p><strong>R1190</strong>:</p>\n<p><strong>Note</strong>: Most of the candidate pairs shown here are located close together in 3D space, except for the pair (39, 60), which are relatively <strong>far apart</strong> despite having high pairing scores.</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4910466%2F8b7ef4647a42af36a344f3af68cd2298%2FScreenshot%202025-03-29%20at%2012.54.10.png?generation=1743220534933611&amp;alt=media\" alt=\"\"></p>",
      "rawMarkdown": "I implemented the idea of **covarying mutation patterns** proposed by @tilii7.\n\nNotebook: https://www.kaggle.com/code/tatamikenn/estimating-covarying-base-pairs-based-on-msa\n\nWithout going to details, I defined a **pairing score** across homologous RNA sequences . Using this score, we can estimate **how likely each pair of bases is to form a structural base-pair**. For more details see the notebook shared above.\n\nA sample pairing score matrix (target_id = R1190) is shown below:\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4910466%2Fdc0a43798f37499730c9e285f1a2b975%2FR1190.png?generation=1743220491874910&alt=media)\n\nIn this matrix, you can clearly see an **anti-diagonal** pattern, which indicates a continuous sequence of base-pairs.  \nThis pattern is likely a result of covarying mutations.\n\nI also visualized the 3D structures using the competition's labeled data, and found that the **predicted base-pairs are actually located close to each other** in 3D space.\n\nThis suggests that @tilii7's original idea really works on actual PDB data: **information about covarying base-pairs can help constrain the RNA 3D structure.**\n\n---\n\n## Current Limitations\n\n1. Estimated co-varying base pairs do not always match with spatially proximal C1'-C1' pairs.  \n   (This might be because the actual base-pairing interaction occurs away from the C1' atoms.)\n2. In some cases, estimated co-varying pairs are even farther apart than other covarying base pairs.\n\n### Possible Explanation of Limitation #2\n\nCovariation reflects evolutionary constraints, not necessarily the current physical configuration. It's possible that such base pairs were close in space **at some point in evolutionary history**, but later diverged due to structural rearrangements or alternative conformations.\n\nIn other words, **covariation signals the potential for pairing**, but doesn't guarantee that those bases are currently interacting in the observed 3D structure.\n\n---\n\n## Sample 3D Structures\n\nBelow are examples of predicted base-pairs (in cyan lines) overlaid on 3D structures.\n\n**1A51_A**:\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4910466%2F127a74cc0a06f0c0cbd3473b090071bc%2FScreenshot%202025-03-29%20at%2012.52.35.png?generation=1743220518911275&alt=media)\n\n**R1190**:\n\n**Note**: Most of the candidate pairs shown here are located close together in 3D space, except for the pair (39, 60), which are relatively **far apart** despite having high pairing scores.\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4910466%2F8b7ef4647a42af36a344f3af68cd2298%2FScreenshot%202025-03-29%20at%2012.54.10.png?generation=1743220534933611&alt=media)",
      "votes": 8
    },
    {
      "id": 3162222,
      "postDate": "2025-03-29T00:36:45.357Z",
      "content": "<p>You did a nice job here to provide a proof of principle for RNA base co-variation.</p>\n<blockquote>\n  <p>Estimated co-varying base pairs do not always match with spatially proximal C1'-C1' pairs. (This might be because the actual base-pairing interaction occurs away from the C1' atoms.)</p>\n</blockquote>\n<p>More likely this is because estimating co-variation is not easy. It requires deep alignments - let's say that means at least 500-1000 sequences that are fairly diverse. Sequences also have to be analyzed so that similar sequences get less weight.</p>\n<p>Strong co-variation does mean that bases are paired, and the exceptions to that are rare. It is likely that some of your calculated co-variations are overly optimistic, which is why they don't correspond to structural patterns from the PDB file. Also, it could be that your observed patterns are real, but there are stronger co-variations involving those bases that you didn't uncover.</p>",
      "rawMarkdown": "You did a nice job here to provide a proof of principle for RNA base co-variation.\n\n> Estimated co-varying base pairs do not always match with spatially proximal C1'-C1' pairs. (This might be because the actual base-pairing interaction occurs away from the C1' atoms.)\n\nMore likely this is because estimating co-variation is not easy. It requires deep alignments - let's say that means at least 500-1000 sequences that are fairly diverse. Sequences also have to be analyzed so that similar sequences get less weight.\n\nStrong co-variation does mean that bases are paired, and the exceptions to that are rare. It is likely that some of your calculated co-variations are overly optimistic, which is why they don't correspond to structural patterns from the PDB file. Also, it could be that your observed patterns are real, but there are stronger co-variations involving those bases that you didn't uncover.",
      "votes": 1,
      "replies": [
        {
          "id": 3162293,
          "postDate": "2025-03-29T04:03:02.950Z",
          "content": "<p><a href=\"https://www.kaggle.com/tilii7\" target=\"_blank\">@tilii7</a> Thank you for the advice. (and also for the original idea!)</p>\n<p>Maybe my first implementation is too naive. I updated my custom paring score with <strong>mutual information (MI)</strong>, and found it improved accuracy of candidates significantly.</p>\n<p>However, some pairs still have large discrepancy even if they have large paring score.</p>\n<p>I believe one of this reason is possibly the <strong>false correlation</strong> detected. More sophisticated paring score calculation (weakening weight for too-close relative RNAs etc.) might solve this problem.</p>",
          "rawMarkdown": "@tilii7 Thank you for the advice. (and also for the original idea!)\n\nMaybe my first implementation is too naive. I updated my custom paring score with **mutual information (MI)**, and found it improved accuracy of candidates significantly.\n\nHowever, some pairs still have large discrepancy even if they have large paring score.\n\nI believe one of this reason is possibly the **false correlation** detected. More sophisticated paring score calculation (weakening weight for too-close relative RNAs etc.) might solve this problem.",
          "votes": 1,
          "replies": [
            {
              "id": 3162355,
              "postDate": "2025-03-29T05:20:44.813Z",
              "content": "<p>I don't think you need to reinvent the wheel here. There are many implementations for co-variance calculation in proteins, and some for RNA as well. It is the same thing conceptually, and this is how AlphaFold functions as well.</p>\n<ul>\n<li><a href=\"https://academic.oup.com/bioinformatics/article/34/19/3308/4987145\" target=\"_blank\">https://academic.oup.com/bioinformatics/article/34/19/3308/4987145</a></li>\n<li><a href=\"https://academic.oup.com/bioinformatics/article/28/2/184/198108\" target=\"_blank\">https://academic.oup.com/bioinformatics/article/28/2/184/198108</a></li>\n<li><a href=\"https://www.cell.com/cell/fulltext/S0092-8674(16)30328-2\" target=\"_blank\">https://www.cell.com/cell/fulltext/S0092-8674(16)30328-2</a></li>\n<li><a href=\"https://journals.plos.org/plosone/article?id=10.1371/journal.pone.0028766\" target=\"_blank\">https://journals.plos.org/plosone/article?id=10.1371/journal.pone.0028766</a></li>\n<li><a href=\"https://www.nature.com/articles/nbt.2419\" target=\"_blank\">https://www.nature.com/articles/nbt.2419</a></li>\n<li><a href=\"https://academic.oup.com/bioinformatics/article/29/22/2933/316439\" target=\"_blank\">https://academic.oup.com/bioinformatics/article/29/22/2933/316439</a></li>\n<li><a href=\"https://www.tandfonline.com/doi/10.4161/rna.25038\" target=\"_blank\">https://www.tandfonline.com/doi/10.4161/rna.25038</a></li>\n</ul>",
              "rawMarkdown": "I don't think you need to reinvent the wheel here. There are many implementations for co-variance calculation in proteins, and some for RNA as well. It is the same thing conceptually, and this is how AlphaFold functions as well.\n\n- https://academic.oup.com/bioinformatics/article/34/19/3308/4987145\n- https://academic.oup.com/bioinformatics/article/28/2/184/198108\n- https://www.cell.com/cell/fulltext/S0092-8674(16)30328-2\n- https://journals.plos.org/plosone/article?id=10.1371/journal.pone.0028766\n- https://www.nature.com/articles/nbt.2419\n- https://academic.oup.com/bioinformatics/article/29/22/2933/316439\n- https://www.tandfonline.com/doi/10.4161/rna.25038\n",
              "votes": 2
            }
          ]
        }
      ]
    },
    {
      "id": 3161756,
      "postDate": "2025-03-28T10:54:08.883Z",
      "rawMarkdown": "",
      "isDeleted": true
    }
  ],
  "comments": [
    {
      "id": 3162222,
      "author_name": "Tilii",
      "author_url": "",
      "post_date": "2025-03-29T00:36:45.357000",
      "content": "<p>You did a nice job here to provide a proof of principle for RNA base co-variation.</p>\n<blockquote>\n  <p>Estimated co-varying base pairs do not always match with spatially proximal C1'-C1' pairs. (This might be because the actual base-pairing interaction occurs away from the C1' atoms.)</p>\n</blockquote>\n<p>More likely this is because estimating co-variation is not easy. It requires deep alignments - let's say that means at least 500-1000 sequences that are fairly diverse. Sequences also have to be analyzed so that similar sequences get less weight.</p>\n<p>Strong co-variation does mean that bases are paired, and the exceptions to that are rare. It is likely that some of your calculated co-variations are overly optimistic, which is why they don't correspond to structural patterns from the PDB file. Also, it could be that your observed patterns are real, but there are stronger co-variations involving those bases that you didn't uncover.</p>",
      "votes": 1,
      "replies": [
        {
          "id": 3162293,
          "author_name": "Bilzard",
          "author_url": "",
          "post_date": "2025-03-29T04:03:02.950000",
          "content": "<p><a href=\"https://www.kaggle.com/tilii7\" target=\"_blank\">@tilii7</a> Thank you for the advice. (and also for the original idea!)</p>\n<p>Maybe my first implementation is too naive. I updated my custom paring score with <strong>mutual information (MI)</strong>, and found it improved accuracy of candidates significantly.</p>\n<p>However, some pairs still have large discrepancy even if they have large paring score.</p>\n<p>I believe one of this reason is possibly the <strong>false correlation</strong> detected. More sophisticated paring score calculation (weakening weight for too-close relative RNAs etc.) might solve this problem.</p>",
          "votes": 1,
          "replies": [
            {
              "id": 3162355,
              "author_name": "Tilii",
              "author_url": "",
              "post_date": "2025-03-29T05:20:44.813000",
              "content": "<p>I don't think you need to reinvent the wheel here. There are many implementations for co-variance calculation in proteins, and some for RNA as well. It is the same thing conceptually, and this is how AlphaFold functions as well.</p>\n<ul>\n<li><a href=\"https://academic.oup.com/bioinformatics/article/34/19/3308/4987145\" target=\"_blank\">https://academic.oup.com/bioinformatics/article/34/19/3308/4987145</a></li>\n<li><a href=\"https://academic.oup.com/bioinformatics/article/28/2/184/198108\" target=\"_blank\">https://academic.oup.com/bioinformatics/article/28/2/184/198108</a></li>\n<li><a href=\"https://www.cell.com/cell/fulltext/S0092-8674(16)30328-2\" target=\"_blank\">https://www.cell.com/cell/fulltext/S0092-8674(16)30328-2</a></li>\n<li><a href=\"https://journals.plos.org/plosone/article?id=10.1371/journal.pone.0028766\" target=\"_blank\">https://journals.plos.org/plosone/article?id=10.1371/journal.pone.0028766</a></li>\n<li><a href=\"https://www.nature.com/articles/nbt.2419\" target=\"_blank\">https://www.nature.com/articles/nbt.2419</a></li>\n<li><a href=\"https://academic.oup.com/bioinformatics/article/29/22/2933/316439\" target=\"_blank\">https://academic.oup.com/bioinformatics/article/29/22/2933/316439</a></li>\n<li><a href=\"https://www.tandfonline.com/doi/10.4161/rna.25038\" target=\"_blank\">https://www.tandfonline.com/doi/10.4161/rna.25038</a></li>\n</ul>",
              "votes": 2,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 3161756,
      "author_name": "",
      "author_url": "",
      "post_date": "2025-03-28T10:54:08.883000",
      "content": "",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "3161738": "I implemented the idea of **covarying mutation patterns** proposed by @tilii7.\n\nNotebook: https://www.kaggle.com/code/tatamikenn/estimating-covarying-base-pairs-based-on-msa\n\nWithout going to details, I defined a **pairing score** across homologous RNA sequences . Using this score, we can estimate **how likely each pair of bases is to form a structural base-pair**. For more details see the notebook shared above.\n\nA sample pairing score matrix (target_id = R1190) is shown below:\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4910466%2Fdc0a43798f37499730c9e285f1a2b975%2FR1190.png?generation=1743220491874910&alt=media)\n\nIn this matrix, you can clearly see an **anti-diagonal** pattern, which indicates a continuous sequence of base-pairs.  \nThis pattern is likely a result of covarying mutations.\n\nI also visualized the 3D structures using the competition's labeled data, and found that the **predicted base-pairs are actually located close to each other** in 3D space.\n\nThis suggests that @tilii7's original idea really works on actual PDB data: **information about covarying base-pairs can help constrain the RNA 3D structure.**\n\n---\n\n## Current Limitations\n\n1. Estimated co-varying base pairs do not always match with spatially proximal C1'-C1' pairs.  \n   (This might be because the actual base-pairing interaction occurs away from the C1' atoms.)\n2. In some cases, estimated co-varying pairs are even farther apart than other covarying base pairs.\n\n### Possible Explanation of Limitation #2\n\nCovariation reflects evolutionary constraints, not necessarily the current physical configuration. It's possible that such base pairs were close in space **at some point in evolutionary history**, but later diverged due to structural rearrangements or alternative conformations.\n\nIn other words, **covariation signals the potential for pairing**, but doesn't guarantee that those bases are currently interacting in the observed 3D structure.\n\n---\n\n## Sample 3D Structures\n\nBelow are examples of predicted base-pairs (in cyan lines) overlaid on 3D structures.\n\n**1A51_A**:\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4910466%2F127a74cc0a06f0c0cbd3473b090071bc%2FScreenshot%202025-03-29%20at%2012.52.35.png?generation=1743220518911275&alt=media)\n\n**R1190**:\n\n**Note**: Most of the candidate pairs shown here are located close together in 3D space, except for the pair (39, 60), which are relatively **far apart** despite having high pairing scores.\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4910466%2F8b7ef4647a42af36a344f3af68cd2298%2FScreenshot%202025-03-29%20at%2012.54.10.png?generation=1743220534933611&alt=media)",
    "3162222": "You did a nice job here to provide a proof of principle for RNA base co-variation.\n\n> Estimated co-varying base pairs do not always match with spatially proximal C1'-C1' pairs. (This might be because the actual base-pairing interaction occurs away from the C1' atoms.)\n\nMore likely this is because estimating co-variation is not easy. It requires deep alignments - let's say that means at least 500-1000 sequences that are fairly diverse. Sequences also have to be analyzed so that similar sequences get less weight.\n\nStrong co-variation does mean that bases are paired, and the exceptions to that are rare. It is likely that some of your calculated co-variations are overly optimistic, which is why they don't correspond to structural patterns from the PDB file. Also, it could be that your observed patterns are real, but there are stronger co-variations involving those bases that you didn't uncover.",
    "3161756": ""
  }
}