{
  "id": 569164,
  "title": "Understanding MSA files ",
  "url": "/competitions/stanford-rna-3d-folding/discussion/569164",
  "author_name": "Ignasi Alemany",
  "post_date": "2025-03-20T11:41:09.117000",
  "votes": 3,
  "comment_count": 6,
  "views": 0,
  "content": "<p>I am currently looking at the MSA files provided in the dataset. For simplicity, I attach some sequences in file '1RY1_E.MSA'. I have no background in this competition and I am trying to understand the MSA files. </p>\n<p>`</p>\n<blockquote>\n  <p>query<br>\n  GGGCCGGGCGCGGUGGCGCGCGCCUGUAGUCCCAGCUACUCGGGAGGCUC<br>\n  ANKR01284863.1/22738-23056/2-47 [subseq from] Myotis brandtii contig284863, whole genome shotgun sequence.<br>\n  ---CCGGGCGCGGUGGCGCGCGCCUGUAGUCCCAGCUACUCGGGAGGCU-<br>\n  DS562880.1/9036401-9036695/2-47 [subseq from] Cavia porcellus supercont2_25 genomic scaffold, whole genome shotgun sequence.<br>\n  ---CCGGGCGCGGUGGCGCACGCCUGUAGUCCCAGCUACUCGGGAGGCU-<br>\n  GL896957.1/7960321-7960609/2-47 [subseq from] Mustela putorius furo unplaced genomic scaffold scaffold00060, whole genome shotgun sequence.<br>\n  c----UGGGCGCGGUGGCGCGCGCCUGUAGUCCCAGCUACUCGGGAGGCU-<br>\n  `</p>\n</blockquote>\n<p>The Multiple Sequence Alignment (MSA) file gives me a bunch of sequences shared between different organisms, due to evolution there are mutations in the DNA and those get reflected in the RNA. We assume that the structure of these RNA will be similar and they clearly share some parts of the sequence. </p>\n<p>Are these files aligned as in maximized to reduce the distance between these conservative regions? If they are not aligned should I use programs like MAFFT or anyone knows which one would be better for non-coding RNAs?</p>",
  "messages": [
    {
      "id": 3154809,
      "postDate": "2025-03-20T12:29:34.460Z",
      "content": "<p>These are aligned. The first sequence (<code>query</code>) is the anchor and everything else is written relative to it. That sequence will have all capital letters and no gaps (no <code>-</code> characters). Gaps in subsequent sequences mean that they miss bases at certain positions. If they have extra bases compared to the query sequences, they will be shown as lowercase letters.</p>\n<p>Maybe it will help you to see the same alignment in a different format.</p>\n<pre><code>                 -GGGCCGGGCGCGGUGGCGCGCGCCUGUAGUCCCAGCUACUCGGGAGGCUC\n./    ----CCGGGCGCGGUGGCGCGCGCCUGUAGUCCCAGCUACUCGGGAGGCU-\n./    ----CCGGGCGCGGUGGCGCACGCCUGUAGUCCCAGCUACUCGGGAGGCU-\n./    c----UGGGCGCGGUGGCGCGCGCCUGUAGUCCCAGCUACUCGGGAGGCU-\n</code></pre>\n<p>By the way, that alignment you used as an example is not right, and it should look like this:</p>\n<pre><code>                 GGGCCGGGCGCGGUGGCGCGCGCCUGUAGUCCCAGCUACUCGGGAGGCUC\n./    ---CCGGGCGCGGUGGCGCGCGCCUGUAGUCCCAGCUACUCGGGAGGCU-\n./    ---CCGGGCGCGGUGGCGCACGCCUGUAGUCCCAGCUACUCGGGAGGCU-\n./    ---CUGGGCGCGGUGGCGCGCGCCUGUAGUCCCAGCUACUCGGGAGGCU-\n</code></pre>",
      "rawMarkdown": "These are aligned. The first sequence (`query`) is the anchor and everything else is written relative to it. That sequence will have all capital letters and no gaps (no `-` characters). Gaps in subsequent sequences mean that they miss bases at certain positions. If they have extra bases compared to the query sequences, they will be shown as lowercase letters.\n\nMaybe it will help you to see the same alignment in a different format.\n\n```\nquery                 -GGGCCGGGCGCGGUGGCGCGCGCCUGUAGUCCCAGCUACUCGGGAGGCUC\nANKR01284863.1/227    ----CCGGGCGCGGUGGCGCGCGCCUGUAGUCCCAGCUACUCGGGAGGCU-\nDS562880.1/9036401    ----CCGGGCGCGGUGGCGCACGCCUGUAGUCCCAGCUACUCGGGAGGCU-\nGL896957.1/7960321    c----UGGGCGCGGUGGCGCGCGCCUGUAGUCCCAGCUACUCGGGAGGCU-\n```\n\nBy the way, that alignment you used as an example is not right, and it should look like this:\n\n```\nquery                 GGGCCGGGCGCGGUGGCGCGCGCCUGUAGUCCCAGCUACUCGGGAGGCUC\nANKR01284863.1/227    ---CCGGGCGCGGUGGCGCGCGCCUGUAGUCCCAGCUACUCGGGAGGCU-\nDS562880.1/9036401    ---CCGGGCGCGGUGGCGCACGCCUGUAGUCCCAGCUACUCGGGAGGCU-\nGL896957.1/7960321    ---CUGGGCGCGGUGGCGCGCGCCUGUAGUCCCAGCUACUCGGGAGGCU-\n```\n",
      "votes": 3,
      "replies": [
        {
          "id": 3154968,
          "postDate": "2025-03-20T15:06:24.537Z",
          "content": "<p>Thank you so much, I tried using RNA.alifold(sequences) from ViennaRNA package but it seems it cannot deal with sequences of diff length. I have looked it up and it seems that aligned sequences should have the same length? I am confused</p>",
          "rawMarkdown": "Thank you so much, I tried using RNA.alifold(sequences) from ViennaRNA package but it seems it cannot deal with sequences of diff length. I have looked it up and it seems that aligned sequences should have the same length? I am confused",
          "replies": [
            {
              "id": 3154975,
              "postDate": "2025-03-20T15:12:45.437Z",
              "rawMarkdown": "",
              "isDeleted": true
            },
            {
              "id": 3154980,
              "postDate": "2025-03-20T15:18:52.007Z",
              "content": "<p>Ok nevermind I just realised that is a specific format called A2M and ViennaRNA does not support this alignment format! I will have to transform it to Clustal or other supported formats or use other packages,any that you rcommend?</p>",
              "rawMarkdown": "Ok nevermind I just realised that is a specific format called A2M and ViennaRNA does not support this alignment format! I will have to transform it to Clustal or other supported formats or use other packages,any that you rcommend?"
            },
            {
              "id": 3155370,
              "postDate": "2025-03-21T01:02:10.863Z",
              "content": "<p>There is a file named <code>reformat.pl</code> in <a href=\"https://github.com/soedinglab/hh-suite/tree/master/scripts\" target=\"_blank\"><strong>this directory</strong></a>. It can convert from the A3M format (which is what we have) into any other common alignment format, including Clustal. It is a standalone Perl script that should work on Kaggle.</p>",
              "rawMarkdown": "There is a file named `reformat.pl` in [**this directory**](https://github.com/soedinglab/hh-suite/tree/master/scripts). It can convert from the A3M format (which is what we have) into any other common alignment format, including Clustal. It is a standalone Perl script that should work on Kaggle.",
              "votes": 3
            },
            {
              "id": 3155784,
              "postDate": "2025-03-21T12:06:11.490Z",
              "content": "<p>Yes exactly sorry is A3M format! It worked just fine, thank you so much!</p>",
              "rawMarkdown": "Yes exactly sorry is A3M format! It worked just fine, thank you so much!"
            }
          ]
        }
      ]
    },
    {
      "id": 3154770,
      "postDate": "2025-03-20T11:41:09.117Z",
      "content": "<p>I am currently looking at the MSA files provided in the dataset. For simplicity, I attach some sequences in file '1RY1_E.MSA'. I have no background in this competition and I am trying to understand the MSA files. </p>\n<p>`</p>\n<blockquote>\n  <p>query<br>\n  GGGCCGGGCGCGGUGGCGCGCGCCUGUAGUCCCAGCUACUCGGGAGGCUC<br>\n  ANKR01284863.1/22738-23056/2-47 [subseq from] Myotis brandtii contig284863, whole genome shotgun sequence.<br>\n  ---CCGGGCGCGGUGGCGCGCGCCUGUAGUCCCAGCUACUCGGGAGGCU-<br>\n  DS562880.1/9036401-9036695/2-47 [subseq from] Cavia porcellus supercont2_25 genomic scaffold, whole genome shotgun sequence.<br>\n  ---CCGGGCGCGGUGGCGCACGCCUGUAGUCCCAGCUACUCGGGAGGCU-<br>\n  GL896957.1/7960321-7960609/2-47 [subseq from] Mustela putorius furo unplaced genomic scaffold scaffold00060, whole genome shotgun sequence.<br>\n  c----UGGGCGCGGUGGCGCGCGCCUGUAGUCCCAGCUACUCGGGAGGCU-<br>\n  `</p>\n</blockquote>\n<p>The Multiple Sequence Alignment (MSA) file gives me a bunch of sequences shared between different organisms, due to evolution there are mutations in the DNA and those get reflected in the RNA. We assume that the structure of these RNA will be similar and they clearly share some parts of the sequence. </p>\n<p>Are these files aligned as in maximized to reduce the distance between these conservative regions? If they are not aligned should I use programs like MAFFT or anyone knows which one would be better for non-coding RNAs?</p>",
      "rawMarkdown": "I am currently looking at the MSA files provided in the dataset. For simplicity, I attach some sequences in file '1RY1_E.MSA'. I have no background in this competition and I am trying to understand the MSA files. \n\n`\n>query\nGGGCCGGGCGCGGUGGCGCGCGCCUGUAGUCCCAGCUACUCGGGAGGCUC\n>ANKR01284863.1/22738-23056/2-47 [subseq from] Myotis brandtii contig284863, whole genome shotgun sequence.\n---CCGGGCGCGGUGGCGCGCGCCUGUAGUCCCAGCUACUCGGGAGGCU-\n>DS562880.1/9036401-9036695/2-47 [subseq from] Cavia porcellus supercont2_25 genomic scaffold, whole genome shotgun sequence.\n---CCGGGCGCGGUGGCGCACGCCUGUAGUCCCAGCUACUCGGGAGGCU-\n>GL896957.1/7960321-7960609/2-47 [subseq from] Mustela putorius furo unplaced genomic scaffold scaffold00060, whole genome shotgun sequence.\nc----UGGGCGCGGUGGCGCGCGCCUGUAGUCCCAGCUACUCGGGAGGCU-\n`\n\nThe Multiple Sequence Alignment (MSA) file gives me a bunch of sequences shared between different organisms, due to evolution there are mutations in the DNA and those get reflected in the RNA. We assume that the structure of these RNA will be similar and they clearly share some parts of the sequence. \n\nAre these files aligned as in maximized to reduce the distance between these conservative regions? If they are not aligned should I use programs like MAFFT or anyone knows which one would be better for non-coding RNAs?\n\n\n",
      "votes": 3
    }
  ],
  "comments": [
    {
      "id": 3154809,
      "author_name": "Tilii",
      "author_url": "",
      "post_date": "2025-03-20T12:29:34.460000",
      "content": "<p>These are aligned. The first sequence (<code>query</code>) is the anchor and everything else is written relative to it. That sequence will have all capital letters and no gaps (no <code>-</code> characters). Gaps in subsequent sequences mean that they miss bases at certain positions. If they have extra bases compared to the query sequences, they will be shown as lowercase letters.</p>\n<p>Maybe it will help you to see the same alignment in a different format.</p>\n<pre><code>                 -GGGCCGGGCGCGGUGGCGCGCGCCUGUAGUCCCAGCUACUCGGGAGGCUC\n./    ----CCGGGCGCGGUGGCGCGCGCCUGUAGUCCCAGCUACUCGGGAGGCU-\n./    ----CCGGGCGCGGUGGCGCACGCCUGUAGUCCCAGCUACUCGGGAGGCU-\n./    c----UGGGCGCGGUGGCGCGCGCCUGUAGUCCCAGCUACUCGGGAGGCU-\n</code></pre>\n<p>By the way, that alignment you used as an example is not right, and it should look like this:</p>\n<pre><code>                 GGGCCGGGCGCGGUGGCGCGCGCCUGUAGUCCCAGCUACUCGGGAGGCUC\n./    ---CCGGGCGCGGUGGCGCGCGCCUGUAGUCCCAGCUACUCGGGAGGCU-\n./    ---CCGGGCGCGGUGGCGCACGCCUGUAGUCCCAGCUACUCGGGAGGCU-\n./    ---CUGGGCGCGGUGGCGCGCGCCUGUAGUCCCAGCUACUCGGGAGGCU-\n</code></pre>",
      "votes": 3,
      "replies": [
        {
          "id": 3154968,
          "author_name": "Ignasi Alemany",
          "author_url": "",
          "post_date": "2025-03-20T15:06:24.537000",
          "content": "<p>Thank you so much, I tried using RNA.alifold(sequences) from ViennaRNA package but it seems it cannot deal with sequences of diff length. I have looked it up and it seems that aligned sequences should have the same length? I am confused</p>",
          "votes": 0,
          "replies": [
            {
              "id": 3154975,
              "author_name": "",
              "author_url": "",
              "post_date": "2025-03-20T15:12:45.437000",
              "content": "",
              "votes": 0,
              "replies": []
            },
            {
              "id": 3154980,
              "author_name": "Ignasi Alemany",
              "author_url": "",
              "post_date": "2025-03-20T15:18:52.007000",
              "content": "<p>Ok nevermind I just realised that is a specific format called A2M and ViennaRNA does not support this alignment format! I will have to transform it to Clustal or other supported formats or use other packages,any that you rcommend?</p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 3155370,
              "author_name": "Tilii",
              "author_url": "",
              "post_date": "2025-03-21T01:02:10.863000",
              "content": "<p>There is a file named <code>reformat.pl</code> in <a href=\"https://github.com/soedinglab/hh-suite/tree/master/scripts\" target=\"_blank\"><strong>this directory</strong></a>. It can convert from the A3M format (which is what we have) into any other common alignment format, including Clustal. It is a standalone Perl script that should work on Kaggle.</p>",
              "votes": 3,
              "replies": []
            },
            {
              "id": 3155784,
              "author_name": "Ignasi Alemany",
              "author_url": "",
              "post_date": "2025-03-21T12:06:11.490000",
              "content": "<p>Yes exactly sorry is A3M format! It worked just fine, thank you so much!</p>",
              "votes": 0,
              "replies": []
            }
          ]
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "3154809": "These are aligned. The first sequence (`query`) is the anchor and everything else is written relative to it. That sequence will have all capital letters and no gaps (no `-` characters). Gaps in subsequent sequences mean that they miss bases at certain positions. If they have extra bases compared to the query sequences, they will be shown as lowercase letters.\n\nMaybe it will help you to see the same alignment in a different format.\n\n```\nquery                 -GGGCCGGGCGCGGUGGCGCGCGCCUGUAGUCCCAGCUACUCGGGAGGCUC\nANKR01284863.1/227    ----CCGGGCGCGGUGGCGCGCGCCUGUAGUCCCAGCUACUCGGGAGGCU-\nDS562880.1/9036401    ----CCGGGCGCGGUGGCGCACGCCUGUAGUCCCAGCUACUCGGGAGGCU-\nGL896957.1/7960321    c----UGGGCGCGGUGGCGCGCGCCUGUAGUCCCAGCUACUCGGGAGGCU-\n```\n\nBy the way, that alignment you used as an example is not right, and it should look like this:\n\n```\nquery                 GGGCCGGGCGCGGUGGCGCGCGCCUGUAGUCCCAGCUACUCGGGAGGCUC\nANKR01284863.1/227    ---CCGGGCGCGGUGGCGCGCGCCUGUAGUCCCAGCUACUCGGGAGGCU-\nDS562880.1/9036401    ---CCGGGCGCGGUGGCGCACGCCUGUAGUCCCAGCUACUCGGGAGGCU-\nGL896957.1/7960321    ---CUGGGCGCGGUGGCGCGCGCCUGUAGUCCCAGCUACUCGGGAGGCU-\n```\n",
    "3154770": "I am currently looking at the MSA files provided in the dataset. For simplicity, I attach some sequences in file '1RY1_E.MSA'. I have no background in this competition and I am trying to understand the MSA files. \n\n`\n>query\nGGGCCGGGCGCGGUGGCGCGCGCCUGUAGUCCCAGCUACUCGGGAGGCUC\n>ANKR01284863.1/22738-23056/2-47 [subseq from] Myotis brandtii contig284863, whole genome shotgun sequence.\n---CCGGGCGCGGUGGCGCGCGCCUGUAGUCCCAGCUACUCGGGAGGCU-\n>DS562880.1/9036401-9036695/2-47 [subseq from] Cavia porcellus supercont2_25 genomic scaffold, whole genome shotgun sequence.\n---CCGGGCGCGGUGGCGCACGCCUGUAGUCCCAGCUACUCGGGAGGCU-\n>GL896957.1/7960321-7960609/2-47 [subseq from] Mustela putorius furo unplaced genomic scaffold scaffold00060, whole genome shotgun sequence.\nc----UGGGCGCGGUGGCGCGCGCCUGUAGUCCCAGCUACUCGGGAGGCU-\n`\n\nThe Multiple Sequence Alignment (MSA) file gives me a bunch of sequences shared between different organisms, due to evolution there are mutations in the DNA and those get reflected in the RNA. We assume that the structure of these RNA will be similar and they clearly share some parts of the sequence. \n\nAre these files aligned as in maximized to reduce the distance between these conservative regions? If they are not aligned should I use programs like MAFFT or anyone knows which one would be better for non-coding RNAs?\n\n\n"
  }
}