{
  "id": 568116,
  "title": "Multiple Sequence Alignments (MSAs) for Targets",
  "url": "/competitions/stanford-rna-3d-folding/discussion/568116",
  "author_name": "Alissa Hummer",
  "post_date": "2025-03-13T21:20:18.860000",
  "votes": 20,
  "comment_count": 10,
  "views": 0,
  "content": "<p>Dear Competitors,</p>\n<p>We would like to inform you that we have uploaded multiple sequence alignments (MSAs) for the target sequences. As some of you have been discussing, MSAs have proven very important for previous ML methods for RNA structure prediction.</p>\n<p>The MSAs were generated using rMSA (<a href=\"https://github.com/pylelab/rMSA)\" target=\"_blank\">https://github.com/pylelab/rMSA)</a>. The MSA depth (number of sequences) ranges greatly from only 1 sequence (the query, i.e. target sequence) to &gt;15,000 sequences. There are some targets for which we have provided placeholder MSAs (query only).</p>\n<p>The files are in the A2M fasta format and can be accessed at MSA/{target_id}.MSA.fasta. MSAs will be available using the same file path format (MSA/{target_id}.MSA.fasta) for targets in the test datasets.</p>\n<p>Happy modeling!</p>\n<p>[Updated 14/03/2025 – clarification on A3M fasta format and availability of MSAs for targets in the test dataset]</p>\n<p>[Updated 21/03/2025 – MSAs updated to be the outputs of rMSA, A2M fasta format; correction on greatest MSA depth (number of sequences not number of lines)]</p>",
  "messages": [
    {
      "id": 3149153,
      "postDate": "2025-03-13T21:20:18.860Z",
      "content": "<p>Dear Competitors,</p>\n<p>We would like to inform you that we have uploaded multiple sequence alignments (MSAs) for the target sequences. As some of you have been discussing, MSAs have proven very important for previous ML methods for RNA structure prediction.</p>\n<p>The MSAs were generated using rMSA (<a href=\"https://github.com/pylelab/rMSA)\" target=\"_blank\">https://github.com/pylelab/rMSA)</a>. The MSA depth (number of sequences) ranges greatly from only 1 sequence (the query, i.e. target sequence) to &gt;15,000 sequences. There are some targets for which we have provided placeholder MSAs (query only).</p>\n<p>The files are in the A2M fasta format and can be accessed at MSA/{target_id}.MSA.fasta. MSAs will be available using the same file path format (MSA/{target_id}.MSA.fasta) for targets in the test datasets.</p>\n<p>Happy modeling!</p>\n<p>[Updated 14/03/2025 – clarification on A3M fasta format and availability of MSAs for targets in the test dataset]</p>\n<p>[Updated 21/03/2025 – MSAs updated to be the outputs of rMSA, A2M fasta format; correction on greatest MSA depth (number of sequences not number of lines)]</p>",
      "rawMarkdown": "Dear Competitors,\n \nWe would like to inform you that we have uploaded multiple sequence alignments (MSAs) for the target sequences. As some of you have been discussing, MSAs have proven very important for previous ML methods for RNA structure prediction.\n \nThe MSAs were generated using rMSA (https://github.com/pylelab/rMSA). The MSA depth (number of sequences) ranges greatly from only 1 sequence (the query, i.e. target sequence) to >15,000 sequences. There are some targets for which we have provided placeholder MSAs (query only).\n \nThe files are in the A2M fasta format and can be accessed at MSA/{target_id}.MSA.fasta. MSAs will be available using the same file path format (MSA/{target_id}.MSA.fasta) for targets in the test datasets.\n\nHappy modeling!\n\n\n[Updated 14/03/2025 – clarification on A3M fasta format and availability of MSAs for targets in the test dataset]\n\n[Updated 21/03/2025 – MSAs updated to be the outputs of rMSA, A2M fasta format; correction on greatest MSA depth (number of sequences not number of lines)]",
      "votes": 20
    },
    {
      "id": 3149988,
      "postDate": "2025-03-14T23:56:25.550Z",
      "content": "<p><a href=\"https://www.kaggle.com/alissahummer\" target=\"_blank\">@alissahummer</a> Thank you for providing these files.</p>\n<blockquote>\n  <p>The MSAs were generated using the parameters and data pipeline as used by AlphaFold3.</p>\n</blockquote>\n<p>Where is this data pipeline described? How does it compare to rMSA?</p>\n<p><a href=\"https://github.com/pylelab/rMSA\" target=\"_blank\">https://github.com/pylelab/rMSA</a></p>\n<blockquote>\n  <p>There are some targets for which we have provided placeholder MSAs (query only).</p>\n</blockquote>\n<p>A quick look through the data makes me think that true MSAs were not provided for most targets, as file sizes seem to be at or smaller than 100 bytes. Any idea when the MSA portion of the dataset will be completed?</p>",
      "rawMarkdown": "@alissahummer Thank you for providing these files.\n\n> The MSAs were generated using the parameters and data pipeline as used by AlphaFold3.\n\nWhere is this data pipeline described? How does it compare to rMSA?\n\nhttps://github.com/pylelab/rMSA\n\n> There are some targets for which we have provided placeholder MSAs (query only).\n\nA quick look through the data makes me think that true MSAs were not provided for most targets, as file sizes seem to be at or smaller than 100 bytes. Any idea when the MSA portion of the dataset will be completed?",
      "votes": 1,
      "replies": [
        {
          "id": 3149993,
          "postDate": "2025-03-15T00:02:43.813Z",
          "content": "<p>i am investigating the quality of MSA at  <br>\n <a href=\"https://www.kaggle.com/competitions/stanford-rna-3d-folding/discussion/568125#3149991\" target=\"_blank\">https://www.kaggle.com/competitions/stanford-rna-3d-folding/discussion/568125#3149991</a><br>\nPlease also advice on my investigation as I am a newbie on RNA 3d prediction.<br>\nThanks!</p>",
          "rawMarkdown": "i am investigating the quality of MSA at  \n https://www.kaggle.com/competitions/stanford-rna-3d-folding/discussion/568125#3149991\nPlease also advice on my investigation as I am a newbie on RNA 3d prediction.\nThanks!",
          "replies": [
            {
              "id": 3150166,
              "postDate": "2025-03-15T06:07:18.610Z",
              "content": "<p>i think kaggle MSA did not include DNA.<br>\nbut when I check other opensource, they align both DNA and RNA even if input is RNA.<br>\nThen in preprocessing, they replace T with U.</p>",
              "rawMarkdown": "i think kaggle MSA did not include DNA.\nbut when I check other opensource, they align both DNA and RNA even if input is RNA.\nThen in preprocessing, they replace T with U."
            }
          ]
        }
      ]
    },
    {
      "id": 3149249,
      "postDate": "2025-03-14T01:36:38.733Z",
      "content": "<p>i think the files are in a3m format.<br>\ne.g. i downloaded the file 1EIY_C.MSA.fasta</p>\n<pre><code>&gt;query\nGCCGAGGUAGCUCAGUUGGUAGAGCAUGCGACUGAAAAUCGCAGUGUCCGCGGUUCGAUUCCGCGCCUCGGCACCA\n\n&gt;|AE/-- [subseq from] Zymomonas mobilis subsp. mobilis ZM4, complete genome.\nc--CCAGGUAGCUCAGUCGGUAGAGCAUGUGACUGAAAAUCACAGUGUCGGCGGUUCGAUUCCGUCCCUGGG-----c\n</code></pre>\n<p>i can see small charcters(insert) and dash(delete) . the extension should have been MSA/{target_id}.MSA.a3m.?</p>\n<p>this can be confusing for beginner, although a3m is derived from fasta (and hence is technically fasta format as well)</p>",
      "rawMarkdown": "i think the files are in a3m format.\ne.g. i downloaded the file 1EIY_C.MSA.fasta\n\n```\n>query\nGCCGAGGUAGCUCAGUUGGUAGAGCAUGCGACUGAAAAUCGCAGUGUCCGCGGUUCGAUUCCGCGCCUCGGCACCA\n\n>003|AE008692.2/1028696-1028624/2-72 [subseq from] Zymomonas mobilis subsp. mobilis ZM4, complete genome.\nc--CCAGGUAGCUCAGUCGGUAGAGCAUGUGACUGAAAAUCACAGUGUCGGCGGUUCGAUUCCGUCCCUGGG-----c\n\n```\n\ni can see small charcters(insert) and dash(delete) . the extension should have been MSA/{target_id}.MSA.a3m.?\n\nthis can be confusing for beginner, although a3m is derived from fasta (and hence is technically fasta format as well)",
      "votes": 2,
      "replies": [
        {
          "id": 3149743,
          "postDate": "2025-03-14T16:26:41.410Z",
          "content": "<p>Hi <a href=\"https://www.kaggle.com/hengck23\" target=\"_blank\">@hengck23</a>, thank you for pointing this out. I have updated the post to clarify that the files are in the A3M fasta format.</p>",
          "rawMarkdown": "Hi @hengck23, thank you for pointing this out. I have updated the post to clarify that the files are in the A3M fasta format.",
          "votes": 2
        },
        {
          "id": 3158621,
          "postDate": "2025-03-24T17:53:11.400Z",
          "rawMarkdown": "",
          "isDeleted": true
        }
      ]
    },
    {
      "id": 3203522,
      "postDate": "2025-05-16T21:42:57.517Z",
      "content": "<p><a href=\"https://www.kaggle.com/alissahummer\" target=\"_blank\">@alissahummer</a> Thank you for the MSA files. </p>\n<p>For sample file MSA_v2/1A51_A.MSA.fasta</p>\n<pre><code>&gt;\nGGCCGAUGGUAGUGUGGGGUCUCCCCAUGCGAGAGUAGGCC\n&gt;CP018029._f/-\nGGCCGAUGGUAGUGUGGGGUUUCCCCAUGUGAGAGUAGGAC\n...\n</code></pre>\n<p>Is the following information about CP018029.1_364898_365020_f/42-82 correct?  I got the following from claude.ai:</p>\n<p><strong>CP018029.1</strong>: This is a GenBank/RefSeq accession number for a genome or genomic region. The \".1\" indicates it's version 1 of this sequence.<br>\n<strong>364898_365020</strong>: These numbers represent genomic coordinates, likely the start and end positions (364,898 to 365,020) of this sequence in the reference genome.<br>\n<strong>f</strong>: This likely indicates the strand orientation, with \"f\" representing the forward strand (as opposed to \"r\" which would be the reverse strand).<br>\n<strong>/42-82</strong>: This part likely specifies the subsequence coordinates within the aligned segment that are being represented in the MSA. This means positions 42 through 82 of the original sequence are what's included in this alignment.</p>\n<p>This format is commonly used in MSA files to provide detailed information about each sequence's origin and position, which is important for phylogenetic analysis and studying evolutionary relationships between sequences.</p>",
      "rawMarkdown": "@alissahummer Thank you for the MSA files. \n\nFor sample file MSA_v2/1A51_A.MSA.fasta\n```\n>query\nGGCCGAUGGUAGUGUGGGGUCUCCCCAUGCGAGAGUAGGCC\n>CP018029.1_364898_365020_f/42-82\nGGCCGAUGGUAGUGUGGGGUUUCCCCAUGUGAGAGUAGGAC\n...\n```\n\nIs the following information about CP018029.1_364898_365020_f/42-82 correct?  I got the following from claude.ai:\n\n**CP018029.1**: This is a GenBank/RefSeq accession number for a genome or genomic region. The \".1\" indicates it's version 1 of this sequence.\n**364898_365020**: These numbers represent genomic coordinates, likely the start and end positions (364,898 to 365,020) of this sequence in the reference genome.\n**f**: This likely indicates the strand orientation, with \"f\" representing the forward strand (as opposed to \"r\" which would be the reverse strand).\n**/42-82**: This part likely specifies the subsequence coordinates within the aligned segment that are being represented in the MSA. This means positions 42 through 82 of the original sequence are what's included in this alignment.\n\nThis format is commonly used in MSA files to provide detailed information about each sequence's origin and position, which is important for phylogenetic analysis and studying evolutionary relationships between sequences.\n"
    },
    {
      "id": 3149468,
      "postDate": "2025-03-14T08:58:52.050Z",
      "content": "<p>Thanks! Will we have MSAs available for all new test samples on inference phase?<br>\n<a href=\"https://www.kaggle.com/alissahummer\" target=\"_blank\">@alissahummer</a> please make it more clear (maybe by putting train/test MSAs in subfolders where public test MSAs will be substituted by public/private LB test sets).</p>",
      "rawMarkdown": "Thanks! Will we have MSAs available for all new test samples on inference phase?\n@alissahummer please make it more clear (maybe by putting train/test MSAs in subfolders where public test MSAs will be substituted by public/private LB test sets).",
      "replies": [
        {
          "id": 3149746,
          "postDate": "2025-03-14T16:27:44.663Z",
          "content": "<p>Hi <a href=\"https://www.kaggle.com/ogurtsov\" target=\"_blank\">@ogurtsov</a>, thank you for asking! MSAs will be available using the same file path format (MSA/{target_id}.MSA.fasta) for targets in the test datasets. I have updated the post to clarify, but please let me know if anything is unclear!</p>",
          "rawMarkdown": "Hi @ogurtsov, thank you for asking! MSAs will be available using the same file path format (MSA/{target_id}.MSA.fasta) for targets in the test datasets. I have updated the post to clarify, but please let me know if anything is unclear!",
          "votes": 4
        }
      ]
    },
    {
      "id": 3149167,
      "postDate": "2025-03-13T21:50:22.737Z",
      "content": "<p>Thanks a lot!</p>",
      "rawMarkdown": "Thanks a lot!"
    }
  ],
  "comments": [
    {
      "id": 3149988,
      "author_name": "Tilii",
      "author_url": "",
      "post_date": "2025-03-14T23:56:25.550000",
      "content": "<p><a href=\"https://www.kaggle.com/alissahummer\" target=\"_blank\">@alissahummer</a> Thank you for providing these files.</p>\n<blockquote>\n  <p>The MSAs were generated using the parameters and data pipeline as used by AlphaFold3.</p>\n</blockquote>\n<p>Where is this data pipeline described? How does it compare to rMSA?</p>\n<p><a href=\"https://github.com/pylelab/rMSA\" target=\"_blank\">https://github.com/pylelab/rMSA</a></p>\n<blockquote>\n  <p>There are some targets for which we have provided placeholder MSAs (query only).</p>\n</blockquote>\n<p>A quick look through the data makes me think that true MSAs were not provided for most targets, as file sizes seem to be at or smaller than 100 bytes. Any idea when the MSA portion of the dataset will be completed?</p>",
      "votes": 1,
      "replies": [
        {
          "id": 3149993,
          "author_name": "hengck23",
          "author_url": "",
          "post_date": "2025-03-15T00:02:43.813000",
          "content": "<p>i am investigating the quality of MSA at  <br>\n <a href=\"https://www.kaggle.com/competitions/stanford-rna-3d-folding/discussion/568125#3149991\" target=\"_blank\">https://www.kaggle.com/competitions/stanford-rna-3d-folding/discussion/568125#3149991</a><br>\nPlease also advice on my investigation as I am a newbie on RNA 3d prediction.<br>\nThanks!</p>",
          "votes": 0,
          "replies": [
            {
              "id": 3150166,
              "author_name": "hengck23",
              "author_url": "",
              "post_date": "2025-03-15T06:07:18.610000",
              "content": "<p>i think kaggle MSA did not include DNA.<br>\nbut when I check other opensource, they align both DNA and RNA even if input is RNA.<br>\nThen in preprocessing, they replace T with U.</p>",
              "votes": 0,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 3149249,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "2025-03-14T01:36:38.733000",
      "content": "<p>i think the files are in a3m format.<br>\ne.g. i downloaded the file 1EIY_C.MSA.fasta</p>\n<pre><code>&gt;query\nGCCGAGGUAGCUCAGUUGGUAGAGCAUGCGACUGAAAAUCGCAGUGUCCGCGGUUCGAUUCCGCGCCUCGGCACCA\n\n&gt;|AE/-- [subseq from] Zymomonas mobilis subsp. mobilis ZM4, complete genome.\nc--CCAGGUAGCUCAGUCGGUAGAGCAUGUGACUGAAAAUCACAGUGUCGGCGGUUCGAUUCCGUCCCUGGG-----c\n</code></pre>\n<p>i can see small charcters(insert) and dash(delete) . the extension should have been MSA/{target_id}.MSA.a3m.?</p>\n<p>this can be confusing for beginner, although a3m is derived from fasta (and hence is technically fasta format as well)</p>",
      "votes": 2,
      "replies": [
        {
          "id": 3149743,
          "author_name": "Alissa Hummer",
          "author_url": "",
          "post_date": "2025-03-14T16:26:41.410000",
          "content": "<p>Hi <a href=\"https://www.kaggle.com/hengck23\" target=\"_blank\">@hengck23</a>, thank you for pointing this out. I have updated the post to clarify that the files are in the A3M fasta format.</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 3158621,
          "author_name": "",
          "author_url": "",
          "post_date": "2025-03-24T17:53:11.400000",
          "content": "",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 3203522,
      "author_name": "Young",
      "author_url": "",
      "post_date": "2025-05-16T21:42:57.517000",
      "content": "<p><a href=\"https://www.kaggle.com/alissahummer\" target=\"_blank\">@alissahummer</a> Thank you for the MSA files. </p>\n<p>For sample file MSA_v2/1A51_A.MSA.fasta</p>\n<pre><code>&gt;\nGGCCGAUGGUAGUGUGGGGUCUCCCCAUGCGAGAGUAGGCC\n&gt;CP018029._f/-\nGGCCGAUGGUAGUGUGGGGUUUCCCCAUGUGAGAGUAGGAC\n...\n</code></pre>\n<p>Is the following information about CP018029.1_364898_365020_f/42-82 correct?  I got the following from claude.ai:</p>\n<p><strong>CP018029.1</strong>: This is a GenBank/RefSeq accession number for a genome or genomic region. The \".1\" indicates it's version 1 of this sequence.<br>\n<strong>364898_365020</strong>: These numbers represent genomic coordinates, likely the start and end positions (364,898 to 365,020) of this sequence in the reference genome.<br>\n<strong>f</strong>: This likely indicates the strand orientation, with \"f\" representing the forward strand (as opposed to \"r\" which would be the reverse strand).<br>\n<strong>/42-82</strong>: This part likely specifies the subsequence coordinates within the aligned segment that are being represented in the MSA. This means positions 42 through 82 of the original sequence are what's included in this alignment.</p>\n<p>This format is commonly used in MSA files to provide detailed information about each sequence's origin and position, which is important for phylogenetic analysis and studying evolutionary relationships between sequences.</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 3149468,
      "author_name": "Ogurtsov",
      "author_url": "",
      "post_date": "2025-03-14T08:58:52.050000",
      "content": "<p>Thanks! Will we have MSAs available for all new test samples on inference phase?<br>\n<a href=\"https://www.kaggle.com/alissahummer\" target=\"_blank\">@alissahummer</a> please make it more clear (maybe by putting train/test MSAs in subfolders where public test MSAs will be substituted by public/private LB test sets).</p>",
      "votes": 0,
      "replies": [
        {
          "id": 3149746,
          "author_name": "Alissa Hummer",
          "author_url": "",
          "post_date": "2025-03-14T16:27:44.663000",
          "content": "<p>Hi <a href=\"https://www.kaggle.com/ogurtsov\" target=\"_blank\">@ogurtsov</a>, thank you for asking! MSAs will be available using the same file path format (MSA/{target_id}.MSA.fasta) for targets in the test datasets. I have updated the post to clarify, but please let me know if anything is unclear!</p>",
          "votes": 4,
          "replies": []
        }
      ]
    },
    {
      "id": 3149167,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "2025-03-13T21:50:22.737000",
      "content": "<p>Thanks a lot!</p>",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "3149153": "Dear Competitors,\n \nWe would like to inform you that we have uploaded multiple sequence alignments (MSAs) for the target sequences. As some of you have been discussing, MSAs have proven very important for previous ML methods for RNA structure prediction.\n \nThe MSAs were generated using rMSA (https://github.com/pylelab/rMSA). The MSA depth (number of sequences) ranges greatly from only 1 sequence (the query, i.e. target sequence) to >15,000 sequences. There are some targets for which we have provided placeholder MSAs (query only).\n \nThe files are in the A2M fasta format and can be accessed at MSA/{target_id}.MSA.fasta. MSAs will be available using the same file path format (MSA/{target_id}.MSA.fasta) for targets in the test datasets.\n\nHappy modeling!\n\n\n[Updated 14/03/2025 – clarification on A3M fasta format and availability of MSAs for targets in the test dataset]\n\n[Updated 21/03/2025 – MSAs updated to be the outputs of rMSA, A2M fasta format; correction on greatest MSA depth (number of sequences not number of lines)]",
    "3149988": "@alissahummer Thank you for providing these files.\n\n> The MSAs were generated using the parameters and data pipeline as used by AlphaFold3.\n\nWhere is this data pipeline described? How does it compare to rMSA?\n\nhttps://github.com/pylelab/rMSA\n\n> There are some targets for which we have provided placeholder MSAs (query only).\n\nA quick look through the data makes me think that true MSAs were not provided for most targets, as file sizes seem to be at or smaller than 100 bytes. Any idea when the MSA portion of the dataset will be completed?",
    "3149249": "i think the files are in a3m format.\ne.g. i downloaded the file 1EIY_C.MSA.fasta\n\n```\n>query\nGCCGAGGUAGCUCAGUUGGUAGAGCAUGCGACUGAAAAUCGCAGUGUCCGCGGUUCGAUUCCGCGCCUCGGCACCA\n\n>003|AE008692.2/1028696-1028624/2-72 [subseq from] Zymomonas mobilis subsp. mobilis ZM4, complete genome.\nc--CCAGGUAGCUCAGUCGGUAGAGCAUGUGACUGAAAAUCACAGUGUCGGCGGUUCGAUUCCGUCCCUGGG-----c\n\n```\n\ni can see small charcters(insert) and dash(delete) . the extension should have been MSA/{target_id}.MSA.a3m.?\n\nthis can be confusing for beginner, although a3m is derived from fasta (and hence is technically fasta format as well)",
    "3203522": "@alissahummer Thank you for the MSA files. \n\nFor sample file MSA_v2/1A51_A.MSA.fasta\n```\n>query\nGGCCGAUGGUAGUGUGGGGUCUCCCCAUGCGAGAGUAGGCC\n>CP018029.1_364898_365020_f/42-82\nGGCCGAUGGUAGUGUGGGGUUUCCCCAUGUGAGAGUAGGAC\n...\n```\n\nIs the following information about CP018029.1_364898_365020_f/42-82 correct?  I got the following from claude.ai:\n\n**CP018029.1**: This is a GenBank/RefSeq accession number for a genome or genomic region. The \".1\" indicates it's version 1 of this sequence.\n**364898_365020**: These numbers represent genomic coordinates, likely the start and end positions (364,898 to 365,020) of this sequence in the reference genome.\n**f**: This likely indicates the strand orientation, with \"f\" representing the forward strand (as opposed to \"r\" which would be the reverse strand).\n**/42-82**: This part likely specifies the subsequence coordinates within the aligned segment that are being represented in the MSA. This means positions 42 through 82 of the original sequence are what's included in this alignment.\n\nThis format is commonly used in MSA files to provide detailed information about each sequence's origin and position, which is important for phylogenetic analysis and studying evolutionary relationships between sequences.\n",
    "3149468": "Thanks! Will we have MSAs available for all new test samples on inference phase?\n@alissahummer please make it more clear (maybe by putting train/test MSAs in subfolders where public test MSAs will be substituted by public/private LB test sets).",
    "3149167": "Thanks a lot!"
  }
}