{
  "id": 569424,
  "title": "MSA Update",
  "url": "/competitions/stanford-rna-3d-folding/discussion/569424",
  "author_name": "Alissa Hummer",
  "post_date": "2025-03-21T18:49:39.197000",
  "votes": 17,
  "comment_count": 23,
  "views": 0,
  "content": "<p>Hi Kagglers,</p>\n<p>Just a quick note that we have updated the multiple sequence alignment (MSA) files to outputs generated using rMSA (<a href=\"https://github.com/pylelab/rMSA)\" target=\"_blank\">https://github.com/pylelab/rMSA)</a>. This method produced deeper MSAs and we used databases with a temporal cutoff more closely aligned with the validation dataset (database versions included below). The rMSA outputs are in the A2M format.</p>\n<p>We have retained the MSA/{target_id}.MSA.fasta file naming.</p>\n<p>rMSA Database Versions/Cutoffs<br>\n• RNAcentral v20.0 (2022-03-28)<br>\n• RFam v14.7 (2021-12-09)<br>\n• NCBI NT (2022-10-03)</p>\n<p>[Updated 20250321 - A2M format clarified]<br>\n[Updated 20250325 - all placeholders have been replaced with rMSA outputs]</p>",
  "messages": [
    {
      "id": 3156129,
      "postDate": "2025-03-21T18:49:39.197Z",
      "content": "<p>Hi Kagglers,</p>\n<p>Just a quick note that we have updated the multiple sequence alignment (MSA) files to outputs generated using rMSA (<a href=\"https://github.com/pylelab/rMSA)\" target=\"_blank\">https://github.com/pylelab/rMSA)</a>. This method produced deeper MSAs and we used databases with a temporal cutoff more closely aligned with the validation dataset (database versions included below). The rMSA outputs are in the A2M format.</p>\n<p>We have retained the MSA/{target_id}.MSA.fasta file naming.</p>\n<p>rMSA Database Versions/Cutoffs<br>\n• RNAcentral v20.0 (2022-03-28)<br>\n• RFam v14.7 (2021-12-09)<br>\n• NCBI NT (2022-10-03)</p>\n<p>[Updated 20250321 - A2M format clarified]<br>\n[Updated 20250325 - all placeholders have been replaced with rMSA outputs]</p>",
      "rawMarkdown": "Hi Kagglers,\n\nJust a quick note that we have updated the multiple sequence alignment (MSA) files to outputs generated using rMSA (https://github.com/pylelab/rMSA). This method produced deeper MSAs and we used databases with a temporal cutoff more closely aligned with the validation dataset (database versions included below). The rMSA outputs are in the A2M format.\n\nWe have retained the MSA/{target_id}.MSA.fasta file naming.\n\nrMSA Database Versions/Cutoffs\n• RNAcentral v20.0 (2022-03-28)\n• RFam v14.7 (2021-12-09)\n• NCBI NT (2022-10-03)\n\n[Updated 20250321 - A2M format clarified]\n[Updated 20250325 - all placeholders have been replaced with rMSA outputs]",
      "votes": 17
    },
    {
      "id": 3156294,
      "postDate": "2025-03-22T02:05:00.203Z",
      "content": "<p>for the newbies:</p>\n<ol>\n<li>A2M: Uses uppercase characters for matches, lowercase for insertions (with gaps aligned to insertions shown as periods), and dashes (-) for deletions. </li>\n<li>A3M: Similar to A2M, but allows gaps aligned to insertions to be omitted(periods), making the format more compact. </li>\n</ol>",
      "rawMarkdown": "for the newbies:\n\n1. A2M: Uses uppercase characters for matches, lowercase for insertions (with gaps aligned to insertions shown as periods), and dashes (-) for deletions. \n1. A3M: Similar to A2M, but allows gaps aligned to insertions to be omitted(periods), making the format more compact. ",
      "votes": 3
    },
    {
      "id": 3156307,
      "postDate": "2025-03-22T03:03:31.283Z",
      "content": "<p><a href=\"https://www.kaggle.com/alissahummer\" target=\"_blank\">@alissahummer</a> <br>\nDo u think u can open source your  rmsa script to make the msa?</p>\n<p>Our team will be making external data and share with the kaggle community. We will be publishing eg af3 3d prediction. We thinking if we want to publish msa of exeternal data as well </p>",
      "rawMarkdown": "@alissahummer \nDo u think u can open source your  rmsa script to make the msa?\n\nOur team will be making external data and share with the kaggle community. We will be publishing eg af3 3d prediction. We thinking if we want to publish msa of exeternal data as well ",
      "votes": 1,
      "replies": [
        {
          "id": 3156312,
          "postDate": "2025-03-22T03:12:19.210Z",
          "content": "<p>Hi <a href=\"https://www.kaggle.com/hengck23\" target=\"_blank\">@hengck23</a>, I am reposting my response to Zacchaeus' question below (posted around the same time as yours):</p>\n<p>We used the rMSA tool with default parameters and database versions/dates as described above. I would recommend looking at the rMSA GitHub repo (<a href=\"https://github.com/pylelab/rMSA)\" target=\"_blank\">https://github.com/pylelab/rMSA)</a>, which has helpful information for installation, database downloads, and MSA generation. To generate the MSAs, we ran <code>./rMSA.pl seq.fasta</code>, as on their GitHub README. We updated the outputs to replace \"T\" nucleotides with \"U\". One thing to be aware of is that the databases for MSA generation are very large (~2 TB) and can take a long time to download/set up (on the order of days).</p>",
          "rawMarkdown": "Hi @hengck23, I am reposting my response to Zacchaeus' question below (posted around the same time as yours):\n\nWe used the rMSA tool with default parameters and database versions/dates as described above. I would recommend looking at the rMSA GitHub repo (https://github.com/pylelab/rMSA), which has helpful information for installation, database downloads, and MSA generation. To generate the MSAs, we ran `./rMSA.pl seq.fasta`, as on their GitHub README. We updated the outputs to replace \"T\" nucleotides with \"U\". One thing to be aware of is that the databases for MSA generation are very large (~2 TB) and can take a long time to download/set up (on the order of days).",
          "votes": 2,
          "replies": [
            {
              "id": 3156381,
              "postDate": "2025-03-22T04:30:30.913Z",
              "content": "<p>yes, i spend i week download then realise that orginal NT server ftp.ncbi.nlm.nih.gov/blast/db/ is slow (and downloaded file corrupted).<br>\nIt is better to look for mirrors in your region (use google)</p>",
              "rawMarkdown": "yes, i spend i week download then realise that orginal NT server ftp.ncbi.nlm.nih.gov/blast/db/ is slow (and downloaded file corrupted).\nIt is better to look for mirrors in your region (use google)"
            },
            {
              "id": 3156659,
              "postDate": "2025-03-22T11:40:17.183Z",
              "content": "<p>nt database is not mandatory, for example, RhoFold was trained with MSAs computed on rnacentral+rfam only. </p>\n<p>And it's straightforward to patch <code>./rMSA.pl</code> this way:</p>\n<blockquote>\n  <p>my $db0    =\"$dbdir/Rfam.cm\";<br>\n  my $db1    =\"$dbdir/rnacentral.fasta\";<br>\n  my $db2    =\"$dbdir/rnacentral.fasta\";<br>\n  my $db0to1 =\"$dbdir/rfam_annotations.tsv.gz\";<br>\n  my $db0to2 =\"$dbdir/rfam_annotations.tsv.gz\";</p>\n</blockquote>",
              "rawMarkdown": "nt database is not mandatory, for example, RhoFold was trained with MSAs computed on rnacentral+rfam only. \n\nAnd it's straightforward to patch `./rMSA.pl` this way:\n\n>my $db0    =\"$dbdir/Rfam.cm\";\nmy $db1    =\"$dbdir/rnacentral.fasta\";\nmy $db2    =\"$dbdir/rnacentral.fasta\";\nmy $db0to1 =\"$dbdir/rfam_annotations.tsv.gz\";\nmy $db0to2 =\"$dbdir/rfam_annotations.tsv.gz\";",
              "votes": 3
            },
            {
              "id": 3157017,
              "postDate": "2025-03-22T20:49:00.213Z",
              "content": "<blockquote>\n  <p>nt database is not mandatory</p>\n</blockquote>\n<p>Exactly. Not only that, but it is a very redundant database (contrary to its name). Sequences that have &gt;80% identity don't help very much in co-variance efforts as they are too similar to each other.</p>",
              "rawMarkdown": "> nt database is not mandatory\n\nExactly. Not only that, but it is a very redundant database (contrary to its name). Sequences that have >80% identity don't help very much in co-variance efforts as they are too similar to each other.",
              "votes": 4
            }
          ]
        }
      ]
    },
    {
      "id": 3156652,
      "postDate": "2025-03-22T11:37:48.013Z",
      "content": "<p>It might be worth considering faster alternatives like JackHMMER and MMseqs2.</p>",
      "rawMarkdown": "It might be worth considering faster alternatives like JackHMMER and MMseqs2.\n",
      "replies": [
        {
          "id": 3158246,
          "postDate": "2025-03-24T11:10:17.640Z",
          "content": "<p>Gpu mmseqs2 is very fast</p>",
          "rawMarkdown": "Gpu mmseqs2 is very fast",
          "votes": 2
        }
      ]
    },
    {
      "id": 3156608,
      "postDate": "2025-03-22T10:23:11.537Z",
      "content": "<p>Hello,<br>\nIt's showing error 404 after accessing the link.<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F25371292%2F25ea360174dda13e475147a74ad155d1%2FScreenshot%202025-03-22%20155234.png?generation=1742638970630946&amp;alt=media\" alt=\"SS\"></p>",
      "rawMarkdown": "Hello,\nIt's showing error 404 after accessing the link.![SS](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F25371292%2F25ea360174dda13e475147a74ad155d1%2FScreenshot%202025-03-22%20155234.png?generation=1742638970630946&alt=media)",
      "replies": [
        {
          "id": 3156888,
          "postDate": "2025-03-22T17:18:05.060Z",
          "content": "<p>Hi <a href=\"https://www.kaggle.com/prnvpwr2612\" target=\"_blank\">@prnvpwr2612</a>, it looks like the URL you tried includes the \")\" at the end. Could you try reloading the page without the \")\"?</p>\n<p><a href=\"https://github.com/pylelab/rMSA\" target=\"_blank\">https://github.com/pylelab/rMSA</a></p>",
          "rawMarkdown": "Hi @prnvpwr2612, it looks like the URL you tried includes the \")\" at the end. Could you try reloading the page without the \")\"?\n\nhttps://github.com/pylelab/rMSA",
          "votes": 1,
          "replies": [
            {
              "id": 3157042,
              "postDate": "2025-03-22T21:40:15.507Z",
              "content": "<p>Thanks a lot it is working now.</p>",
              "rawMarkdown": "Thanks a lot it is working now."
            }
          ]
        }
      ]
    },
    {
      "id": 3156305,
      "postDate": "2025-03-22T02:59:48.620Z",
      "content": "<p>Would you mind sharing the pipeline generating MSAs? It could be helpful that we adopt same pipeline in our experiments.</p>",
      "rawMarkdown": "Would you mind sharing the pipeline generating MSAs? It could be helpful that we adopt same pipeline in our experiments.",
      "replies": [
        {
          "id": 3156310,
          "postDate": "2025-03-22T03:08:38.947Z",
          "content": "<p>Hi <a href=\"https://www.kaggle.com/zacchaeus\" target=\"_blank\">@zacchaeus</a>, we used the rMSA tool with default parameters and database versions/dates as described above. I would recommend looking at the rMSA GitHub repo (<a href=\"https://github.com/pylelab/rMSA)\" target=\"_blank\">https://github.com/pylelab/rMSA)</a>, which has helpful information for installation, database downloads, and MSA generation. To generate the MSAs, we ran <code>./rMSA.pl seq.fasta</code> as on their GitHub README. We updated the outputs to replace \"T\" nucleotides with \"U\". One thing to be aware of is that the databases for MSA generation are very large (~2 TB) and can take a long time to download/set up (on the order of days).</p>",
          "rawMarkdown": "Hi @zacchaeus, we used the rMSA tool with default parameters and database versions/dates as described above. I would recommend looking at the rMSA GitHub repo (https://github.com/pylelab/rMSA), which has helpful information for installation, database downloads, and MSA generation. To generate the MSAs, we ran `./rMSA.pl seq.fasta` as on their GitHub README. We updated the outputs to replace \"T\" nucleotides with \"U\". One thing to be aware of is that the databases for MSA generation are very large (~2 TB) and can take a long time to download/set up (on the order of days).",
          "votes": 1
        }
      ]
    },
    {
      "id": 3156271,
      "postDate": "2025-03-22T00:38:34.077Z",
      "content": "<p><a href=\"https://www.kaggle.com/alissahummer\" target=\"_blank\">@alissahummer</a> Good to have these files, thanks.</p>\n<p>It seems that we now have straight-up aligned FASTA files, with no insert states in lowercase letters. If so, that should be specified for those who have scripts expecting A3M files. I don't think what you did with this update fulfills the criterion <code>so that all notebooks that use the files are compatible with the update, without any changes required.</code></p>\n<p>Same question as before: when do you expect that all alignments will be available?</p>",
      "rawMarkdown": "@alissahummer Good to have these files, thanks.\n\nIt seems that we now have straight-up aligned FASTA files, with no insert states in lowercase letters. If so, that should be specified for those who have scripts expecting A3M files. I don't think what you did with this update fulfills the criterion `so that all notebooks that use the files are compatible with the update, without any changes required.`\n\nSame question as before: when do you expect that all alignments will be available?",
      "replies": [
        {
          "id": 3156292,
          "postDate": "2025-03-22T01:58:52.893Z",
          "content": "<p>Hi <a href=\"https://www.kaggle.com/tilii7\" target=\"_blank\">@tilii7</a>, thank you for pointing this out and my apologies for any confusion. I have updated the post to clarify the A2M format.</p>\n<p>We hope to have all alignments available next week, but please note that this timeline might change.</p>",
          "rawMarkdown": "Hi @tilii7, thank you for pointing this out and my apologies for any confusion. I have updated the post to clarify the A2M format.\n\nWe hope to have all alignments available next week, but please note that this timeline might change.",
          "votes": 1
        },
        {
          "id": 3159516,
          "postDate": "2025-03-25T17:13:26.750Z",
          "content": "<p>Hi <a href=\"https://www.kaggle.com/tilii7\" target=\"_blank\">@tilii7</a>, I wanted to let you know that all placeholders have now been updated with the rMSA outputs.</p>",
          "rawMarkdown": "Hi @tilii7, I wanted to let you know that all placeholders have now been updated with the rMSA outputs.",
          "votes": 2,
          "replies": [
            {
              "id": 3159525,
              "postDate": "2025-03-25T17:30:11.743Z",
              "content": "<p>Thank you, that was fast.</p>\n<p>For some of short RNAs there are still no additional sequences, and that can't be because they don't have homologs in the database (<em>e.g.,</em> 1AFX and 1ESH). I suspect it is because the search procedure is not optimized to work with very short sequences. I will try to generate MSAs for them on my own and will share them if there is a global solution.</p>",
              "rawMarkdown": "Thank you, that was fast.\n\nFor some of short RNAs there are still no additional sequences, and that can't be because they don't have homologs in the database (*e.g.,* 1AFX and 1ESH). I suspect it is because the search procedure is not optimized to work with very short sequences. I will try to generate MSAs for them on my own and will share them if there is a global solution.",
              "votes": 2
            }
          ]
        }
      ]
    },
    {
      "id": 3156228,
      "postDate": "2025-03-21T22:55:47.617Z",
      "content": "<p>Thanks! By the way, what does it mean for some query-only files to be \"placeholders\"? Would these be the ones in the test spplit and they are then replaced with the real MSAs?</p>",
      "rawMarkdown": "Thanks! By the way, what does it mean for some query-only files to be \"placeholders\"? Would these be the ones in the test spplit and they are then replaced with the real MSAs?",
      "replies": [
        {
          "id": 3156287,
          "postDate": "2025-03-22T01:41:41.033Z",
          "content": "<p>Hi <a href=\"https://www.kaggle.com/asarvazyan\" target=\"_blank\">@asarvazyan</a>, thanks for the question! The placeholder MSAs are only for some targets in the training dataset, where MSA generation is taking a long time to complete. We did not want to hold up the release of the majority of the MSAs waiting for these. We plan to update the placeholders with the rMSA outputs once the jobs have finished.</p>",
          "rawMarkdown": "Hi @asarvazyan, thanks for the question! The placeholder MSAs are only for some targets in the training dataset, where MSA generation is taking a long time to complete. We did not want to hold up the release of the majority of the MSAs waiting for these. We plan to update the placeholders with the rMSA outputs once the jobs have finished.",
          "replies": [
            {
              "id": 3156293,
              "postDate": "2025-03-22T02:03:07.720Z",
              "content": "<p>Got it, thanks!</p>",
              "rawMarkdown": "Got it, thanks!"
            }
          ]
        }
      ]
    },
    {
      "id": 3158169,
      "postDate": "2025-03-24T09:30:45.823Z",
      "rawMarkdown": "",
      "isDeleted": true
    },
    {
      "id": 3156226,
      "postDate": "2025-03-21T22:43:48.767Z",
      "content": "<p>thanks a lot again!</p>",
      "rawMarkdown": "thanks a lot again!",
      "votes": 1
    },
    {
      "id": 3158337,
      "postDate": "2025-03-24T13:09:34.423Z",
      "content": "<p>Thanks for the update </p>",
      "rawMarkdown": "Thanks for the update "
    }
  ],
  "comments": [
    {
      "id": 3156294,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "2025-03-22T02:05:00.203000",
      "content": "<p>for the newbies:</p>\n<ol>\n<li>A2M: Uses uppercase characters for matches, lowercase for insertions (with gaps aligned to insertions shown as periods), and dashes (-) for deletions. </li>\n<li>A3M: Similar to A2M, but allows gaps aligned to insertions to be omitted(periods), making the format more compact. </li>\n</ol>",
      "votes": 3,
      "replies": []
    },
    {
      "id": 3156307,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "2025-03-22T03:03:31.283000",
      "content": "<p><a href=\"https://www.kaggle.com/alissahummer\" target=\"_blank\">@alissahummer</a> <br>\nDo u think u can open source your  rmsa script to make the msa?</p>\n<p>Our team will be making external data and share with the kaggle community. We will be publishing eg af3 3d prediction. We thinking if we want to publish msa of exeternal data as well </p>",
      "votes": 1,
      "replies": [
        {
          "id": 3156312,
          "author_name": "Alissa Hummer",
          "author_url": "",
          "post_date": "2025-03-22T03:12:19.210000",
          "content": "<p>Hi <a href=\"https://www.kaggle.com/hengck23\" target=\"_blank\">@hengck23</a>, I am reposting my response to Zacchaeus' question below (posted around the same time as yours):</p>\n<p>We used the rMSA tool with default parameters and database versions/dates as described above. I would recommend looking at the rMSA GitHub repo (<a href=\"https://github.com/pylelab/rMSA)\" target=\"_blank\">https://github.com/pylelab/rMSA)</a>, which has helpful information for installation, database downloads, and MSA generation. To generate the MSAs, we ran <code>./rMSA.pl seq.fasta</code>, as on their GitHub README. We updated the outputs to replace \"T\" nucleotides with \"U\". One thing to be aware of is that the databases for MSA generation are very large (~2 TB) and can take a long time to download/set up (on the order of days).</p>",
          "votes": 2,
          "replies": [
            {
              "id": 3156381,
              "author_name": "hengck23",
              "author_url": "",
              "post_date": "2025-03-22T04:30:30.913000",
              "content": "<p>yes, i spend i week download then realise that orginal NT server ftp.ncbi.nlm.nih.gov/blast/db/ is slow (and downloaded file corrupted).<br>\nIt is better to look for mirrors in your region (use google)</p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 3156659,
              "author_name": "Ogurtsov",
              "author_url": "",
              "post_date": "2025-03-22T11:40:17.183000",
              "content": "<p>nt database is not mandatory, for example, RhoFold was trained with MSAs computed on rnacentral+rfam only. </p>\n<p>And it's straightforward to patch <code>./rMSA.pl</code> this way:</p>\n<blockquote>\n  <p>my $db0    =\"$dbdir/Rfam.cm\";<br>\n  my $db1    =\"$dbdir/rnacentral.fasta\";<br>\n  my $db2    =\"$dbdir/rnacentral.fasta\";<br>\n  my $db0to1 =\"$dbdir/rfam_annotations.tsv.gz\";<br>\n  my $db0to2 =\"$dbdir/rfam_annotations.tsv.gz\";</p>\n</blockquote>",
              "votes": 3,
              "replies": []
            },
            {
              "id": 3157017,
              "author_name": "Tilii",
              "author_url": "",
              "post_date": "2025-03-22T20:49:00.213000",
              "content": "<blockquote>\n  <p>nt database is not mandatory</p>\n</blockquote>\n<p>Exactly. Not only that, but it is a very redundant database (contrary to its name). Sequences that have &gt;80% identity don't help very much in co-variance efforts as they are too similar to each other.</p>",
              "votes": 4,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 3156652,
      "author_name": "Ogurtsov",
      "author_url": "",
      "post_date": "2025-03-22T11:37:48.013000",
      "content": "<p>It might be worth considering faster alternatives like JackHMMER and MMseqs2.</p>",
      "votes": 0,
      "replies": [
        {
          "id": 3158246,
          "author_name": "hengck23",
          "author_url": "",
          "post_date": "2025-03-24T11:10:17.640000",
          "content": "<p>Gpu mmseqs2 is very fast</p>",
          "votes": 2,
          "replies": []
        }
      ]
    },
    {
      "id": 3156608,
      "author_name": "Pranav Pawar",
      "author_url": "",
      "post_date": "2025-03-22T10:23:11.537000",
      "content": "<p>Hello,<br>\nIt's showing error 404 after accessing the link.<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F25371292%2F25ea360174dda13e475147a74ad155d1%2FScreenshot%202025-03-22%20155234.png?generation=1742638970630946&amp;alt=media\" alt=\"SS\"></p>",
      "votes": 0,
      "replies": [
        {
          "id": 3156888,
          "author_name": "Alissa Hummer",
          "author_url": "",
          "post_date": "2025-03-22T17:18:05.060000",
          "content": "<p>Hi <a href=\"https://www.kaggle.com/prnvpwr2612\" target=\"_blank\">@prnvpwr2612</a>, it looks like the URL you tried includes the \")\" at the end. Could you try reloading the page without the \")\"?</p>\n<p><a href=\"https://github.com/pylelab/rMSA\" target=\"_blank\">https://github.com/pylelab/rMSA</a></p>",
          "votes": 1,
          "replies": [
            {
              "id": 3157042,
              "author_name": "Pranav Pawar",
              "author_url": "",
              "post_date": "2025-03-22T21:40:15.507000",
              "content": "<p>Thanks a lot it is working now.</p>",
              "votes": 0,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 3156305,
      "author_name": "Zacchaeus",
      "author_url": "",
      "post_date": "2025-03-22T02:59:48.620000",
      "content": "<p>Would you mind sharing the pipeline generating MSAs? It could be helpful that we adopt same pipeline in our experiments.</p>",
      "votes": 0,
      "replies": [
        {
          "id": 3156310,
          "author_name": "Alissa Hummer",
          "author_url": "",
          "post_date": "2025-03-22T03:08:38.947000",
          "content": "<p>Hi <a href=\"https://www.kaggle.com/zacchaeus\" target=\"_blank\">@zacchaeus</a>, we used the rMSA tool with default parameters and database versions/dates as described above. I would recommend looking at the rMSA GitHub repo (<a href=\"https://github.com/pylelab/rMSA)\" target=\"_blank\">https://github.com/pylelab/rMSA)</a>, which has helpful information for installation, database downloads, and MSA generation. To generate the MSAs, we ran <code>./rMSA.pl seq.fasta</code> as on their GitHub README. We updated the outputs to replace \"T\" nucleotides with \"U\". One thing to be aware of is that the databases for MSA generation are very large (~2 TB) and can take a long time to download/set up (on the order of days).</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 3156271,
      "author_name": "Tilii",
      "author_url": "",
      "post_date": "2025-03-22T00:38:34.077000",
      "content": "<p><a href=\"https://www.kaggle.com/alissahummer\" target=\"_blank\">@alissahummer</a> Good to have these files, thanks.</p>\n<p>It seems that we now have straight-up aligned FASTA files, with no insert states in lowercase letters. If so, that should be specified for those who have scripts expecting A3M files. I don't think what you did with this update fulfills the criterion <code>so that all notebooks that use the files are compatible with the update, without any changes required.</code></p>\n<p>Same question as before: when do you expect that all alignments will be available?</p>",
      "votes": 0,
      "replies": [
        {
          "id": 3156292,
          "author_name": "Alissa Hummer",
          "author_url": "",
          "post_date": "2025-03-22T01:58:52.893000",
          "content": "<p>Hi <a href=\"https://www.kaggle.com/tilii7\" target=\"_blank\">@tilii7</a>, thank you for pointing this out and my apologies for any confusion. I have updated the post to clarify the A2M format.</p>\n<p>We hope to have all alignments available next week, but please note that this timeline might change.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 3159516,
          "author_name": "Alissa Hummer",
          "author_url": "",
          "post_date": "2025-03-25T17:13:26.750000",
          "content": "<p>Hi <a href=\"https://www.kaggle.com/tilii7\" target=\"_blank\">@tilii7</a>, I wanted to let you know that all placeholders have now been updated with the rMSA outputs.</p>",
          "votes": 2,
          "replies": [
            {
              "id": 3159525,
              "author_name": "Tilii",
              "author_url": "",
              "post_date": "2025-03-25T17:30:11.743000",
              "content": "<p>Thank you, that was fast.</p>\n<p>For some of short RNAs there are still no additional sequences, and that can't be because they don't have homologs in the database (<em>e.g.,</em> 1AFX and 1ESH). I suspect it is because the search procedure is not optimized to work with very short sequences. I will try to generate MSAs for them on my own and will share them if there is a global solution.</p>",
              "votes": 2,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 3156228,
      "author_name": "asarvazyan",
      "author_url": "",
      "post_date": "2025-03-21T22:55:47.617000",
      "content": "<p>Thanks! By the way, what does it mean for some query-only files to be \"placeholders\"? Would these be the ones in the test spplit and they are then replaced with the real MSAs?</p>",
      "votes": 0,
      "replies": [
        {
          "id": 3156287,
          "author_name": "Alissa Hummer",
          "author_url": "",
          "post_date": "2025-03-22T01:41:41.033000",
          "content": "<p>Hi <a href=\"https://www.kaggle.com/asarvazyan\" target=\"_blank\">@asarvazyan</a>, thanks for the question! The placeholder MSAs are only for some targets in the training dataset, where MSA generation is taking a long time to complete. We did not want to hold up the release of the majority of the MSAs waiting for these. We plan to update the placeholders with the rMSA outputs once the jobs have finished.</p>",
          "votes": 0,
          "replies": [
            {
              "id": 3156293,
              "author_name": "asarvazyan",
              "author_url": "",
              "post_date": "2025-03-22T02:03:07.720000",
              "content": "<p>Got it, thanks!</p>",
              "votes": 0,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 3158169,
      "author_name": "",
      "author_url": "",
      "post_date": "2025-03-24T09:30:45.823000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 3156226,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "2025-03-21T22:43:48.767000",
      "content": "<p>thanks a lot again!</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 3158337,
      "author_name": "",
      "author_url": "",
      "post_date": "2025-03-24T13:09:34.423000",
      "content": "<p>Thanks for the update </p>",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "3156129": "Hi Kagglers,\n\nJust a quick note that we have updated the multiple sequence alignment (MSA) files to outputs generated using rMSA (https://github.com/pylelab/rMSA). This method produced deeper MSAs and we used databases with a temporal cutoff more closely aligned with the validation dataset (database versions included below). The rMSA outputs are in the A2M format.\n\nWe have retained the MSA/{target_id}.MSA.fasta file naming.\n\nrMSA Database Versions/Cutoffs\n• RNAcentral v20.0 (2022-03-28)\n• RFam v14.7 (2021-12-09)\n• NCBI NT (2022-10-03)\n\n[Updated 20250321 - A2M format clarified]\n[Updated 20250325 - all placeholders have been replaced with rMSA outputs]",
    "3156294": "for the newbies:\n\n1. A2M: Uses uppercase characters for matches, lowercase for insertions (with gaps aligned to insertions shown as periods), and dashes (-) for deletions. \n1. A3M: Similar to A2M, but allows gaps aligned to insertions to be omitted(periods), making the format more compact. ",
    "3156307": "@alissahummer \nDo u think u can open source your  rmsa script to make the msa?\n\nOur team will be making external data and share with the kaggle community. We will be publishing eg af3 3d prediction. We thinking if we want to publish msa of exeternal data as well ",
    "3156652": "It might be worth considering faster alternatives like JackHMMER and MMseqs2.\n",
    "3156608": "Hello,\nIt's showing error 404 after accessing the link.![SS](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F25371292%2F25ea360174dda13e475147a74ad155d1%2FScreenshot%202025-03-22%20155234.png?generation=1742638970630946&alt=media)",
    "3156305": "Would you mind sharing the pipeline generating MSAs? It could be helpful that we adopt same pipeline in our experiments.",
    "3156271": "@alissahummer Good to have these files, thanks.\n\nIt seems that we now have straight-up aligned FASTA files, with no insert states in lowercase letters. If so, that should be specified for those who have scripts expecting A3M files. I don't think what you did with this update fulfills the criterion `so that all notebooks that use the files are compatible with the update, without any changes required.`\n\nSame question as before: when do you expect that all alignments will be available?",
    "3156228": "Thanks! By the way, what does it mean for some query-only files to be \"placeholders\"? Would these be the ones in the test spplit and they are then replaced with the real MSAs?",
    "3158169": "",
    "3156226": "thanks a lot again!",
    "3158337": "Thanks for the update "
  }
}