{
  "id": 570704,
  "title": "[draft] how to finetune proteinx (AF3 clone) for MSA RNA finetunning",
  "url": "/competitions/stanford-rna-3d-folding/discussion/570704",
  "author_name": "",
  "post_date": "2025-03-30T07:16:23.017812500Z",
  "votes": 21,
  "comment_count": 8,
  "views": 0,
  "content": "<h2>background</h2>\n<ul>\n<li>As mentioned in the other discussions, all AF3 clones (proteinX, Boltz,Chai-1) does not support RNA MSA.</li>\n<li>Host showed that orginal alphafold3 improved results when kaggle rMSA is used (this could mean that there are useful information in MSA)</li>\n<li>Many of the kagglers are newble (including me) in RNA 3d structure prediction, MSA search, let's help each other!</li>\n<li>proteinX is chosen becuase its API does support RNA MSA, but it is not used. </li>\n</ul>\n<h2>We follow orginal AF3 paper and proteinX paper first:</h2>\n<p>-Supplementary information:<br>\n\"Accurate structure prediction of biomolecular interactions with AlphaFold 3\"</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2Faa721021309efc354d5af37edf3a2a6f%2FSelection_173.png?generation=1743318929775451&amp;alt=media\" alt=\"\"></p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2F83d610199cf7cf3248824295b7963e9e%2FSelection_172.png?generation=1743318971095569&amp;alt=media\" alt=\"\"></p>\n<hr>\n<p>note: this is still in experiment. so maybe the parameters may not be correct (e.g. msa search parameters)</p>",
  "messages": [
    {
      "id": "3163015",
      "postDate": "03/30/2025 07:16:23",
      "content": "<h2>background</h2>\n<ul>\n<li>As mentioned in the other discussions, all AF3 clones (proteinX, Boltz,Chai-1) does not support RNA MSA.</li>\n<li>Host showed that orginal alphafold3 improved results when kaggle rMSA is used (this could mean that there are useful information in MSA)</li>\n<li>Many of the kagglers are newble (including me) in RNA 3d structure prediction, MSA search, let's help each other!</li>\n<li>proteinX is chosen becuase its API does support RNA MSA, but it is not used. </li>\n</ul>\n<h2>We follow orginal AF3 paper and proteinX paper first:</h2>\n<p>-Supplementary information:<br>\n\"Accurate structure prediction of biomolecular interactions with AlphaFold 3\"</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2Faa721021309efc354d5af37edf3a2a6f%2FSelection_173.png?generation=1743318929775451&amp;alt=media\" alt=\"\"></p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2F83d610199cf7cf3248824295b7963e9e%2FSelection_172.png?generation=1743318971095569&amp;alt=media\" alt=\"\"></p>\n<hr>\n<p>note: this is still in experiment. so maybe the parameters may not be correct (e.g. msa search parameters)</p>",
      "rawMarkdown": "##background\n- As mentioned in the other discussions, all AF3 clones (proteinX, Boltz,Chai-1) does not support RNA MSA.\n- Host showed that orginal alphafold3 improved results when kaggle rMSA is used (this could mean that there are useful information in MSA)\n- Many of the kagglers are newble (including me) in RNA 3d structure prediction, MSA search, let's help each other!\n- proteinX is chosen becuase its API does support RNA MSA, but it is not used. \n\n\n##We follow orginal AF3 paper and proteinX paper first:\n-Supplementary information:\n\"Accurate structure prediction of biomolecular interactions with AlphaFold 3\"\n\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2Faa721021309efc354d5af37edf3a2a6f%2FSelection_173.png?generation=1743318929775451&alt=media)\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2F83d610199cf7cf3248824295b7963e9e%2FSelection_172.png?generation=1743318971095569&alt=media)\n\n\n---\n\nnote: this is still in experiment. so maybe the parameters may not be correct (e.g. msa search parameters)",
      "votes": null
    },
    {
      "id": "3163018",
      "postDate": "03/30/2025 07:20:39",
      "content": "<p>step.0: setup</p>\n<p>Here we how example for RNACentral.</p>\n<ul>\n<li>download RNACentral</li>\n<li>install HMMER</li>\n<li>install mmseq2<br>\n(these are quite standard installations. one can just follows instruction from their repo, so i won't go over the details. in doubt, chatgpt can help)</li>\n</ul>\n<p>here are my downloaded RNACentral files:<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2F22f5826e6a2870fa0c695d497bc84df3%2FSelection_174.png?generation=1743319262803448&amp;alt=media\" alt=\"\"></p>\n<p>NOTE: i have not followed the cutoff date requirement for early prize, etc … (since i am still experimenting). Plesae adjust for your case accordingly</p>",
      "rawMarkdown": "step.0: setup\n\nHere we how example for RNACentral.\n- download RNACentral\n- install HMMER\n- install mmseq2\n(these are quite standard installations. one can just follows instruction from their repo, so i won't go over the details. in doubt, chatgpt can help)\n\nhere are my downloaded RNACentral files:\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2F22f5826e6a2870fa0c695d497bc84df3%2FSelection_174.png?generation=1743319262803448&alt=media)\n\nNOTE: i have not followed the cutoff date requirement for early prize, etc ... (since i am still experimenting). Plesae adjust for your case accordingly",
      "votes": null
    },
    {
      "id": "3163024",
      "postDate": "03/30/2025 07:37:35",
      "content": "<p>step.1 clustering with mmseqs</p>\n<pre><code>import , sys\n\nSSD_DIR = \nHMMER_DIR =f\nMMSEQS_DIR =f\ndb_file = \\\n    f\n    #f #smaller file  \n    #seqkit -n  ...\n\n : \n    .makedirs(, exist_ok=True)\n\n    cmd = .join([\n        f,\n        f,\n        f,\n        f,\n        f,\n        f,\n        f,\n        f,\n        f,\n        f,\n    ])\n     = .(cmd).()\n    ()\n</code></pre>\n<p>chatgpt:<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2Fd5b47503eb5879d60ee6c6f985b73433%2FSelection_175.png?generation=1743320248195126&amp;alt=media\" alt=\"\"></p>\n<hr>\n<p>output:</p>\n<p>Number of clusters: 16972265<br>\nThis means MMseqs grouped ~37.9 million RNAcentral sequences down to ~17 million representative</p>",
      "rawMarkdown": "step.1 clustering with mmseqs\n\n```\n\nimport os, sys\n\nSSD_DIR = '/media/hp/xxx-xxx-xxx-xxx'\nHMMER_DIR =f'{SSD_DIR}/my-msa-server/tool/hmmer/binary/bin'\nMMSEQS_DIR =f'{SSD_DIR}/my-msa-server/tool/mmseqs/mmseqs/bin/'\ndb_file = \\\n    f'{SSD_DIR}/my-msa-server/database/RNAcentral/rnacentral_species_specific_ids.fasta'\n    #f'{SSD_DIR}/my-msa-server/database/RNAcentral/mini_rnacentral.fasta' #smaller file for debug\n    #seqkit -n 1000 ...\n\nif 1: \n    os.makedirs('mmseqs_tmp', exist_ok=True)\n\n    cmd = \" \".join([\n        f'{MMSEQS_DIR}/mmseqs',\n        f'easy-linclust',\n        f'{db_file}',\n        f'rnacentral_linclust',\n        f'mmseqs_tmp',\n        f'--min-seq-id 0.9',\n        f'-c 0.8',\n        f'--cov-mode 0',\n        f'--kmer-per-seq-scale 0.3',\n        f'--threads 72',\n    ])\n    output = os.popen(cmd).read()\n    print(output)\n\n\n```\n\nchatgpt:\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2Fd5b47503eb5879d60ee6c6f985b73433%2FSelection_175.png?generation=1743320248195126&alt=media)\n\n\n---\n\noutput:\n\nNumber of clusters: 16972265\nThis means MMseqs grouped ~37.9 million RNAcentral sequences down to ~17 million representative",
      "votes": null
    },
    {
      "id": "3163045",
      "postDate": "03/30/2025 08:22:35",
      "content": "<p>step.2. search with hmmer</p>\n<pre><code>    query_file = \n    \n\n    cluster_db_file = \\\n    \n    cmd = .join([\n        ,\n        ,\n        ,\n        ,\n        ,\n        ,\n        ,\n        ,\n        ,\n        ,\n    ])\n    output = os.popen(cmd).read()\n    (output)\n</code></pre>\n<p>Query model(s):                            1  (69 nodes)<br>\nTarget sequences:                   16972265  (14294762522 residues searched)<br>\nResidues passing SSV filter:       183030293  (0.0128); expected (0.02)<br>\nResidues passing bias filter:      176205600  (0.0123); expected (0.02)<br>\nResidues passing Vit filter:        14003806  (0.00098); expected (0.003)<br>\nResidues passing Fwd filter:          165886  (1.16e-05); expected (5e-05)<br>\nTotal number of hits:                      4  (1.87e-08)<br>\nElapsed: 00:01:10.64</p>\n<p>protineX is reading in *.sto</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2Fc05bfd930381eff9c31c80c6fe2de1a3%2FSelection_177.png?generation=1743323280546390&amp;alt=media\" alt=\"\"></p>",
      "rawMarkdown": "step.2. search with hmmer\n\n```\n    query_file = f'{SSD_DIR}/2025/kaggle/stanford-rna-3d-folding/data/my-data/casp15/R1107.fasta'\n    # note seq length.flag F3 is 0.02 or 0.00005 dependent on seq length \n \n    cluster_db_file = \\\n    f'{SSD_DIR}/my-msa-server/database/RNAcentral/clustered/rnacentral_linclust_rep_seq.fasta'\n    cmd = \" \".join([\n        f'{HMMER_DIR}/nhmmer',\n        f'--rna',\n        f'--cpu 16',\n        f'-E 0.001',\n        f'--incE 0.001',\n        f'--watson',\n        f'--F3 0.00005',\n        f'--tblout nhmmer_hits.tbl',\n        f'-A nhmmer_hits.sto',\n        f'{query_file} {cluster_db_file}',\n    ])\n    output = os.popen(cmd).read()\n    print(output)\n```\nQuery model(s):                            1  (69 nodes)\nTarget sequences:                   16972265  (14294762522 residues searched)\nResidues passing SSV filter:       183030293  (0.0128); expected (0.02)\nResidues passing bias filter:      176205600  (0.0123); expected (0.02)\nResidues passing Vit filter:        14003806  (0.00098); expected (0.003)\nResidues passing Fwd filter:          165886  (1.16e-05); expected (5e-05)\nTotal number of hits:                      4  (1.87e-08)\nElapsed: 00:01:10.64\n\nprotineX is reading in *.sto\n\n ![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2Fc05bfd930381eff9c31c80c6fe2de1a3%2FSelection_177.png?generation=1743323280546390&alt=media)",
      "votes": null
    },
    {
      "id": "3163050",
      "postDate": "03/30/2025 08:29:05",
      "content": "<p>step.3: check if *.sto can be read by proteinX?</p>\n<p>reference:<br>\n[1] <a href=\"https://github.com/bytedance/Protenix/issues/7\" target=\"_blank\">https://github.com/bytedance/Protenix/issues/7</a><br>\nHow to get seq_to_pdb_index.json when fine-tuning #7</p>\n<p>[2] rna msa related source code<br>\n/Protenix-main/protenix/data/msa_featurizer.py<br>\n/Protenix-main/protenix/data/msa_utils.py</p>\n<pre><code>\n\n\n protenix.data.msa_featurizer import *\ntarget_id = \nseq = \nraw_msa_paths = [\n       #result  hmmer search\n]\n\n\nr = process_single_sequence(\n    =target_id,\n    =seq,\n    =raw_msa_paths, \n    =SEQ_LIMITS, # Optional[list[str]],\n    msa_entity_type =,\n    msa_type = ,\n)\n k,v  r.items():\n    (,k)\n    (v)\n\n\n\n\n\n\nmsa_data = parse_rna_msa_data(.)\n\n k,v  msa_data.items():\n    (,k)\n    (v)\n\n** /media/hp/c30d34ed-0d55-4077-82dc-b56cd13dd548/2025/kaggle/stanford-rna-3d-folding/code/proteinx01/000/DUMMY_MSA/0/nhmmer_hits.sto\n\nMsa(sequences=[, , , , ], \n\ndeletion_matrix=[[0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0], [0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0], [0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0], [0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0], [0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0]], \n\ndescriptions=[, , , , ])\n</code></pre>\n<p>a natural question is can kaggle a2m/a3m file be read by proteinX?<br>\nanswer is yes and not</p>\n<ul>\n<li>if you check the code, for protein input, proteinX accept  a3m file (and sto file)</li>\n<li>however, for rna input, it read only sto file</li>\n</ul>\n<p>so you can write rna a3m reader yourself or convert kaggle a3m file to sto</p>\n<hr>\n<p>then you may ask why no rna a3m reader? because proteinX follows alphafold3 paper. only HMMER( i.e.  <em>.sto) for RNA. Protein uses jackhammer(</em>a3m) and HHBlits(*.a3m)</p>",
      "rawMarkdown": "step.3: check if *.sto can be read by proteinX?\n\nreference:\n[1] https://github.com/bytedance/Protenix/issues/7\nHow to get seq_to_pdb_index.json when fine-tuning #7\n\n[2] rna msa related source code\n/Protenix-main/protenix/data/msa_featurizer.py\n/Protenix-main/protenix/data/msa_utils.py\n\n````\n\n#first test this function\n\n\nfrom protenix.data.msa_featurizer import *\ntarget_id = 'R1107'\nseq = 'GGGGGCCACAGCAGAAGCGUUCACGUCGCAGCCCCUGUCAGCCAUUGCACUCCGGCUGCGAAUUCUGCU'\nraw_msa_paths = [\n    '/000/DUMMY_MSA/0/nhmmer_hits.sto'   #result from hmmer search\n]\n\n\nr = process_single_sequence(\n    pdb_name=target_id,\n    sequence=seq,\n    raw_msa_paths=raw_msa_paths, \n    seq_limits=SEQ_LIMITS, # Optional[list[str]],\n    msa_entity_type =\"rna\",\n    msa_type = \"non_pairing\",\n)\nfor k,v in r.items():\n    print('**',k)\n    print(v)\n\n#-----\n\n#process_single_sequence() will end up calling parse_rna_msa_data() in msa_utils.py\n#e.g. \n\nmsa_data = parse_rna_msa_data(...)\n\nfor k,v in msa_data.items():\n    print('**',k)\n    print(v)\n    \n** /media/hp/c30d34ed-0d55-4077-82dc-b56cd13dd548/2025/kaggle/stanford-rna-3d-folding/code/proteinx01/000/DUMMY_MSA/0/nhmmer_hits.sto\n\nMsa(sequences=['GGGGGCCACAGCAGAAGCGUUCACGUCGCAGCCCCUGUCAGCCAUUGCACUCCGGCUGCGAAUUCUGCU', 'GGGGGCCACAGCAGAAGCGUUCACGUCGCGGCCCCUGUCAGCCAUUGCACUCCGGCUGCGAAUUCUGCU', 'GGGGGCCAUAGCAGAAGCGUUCACGUCGCAGCCCCUGUCAGAUUCU--UACGAACCUGCGAAUUCUGCU', '-GGGGCCACAGCAGAAGCGUUCACGUCGCGGCCCCUGUCAGAUUCUG--GUGAAUCUGCGAAUUCUGCU', '-GGCGCCAUAGCAGAAGCGUUCACGUCGCAGCCCCUGUCAGAUUCU--UACGAAUCUGCGAAUUCUGC-'], \n\ndeletion_matrix=[[0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0], [0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0], [0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0], [0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0], [0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0]], \n\ndescriptions=['query', 'URS0002617738_9598/1-69', 'URS00006F3FF9_7897/12-78', 'URS0000626D9E_230844/13-78', 'URS0000A93AF3_7897/13-77'])\n````\n\n\na natural question is can kaggle a2m/a3m file be read by proteinX?\nanswer is yes and not\n- if you check the code, for protein input, proteinX accept  a3m file (and sto file)\n- however, for rna input, it read only sto file\n\nso you can write rna a3m reader yourself or convert kaggle a3m file to sto\n\n---\n\nthen you may ask why no rna a3m reader? because proteinX follows alphafold3 paper. only HMMER( i.e.  *.sto) for RNA. Protein uses jackhammer(*a3m) and HHBlits(*.a3m)",
      "votes": null
    },
    {
      "id": "3163128",
      "postDate": "03/30/2025 10:39:50",
      "content": "<p>step.4 setup a training database for rna msa<br>\nreference:</p>\n<p>[1] <a href=\"https://github.com/bytedance/Protenix/issues/7\" target=\"_blank\">https://github.com/bytedance/Protenix/issues/7</a><br>\nHow to get seq_to_pdb_index.json when fine-tuning #7</p>\n<p>[2] training guide<br>\n<a href=\"https://github.com/bytedance/Protenix/blob/main/docs/msa_pipeline.md\" target=\"_blank\">https://github.com/bytedance/Protenix/blob/main/docs/msa_pipeline.md</a><br>\n<a href=\"https://github.com/bytedance/Protenix/blob/main/docs/training.md\" target=\"_blank\">https://github.com/bytedance/Protenix/blob/main/docs/training.md</a></p>\n<p>[3] source code<br>\n/Protenix-main/protenix/data/msa_featurizer.py</p>\n<hr>\n<p>basically, the pytorch dataset object will create MSAFeaturizer, which will in turn create RNAMSAFeaturizer.<br>\nPlease check the code of these classes.</p>\n<p>the information below may not be 100%. it can only be confirmed after I use the dataset object for training. nevertheless:</p>\n<pre><code>   strcuture\nDUMMY_MSA\n  |-R1107\n  |    |-rnacentral.sto  #your result from hmmer \n  |\n  |-R1156\n  |    |-rnacentral.sto\n  |\n  |-seq_to_pdb_index.json\n\n\nseq_to_pdb_index.json:\n{\n    : [],\n    : []\n}\n\nnote that [] will  the folder    *sto \n</code></pre>\n<p>then try the test code:</p>\n<pre><code>rna_msa_dir = \nseq_to_pdb_idx_path = \n\n\nf = RNAMSAFeaturizer(\n    seq_to_pdb_idx_path = seq_to_pdb_index_file,  \n    indexing_method = , \n    merge_method = ,\n    seq_limits = SEQ_LIMITS,\n    max_size = ,\n    rna_msa_dir = rna_msa_dir, \n    \n)\n( f.seq_to_pdb_idx )\n\nr = f.process_single_sequence(\n    pdb_name= , \n    sequence=seq,\n    pdb_id=target_id,\n    is_homomer_or_monomer=,\n)\n\n k,v  r.items():\n    (,k)\n    (v)\n</code></pre>\n<p>modify the code to print raw_msa_paths<br>\nraw_msa_paths [DUMMY_MSA/R1107/rnacentral.sto']</p>\n<pre><code> ():\n...\n ()\n  \n\n     (</code></pre>",
      "rawMarkdown": "step.4 setup a training database for rna msa\nreference:\n\n[1] https://github.com/bytedance/Protenix/issues/7\nHow to get seq_to_pdb_index.json when fine-tuning #7\n\n[2] training guide\nhttps://github.com/bytedance/Protenix/blob/main/docs/msa_pipeline.md\nhttps://github.com/bytedance/Protenix/blob/main/docs/training.md\n\n[3] source code\n/Protenix-main/protenix/data/msa_featurizer.py\n\n----\n\nbasically, the pytorch dataset object will create MSAFeaturizer, which will in turn create RNAMSAFeaturizer.\nPlease check the code of these classes.\n\nthe information below may not be 100%. it can only be confirmed after I use the dataset object for training. nevertheless:\n\n```\nset up file strcuture\nDUMMY_MSA\n  |-R1107\n  |    |-rnacentral.sto  #your result from hmmer search\n  |\n  |-R1156\n  |    |-rnacentral.sto\n  |\n  |-seq_to_pdb_index.json\n\n\nseq_to_pdb_index.json:\n{\n    \"GGGGGCCACAGCAGAAGCGUUCACGUCGCAGCCCCUGUCAGCCAUUGCACUCCGGCUGCGAAUUCUGCU\": [\"R1107\"],\n    \"GGAGCAUCGUGUCUCAAGUGCUUCACGGUCACAAUAUACCGUUUCGUCGGGUGCGUGGCAAUUCGGUGCACAUCAUGUCUUUCGUGGCUGGUGUGGCUCCUCAAGGUGCGAGGGGCAAGUAUAGAGCAGAGCUCC\": [\"R1156\"]\n}\n\nnote that [\"R1107\"] will be the folder to search for *sto file\n```\n\nthen try the test code:\n\n```\nrna_msa_dir = 'DUMMY_MSA'\nseq_to_pdb_idx_path = 'DUMMY_MSA/seq_to_pdb_index.json'\n\n\nf = RNAMSAFeaturizer(\n    seq_to_pdb_idx_path = seq_to_pdb_index_file,  \n    indexing_method = 'sequence', #'pdb_id_entity_id', #\"sequence\",\n    merge_method = 'dense_max',\n    seq_limits = SEQ_LIMITS,\n    max_size = 16384,\n    rna_msa_dir = rna_msa_dir, \n    #**kwargs,\n)\nprint( f.seq_to_pdb_idx )\n\nr = f.process_single_sequence(\n    pdb_name=f'{target_id}_1' , #pdb_id_entity_id\n    sequence=seq,\n    pdb_id=target_id,\n    is_homomer_or_monomer=True,\n)\n\nfor k,v in r.items():\n    print('**',k)\n    print(v)\n\n```\nmodify the code to print raw_msa_paths\nraw_msa_paths [DUMMY_MSA/R1107/rnacentral.sto']\n\n\n```\nclass RNAMSAFeaturizer(BaseMSAFeaturizer):\n...\ndef get_msa_path(\n        self, db_name: str, sequence: str, pdb_id_entity_id: str, reduced: bool = False\n    )\n  ##!!!! reduced must be patch to False  !!!!\n\n    def process_single_sequence(\n        self,\n        pdb_name: str,\n ...\n\n        raw_msa_paths, seq_limits = [], []\n        for db_name in self.non_pairing_db:\n            if opexists(\n                path := self.get_msa_path(db_name, sequence, pdb_name)\n            ) and path.endswith(\".sto\"):\n                raw_msa_paths.append(path)\n                seq_limits.append(self.seq_limits.get(db_name, SEQ_LIMITS[db_name]))\n        print('HCK!!!! raw_msa_paths', raw_msa_paths)\n```",
      "votes": null
    },
    {
      "id": "3163318",
      "postDate": "03/30/2025 16:07:07",
      "content": "<p>let's try if we can do infrerence with rna msa.</p>\n<p>update the configs_data.py used by runner/inference.py:</p>\n<pre><code> = \n\n\n = ...\n = ...\n = ...\n</code></pre>\n<p>trick to run inference without bash file sh:</p>\n<p>copy inference.py and make the following change:</p>\n<pre><code>\n os\nos.environ[] = \nos.environ[] = \nos.environ[] = \n\n sys\nsys.path.insert(, )\n\n....\n\n __name__ == :\n    N_sample = \n    N_step = \n    N_cycle = \n    seed = \n    use_deepspeed_evo_attention=\n    input_json_path=\n    dump_dir=\n\n   \n    sys.argv +=[\n        , ,\n        , ,\n        , ,\n        , ,\n        , ,\n        , ,\n    ]\n</code></pre>\n<p>to see how rna mas is loaded, we can check the function get_inference_dataloader() of infer_data_pipeline.py. it leads us to:</p>\n<pre><code> (object):\n    \n\nwe need  modify this  !!!!!\n...  be continued ...\n</code></pre>",
      "rawMarkdown": "let's try if we can do infrerence with rna msa.\n\nupdate the configs_data.py used by runner/inference.py:\n```\nmsa.enable_rna_msa = True\n\n#used by training ... but we updated these just in case ...\nrna.seq_to_pdb_idx_path = ...\nrna.rna_msa_dir = ...\nrna.indexing_method = ...\n\n```\n\n\ntrick to run inference without bash file sh:\n\ncopy inference.py and make the following change:\n\n```\n#top of inference.py:\nimport os\nos.environ[\"CUTLASS_PATH\"] = \"....your path.../cutlass/cutlass\"\nos.environ[\"LAYERNORM_TYPE\"] = \"fast_layernorm\"\nos.environ[\"USE_DEEPSPEED_EVO_ATTENTION\"] = \"true\"\n\nimport sys\nsys.path.insert(0, '/.... run local opy instead of pip installed site-package ..../Protenix-main')\n\n....\n\nif __name__ == \"__main__\":\n    N_sample = 5\n    N_step = 200\n    N_cycle = 10\n    seed = 101\n    use_deepspeed_evo_attention=True\n    input_json_path=\"... your path ..../casp15-msa.json\"\n    dump_dir=\"... your path ..../casp15-msa\"\n\n   #fake command line arguments\n    sys.argv +=[\n        '--seeds', f'{seed}',\n        '--dump_dir', f'{dump_dir}',\n        '--input_json_path', f'{input_json_path}',\n        '--model.N_cycle', f'{N_cycle}',\n        '--sample_diffusion.N_sample', f'{N_sample}',\n        '--sample_diffusion.N_step', f'{N_step}',\n    ]\n```\n\n\nto see how rna mas is loaded, we can check the function get_inference_dataloader() of infer_data_pipeline.py. it leads us to:\n```\n\nclass InferenceMSAFeaturizer(object):\n    # Now we only support protein msa in inference\n\nwe need to modify this function !!!!!\n... to be continued ...\n\n```",
      "votes": null
    },
    {
      "id": "3175261",
      "postDate": "04/10/2025 00:47:35",
      "content": "<p>Thank you for sharing! I am wondering about the results that you have done inferencing with msa or training it! Can you please share it with us?</p>",
      "rawMarkdown": "Thank you for sharing! I am wondering about the results that you have done inferencing with msa or training it! Can you please share it with us?",
      "votes": null
    },
    {
      "id": "3177791",
      "postDate": "04/13/2025 09:30:12",
      "content": "<p>I appreciate your input; it will be very helpful to me.</p>",
      "rawMarkdown": "I appreciate your input; it will be very helpful to me.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 3163018,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "03/30/2025 07:20:39",
      "content": "<p>step.0: setup</p>\n<p>Here we how example for RNACentral.</p>\n<ul>\n<li>download RNACentral</li>\n<li>install HMMER</li>\n<li>install mmseq2<br>\n(these are quite standard installations. one can just follows instruction from their repo, so i won't go over the details. in doubt, chatgpt can help)</li>\n</ul>\n<p>here are my downloaded RNACentral files:<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2F22f5826e6a2870fa0c695d497bc84df3%2FSelection_174.png?generation=1743319262803448&amp;alt=media\" alt=\"\"></p>\n<p>NOTE: i have not followed the cutoff date requirement for early prize, etc … (since i am still experimenting). Plesae adjust for your case accordingly</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 3163024,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "03/30/2025 07:37:35",
      "content": "<p>step.1 clustering with mmseqs</p>\n<pre><code>import , sys\n\nSSD_DIR = \nHMMER_DIR =f\nMMSEQS_DIR =f\ndb_file = \\\n    f\n    #f #smaller file  \n    #seqkit -n  ...\n\n : \n    .makedirs(, exist_ok=True)\n\n    cmd = .join([\n        f,\n        f,\n        f,\n        f,\n        f,\n        f,\n        f,\n        f,\n        f,\n        f,\n    ])\n     = .(cmd).()\n    ()\n</code></pre>\n<p>chatgpt:<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2Fd5b47503eb5879d60ee6c6f985b73433%2FSelection_175.png?generation=1743320248195126&amp;alt=media\" alt=\"\"></p>\n<hr>\n<p>output:</p>\n<p>Number of clusters: 16972265<br>\nThis means MMseqs grouped ~37.9 million RNAcentral sequences down to ~17 million representative</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 3163045,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "03/30/2025 08:22:35",
      "content": "<p>step.2. search with hmmer</p>\n<pre><code>    query_file = \n    \n\n    cluster_db_file = \\\n    \n    cmd = .join([\n        ,\n        ,\n        ,\n        ,\n        ,\n        ,\n        ,\n        ,\n        ,\n        ,\n    ])\n    output = os.popen(cmd).read()\n    (output)\n</code></pre>\n<p>Query model(s):                            1  (69 nodes)<br>\nTarget sequences:                   16972265  (14294762522 residues searched)<br>\nResidues passing SSV filter:       183030293  (0.0128); expected (0.02)<br>\nResidues passing bias filter:      176205600  (0.0123); expected (0.02)<br>\nResidues passing Vit filter:        14003806  (0.00098); expected (0.003)<br>\nResidues passing Fwd filter:          165886  (1.16e-05); expected (5e-05)<br>\nTotal number of hits:                      4  (1.87e-08)<br>\nElapsed: 00:01:10.64</p>\n<p>protineX is reading in *.sto</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2Fc05bfd930381eff9c31c80c6fe2de1a3%2FSelection_177.png?generation=1743323280546390&amp;alt=media\" alt=\"\"></p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 3163050,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "03/30/2025 08:29:05",
      "content": "<p>step.3: check if *.sto can be read by proteinX?</p>\n<p>reference:<br>\n[1] <a href=\"https://github.com/bytedance/Protenix/issues/7\" target=\"_blank\">https://github.com/bytedance/Protenix/issues/7</a><br>\nHow to get seq_to_pdb_index.json when fine-tuning #7</p>\n<p>[2] rna msa related source code<br>\n/Protenix-main/protenix/data/msa_featurizer.py<br>\n/Protenix-main/protenix/data/msa_utils.py</p>\n<pre><code>\n\n\n protenix.data.msa_featurizer import *\ntarget_id = \nseq = \nraw_msa_paths = [\n       #result  hmmer search\n]\n\n\nr = process_single_sequence(\n    =target_id,\n    =seq,\n    =raw_msa_paths, \n    =SEQ_LIMITS, # Optional[list[str]],\n    msa_entity_type =,\n    msa_type = ,\n)\n k,v  r.items():\n    (,k)\n    (v)\n\n\n\n\n\n\nmsa_data = parse_rna_msa_data(.)\n\n k,v  msa_data.items():\n    (,k)\n    (v)\n\n** /media/hp/c30d34ed-0d55-4077-82dc-b56cd13dd548/2025/kaggle/stanford-rna-3d-folding/code/proteinx01/000/DUMMY_MSA/0/nhmmer_hits.sto\n\nMsa(sequences=[, , , , ], \n\ndeletion_matrix=[[0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0], [0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0], [0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0], [0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0], [0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0]], \n\ndescriptions=[, , , , ])\n</code></pre>\n<p>a natural question is can kaggle a2m/a3m file be read by proteinX?<br>\nanswer is yes and not</p>\n<ul>\n<li>if you check the code, for protein input, proteinX accept  a3m file (and sto file)</li>\n<li>however, for rna input, it read only sto file</li>\n</ul>\n<p>so you can write rna a3m reader yourself or convert kaggle a3m file to sto</p>\n<hr>\n<p>then you may ask why no rna a3m reader? because proteinX follows alphafold3 paper. only HMMER( i.e.  <em>.sto) for RNA. Protein uses jackhammer(</em>a3m) and HHBlits(*.a3m)</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 3163128,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "03/30/2025 10:39:50",
      "content": "<p>step.4 setup a training database for rna msa<br>\nreference:</p>\n<p>[1] <a href=\"https://github.com/bytedance/Protenix/issues/7\" target=\"_blank\">https://github.com/bytedance/Protenix/issues/7</a><br>\nHow to get seq_to_pdb_index.json when fine-tuning #7</p>\n<p>[2] training guide<br>\n<a href=\"https://github.com/bytedance/Protenix/blob/main/docs/msa_pipeline.md\" target=\"_blank\">https://github.com/bytedance/Protenix/blob/main/docs/msa_pipeline.md</a><br>\n<a href=\"https://github.com/bytedance/Protenix/blob/main/docs/training.md\" target=\"_blank\">https://github.com/bytedance/Protenix/blob/main/docs/training.md</a></p>\n<p>[3] source code<br>\n/Protenix-main/protenix/data/msa_featurizer.py</p>\n<hr>\n<p>basically, the pytorch dataset object will create MSAFeaturizer, which will in turn create RNAMSAFeaturizer.<br>\nPlease check the code of these classes.</p>\n<p>the information below may not be 100%. it can only be confirmed after I use the dataset object for training. nevertheless:</p>\n<pre><code>   strcuture\nDUMMY_MSA\n  |-R1107\n  |    |-rnacentral.sto  #your result from hmmer \n  |\n  |-R1156\n  |    |-rnacentral.sto\n  |\n  |-seq_to_pdb_index.json\n\n\nseq_to_pdb_index.json:\n{\n    : [],\n    : []\n}\n\nnote that [] will  the folder    *sto \n</code></pre>\n<p>then try the test code:</p>\n<pre><code>rna_msa_dir = \nseq_to_pdb_idx_path = \n\n\nf = RNAMSAFeaturizer(\n    seq_to_pdb_idx_path = seq_to_pdb_index_file,  \n    indexing_method = , \n    merge_method = ,\n    seq_limits = SEQ_LIMITS,\n    max_size = ,\n    rna_msa_dir = rna_msa_dir, \n    \n)\n( f.seq_to_pdb_idx )\n\nr = f.process_single_sequence(\n    pdb_name= , \n    sequence=seq,\n    pdb_id=target_id,\n    is_homomer_or_monomer=,\n)\n\n k,v  r.items():\n    (,k)\n    (v)\n</code></pre>\n<p>modify the code to print raw_msa_paths<br>\nraw_msa_paths [DUMMY_MSA/R1107/rnacentral.sto']</p>\n<pre><code> ():\n...\n ()\n  \n\n     (</code></pre>",
      "votes": null,
      "replies": [
        {
          "id": 3163318,
          "author_name": "hengck23",
          "author_url": "",
          "post_date": "03/30/2025 16:07:07",
          "content": "<p>let's try if we can do infrerence with rna msa.</p>\n<p>update the configs_data.py used by runner/inference.py:</p>\n<pre><code> = \n\n\n = ...\n = ...\n = ...\n</code></pre>\n<p>trick to run inference without bash file sh:</p>\n<p>copy inference.py and make the following change:</p>\n<pre><code>\n os\nos.environ[] = \nos.environ[] = \nos.environ[] = \n\n sys\nsys.path.insert(, )\n\n....\n\n __name__ == :\n    N_sample = \n    N_step = \n    N_cycle = \n    seed = \n    use_deepspeed_evo_attention=\n    input_json_path=\n    dump_dir=\n\n   \n    sys.argv +=[\n        , ,\n        , ,\n        , ,\n        , ,\n        , ,\n        , ,\n    ]\n</code></pre>\n<p>to see how rna mas is loaded, we can check the function get_inference_dataloader() of infer_data_pipeline.py. it leads us to:</p>\n<pre><code> (object):\n    \n\nwe need  modify this  !!!!!\n...  be continued ...\n</code></pre>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 3175261,
      "author_name": "doheon114",
      "author_url": "",
      "post_date": "04/10/2025 00:47:35",
      "content": "<p>Thank you for sharing! I am wondering about the results that you have done inferencing with msa or training it! Can you please share it with us?</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 3177791,
      "author_name": "hvanphucs112",
      "author_url": "",
      "post_date": "04/13/2025 09:30:12",
      "content": "<p>I appreciate your input; it will be very helpful to me.</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "3163015": "##background\n- As mentioned in the other discussions, all AF3 clones (proteinX, Boltz,Chai-1) does not support RNA MSA.\n- Host showed that orginal alphafold3 improved results when kaggle rMSA is used (this could mean that there are useful information in MSA)\n- Many of the kagglers are newble (including me) in RNA 3d structure prediction, MSA search, let's help each other!\n- proteinX is chosen becuase its API does support RNA MSA, but it is not used. \n\n\n##We follow orginal AF3 paper and proteinX paper first:\n-Supplementary information:\n\"Accurate structure prediction of biomolecular interactions with AlphaFold 3\"\n\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2Faa721021309efc354d5af37edf3a2a6f%2FSelection_173.png?generation=1743318929775451&alt=media)\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2F83d610199cf7cf3248824295b7963e9e%2FSelection_172.png?generation=1743318971095569&alt=media)\n\n\n---\n\nnote: this is still in experiment. so maybe the parameters may not be correct (e.g. msa search parameters)",
    "3163018": "step.0: setup\n\nHere we how example for RNACentral.\n- download RNACentral\n- install HMMER\n- install mmseq2\n(these are quite standard installations. one can just follows instruction from their repo, so i won't go over the details. in doubt, chatgpt can help)\n\nhere are my downloaded RNACentral files:\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2F22f5826e6a2870fa0c695d497bc84df3%2FSelection_174.png?generation=1743319262803448&alt=media)\n\nNOTE: i have not followed the cutoff date requirement for early prize, etc ... (since i am still experimenting). Plesae adjust for your case accordingly",
    "3163024": "step.1 clustering with mmseqs\n\n```\n\nimport os, sys\n\nSSD_DIR = '/media/hp/xxx-xxx-xxx-xxx'\nHMMER_DIR =f'{SSD_DIR}/my-msa-server/tool/hmmer/binary/bin'\nMMSEQS_DIR =f'{SSD_DIR}/my-msa-server/tool/mmseqs/mmseqs/bin/'\ndb_file = \\\n    f'{SSD_DIR}/my-msa-server/database/RNAcentral/rnacentral_species_specific_ids.fasta'\n    #f'{SSD_DIR}/my-msa-server/database/RNAcentral/mini_rnacentral.fasta' #smaller file for debug\n    #seqkit -n 1000 ...\n\nif 1: \n    os.makedirs('mmseqs_tmp', exist_ok=True)\n\n    cmd = \" \".join([\n        f'{MMSEQS_DIR}/mmseqs',\n        f'easy-linclust',\n        f'{db_file}',\n        f'rnacentral_linclust',\n        f'mmseqs_tmp',\n        f'--min-seq-id 0.9',\n        f'-c 0.8',\n        f'--cov-mode 0',\n        f'--kmer-per-seq-scale 0.3',\n        f'--threads 72',\n    ])\n    output = os.popen(cmd).read()\n    print(output)\n\n\n```\n\nchatgpt:\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2Fd5b47503eb5879d60ee6c6f985b73433%2FSelection_175.png?generation=1743320248195126&alt=media)\n\n\n---\n\noutput:\n\nNumber of clusters: 16972265\nThis means MMseqs grouped ~37.9 million RNAcentral sequences down to ~17 million representative",
    "3163045": "step.2. search with hmmer\n\n```\n    query_file = f'{SSD_DIR}/2025/kaggle/stanford-rna-3d-folding/data/my-data/casp15/R1107.fasta'\n    # note seq length.flag F3 is 0.02 or 0.00005 dependent on seq length \n \n    cluster_db_file = \\\n    f'{SSD_DIR}/my-msa-server/database/RNAcentral/clustered/rnacentral_linclust_rep_seq.fasta'\n    cmd = \" \".join([\n        f'{HMMER_DIR}/nhmmer',\n        f'--rna',\n        f'--cpu 16',\n        f'-E 0.001',\n        f'--incE 0.001',\n        f'--watson',\n        f'--F3 0.00005',\n        f'--tblout nhmmer_hits.tbl',\n        f'-A nhmmer_hits.sto',\n        f'{query_file} {cluster_db_file}',\n    ])\n    output = os.popen(cmd).read()\n    print(output)\n```\nQuery model(s):                            1  (69 nodes)\nTarget sequences:                   16972265  (14294762522 residues searched)\nResidues passing SSV filter:       183030293  (0.0128); expected (0.02)\nResidues passing bias filter:      176205600  (0.0123); expected (0.02)\nResidues passing Vit filter:        14003806  (0.00098); expected (0.003)\nResidues passing Fwd filter:          165886  (1.16e-05); expected (5e-05)\nTotal number of hits:                      4  (1.87e-08)\nElapsed: 00:01:10.64\n\nprotineX is reading in *.sto\n\n ![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2Fc05bfd930381eff9c31c80c6fe2de1a3%2FSelection_177.png?generation=1743323280546390&alt=media)",
    "3163050": "step.3: check if *.sto can be read by proteinX?\n\nreference:\n[1] https://github.com/bytedance/Protenix/issues/7\nHow to get seq_to_pdb_index.json when fine-tuning #7\n\n[2] rna msa related source code\n/Protenix-main/protenix/data/msa_featurizer.py\n/Protenix-main/protenix/data/msa_utils.py\n\n````\n\n#first test this function\n\n\nfrom protenix.data.msa_featurizer import *\ntarget_id = 'R1107'\nseq = 'GGGGGCCACAGCAGAAGCGUUCACGUCGCAGCCCCUGUCAGCCAUUGCACUCCGGCUGCGAAUUCUGCU'\nraw_msa_paths = [\n    '/000/DUMMY_MSA/0/nhmmer_hits.sto'   #result from hmmer search\n]\n\n\nr = process_single_sequence(\n    pdb_name=target_id,\n    sequence=seq,\n    raw_msa_paths=raw_msa_paths, \n    seq_limits=SEQ_LIMITS, # Optional[list[str]],\n    msa_entity_type =\"rna\",\n    msa_type = \"non_pairing\",\n)\nfor k,v in r.items():\n    print('**',k)\n    print(v)\n\n#-----\n\n#process_single_sequence() will end up calling parse_rna_msa_data() in msa_utils.py\n#e.g. \n\nmsa_data = parse_rna_msa_data(...)\n\nfor k,v in msa_data.items():\n    print('**',k)\n    print(v)\n    \n** /media/hp/c30d34ed-0d55-4077-82dc-b56cd13dd548/2025/kaggle/stanford-rna-3d-folding/code/proteinx01/000/DUMMY_MSA/0/nhmmer_hits.sto\n\nMsa(sequences=['GGGGGCCACAGCAGAAGCGUUCACGUCGCAGCCCCUGUCAGCCAUUGCACUCCGGCUGCGAAUUCUGCU', 'GGGGGCCACAGCAGAAGCGUUCACGUCGCGGCCCCUGUCAGCCAUUGCACUCCGGCUGCGAAUUCUGCU', 'GGGGGCCAUAGCAGAAGCGUUCACGUCGCAGCCCCUGUCAGAUUCU--UACGAACCUGCGAAUUCUGCU', '-GGGGCCACAGCAGAAGCGUUCACGUCGCGGCCCCUGUCAGAUUCUG--GUGAAUCUGCGAAUUCUGCU', '-GGCGCCAUAGCAGAAGCGUUCACGUCGCAGCCCCUGUCAGAUUCU--UACGAAUCUGCGAAUUCUGC-'], \n\ndeletion_matrix=[[0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0], [0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0], [0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0], [0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0], [0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0]], \n\ndescriptions=['query', 'URS0002617738_9598/1-69', 'URS00006F3FF9_7897/12-78', 'URS0000626D9E_230844/13-78', 'URS0000A93AF3_7897/13-77'])\n````\n\n\na natural question is can kaggle a2m/a3m file be read by proteinX?\nanswer is yes and not\n- if you check the code, for protein input, proteinX accept  a3m file (and sto file)\n- however, for rna input, it read only sto file\n\nso you can write rna a3m reader yourself or convert kaggle a3m file to sto\n\n---\n\nthen you may ask why no rna a3m reader? because proteinX follows alphafold3 paper. only HMMER( i.e.  *.sto) for RNA. Protein uses jackhammer(*a3m) and HHBlits(*.a3m)",
    "3163128": "step.4 setup a training database for rna msa\nreference:\n\n[1] https://github.com/bytedance/Protenix/issues/7\nHow to get seq_to_pdb_index.json when fine-tuning #7\n\n[2] training guide\nhttps://github.com/bytedance/Protenix/blob/main/docs/msa_pipeline.md\nhttps://github.com/bytedance/Protenix/blob/main/docs/training.md\n\n[3] source code\n/Protenix-main/protenix/data/msa_featurizer.py\n\n----\n\nbasically, the pytorch dataset object will create MSAFeaturizer, which will in turn create RNAMSAFeaturizer.\nPlease check the code of these classes.\n\nthe information below may not be 100%. it can only be confirmed after I use the dataset object for training. nevertheless:\n\n```\nset up file strcuture\nDUMMY_MSA\n  |-R1107\n  |    |-rnacentral.sto  #your result from hmmer search\n  |\n  |-R1156\n  |    |-rnacentral.sto\n  |\n  |-seq_to_pdb_index.json\n\n\nseq_to_pdb_index.json:\n{\n    \"GGGGGCCACAGCAGAAGCGUUCACGUCGCAGCCCCUGUCAGCCAUUGCACUCCGGCUGCGAAUUCUGCU\": [\"R1107\"],\n    \"GGAGCAUCGUGUCUCAAGUGCUUCACGGUCACAAUAUACCGUUUCGUCGGGUGCGUGGCAAUUCGGUGCACAUCAUGUCUUUCGUGGCUGGUGUGGCUCCUCAAGGUGCGAGGGGCAAGUAUAGAGCAGAGCUCC\": [\"R1156\"]\n}\n\nnote that [\"R1107\"] will be the folder to search for *sto file\n```\n\nthen try the test code:\n\n```\nrna_msa_dir = 'DUMMY_MSA'\nseq_to_pdb_idx_path = 'DUMMY_MSA/seq_to_pdb_index.json'\n\n\nf = RNAMSAFeaturizer(\n    seq_to_pdb_idx_path = seq_to_pdb_index_file,  \n    indexing_method = 'sequence', #'pdb_id_entity_id', #\"sequence\",\n    merge_method = 'dense_max',\n    seq_limits = SEQ_LIMITS,\n    max_size = 16384,\n    rna_msa_dir = rna_msa_dir, \n    #**kwargs,\n)\nprint( f.seq_to_pdb_idx )\n\nr = f.process_single_sequence(\n    pdb_name=f'{target_id}_1' , #pdb_id_entity_id\n    sequence=seq,\n    pdb_id=target_id,\n    is_homomer_or_monomer=True,\n)\n\nfor k,v in r.items():\n    print('**',k)\n    print(v)\n\n```\nmodify the code to print raw_msa_paths\nraw_msa_paths [DUMMY_MSA/R1107/rnacentral.sto']\n\n\n```\nclass RNAMSAFeaturizer(BaseMSAFeaturizer):\n...\ndef get_msa_path(\n        self, db_name: str, sequence: str, pdb_id_entity_id: str, reduced: bool = False\n    )\n  ##!!!! reduced must be patch to False  !!!!\n\n    def process_single_sequence(\n        self,\n        pdb_name: str,\n ...\n\n        raw_msa_paths, seq_limits = [], []\n        for db_name in self.non_pairing_db:\n            if opexists(\n                path := self.get_msa_path(db_name, sequence, pdb_name)\n            ) and path.endswith(\".sto\"):\n                raw_msa_paths.append(path)\n                seq_limits.append(self.seq_limits.get(db_name, SEQ_LIMITS[db_name]))\n        print('HCK!!!! raw_msa_paths', raw_msa_paths)\n```",
    "3163318": "let's try if we can do infrerence with rna msa.\n\nupdate the configs_data.py used by runner/inference.py:\n```\nmsa.enable_rna_msa = True\n\n#used by training ... but we updated these just in case ...\nrna.seq_to_pdb_idx_path = ...\nrna.rna_msa_dir = ...\nrna.indexing_method = ...\n\n```\n\n\ntrick to run inference without bash file sh:\n\ncopy inference.py and make the following change:\n\n```\n#top of inference.py:\nimport os\nos.environ[\"CUTLASS_PATH\"] = \"....your path.../cutlass/cutlass\"\nos.environ[\"LAYERNORM_TYPE\"] = \"fast_layernorm\"\nos.environ[\"USE_DEEPSPEED_EVO_ATTENTION\"] = \"true\"\n\nimport sys\nsys.path.insert(0, '/.... run local opy instead of pip installed site-package ..../Protenix-main')\n\n....\n\nif __name__ == \"__main__\":\n    N_sample = 5\n    N_step = 200\n    N_cycle = 10\n    seed = 101\n    use_deepspeed_evo_attention=True\n    input_json_path=\"... your path ..../casp15-msa.json\"\n    dump_dir=\"... your path ..../casp15-msa\"\n\n   #fake command line arguments\n    sys.argv +=[\n        '--seeds', f'{seed}',\n        '--dump_dir', f'{dump_dir}',\n        '--input_json_path', f'{input_json_path}',\n        '--model.N_cycle', f'{N_cycle}',\n        '--sample_diffusion.N_sample', f'{N_sample}',\n        '--sample_diffusion.N_step', f'{N_step}',\n    ]\n```\n\n\nto see how rna mas is loaded, we can check the function get_inference_dataloader() of infer_data_pipeline.py. it leads us to:\n```\n\nclass InferenceMSAFeaturizer(object):\n    # Now we only support protein msa in inference\n\nwe need to modify this function !!!!!\n... to be continued ...\n\n```",
    "3175261": "Thank you for sharing! I am wondering about the results that you have done inferencing with msa or training it! Can you please share it with us?",
    "3177791": "I appreciate your input; it will be very helpful to me."
  },
  "source": "meta"
}