{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"name":"python","version":"3.10.12","mimetype":"text/x-python","codemirror_mode":{"name":"ipython","version":3},"pygments_lexer":"ipython3","nbconvert_exporter":"python","file_extension":".py"},"kaggle":{"accelerator":"none","dataSources":[{"sourceId":87793,"databundleVersionId":12276181,"sourceType":"competition"}],"dockerImageVersionId":30918,"isInternetEnabled":true,"language":"python","sourceType":"notebook","isGpuEnabled":false}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"![Header](https://raw.githubusercontent.com/ammomahdavikia/asset-holding/main/stanford-rna-3d.gif)","metadata":{}},{"cell_type":"markdown","source":"### Why is RNA important?\n\nMany of us recognize RNA as the intermediary between DNA and proteins, following the **Central Dogma of Life**—the flow of genetic information from **DNA → RNA → Protein**.\n\nHowever, RNA is far more than just a messenger. It plays diverse and essential roles in biological systems, from **transporting amino acids for protein synthesis** (tRNA) to **regulating gene expression** (miRNA) and even **catalyzing biochemical reactions** (ribozymes).\n\n![RNA Types](https://raw.githubusercontent.com/ammomahdavikia/asset-holding/main/rna_types.png)\n\nWhat makes RNA even more fascinating is its potential as the earliest biomolecule of life. Unlike DNA, RNA can both store genetic information and catalyze its own replication, supporting the **RNA World Hypothesis**—the idea that life may have originated with self-replicating RNA molecules.\n\nThe critical functional roles of RNAs make them a new type of **drug target**. Since many biological functions depend on the specifc tertiary structures of RNAs, it is imperative to determine the 3D structures of RNAs in order to facilitate RNA-based function annotation and drug discovery.","metadata":{}},{"cell_type":"markdown","source":"### ↓ Libraries","metadata":{}},{"cell_type":"code","source":"!pip install biopython\n!pip install ViennaRNA\n!pip install py3Dmol\n!pip install forgi","metadata":{"_uuid":"8f2839f25d086af736a60e9eeb907d3b93b6e0e5","_cell_guid":"b1076dfc-b9ad-4769-8c92-a6c4dae69d19","trusted":true,"_kg_hide-output":true,"execution":{"iopub.status.busy":"2025-05-20T10:40:14.473732Z","iopub.execute_input":"2025-05-20T10:40:14.474112Z","iopub.status.idle":"2025-05-20T10:40:40.341331Z","shell.execute_reply.started":"2025-05-20T10:40:14.474081Z","shell.execute_reply":"2025-05-20T10:40:40.340226Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"### ↓ Imports","metadata":{}},{"cell_type":"code","source":"import os\nimport re\nimport requests\nimport zipfile\nimport pandas as pd\nimport numpy as np\nimport matplotlib.pyplot as plt\nimport matplotlib.font_manager as fm\nfrom matplotlib.colors import ListedColormap\nimport seaborn as sns\nfrom scipy.signal import resample\nfrom Bio import PDB, SeqIO\nfrom Bio.PDB import MMCIFParser\nfrom io import StringIO\nimport RNA\nimport py3Dmol\nimport forgi.graph.bulge_graph as fgb\nimport forgi.visual.mplotlib as fvm\nfrom collections import Counter\nfrom tqdm import tqdm\nfrom joblib import Parallel, delayed\nfrom IPython.display import display_html, display\nimport ipywidgets as widgets\nimport warnings\n\n\nwarnings.simplefilter(action='ignore', category=FutureWarning)\n\nclass clr:\n    B = '\\033[1m'\n    S = '\\033[1m' + '\\033[94m'\n    L = '\\033[1m' + '\\033[91m'\n    E = '\\033[0m'\n\nprint(clr.B+'\\n----- Font -----\\n'+clr.E)\nfont_url = 'https://ifonts.xyz/core/ifonts-files/downloads/472336/carbon-plus-font.zip'\noutput_path = 'carbon-plus.zip'\n\nresponse = requests.get(font_url, stream=True)\n\nif response.status_code == 200:\n    with open(output_path, 'wb') as f:\n        for chunk in response.iter_content(1024):\n            f.write(chunk)\n    print(f'[Success] downloaded: {output_path}')\nelse:\n    print('[Error] Failed to download the font. Check the URL.')\n\nwith zipfile.ZipFile(output_path, 'r') as zip_ref:\n    zip_ref.extractall('carbon_plus_font')\n    print('[Success] extracted.')\n    os.system(f'rm {output_path}')\n\nfont_path = '/kaggle/working/carbon_plus_font/carbonplus-regular-bl.otf'\ntry:\n    fm.fontManager.addfont(font_path)\n    primary_font = fm.FontProperties(fname=font_path).get_name()\n    plt.rcParams['font.family'] = [primary_font, \"DejaVu Sans\"]\n    print('[Success] loaded.')\nexcept:\n    print('[Error] Failed to load the font. Check the address.')\n\ncolors = ['#780000', '#c1121f', '#fdf0d5', '#02c39a', '#669bbc', '#003049']\nnucleotides = ['X', 'C', 'U', 'G', 'A', '-']\nnt_clr = {nt: colors[idx] for idx, nt in enumerate(nucleotides)}\nprint(clr.B+'\\n----- Color -----\\n'+clr.E)\nsns.palplot(sns.color_palette(colors))","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-05-20T10:40:40.342805Z","iopub.execute_input":"2025-05-20T10:40:40.343219Z","iopub.status.idle":"2025-05-20T10:40:42.603861Z","shell.execute_reply.started":"2025-05-20T10:40:40.343174Z","shell.execute_reply":"2025-05-20T10:40:42.602979Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"### ↓ Helpers","metadata":{}},{"cell_type":"code","source":"def interactive_subplot(plot_func, dfs_dict):\n    output = widgets.Output()\n\n    def on_select(change):\n        with output:\n            output.clear_output(wait=True)\n            plot_func(dfs_dict[change.new], change.new)\n            plt.show()\n    \n    dropdown = widgets.Dropdown(\n        options=dfs_dict.keys(),\n        description='Dataset:',\n        style={'description_width': 'initial'}\n    )\n    \n    dropdown.observe(on_select, names='value')\n    display(dropdown, output)\n    \n    first_key = next(iter(dfs_dict.keys()), None)\n    if first_key:\n        with output:\n            plot_func(dfs_dict[first_key], first_key)\n            plt.show()\n\ndef parallel_apply(series, func, desc, n_jobs=-1):\n    \"\"\"Parallel apply with tqdm.\"\"\"\n    results = Parallel(n_jobs=n_jobs)(\n        delayed(func)(x) for x in tqdm(series, desc=desc)\n    )\n    return results","metadata":{"trusted":true,"_kg_hide-input":true,"execution":{"iopub.status.busy":"2025-05-20T10:40:42.605462Z","iopub.execute_input":"2025-05-20T10:40:42.605979Z","iopub.status.idle":"2025-05-20T10:40:42.613285Z","shell.execute_reply.started":"2025-05-20T10:40:42.605939Z","shell.execute_reply":"2025-05-20T10:40:42.612212Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"# <div style=\"text-align:center; border-radius:10px; color:black; margin:0; font-size:100%; font-family:Carbon Plus; background-color:white; overflow:hidden\"><b>Loading the data</b></div>","metadata":{}},{"cell_type":"code","source":"class Config():\n    def __init__(self):\n        self.PATH = '/kaggle/input/stanford-rna-3d-folding'\n\nconfig = Config()\nos.makedirs('Figures', exist_ok=True)","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-05-20T10:40:42.614868Z","iopub.execute_input":"2025-05-20T10:40:42.615240Z","iopub.status.idle":"2025-05-20T10:40:42.634776Z","shell.execute_reply.started":"2025-05-20T10:40:42.615209Z","shell.execute_reply":"2025-05-20T10:40:42.633484Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"def load_data():\n    \n    train_seq = pd.read_csv(os.path.join(config.PATH, 'train_sequences.csv'))\n    train_v2_seq = pd.read_csv(os.path.join(config.PATH, 'train_sequences.v2.csv'))\n    valid_seq = pd.read_csv(os.path.join(config.PATH, 'validation_sequences.csv'))\n    test_seq = pd.read_csv(os.path.join(config.PATH, 'test_sequences.csv'))\n\n    train_labels = pd.read_csv(os.path.join(config.PATH, 'train_labels.csv'))\n    train_v2_labels = pd.read_csv(os.path.join(config.PATH, 'train_labels.v2.csv'))\n    valid_labels = pd.read_csv(os.path.join(config.PATH, 'validation_labels.csv'))\n\n    return train_seq, train_v2_seq, valid_seq, test_seq, train_labels, train_v2_labels, valid_labels\n\ndef summarize(df, desc='Summary'):\n\n    start = clr.S if 'Sequence' in desc else clr.L\n    \n    print(start+f'\\n----- {desc} -----\\n'+clr.E)\n    print(f'Shape: {df.shape}')\n    print(f'Missing: {df.isna().sum().sum()}')\n    print(f'Columns: {df.columns.to_list()}\\n')\n    display_html(df.head(3))\n    print('\\n')","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-05-20T10:40:42.635810Z","iopub.execute_input":"2025-05-20T10:40:42.636194Z","iopub.status.idle":"2025-05-20T10:40:42.645037Z","shell.execute_reply.started":"2025-05-20T10:40:42.636155Z","shell.execute_reply":"2025-05-20T10:40:42.643954Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"train_seq, train_v2_seq, valid_seq, test_seq, train_labels, train_v2_labels, valid_labels = load_data()\n\nfor df, desc in zip([train_seq, train_v2_seq, valid_seq, test_seq, train_labels, train_v2_labels, valid_labels],\n                    ['Train Sequence', 'Train Sequence 2', 'Validation Sequence','Test Sequence',\n                     'Train Label', 'Train Label 2', 'Validation Label']):\n    summarize(df, desc)","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-05-20T10:40:42.645944Z","iopub.execute_input":"2025-05-20T10:40:42.646248Z","iopub.status.idle":"2025-05-20T10:40:52.220193Z","shell.execute_reply.started":"2025-05-20T10:40:42.646223Z","shell.execute_reply":"2025-05-20T10:40:52.219177Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"**Training data update**:\n\nAs announced by the organizers, a new set of training data (`train_sequences.v2.csv`) with more relaxed filters has been released.\n1. They downloaded pdbs from the protein data bank with full text search for keyword RNA\n2. They relaxed filter for unstructured RNAs based on pairwise C1' distances, where 20% of residues have to be close to some other residue that is over 4 bases apart\n\nLet's see how much are the new sequences similar to the previous training set:","metadata":{}},{"cell_type":"code","source":"from matplotlib_venn import venn2\n\ntrain_ids = set(train_seq.target_id)\ntrain_v2_ids = set(train_v2_seq.target_id)\n\nplt.figure(figsize=(8, 8))\nv = venn2([train_ids, train_v2_ids], set_labels=('', ''))\nv.get_label_by_id('10').set_text(f'v1\\n({len(train_ids - train_v2_ids)})')\nv.get_label_by_id('01').set_text(f'v2\\n({len(train_v2_ids - train_ids)})')\nv.get_label_by_id('11').set_text(f'v1 ∩ v2\\n({len(train_ids & train_v2_ids)})')\nplt.title(\"Training data overlap\")\nplt.show()","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-05-20T10:40:52.221043Z","iopub.execute_input":"2025-05-20T10:40:52.221393Z","iopub.status.idle":"2025-05-20T10:40:52.432289Z","shell.execute_reply.started":"2025-05-20T10:40:52.221366Z","shell.execute_reply":"2025-05-20T10:40:52.431253Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"# <div style=\"text-align:center; border-radius:10px; color:black; margin:0; font-size:100%; font-family:Carbon Plus; background-color:white; overflow:hidden\"><b>Road Map</b></div>\n\nThere are four different types of RNA data structures commonly used for RNA 3D folding:\n1. [**Raw Sequence**](#primary-structure): The linear string of nucleotides (A, U, C, G) that make up the RNA molecule. This primary sequence serves as the foundation for predicting higher-order structures and functions.\n2. [**Secondary Structure**](#secondary-structure): Represents the base-pairing interactions within an RNA sequence, typically visualized as dot-bracket notation or base-pair probability matrices. It captures elements like stems, loops, and bulges, which are crucial for understanding RNA folding and function.\n3. [**Tertiary Structure**](#tertiary-structure): The 3D spatial conformation of the RNA molecule, describing how it folds in real space. This includes complex interactions like pseudoknots and long-range contacts, often derived from experimental techniques or computational modeling.\n4. [**Multiple Sequence Alignment (MSA)**](#multiple-sequence-alignment): An alignment of homologous RNA sequences from different organisms. MSAs highlight conserved regions and covariation patterns, offering evolutionary insights that can improve structural and functional predictions.\n\n![RNA Data Structures](https://raw.githubusercontent.com/ammomahdavikia/asset-holding/main/rna_data_structures.png)\n\nWe will explore each of these data structures in detail in the following sections.","metadata":{}},{"cell_type":"markdown","source":"# <div style=\"text-align:center; border-radius:10px; color:black; margin:0; font-size:100%; font-family:Carbon Plus; background-color:white; overflow:hidden\"><b>Primary Structure</b></div>\n\nNucleic acids are macromolecules that exist as polymers called **polynucleotides**. As indicated by the name, each polynucleotide consists of monomers called **nucleotides**. A nucleotide, in general, is composed of three parts: \n+ A five-carbon sugar (a pentose)\n+ A nitrogen-containing (nitrogenous) base\n+ One to three phosphate groups\n\n![RNA Types](https://raw.githubusercontent.com/ammomahdavikia/asset-holding/main/rna_backbone.png)\n\nTo understand the structure of a single nucleotide, let’s first consider the **nitrogenous bases**. Each nitrogenous base has one or two rings that include nitrogen atoms. There are two families of nitrogenous bases: \n+ **Pyrimidines**: A pyrimidine has one six-membered ring of carbon and nitrogen atoms.\n    + Cytosine `C`\n    + Thymine `T`\n    + Uracil `U`\n+ **Purines**: Purine is larger, with a six-membered ring fused to a five-membered ring.\n    + Adenine `A`\n    + Guanine `G`\n\nAdenine, guanine, and cytosine are found in both DNA and RNA; thymine is found only in DNA, and uracil only in RNA. Let's get our hands dirty and explore the dataset a little bit:","metadata":{}},{"cell_type":"code","source":"def nucleotide_component(df, desc, ax=None):\n    print(clr.S+f'\\n----- {desc} components -----\\n'+clr.E)\n    display_html(df.sequence.apply(lambda x: set(x)).value_counts().reset_index())\n\n\nfor df, desc in zip([train_seq, train_v2_seq, valid_seq],\n                    ['Train Sequence', 'Train Sequence V2', 'Validation Sequence']):\n    nucleotide_component(df, desc)","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-05-20T10:40:52.434903Z","iopub.execute_input":"2025-05-20T10:40:52.435213Z","iopub.status.idle":"2025-05-20T10:40:52.544680Z","shell.execute_reply.started":"2025-05-20T10:40:52.435186Z","shell.execute_reply":"2025-05-20T10:40:52.543805Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"As you can see, they contain up to 6 different types of characters based on [FASTA format](https://en.wikipedia.org/wiki/FASTA_format). But don't worry about the addition `X` and `-`, the organizers announced for `test_sequences.csv`, this is guaranteed to be a string of `A`, `C`, `G`, and `U`.\n\n| Symbol | Meaning                 |\n|--------|-------------------------|\n| A      | Adenine                 |\n| U      | Uracil                  |\n| C      | Cytosine                |\n| G      | Guanine                 |\n| X      | Any                     |\n| -      | Gap or missing base     |","metadata":{}},{"cell_type":"code","source":"def plot_nucleotide_frequency(df, desc, ax=None):\n    nucleotide_counts = Counter(''.join(df['sequence']))\n    c = {nt: [count, nt_clr[nt]] for nt, count in nucleotide_counts.items()}\n    c = pd.DataFrame(c).T.reset_index().rename(columns={'index': 'nt', 0: 'count', 1: 'color'})\n    c['frequency'] = c['count'] / c['count'].sum()\n\n    if ax is None:\n        fig, ax = plt.subplots(figsize=(8, 5))\n\n    sns.barplot(data=c, x='nt', y='frequency', palette=c.color.to_list(), ax=ax)\n    ax.set_title(f'Nucleotide frequency - {desc}')\n    ax.set_xlabel('')\n    if fig:\n        plt.savefig(f'Figures/{desc}_plot_nucleotide_frequency.png', dpi=300, bbox_inches='tight')","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-05-20T10:40:52.546785Z","iopub.execute_input":"2025-05-20T10:40:52.547114Z","iopub.status.idle":"2025-05-20T10:40:52.553751Z","shell.execute_reply.started":"2025-05-20T10:40:52.547087Z","shell.execute_reply":"2025-05-20T10:40:52.552617Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"plot_nucleotide_frequency(train_seq, 'Train')","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-05-20T10:40:52.554668Z","iopub.execute_input":"2025-05-20T10:40:52.555022Z","iopub.status.idle":"2025-05-20T10:40:53.058875Z","shell.execute_reply.started":"2025-05-20T10:40:52.554995Z","shell.execute_reply":"2025-05-20T10:40:53.057884Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"plot_nucleotide_frequency(train_v2_seq, 'Train V2')","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-05-20T10:40:53.059946Z","iopub.execute_input":"2025-05-20T10:40:53.060237Z","iopub.status.idle":"2025-05-20T10:40:53.763218Z","shell.execute_reply.started":"2025-05-20T10:40:53.060212Z","shell.execute_reply":"2025-05-20T10:40:53.762284Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"The higher the percent of **G:C base pairs** in the nucleic acids (and hence the lower the content of A:T base pairs), the higher is the melting point. How do we explain this behavior? \n\nG:C base pairs contribute more to the stability of nucleic acids than do A:T base pairs because of the greater number of hydrogen bonds for the former (three in a G:C base pair vs. two for A:T), but also, importantly, because the stacking interactions of G:C base pairs with adjacent base pairs are more favorable than the corresponding interactions of A:T base pairs with their neighboring base pairs.","metadata":{}},{"cell_type":"code","source":"def CG_ratio(seq):\n    nucleotide_counts = Counter(''.join(seq))\n    return (nucleotide_counts['C'] + nucleotide_counts['G']) / len(seq)\n\ndef plot_CG_ratio(df, desc, ax=None):\n    if ax is None:\n        fig, ax = plt.subplots(figsize=(8, 5))\n    sns.histplot(df.sequence.apply(CG_ratio), kde=True, ax=ax, color=colors[-1])\n    ax.set_title(f'C+G Ratio frequency - {desc}')\n    ax.set_xlabel('C+G Ratio')\n    ax.set_xlim(0, 1)\n    if fig:\n        plt.savefig(f'Figures/{desc}_plot_CG_ratio.png', dpi=300, bbox_inches='tight')","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-05-20T10:40:53.764197Z","iopub.execute_input":"2025-05-20T10:40:53.764572Z","iopub.status.idle":"2025-05-20T10:40:53.771115Z","shell.execute_reply.started":"2025-05-20T10:40:53.764537Z","shell.execute_reply":"2025-05-20T10:40:53.770015Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"plot_CG_ratio(train_seq, 'Train')","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-05-20T10:40:53.772017Z","iopub.execute_input":"2025-05-20T10:40:53.772276Z","iopub.status.idle":"2025-05-20T10:40:54.429172Z","shell.execute_reply.started":"2025-05-20T10:40:53.772253Z","shell.execute_reply":"2025-05-20T10:40:54.428220Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"Now let’s add the sugar to which the nitrogenous base is attached. In DNA the sugar is **deoxyribose**; in RNA it is **ribose**. The only difference between these two sugars is that deoxyribose lacks an oxygen atom on the second carbon in the ring, hence the name deoxyribose.\n\nSo far, we have built a nucleoside (base plus sugar). To complete the construction of a nucleotide, we attach one to three **phosphate groups** to the 5' carbon of the sugar (the carbon numbers in the sugar include `'`, the prime symbol); With one phosphate, this is a nucleoside monophosphate, more often called a nucleotide.","metadata":{}},{"cell_type":"code","source":"def plot_backbone(pdb_id):\n    \n    cif_filename = os.path.join(config.PATH, 'PDB_RNA', f\"{pdb_id.lower()}.cif\")\n    \n    with open(cif_filename, 'r') as file:\n        cif_string = file.read()\n    \n    viewer = py3Dmol.view(width=600, height=400)\n    viewer.addModel(cif_string, 'cif')\n    \n    viewer.setStyle({'cartoon': {'colorscheme': 'spectrum'}})\n    viewer.setStyle({'atom': ['P']}, {'sphere': {'color': colors[1], 'radius': 0.8}})\n    viewer.setStyle({'atom': [\"C1'\", \"C2'\", \"C3'\", \"C4'\", \"O4'\"]}, {'stick': {'color': 'spectrum'}})\n    \n    viewer.zoomTo()\n    viewer.show()","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-05-20T10:40:54.430063Z","iopub.execute_input":"2025-05-20T10:40:54.430323Z","iopub.status.idle":"2025-05-20T10:40:54.436584Z","shell.execute_reply.started":"2025-05-20T10:40:54.430302Z","shell.execute_reply":"2025-05-20T10:40:54.435561Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"# Feel free to change the PDB ID\nsample_pdb_id = '1HS3'\nplot_backbone(sample_pdb_id)","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-05-20T10:40:54.437530Z","iopub.execute_input":"2025-05-20T10:40:54.437866Z","iopub.status.idle":"2025-05-20T10:40:54.469804Z","shell.execute_reply.started":"2025-05-20T10:40:54.437816Z","shell.execute_reply":"2025-05-20T10:40:54.468597Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"In this visualiztion, phosphate group and sugar are represented by red sphere and pentagon repectively. Now you might ask, Do we have to predict the location of each atom of the nucleotide? The short answer is **NO**, not in this competition! Here, our goal is to predict the location of each C'1 atom (first-red carbon sugar backbone) of each nucleotide in RNA sequence. \n\nWe will come back to the details of the conformation later in the notebook.","metadata":{}},{"cell_type":"code","source":"for df, desc in zip([train_seq, train_v2_seq, valid_seq, test_seq],\n                    ['Train Sequence', 'Train Sequence V2', 'Validation Sequence','Test Sequence']):\n    \n    df['length'] = df.sequence.apply(lambda x: len(x))\n    print(clr.S+f'\\n----- {desc} Length -----\\n'+clr.E)\n    print(f'[MIN]: {df.length.min()}')\n    print(f'[MAX]: {df.length.max()}')\n    print(f'[AVG]: {df.length.mean():.2f}')\n    print(f'[STD]: {df.length.std():.2f}')\n    print(f'[Q1]: {df.length.quantile([0.25]).iloc[0]:.2f}')\n    print(f'[Q2]: {df.length.quantile([0.5]).iloc[0]:.2f}')\n    print(f'[Q3]: {df.length.quantile([0.75]).iloc[0]:.2f}')","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-05-20T10:40:54.470941Z","iopub.execute_input":"2025-05-20T10:40:54.471214Z","iopub.status.idle":"2025-05-20T10:40:54.507824Z","shell.execute_reply.started":"2025-05-20T10:40:54.471191Z","shell.execute_reply":"2025-05-20T10:40:54.506751Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"def plot_length_histogram(df, desc, ax=None):\n    if ax is None:\n        fig, ax = plt.subplots(figsize=(8, 5))\n    sns.histplot(df.length, kde=True, ax=ax, color=colors[-1])\n    ax.set_title(f'Sequence Length Frequency - {desc}')\n    ax.set_xlabel('Length')\n    if fig:\n        plt.savefig(f'Figures/{desc}_plot_length_histogram.png', dpi=300, bbox_inches='tight')","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-05-20T10:40:54.509015Z","iopub.execute_input":"2025-05-20T10:40:54.509392Z","iopub.status.idle":"2025-05-20T10:40:54.515163Z","shell.execute_reply.started":"2025-05-20T10:40:54.509356Z","shell.execute_reply":"2025-05-20T10:40:54.514012Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"plot_length_histogram(train_seq, 'Train')","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-05-20T10:40:54.516287Z","iopub.execute_input":"2025-05-20T10:40:54.516660Z","iopub.status.idle":"2025-05-20T10:40:55.610805Z","shell.execute_reply.started":"2025-05-20T10:40:54.516629Z","shell.execute_reply":"2025-05-20T10:40:55.609849Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"plot_length_histogram(train_v2_seq, 'Train V2')","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-05-20T10:40:55.611938Z","iopub.execute_input":"2025-05-20T10:40:55.612231Z","iopub.status.idle":"2025-05-20T10:40:56.193676Z","shell.execute_reply.started":"2025-05-20T10:40:55.612205Z","shell.execute_reply":"2025-05-20T10:40:56.192771Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"# <div style=\"text-align:center; border-radius:10px; color:black; margin:0; font-size:100%; font-family:Carbon Plus; background-color:white; overflow:hidden\"><b>Exploration</b></div>","metadata":{}},{"cell_type":"code","source":"for df, desc in zip([train_seq, train_v2_seq, valid_seq, test_seq],\n                    ['Train Sequence', 'Train Sequence V2', 'Validation Sequence','Test Sequence']):\n    \n    df['temporal_cutoff'] = pd.to_datetime(df['temporal_cutoff'], errors='coerce')\n    print(clr.S+f'\\n----- {desc} Cutoff -----\\n'+clr.E)\n    print(f'[MIN]: {df.temporal_cutoff.min()}')\n    print(f'[MAX]: {df.temporal_cutoff.max()}')","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-05-20T10:40:56.194697Z","iopub.execute_input":"2025-05-20T10:40:56.195041Z","iopub.status.idle":"2025-05-20T10:40:56.215116Z","shell.execute_reply.started":"2025-05-20T10:40:56.195014Z","shell.execute_reply":"2025-05-20T10:40:56.214041Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"def plot_cutoff_histogram(df, desc, axs=None):\n    if axs is None:\n        fig, axs = plt.subplots(nrows=1, ncols=2, figsize=(12, 6), gridspec_kw={'width_ratios': [3, 1]})\n    else:\n        fig = None\n    \n    years = df.temporal_cutoff.dt.to_period('Y')\n    year_strs = years.astype(str)\n\n    year_order = sorted(year_strs.unique())\n    year_cat = pd.Categorical(year_strs, categories=year_order, ordered=True)\n\n    bars = sns.histplot(year_cat, kde=True, ax=axs[0], color=colors[-1])\n    axs[0].set_title('Temporal Cutoff Frequency')\n    axs[0].set_xlabel('Date')\n    plt.setp(axs[0].xaxis.get_majorticklabels(), rotation=45)\n    \n    for year, bar in zip(year_order ,bars.patches):\n        if int(year) > 2021:\n            bar.set_facecolor(colors[1])\n\n    valid_cutoff = (df.temporal_cutoff < valid_seq.temporal_cutoff.min()).value_counts().reset_index()\n    sizes = valid_cutoff['count'].to_list()\n    labels = valid_cutoff['temporal_cutoff'].to_list()\n    valid_map = {True: ['Valid', colors[3]], False: ['Invalid', colors[1]]}\n    validity = [f'{valid_map[label][0]} ({size})' for size, label in zip(sizes, labels)]\n    _colors = [valid_map[label][1] for label in labels]\n    axs[1].pie(sizes, labels=validity, colors=_colors)\n    axs[1].set_title('Cutoff Validity')\n\n    if fig:\n        plt.tight_layout()\n        plt.savefig(f'Figures/{desc}_plot_cutoff_histogram.png', dpi=300, bbox_inches='tight')","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-05-20T10:40:56.216302Z","iopub.execute_input":"2025-05-20T10:40:56.216679Z","iopub.status.idle":"2025-05-20T10:40:56.239168Z","shell.execute_reply.started":"2025-05-20T10:40:56.216642Z","shell.execute_reply":"2025-05-20T10:40:56.238147Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"plot_cutoff_histogram(train_seq, 'Train')","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-05-20T10:40:56.240351Z","iopub.execute_input":"2025-05-20T10:40:56.240741Z","iopub.status.idle":"2025-05-20T10:40:57.570071Z","shell.execute_reply.started":"2025-05-20T10:40:56.240690Z","shell.execute_reply":"2025-05-20T10:40:57.569016Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"plot_cutoff_histogram(train_v2_seq, 'Train V2')","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-05-20T10:40:57.574418Z","iopub.execute_input":"2025-05-20T10:40:57.574761Z","iopub.status.idle":"2025-05-20T10:40:59.043704Z","shell.execute_reply.started":"2025-05-20T10:40:57.574729Z","shell.execute_reply":"2025-05-20T10:40:59.042663Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"If you have decided to go for early sharing prizes, make sure that you train only on **train_sequences.csv** and **train_sequences.v2.csv** that have `temporal_cutoff` before the test_sequences (`2022-05-27` is a safe date).","metadata":{}},{"cell_type":"code","source":"def count_chains(fasta_string):\n    try:\n        chains = len(fasta_string.split('\\n')) // 2\n    except:\n        chains = 0\n    return chains == 1\n\ndef plot_chains(df, desc, ax=None):\n    fig = None\n    if ax is None:\n        fig, ax = plt.subplots(figsize=(6, 6))\n\n    palette = {True: colors[2], False: colors[-1]}\n    sns.histplot(df.all_sequences.apply(count_chains), ax=ax, discrete=True, color=colors[-1])\n\n    ax.set_title(f'Unique Chains - {desc}')\n    ax.set_xticks([0, 1])\n    ax.set_xticklabels(['Poly', 'Mono'])\n    ax.set_xlabel('')\n\n    if fig:\n        plt.savefig(f'Figures/{desc}_plot_chains.png', dpi=300, bbox_inches='tight')","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-05-20T10:40:59.045624Z","iopub.execute_input":"2025-05-20T10:40:59.045986Z","iopub.status.idle":"2025-05-20T10:40:59.052801Z","shell.execute_reply.started":"2025-05-20T10:40:59.045956Z","shell.execute_reply":"2025-05-20T10:40:59.051624Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"import matplotlib.pyplot as plt\nimport seaborn as sns\nimport numpy as np\n\ndef count_chains(fasta_string):\n    try:\n        chains = len(fasta_string.split('\\n')) // 2\n    except:\n        chains = 0\n    return chains == 1\n\ndef plot_chains(df, desc, ax=None):\n    fig = None\n    if ax is None:\n        fig, ax = plt.subplots(figsize=(6, 6))\n\n    chain_counts = df.all_sequences.apply(count_chains).astype(int)\n\n    sns.histplot(chain_counts, ax=ax, stat='probability', discrete=True, color=colors[-1])\n\n    ax.set_title(desc)\n    ax.set_xticks([0, 1])\n    ax.set_xticklabels(['Multi', 'Mono'])\n    ax.set_xlabel('')\n    ax.set_ylabel('Frequency')\n    ax.set_ylim(0., 1.)\n\n    if fig:\n        plt.savefig(f'Figures/{desc}_plot_chains.png', dpi=300, bbox_inches='tight')","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-05-20T10:40:59.054144Z","iopub.execute_input":"2025-05-20T10:40:59.054491Z","iopub.status.idle":"2025-05-20T10:40:59.073782Z","shell.execute_reply.started":"2025-05-20T10:40:59.054451Z","shell.execute_reply":"2025-05-20T10:40:59.072764Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"fig, axs = plt.subplots(nrows=1, ncols=3, figsize=(12, 4))\n\nfor idx, (df, desc) in enumerate(zip([train_seq, train_v2_seq, valid_seq],\n                                     ['Train', 'Train V2', 'Validation'])):\n    plt.suptitle('Unique Chains')\n    plot_chains(df, desc, ax=axs[idx])","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-05-20T10:40:59.074810Z","iopub.execute_input":"2025-05-20T10:40:59.075160Z","iopub.status.idle":"2025-05-20T10:40:59.673988Z","shell.execute_reply.started":"2025-05-20T10:40:59.075124Z","shell.execute_reply":"2025-05-20T10:40:59.672911Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"As you can see, most of the samples in the training data version 2 are multi-stranded which makes sense. Because they are downloaded pdbs from the protein data bank with full text search for keyword **\"RNA\"**. Keep in mind that we only count unique chains. There are molecules with multiple similar chains, we count them as mono-stranded here.","metadata":{}},{"cell_type":"code","source":"for df, desc in zip([train_labels, train_v2_labels, valid_labels],\n                    ['Train Label', 'Train Label V2', 'Validation Label']):\n    \n    df['chain'] = df.ID.apply(lambda x: x.split('_')[1])\n    df['pdb_id'] = df.ID.apply(lambda x: x.split('_')[0])\n    df['target_id'] = df.ID.apply(lambda x: re.sub(r'_(\\d+)$', '', x))\n    missing_count = df.groupby('target_id').x_1.apply(lambda x: x.isna().sum())\n\n    if 'V2' in desc:\n        o_missing_count = pd.merge(missing_count.reset_index(), train_v2_seq, on='target_id', how='right')['x_1']\n        train_v2_seq['missing_label'] = o_missing_count\n        train_v2_seq['missing_ratio'] = train_v2_seq['missing_label'] / train_v2_seq['length']\n    elif 'Train' in desc:\n        o_missing_count = pd.merge(missing_count.reset_index(), train_seq, on='target_id', how='right')['x_1']\n        train_seq['missing_label'] = o_missing_count\n        train_seq['missing_ratio'] = train_seq['missing_label'] / train_seq['length']\n    else:\n        o_missing_count = pd.merge(missing_count.reset_index(), valid_seq, on='target_id', how='right')['x_1']\n        valid_seq['missing_label'] = o_missing_count\n        valid_seq['missing_ratio'] = valid_seq['missing_label'] / valid_seq['length']\n\n    print(clr.L+f'\\n----- {desc} Missing -----\\n'+clr.E)\n    print(f'[COUNT]: {missing_count[missing_count>0].count()}')\n    print(f'[MIN]: {missing_count.min()}')\n    print(f'[MAX]: {missing_count.max()}')\n    print(f'[AVG]: {missing_count.mean():.2f}')","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-05-20T10:40:59.675008Z","iopub.execute_input":"2025-05-20T10:40:59.675377Z","iopub.status.idle":"2025-05-20T10:41:08.498312Z","shell.execute_reply.started":"2025-05-20T10:40:59.675347Z","shell.execute_reply":"2025-05-20T10:41:08.497244Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"def plot_missing_label(df, desc, axs=None):\n    \n    if axs is None:\n        fig, axs = plt.subplots(nrows=1, ncols=2, figsize=(12, 6), gridspec_kw={'width_ratios': [1, 4]})\n    else:\n        fig = None\n\n    s=df['missing_label'] > 0\n    sns.countplot(\n        x=s,\n        ax=axs[0],\n        palette=[colors[4], colors[-1]]\n    )\n    axs[0].set_title('Any coordinate missing?')\n    axs[0].set_xlabel('')\n    ticks = axs[0].get_xticks()\n    _labels = ['No', 'Yes']\n    axs[0].set_xticks(ticks=ticks, labels=_labels[:len(ticks)])\n    axs[0].set_ylabel('Count')\n    ylim = axs[0].get_ylim()\n\n    sns.histplot(\n        df.loc[df.missing_label > 0, 'missing_ratio'], \n        ax=axs[1],\n        color=colors[-1],\n        kde=True\n    )\n\n    axs[1].set_title('Missing Ratio')\n    axs[1].set_xlabel('Ratio')\n    axs[1].set_ylabel('')\n    axs[1].set_ylim(ylim)\n\n    if fig:\n        plt.tight_layout()\n        plt.savefig(f'Figures/{desc}_plot_missing_label.png', dpi=300, bbox_inches='tight')","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-05-20T10:41:08.499375Z","iopub.execute_input":"2025-05-20T10:41:08.499680Z","iopub.status.idle":"2025-05-20T10:41:08.508054Z","shell.execute_reply.started":"2025-05-20T10:41:08.499651Z","shell.execute_reply":"2025-05-20T10:41:08.506990Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"plot_missing_label(train_seq, 'Train')","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-05-20T10:41:08.509093Z","iopub.execute_input":"2025-05-20T10:41:08.509446Z","iopub.status.idle":"2025-05-20T10:41:09.534557Z","shell.execute_reply.started":"2025-05-20T10:41:08.509417Z","shell.execute_reply":"2025-05-20T10:41:09.533480Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"plot_missing_label(train_v2_seq, 'Train V2')","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-05-20T10:41:09.535551Z","iopub.execute_input":"2025-05-20T10:41:09.535953Z","iopub.status.idle":"2025-05-20T10:41:10.645878Z","shell.execute_reply.started":"2025-05-20T10:41:09.535913Z","shell.execute_reply":"2025-05-20T10:41:10.644857Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"def resample_array(arr, target_length=1024):\n    return np.round(resample(arr, target_length)).astype(int)\n\ndef binary_sequence(target_id):\n    sequence = train_seq[train_seq.target_id == target_id].sequence\n    b_sequence = np.array(train_labels[train_labels.target_id == target_id].x_1.notna().astype(int))\n    resampled = resample_array(b_sequence)\n    return resampled\n\ndef plot_missing_label_pattern(df, desc, ax=None):\n    if ax is None:\n        fig, ax = plt.subplots(figsize=(12, 6))\n\n    concat_matrix = []\n    selected_df = df.sort_values('missing_label', ascending=False).iloc[:30]\n    for idx, rna in selected_df.iterrows():\n        b_seq = binary_sequence(rna.target_id)\n        concat_matrix.append(b_seq)\n    sns.heatmap(np.array(concat_matrix), cmap='viridis', ax=ax, cbar=False)\n    ax.set_title('Missing Label Pattern')\n    ax.set_yticklabels(labels=selected_df.target_id.to_list())\n    ax.xaxis.set_visible(False)\n    plt.setp(ax.yaxis.get_majorticklabels(), rotation=0)\n    if fig:\n        plt.savefig(f'Figures/{desc}_plot_missing_label_pattern.png', dpi=300, bbox_inches='tight')","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-05-20T10:41:10.647063Z","iopub.execute_input":"2025-05-20T10:41:10.647336Z","iopub.status.idle":"2025-05-20T10:41:10.655372Z","shell.execute_reply.started":"2025-05-20T10:41:10.647313Z","shell.execute_reply":"2025-05-20T10:41:10.654232Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"plot_missing_label_pattern(train_seq, 'Train')","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-05-20T10:41:10.656400Z","iopub.execute_input":"2025-05-20T10:41:10.656815Z","iopub.status.idle":"2025-05-20T10:41:12.658581Z","shell.execute_reply.started":"2025-05-20T10:41:10.656777Z","shell.execute_reply":"2025-05-20T10:41:12.657524Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"### Quick Takeaways:\n\n+ `[Train Dataset]`:\n    + 90% of sequences contain fewer than 160 nucleotides.\n    + Only 34 sequences exceed 1000 nucleotides.\n    + 152 sequences are **invalid for Phase 1 training** due to a temporal cutoff after `2022-05-27`.\n    + 238 sequences **lack at least one residue coordinate**.\n    + 46 sequences are **missing coordinate labels entirely**.\n    + No clear pattern is observed in the missing coordinate labels.","metadata":{}},{"cell_type":"markdown","source":"# <div style=\"text-align:center; border-radius:10px; color:black; margin:0; font-size:100%; font-family:Carbon Plus; background-color:white; overflow:hidden\"><b>Secondary Structure</b></div>","metadata":{}},{"cell_type":"markdown","source":"Despite being single-stranded, RNA molecules often exhibit a great deal of **double-helical** character. This is because RNA chains frequently fold back on themselves to form base-paired segments between short stretches of complementary sequences.\n\n![Secondary Structure](https://raw.githubusercontent.com/ammomahdavikia/asset-holding/main/rna_secondary_structure.png)\n\nIf the two stretches of complementary sequence are near each other, the RNA may adopt a **stem-loop structure** in which the intervening RNA is looped out from the end of the double-helical segment. \n\nStretches of double-helical RNA may also exhibit **internal loops** (unpaired nucleotides on either side of the stem), **bulges** (anunpaired nucleotide on one side of the bulge), or **junctions**.","metadata":{}},{"cell_type":"markdown","source":"### How Can We Predict RNA Secondary Structure?\n\nUltimately, a true secondary structure can be verifed when the 3D structure is determined. Crystallography, NMR, and now high-resolution cryo-electron microscopy (cryo-EM) represent the _crème de la crème_, with NMR and cryo-EM also having the ability to potentially detect conformational ensembles and dynamics.\n\n![Secondary Structure](https://raw.githubusercontent.com/ammomahdavikia/asset-holding/main/rna_secondary_structure_prediction.png)\n\nHowever, for large-scale analysis, computational prediction remains the most practical approach. For now, we will use automatic **thermodynamics-based** algorithm (**RNAfold**, part of the **ViennaRNA**) to predict the secondary structure for each sequence. ","metadata":{}},{"cell_type":"code","source":"def base_pairing_probability(seq):\n    \"\"\"\n    Computes the base pairing probability matrix (BPPM) for a given RNA sequence.\n    \"\"\"\n    fc = RNA.fold_compound(seq)\n    fc.pf()\n    bpp = fc.bpp()\n    bppm = np.array([list(bp) for bp in bpp])\n    bppm = bppm[1:, 1:]  # Remove 0th dummy row and column\n    return bppm\n\ndef dot_bracket(seq):\n    \"\"\"\n    Predicts the minimum free energy (MFE) secondary structure of an RNA sequence in dot-bracket notation.\n    \"\"\"\n    fc = RNA.fold_compound(seq)\n    return fc.mfe()\n\ndef predict_secondary_structure(df, df_name='dataset'):\n    print(clr.S+f\"\\n--- Dot-Bracket {df_name} ---\\n\"+clr.E)\n\n    ss_results = parallel_apply(df['sequence'], dot_bracket, desc=f'{df_name} - DB', n_jobs=-1)\n    df[['secondary_structure', 'mfe']] = pd.DataFrame(ss_results, index=df.index)\n\n    return df","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-05-20T10:41:12.659544Z","iopub.execute_input":"2025-05-20T10:41:12.659852Z","iopub.status.idle":"2025-05-20T10:41:12.666246Z","shell.execute_reply.started":"2025-05-20T10:41:12.659809Z","shell.execute_reply":"2025-05-20T10:41:12.665291Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"for df, desc in zip([train_seq, valid_seq, test_seq],\n                    ['Train Sequence', 'Validation Sequence','Test Sequence']):\n    \n    df = predict_secondary_structure(df, desc)","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-05-20T10:41:12.667352Z","iopub.execute_input":"2025-05-20T10:41:12.667683Z","iopub.status.idle":"2025-05-20T10:45:49.577641Z","shell.execute_reply.started":"2025-05-20T10:41:12.667651Z","shell.execute_reply":"2025-05-20T10:45:49.576511Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"The **dot-bracket** notation is a simplified way to represent RNA secondary structures, where dots `.` indicate unpaired nucleotides, and matching parentheses `( )` represent base pairs (e.g., ”((…))” for a hairpin loop). ","metadata":{}},{"cell_type":"code","source":"train_seq[['target_id', 'sequence', 'secondary_structure']].head()","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-05-20T10:45:49.579139Z","iopub.execute_input":"2025-05-20T10:45:49.579563Z","iopub.status.idle":"2025-05-20T10:45:49.597429Z","shell.execute_reply.started":"2025-05-20T10:45:49.579521Z","shell.execute_reply":"2025-05-20T10:45:49.596400Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"It doesn't seem intuitive, right? Let's visualize better:","metadata":{}},{"cell_type":"code","source":"def plot_dot_bracket(seq, dot_bracket, desc=None, max_length=50):\n    fig, ax = plt.subplots(figsize=(4, 4))\n    bg = fgb.BulgeGraph.from_dotbracket(dot_bracket, seq)\n    fvm.plot_rna(bg, ax=ax, text_kwargs={'fontsize': 6, 'color': 'black'})\n    ax.set_title(f'(Potential) Secondary Structure - {desc}')\n    if fig:\n        plt.savefig(f'Figures/{desc}_plot_dot_bracket.png', dpi=300, bbox_inches='tight')","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-05-20T10:45:49.598467Z","iopub.execute_input":"2025-05-20T10:45:49.598886Z","iopub.status.idle":"2025-05-20T10:45:49.614258Z","shell.execute_reply.started":"2025-05-20T10:45:49.598826Z","shell.execute_reply":"2025-05-20T10:45:49.612963Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"sample_idx = np.random.randint(len(train_seq))\nsample_seq = train_seq.sequence.iloc[sample_idx]\nsample_target_id = train_seq.target_id.iloc[sample_idx]\nsample_dot_bracket = train_seq.secondary_structure.iloc[sample_idx]\nplot_dot_bracket(sample_seq, sample_dot_bracket, desc=sample_target_id)","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-05-20T10:45:49.615162Z","iopub.execute_input":"2025-05-20T10:45:49.615547Z","iopub.status.idle":"2025-05-20T10:45:51.693344Z","shell.execute_reply.started":"2025-05-20T10:45:49.615510Z","shell.execute_reply":"2025-05-20T10:45:51.692175Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"### How can we use it for prediction?\n\nWe need better representation, **Base Pair Probability Matrix (BPPM)**. This matrix represents the likelihood of each nucleotide pairing with another in an RNA secondary structure, where each entry $(i, j)$ contains the probability that nucleotide $i$ forms a base pair with nucleotide $j$.","metadata":{}},{"cell_type":"code","source":"def plot_bppm(seq, dot_bracket, desc=None, max_length=50):\n    \n    bg = fgb.BulgeGraph.from_dotbracket(dot_bracket, seq)\n    if len(seq) > max_length:\n        seq = seq[:max_length]\n    \n    bppm = base_pairing_probability(seq)\n    fig, axs = plt.subplots(nrows=1, ncols=2, figsize=(12, 6), gridspec_kw={'width_ratios': [2, 1]})\n    \n    sns.heatmap(bppm, \n                cmap='viridis',\n                ax=axs[0],\n                annot=False)\n    axs[0].set_xlabel('Position in sequence')\n    axs[0].set_ylabel('Position in sequence')\n    axs[0].set_title(f'Base Pairing Probability Matrix - {desc}')\n\n    \n    fvm.plot_rna(bg, ax=axs[1], text_kwargs={'fontsize': 8, 'color': 'black'})\n    axs[1].set_title(f'Secondary Structure - {desc}')\n    if fig:\n        plt.savefig(f'Figures/{desc}_plot_bppm.png', dpi=300, bbox_inches='tight')","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-05-20T10:45:51.694336Z","iopub.execute_input":"2025-05-20T10:45:51.694695Z","iopub.status.idle":"2025-05-20T10:45:51.701567Z","shell.execute_reply.started":"2025-05-20T10:45:51.694663Z","shell.execute_reply":"2025-05-20T10:45:51.700474Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"plot_bppm(sample_seq, sample_dot_bracket, desc=sample_target_id)","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-05-20T10:45:51.702658Z","iopub.execute_input":"2025-05-20T10:45:51.703102Z","iopub.status.idle":"2025-05-20T10:45:54.512855Z","shell.execute_reply.started":"2025-05-20T10:45:51.703056Z","shell.execute_reply":"2025-05-20T10:45:54.511631Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"### Pitfall?\n\nThe danger of only considering the lowest free-energy structure and showcasing it as **“The”** secondary structure is the major pitfall here. Thermodynamics-based algorithms consider all nucleotides as equally likely to be involved in secondary structure elements, which often leads to erroneous assumptions about base pairs.\n\nA feature of RNA that adds to its propensity to form double-helical structures is additional **non-Watson-Crick** base pairs. One such example is the G:U base pair. Non-Watson-Crick base pairs can be found in all\ncombinations in RNA (GA and GU are the most abundant in ribosomal RNA).\n\n![Base Pair Types](https://raw.githubusercontent.com/ammomahdavikia/asset-holding/main/base_pair_types.png)\n\nBecause such non-Watson-Crick base pairs can occur as well as the two conventional Watson-Crick base pairs, RNA chains have an enhanced capacity for self-complementarity.","metadata":{}},{"cell_type":"code","source":"def plot_mfe_length(df, desc, ax=None):\n    sns.jointplot(df, x='mfe', y='length', kind='reg', color=colors[-1])","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-05-20T10:45:54.513766Z","iopub.execute_input":"2025-05-20T10:45:54.514061Z","iopub.status.idle":"2025-05-20T10:45:54.518579Z","shell.execute_reply.started":"2025-05-20T10:45:54.514036Z","shell.execute_reply":"2025-05-20T10:45:54.517633Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"plot_mfe_length(train_seq, 'Train')","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-05-20T10:45:54.519485Z","iopub.execute_input":"2025-05-20T10:45:54.519772Z","iopub.status.idle":"2025-05-20T10:45:56.412011Z","shell.execute_reply.started":"2025-05-20T10:45:54.519746Z","shell.execute_reply":"2025-05-20T10:45:56.411000Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"# <div style=\"text-align:center; border-radius:10px; color:black; margin:0; font-size:100%; font-family:Carbon Plus; background-color:white; overflow:hidden\"><b>🎯 Tertiary Structure 🎯</b></div>","metadata":{}},{"cell_type":"code","source":"def plot_xyz_frequency(df, desc, axs=None):\n\n    if axs is None:\n        fig, axs = plt.subplots(nrows=1, ncols=3, figsize=(12, 4))\n    else:\n        fig = None\n\n    sns.histplot(df.x_1.replace(-1.000000e+18, np.nan), kde=True, ax=axs[0], color=colors[-1])\n    axs[0].set_title('X coordinate')\n    axs[0].set_xlabel('')\n    axs[0].set_ylabel('Count')\n    axs[0].set_xlim([-1000, 1000])\n    axs[0].set_ylim([0, 6000])\n\n    sns.histplot(df.y_1.replace(-1.000000e+18, np.nan), kde=True, ax=axs[1], color=colors[-1])\n    axs[1].set_title('Y coordinate')\n    axs[1].set_xlabel('')\n    axs[1].set_ylabel('Count')\n    axs[1].set_xlim([-1000, 1000])\n    axs[1].set_ylim([0, 6000])\n\n    sns.histplot(df.z_1.replace(-1.000000e+18, np.nan), kde=True, ax=axs[2], color=colors[-1])\n    axs[2].set_title('Z coordinate')\n    axs[2].set_xlabel('')\n    axs[2].set_ylabel('Count')\n    axs[2].set_xlim([-1000, 1000])\n    axs[2].set_ylim([0, 6000])\n    \n    if fig:\n        plt.tight_layout()\n        plt.savefig(f'Figures/{desc}_plot_xyz_frequency.png', dpi=300, bbox_inches='tight')","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-05-20T10:45:56.412966Z","iopub.execute_input":"2025-05-20T10:45:56.413262Z","iopub.status.idle":"2025-05-20T10:45:56.421899Z","shell.execute_reply.started":"2025-05-20T10:45:56.413235Z","shell.execute_reply":"2025-05-20T10:45:56.420942Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"plot_xyz_frequency(train_labels, 'Train')","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-05-20T10:45:56.422984Z","iopub.execute_input":"2025-05-20T10:45:56.423358Z","iopub.status.idle":"2025-05-20T10:46:00.669687Z","shell.execute_reply.started":"2025-05-20T10:45:56.423320Z","shell.execute_reply":"2025-05-20T10:46:00.668582Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"def plot_3d_strcutrue(pdb_id):\n    rna = train_labels[train_labels.pdb_id == pdb_id]\n    chains = sorted(set(rna.chain))\n    num_chains = len(chains)\n\n    def interpolate_gray(index, total):\n        ratio = index / max(total - 1, 1)\n        gray_value = int(200 - (150 * ratio))\n        return f'rgb({gray_value},{gray_value},{gray_value})'\n    \n    chain_color_map = {chain: interpolate_gray(i, num_chains) for i, chain in enumerate(chains)}\n    \n    pdb_data = \"\"\n    for _, residue in rna.iterrows():\n        pdb_data += f'ATOM  {residue.resid:5d}  P   {residue.resname} {residue.chain} {residue.resid:4d}    '\\\n                    f'{residue.x_1:8.3f}{residue.y_1:8.3f}{residue.z_1:8.3f}\\n'\n    \n    view = py3Dmol.view(width=600, height=400)\n    view.addModel(pdb_data, 'pdb')\n    \n    for _, residue in rna.iterrows():\n        view.addSphere({\n            'center': {'x': residue.x_1, 'y': residue.y_1, 'z': residue.z_1},\n            'radius': 1.2,\n            'color': nt_clr.get(residue.resname, 'white')\n        })\n        view.addLabel(f'{residue.resname}{residue.resid}', {\n            'position': {'x': residue.x_1, 'y': residue.y_1, 'z': residue.z_1},\n            'backgroundColor': chain_color_map[residue.chain],\n            'fontColor': 'white',\n            'fontSize': 10\n        })\n\n    cif_filename = os.path.join(config.PATH, 'PDB_RNA', f\"{pdb_id.lower()}.cif\")\n    \n    parser = MMCIFParser(QUIET=True)\n    structure = parser.get_structure(pdb_id, cif_filename)\n    \n    output_pdb = StringIO()\n    first_model = next(structure.get_models())\n    \n    for chain in first_model:\n        if chain.id in chains:  # Include only relevant chains\n            for residue in chain:\n                for atom in residue:\n                    output_pdb.write(f\"ATOM  {atom.serial_number:5d}  {atom.name:<4} {residue.resname} {chain.id} {residue.id[1]:4d}    \"\n                                     f\"{atom.coord[0]:8.3f}{atom.coord[1]:8.3f}{atom.coord[2]:8.3f}\\n\")\n    \n    target_pdb_string = output_pdb.getvalue()\n    \n    view.addModel(target_pdb_string, 'pdb')\n    view.setStyle({'cartoon': {'color': 'spectrum'}})\n    view.zoomTo()\n    view.show()","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-05-20T10:46:00.670693Z","iopub.execute_input":"2025-05-20T10:46:00.671068Z","iopub.status.idle":"2025-05-20T10:46:00.681407Z","shell.execute_reply.started":"2025-05-20T10:46:00.671034Z","shell.execute_reply":"2025-05-20T10:46:00.680217Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"plot_3d_strcutrue(sample_pdb_id)","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-05-20T10:46:00.682504Z","iopub.execute_input":"2025-05-20T10:46:00.682801Z","iopub.status.idle":"2025-05-20T10:46:00.754863Z","shell.execute_reply.started":"2025-05-20T10:46:00.682749Z","shell.execute_reply":"2025-05-20T10:46:00.753806Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"In this representation, the spheres correspond to XYZ coordinates of C1' atom from `[]_labels.csv`, while the ribbon depicts the actual 3D structure obtained from the PDB.","metadata":{}},{"cell_type":"markdown","source":"# <div style=\"text-align:center; border-radius:10px; color:black; margin:0; font-size:100%; font-family:Carbon Plus; background-color:white; overflow:hidden\"><b>Multiple Sequence Alignment</b></div>","metadata":{}},{"cell_type":"markdown","source":"**Multiple Sequence Alignment (MSA)** is a technique to align structurally similar sequences in order to extract evolutionary information about the sequence structure and potentialy its function. These similarities can reveal evolutionary relationships, conserved functional domains, and structural motifs.\n\nRNAs with similar sequences often fold into similar structures, so aligning homologous sequences helps identify conserved residues and co-evolutionary patterns. These patterns are essential for methods like **AlphaFold** and **RoseTTAFold**.\n\nOne can extract MSAs for each sequence of dataset by retrieving the related datasets (NCBI, Rfam, ...) for similar patterns, But here we will use tha kaggle provided MSA dataset:","metadata":{}},{"cell_type":"code","source":"def msa_count(target_id, suffix=''):\n    msa_file = os.path.join(config.PATH, 'MSA'+suffix, f'{target_id}.MSA.fasta')\n    try:\n        return sum(1 for _ in SeqIO.parse(msa_file, 'fasta')) - 1\n    except FileNotFoundError:\n        return np.nan\n\nfor df, desc in zip([train_seq, train_v2_seq, valid_seq, test_seq],\n                    ['Train Sequence', 'Train Sequence V2', 'Validation Sequence','Test Sequence']):\n\n    suffix = ''\n    if desc == 'Train Sequence V2':\n        suffix = '_v2'\n    df['msa_count'] = df['target_id'].apply(lambda tid: msa_count(tid, suffix))\n    df['msa_presence'] = pd.Series(df['msa_count']).gt(0)\n    print(clr.S+f'\\n----- {desc} MSA -----\\n'+clr.E)\n    print(f'[HIT]: {df.msa_presence.sum()}')\n    print(f'[MIN]: {df.msa_count.min()}')\n    print(f'[MAX]: {df.msa_count.max()}')\n    print(f'[AVG]: {df.msa_count.mean():.2f}')\n    df['msa_presence'] = df['msa_presence'].map({True: 'Hit', False: 'No Hit'})","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-05-20T10:46:00.755760Z","iopub.execute_input":"2025-05-20T10:46:00.756075Z","iopub.status.idle":"2025-05-20T10:47:45.858590Z","shell.execute_reply.started":"2025-05-20T10:46:00.756046Z","shell.execute_reply":"2025-05-20T10:47:45.857547Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"def plot_msa_analytics(df, desc, axs=None):\n    if axs is None:\n        fig, axs = plt.subplots(nrows=1, ncols=2, figsize=(12, 6), gridspec_kw={'width_ratios': [1, 4]})\n    else:\n        fig = None\n\n    sns.countplot(\n        data=df,\n        x='msa_presence',\n        ax=axs[0],\n        palette=[colors[3], colors[1]]\n    )\n    axs[0].set_title('MSA Presence')\n    axs[0].set_xlabel('')\n    axs[0].set_ylabel('Count')\n    ylim = axs[0].get_ylim()\n\n    sns.histplot(\n        df.loc[df.msa_count > 0, 'msa_count'].dropna(), \n        ax=axs[1],\n        color=colors[-1]\n    )\n\n    axs[1].set_title('Frequency of MSA Hits')\n    axs[1].set_xlabel('Hits')\n    axs[1].set_ylabel('')\n    axs[1].set_ylim(ylim)\n\n    if fig:\n        plt.tight_layout()\n        plt.savefig(f'Figures/{desc}_plot_msa_analytics.png', dpi=300, bbox_inches='tight')","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-05-20T10:47:45.859567Z","iopub.execute_input":"2025-05-20T10:47:45.859882Z","iopub.status.idle":"2025-05-20T10:47:45.867482Z","shell.execute_reply.started":"2025-05-20T10:47:45.859831Z","shell.execute_reply":"2025-05-20T10:47:45.866324Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"plot_msa_analytics(train_seq, 'Train')","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-05-20T10:47:45.868566Z","iopub.execute_input":"2025-05-20T10:47:45.868994Z","iopub.status.idle":"2025-05-20T10:47:46.958631Z","shell.execute_reply.started":"2025-05-20T10:47:45.868964Z","shell.execute_reply":"2025-05-20T10:47:46.957632Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"def plot_msa(target_id, max_sequence=20, max_position=40):\n    msa_file = os.path.join(config.PATH, 'MSA', f'{target_id}.MSA.fasta')\n    \n    sequences = []\n    try:\n        for record in SeqIO.parse(msa_file, 'fasta'):\n            sequences.append(list(str(record.seq)))\n    except FileNotFoundError:\n        print(np.nan)\n    \n    msa_df = pd.DataFrame(sequences)\n    \n    present_nucleotides = set(msa_df.values.flatten())\n    _nt_clr = {nt: nt_clr[nt] for nt in present_nucleotides}\n    nt_to_num = {nt: idx for idx, nt in enumerate(_nt_clr.keys())}\n    cmap = ListedColormap([_nt_clr[nt] for nt in _nt_clr.keys()])\n    \n    msa_numeric = msa_df.replace(nt_to_num)\n    if msa_numeric.shape[0] > max_sequence:\n        msa_numeric = msa_numeric[:max_sequence]\n    if msa_numeric.shape[1] > max_position:\n        msa_numeric = msa_numeric[:, :max_position]    \n\n    fig, ax = plt.subplots(figsize=(12, 6))\n    sns.heatmap(msa_numeric, \n                cmap=cmap, \n                cbar=False, \n                xticklabels=False,\n                yticklabels=False, \n                ax=ax,\n                annot=False)\n    \n    for i in range(msa_numeric.shape[0]):\n        for j in range(msa_numeric.shape[1]):\n            nucleotide = msa_df.iloc[i, j]\n            ax.text(j + 0.5, i + 0.5, \n                    nucleotide, \n                    ha='center', \n                    va='center', \n                    fontsize=8, \n                    color='black', \n                    fontweight='bold')\n    \n    _labels = ['Original'] + [str(idx+1) for idx in range(msa_numeric.shape[0]-1)]\n    ax.set_yticks(np.arange(msa_numeric.shape[0]) + 0.5)\n    ax.set_yticklabels(_labels)\n    plt.xlabel('Position in sequence')\n    plt.title(f'MSA - {target_id}')\n\n    if fig:\n        plt.savefig(f'Figures/{target_id}_plot_msa.png', dpi=300, bbox_inches='tight')\n    plt.show()","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-05-20T10:47:46.959613Z","iopub.execute_input":"2025-05-20T10:47:46.959969Z","iopub.status.idle":"2025-05-20T10:47:46.969777Z","shell.execute_reply.started":"2025-05-20T10:47:46.959942Z","shell.execute_reply":"2025-05-20T10:47:46.968856Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"plot_msa(target_id='1EBQ_A') # Feel free to change","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-05-20T10:47:46.970726Z","iopub.execute_input":"2025-05-20T10:47:46.971036Z","iopub.status.idle":"2025-05-20T10:47:49.581714Z","shell.execute_reply.started":"2025-05-20T10:47:46.971001Z","shell.execute_reply":"2025-05-20T10:47:49.580700Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"# <div style=\"text-align:center; border-radius:10px; color:black; margin:0; font-size:100%; font-family:Carbon Plus; background-color:white; overflow:hidden\"><b>State-of-the-RNArt</b></div>\n\nI was planning to provide a PyTorch baseline but there was a trade-off between the educational purpose of the notebook and technical complexity it could provide. So far, It is enough both for the content and the complexity. I will just append the state-of-the-art methods for RNA 3D folding in the time competition is happening. It is inspired by State-of-the-RNArt paper. You can trace the path back to the roots by the links provided.\n\nLet's start with the methods, then we will see how can one use and ensemble these methods to attain competitive results.\n\nMethods can be classified as ab initio, template-based or deep learning-based. \n+ **Ab initio** (or prediction-based) methods tend to simulate the physics of the system.\n+ **Template-based** (or fragment-assembly) approaches rely on the fact that molecules that have evolution similitude adopt similar structures. A database of known RNA structures is used as a reference. \n+ **Deep learning** approaches use available data to create a neural network architecture that predicts RNA 3D structures from different data sources.\n\n\n![State-of-the-RNArt](https://raw.githubusercontent.com/ammomahdavikia/asset-holding/main/State-of-the-RNArt.png)\n\nSome of the open-source methods have been implemented by the kaggle community:\n| Method              | Kaggle                                                                                                 | GitHub                                                                 |\n|---------------------|--------------------------------------------------------------------------------------------------------|------------------------------------------------------------------------|\n| **DRFold2**         | [![View on Kaggle](https://img.shields.io/badge/Kaggle-View-blue?logo=kaggle)](https://www.kaggle.com/code/hengck23/lb0-321-simple-drfold-no-msa) | [![GitHub](https://img.shields.io/badge/GitHub-Repo-black?logo=github)](https://github.com/leeyang/DRfold2) |\n| **NuFold**          | [![View on Kaggle](https://img.shields.io/badge/Kaggle-View-blue?logo=kaggle)](https://www.kaggle.com/code/amirmmahdavikia/stanford-rna-3d-folding-nufold-inference) | [![GitHub](https://img.shields.io/badge/GitHub-Repo-black?logo=github)](https://github.com/kiharalab/NuFold) |\n| **RhoFold**         | [![View on Kaggle](https://img.shields.io/badge/Kaggle-View-blue?logo=kaggle)](https://www.kaggle.com/code/ogurtsov/rhofold-ribonanzanet-msas-lb-0-215)<br>[![View on Kaggle](https://img.shields.io/badge/Kaggle-View-blue?logo=kaggle)](https://www.kaggle.com/code/hengck23/demo-for-rhofold-plus-with-kaggle-msa) | [![GitHub](https://img.shields.io/badge/GitHub-Repo-black?logo=github)](https://github.com/ml4bio/RhoFold) |\n| **RibonanzaNet**    | [![View on Kaggle](https://img.shields.io/badge/Kaggle-View-blue?logo=kaggle)](https://www.kaggle.com/code/shujun717/ribonanzanet-3d-inference) | [![GitHub](https://img.shields.io/badge/GitHub-Repo-black?logo=github)](https://github.com/Shujun-He/RibonanzaNet) |\n| **Protenix (AF3)**  | [![View on Kaggle](https://img.shields.io/badge/Kaggle-View-blue?logo=kaggle)](https://www.kaggle.com/code/geraseva/protenix) | [![GitHub](https://img.shields.io/badge/GitHub-Repo-black?logo=github)](https://github.com/bytedance/Protenix) |\n| **Boltz-1 (AF3)**   | [![View on Kaggle](https://img.shields.io/badge/Kaggle-View-blue?logo=kaggle)](https://www.kaggle.com/code/youhanlee/boltz-1-inference-submission) | [![GitHub](https://img.shields.io/badge/GitHub-Repo-black?logo=github)](https://github.com/jwohlwend/boltz) |\n| **trRosettaRNA**    | [![View on Kaggle](https://img.shields.io/badge/Kaggle-View-blue?logo=kaggle)](https://www.kaggle.com/code/amirmmahdavikia/stanford-rna-3d-folding-trrosettarna-inference) | [![GitHub](https://img.shields.io/badge/GitHub-Repo-black?logo=github)](https://yanglab.qd.sdu.edu.cn/trRosettaRNA/download/) |","metadata":{}},{"cell_type":"markdown","source":"# <div style=\"text-align:center; border-radius:10px; color:black; margin:0; font-size:100%; font-family:Carbon Plus; background-color:white; overflow:hidden\"><b>CASP-16</b></div>\n\n**CASP (Critical Assessment of Structure Prediction)** is a global competition where researchers test how well their algorithms can predict the 3D shape of biological molecules. In CASP16, held in 2024, RNA structure prediction was included as a major focus.\nI was curious to see how did CASP16 top competitors performed on RNA structures. \n\n| Team           | Core Model/Framework                                                                                  | MSA | Refinement Method                                                |\n|----------------|--------------------------------------------------------------------------------------------------------|-----|------------------------------------------------------------------|\n| **Vfold**      | Vfold-Pipeline, IsRNA, RNAJP                                                                           | ✘   | SimRNA + Filtering + Templates + Literature information for ranking |\n| **Guangzhou**  | Guijunlab-complex, AlphaFold2, AlphaFold3                                                              | ✔   | Custom model quality assessment (GuijunLab-QA:QA4) + DeepAssembly |\n| **Kiharalab**  | NuFold, DeepFoldRNA, DRfold, RosettaFold2NA, RhoFold, trRosettaRNA, AlphaFold3                        | ✔   | Ranked by combining Rosetta score and ARES                       |\n| **Yang-server**| trRosettaRNA                                                                                           | ✔   | Restraint-based GD (PyRosetta-based)                             |\n| **GeneSilico** | Main: (ModeRNA + SimRNA) / Additional: (RNAComposer, RNA-BRiQ, RNAJP, trRosettaRNA, AlphaFold3)        | ✔   | RMSD scoring + Template refinement (QRNAS)                       |","metadata":{}},{"cell_type":"markdown","source":"# <div style=\"text-align:center; border-radius:10px; color:black; margin:0; font-size:100%; font-family:Carbon Plus; background-color:white; overflow:hidden\"><b>Summary</b></div>\n\n![Base Pair Types](https://raw.githubusercontent.com/ammomahdavikia/asset-holding/main/stanford_rna_3d_folding_summary.png)","metadata":{}},{"cell_type":"markdown","source":"### Resources\n\n1.\tWatson, J. D., Baker, T. A., Bell, S. P., Gann, A., Levine, M., & Losick, R. (2013). Molecular Biology of the Gene (7th ed.). Pearson Education. ISBN: 978-0321762436.\n2. Nelson, D. L., & Cox, M. M. (2017). Lehninger Principles of Biochemistry (7th ed.). W.H. Freeman. ISBN: 978-1464126116.\n3. Berman, H. M., Westbrook, J., Feng, Z., Gilliland, G., Bhat, T. N., Weissig, H., … & Bourne, P. E. (2000). The Protein Data Bank. Nucleic Acids Research, 28(1), 235-242. DOI: 10.1093/nar/28.1.235.\n4. Urry, L. A., Cain, M. L., Wasserman, S. A., Minorsky, P. V., & Reece, J. B. (2020). Campbell Biology (12th ed.). Pearson. ISBN: 978-0135188743.\n5. Jumper, J., Evans, R., Pritzel, A., et al. (2021). Highly accurate protein structure prediction with AlphaFold. Nature, 596, 583–589. DOI: 10.1038/s41586-021-03819-2.\n6. Weeks, K. M. (2021). Thoughts on how to think (and talk) about RNA structure. Proceedings of the National Academy of Sciences, 118(37), e2112677119. DOI: 10.1073/pnas.2112677119.\n7. Jumper, J., et al. (2021). Highly accurate protein structure prediction with AlphaFold. Nature, 596, 583–589. DOI: 10.1038/s41586-021-03819-2.\n8. Abramson, J., et al. (2024). Accurate structure prediction of biomolecular interactions with AlphaFold 3. Nature, 630, 493–500. DOI: 10.1038/s41586-024-07487-w.\n9. Wang, W., Feng, C., Han, R., Wang, Z., Ye, L., Du, Z., Wei, H., Zhang, F., Peng, Z., & Yang, J. (2023). trRosettaRNA: automated prediction of RNA 3D structure with transformer network. Nature Communications, 14, 7266. DOI: 10.1038/s41467-023-42528-4.\n10. Kagaya, Y. (2025). NuFold: end-to-end approach for RNA tertiary structure prediction with flexible nucleobase center representation. Nature Communications, 16, 881. DOI: 10.1038/s41467-025-56261-7.\n11. Li, Y. (2025). Ab initio RNA structure prediction with composite language model and denoised end-to-end learning. bioRxiv. DOI:10.1101/2025.03.05.641632.\n12. He, S. (2024). Ribonanza: deep learning of RNA structure through dual crowdsourcing. bioRxiv. DOI:10.1101/2024.02.24.581671.\n13. Shen, T. (2024). Accurate RNA 3D structure prediction using a language model-based deep learning approach. Nature Methods, 21, 2287–2298. DOI: 10.1038/s41592-024-02487-0.\n14. Wohlwend, J. (2024). Boltz-1: Democratizing biomolecular interaction modeling. bioRxiv. DOI: 10.1101/2024.11.19.624167\n15. Chen, X. (2025). Protenix: Advancing structure prediction through a comprehensive AlphaFold3 reproduction. bioRxiv. DOI: 10.1101/2025.01.08.631967","metadata":{}},{"cell_type":"markdown","source":"### <div style=\"text-align:center; border-radius:10px; color:black; margin:0; font-size:100%; font-family:Carbon Plus; background-color:white; overflow:hidden\"><b>🔼 Your Welcome 🔼</b></div>","metadata":{}}]}