{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"name":"python","version":"3.10.12","mimetype":"text/x-python","codemirror_mode":{"name":"ipython","version":3},"pygments_lexer":"ipython3","nbconvert_exporter":"python","file_extension":".py"},"kaggle":{"accelerator":"none","dataSources":[{"sourceId":87793,"databundleVersionId":12024591,"sourceType":"competition"},{"sourceId":11172002,"sourceType":"datasetVersion","datasetId":6972289}],"dockerImageVersionId":30918,"isInternetEnabled":true,"language":"python","sourceType":"notebook","isGpuEnabled":false}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"![Header](https://raw.githubusercontent.com/ammomahdavikia/asset-holding/main/stanford-rna-3d.gif)","metadata":{}},{"cell_type":"markdown","source":"### Why is RNA important?\n\nMany of us recognize RNA as the intermediary between DNA and proteins, following the **Central Dogma of Life**—the flow of genetic information from **DNA → RNA → Protein**.\n\nHowever, RNA is far more than just a messenger. It plays diverse and essential roles in biological systems, from **transporting amino acids for protein synthesis** (tRNA) to **regulating gene expression** (miRNA) and even **catalyzing biochemical reactions** (ribozymes).\n\n**[Import: RNA Types image]**\n![RNA Types](https://raw.githubusercontent.com/ammomahdavikia/asset-holding/main/rna_types.png)\n\nWhat makes RNA even more fascinating is its potential as the earliest biomolecule of life. Unlike DNA, RNA can both store genetic information and catalyze its own replication, supporting the **RNA World Hypothesis**—the idea that life may have originated with self-replicating RNA molecules.\n\nThe critical functional roles of RNAs make them a new type of **drug target**. Since many biological functions depend on the specifc tertiary structures of RNAs, it is imperative to determine the 3D structures of RNAs in order to facilitate RNA-based function annotation and drug discovery.","metadata":{}},{"cell_type":"markdown","source":"### ↓ Libraries","metadata":{}},{"cell_type":"code","source":"!pip install biopython  # 생물학 관련 파이썬 라이브리리\n!pip install ViennaRNA  # RNA의 2차원 구조를 예측하는 라이브러리\n!pip install py3Dmol # WebGL 기반의 3D 단백질, 분자 구조 시각화 라이브러리\n!pip install forgi","metadata":{"_uuid":"8f2839f25d086af736a60e9eeb907d3b93b6e0e5","_cell_guid":"b1076dfc-b9ad-4769-8c92-a6c4dae69d19","trusted":true,"_kg_hide-output":true,"execution":{"iopub.status.busy":"2025-04-28T08:03:28.247318Z","iopub.execute_input":"2025-04-28T08:03:28.247807Z","iopub.status.idle":"2025-04-28T08:04:00.427966Z","shell.execute_reply.started":"2025-04-28T08:03:28.247762Z","shell.execute_reply":"2025-04-28T08:04:00.426620Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"### ↓ Imports","metadata":{}},{"cell_type":"code","source":"import os\nimport re\nimport requests\nimport zipfile  # 압축 파일을 다루는 모델, python 3.13 이상\nimport pandas as pd\nimport numpy as np\nimport matplotlib.pyplot as plt\nimport matplotlib.font_manager as fm\nfrom matplotlib.colors import ListedColormap\nimport seaborn as sns\nfrom scipy.signal import resample\nfrom Bio import PDB, SeqIO\nfrom Bio.PDB import MMCIFParser\nfrom io import StringIO\nimport RNA\nimport py3Dmol\nimport forgi.graph.bulge_graph as fgb\nimport forgi.visual.mplotlib as fvm\nfrom collections import Counter\nfrom tqdm import tqdm\nfrom typing import Callable\nfrom joblib import Parallel, delayed  # 병렬 처리를 도와주는 모듈\nfrom IPython.display import display_html, display  # 콘솔 창에서 print보다 df를 더 보기 쉽게 해준다.\nimport ipywidgets as widgets  # 인터랙티브 위젯\nimport warnings","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-04-28T08:04:00.429869Z","iopub.execute_input":"2025-04-28T08:04:00.430328Z","iopub.status.idle":"2025-04-28T08:04:02.510669Z","shell.execute_reply.started":"2025-04-28T08:04:00.430284Z","shell.execute_reply":"2025-04-28T08:04:02.509333Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"warnings.simplefilter(action='ignore', category=FutureWarning)\n\nclass clr:\n    \"\"\"\n    터미널에서 텍스트의 스타일과 색상을 설정하기 위한 ANSI 이스케이프 코드를 정의합니다.\n    \"\"\"\n    B = '\\033[1m'  # 굵은 글씨\n    S = '\\033[1m' + '\\033[94m'  # 굵은 글씨 + 파란색\n    L = '\\033[1m' + '\\033[91m'  # 굵은 글씨 + 빨간색\n    E = '\\033[0m'  # 모든 스타일 초기화\n\nprint(clr.B+'\\n----- Font -----\\n'+clr.E)\n# 폰트 파일 다운로드\nfont_url = 'https://ifonts.xyz/core/ifonts-files/downloads/472336/carbon-plus-font.zip'\noutput_path = 'carbon-plus.zip'\n\nresponse = requests.get(font_url, stream=True)\n\nif response.status_code == 200:\n    with open(output_path, 'wb') as f:\n        for chunk in response.iter_content(1024):  # 1024바이트씩 읽어와서 저장\n            f.write(chunk)\n    print(f'[Success] downloaded: {output_path}')\nelse:\n    print('[Error] Failed to download the font. Check the URL.')\n\n## 압축 해제\n# with zipfile.ZipFile(output_path, 'r') as zip_ref:\n#     zip_ref.extractall('carbon_plus_font') # ZIP 파일의 내용을 'carbon_plus_font' 폴더에 압축 해제\n#     print('[Success] extracted.')\n\n# # 폰트 로드 및 설정\n# font_path = '/kaggle/working/carbon_plus_font/carbonplus-regular-bl.otf'\n# try:\n#     fm.fontManager.addfont(font_path)  # 폰트 매니저에 폰트 추가\n#     primary_font = fm.FontProperties(fname=font_path).get_name()  # 추가한 폰트의 이름 가져오기\n#     plt.rcParams['font.family'] = [primary_font, \"DejaVu Sans\"]  # plt의 기본 폰트를 새로 추가한 폰트와 dejavu sans로 설정\n#     print('[Success] loaded.')\n# except:\n#     print('[Error] Failed to load the font. Check the address.')","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-04-28T08:04:02.512302Z","iopub.execute_input":"2025-04-28T08:04:02.512842Z","iopub.status.idle":"2025-04-28T08:04:02.742439Z","shell.execute_reply.started":"2025-04-28T08:04:02.512811Z","shell.execute_reply":"2025-04-28T08:04:02.741204Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"colors = ['#780000', '#c1121f', '#fdf0d5', '#02c39a', '#669bbc', '#003049']\nnucleotides = ['X', 'C', 'U', 'G', 'A', '-']\nnt_clr = {nt: colors[idx] for idx, nt in enumerate(nucleotides)}\nprint(clr.B+'\\n----- Color -----\\n'+clr.E)\nsns.palplot(sns.color_palette(colors))","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-04-28T08:04:02.744407Z","iopub.execute_input":"2025-04-28T08:04:02.744723Z","iopub.status.idle":"2025-04-28T08:04:02.920756Z","shell.execute_reply.started":"2025-04-28T08:04:02.744697Z","shell.execute_reply":"2025-04-28T08:04:02.919642Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"### ↓ Helpers","metadata":{}},{"cell_type":"code","source":"def interactive_subplot(plot_func, dfs_dict):\n    \"\"\"\"\"\"\n    output = widgets.Output()\n\n    def on_select(change):\n        \"\"\"\"\"\"\n        with output:\n            output.clear_output(wait=True)\n            plot_func(dfs_dict[change.new], change.new)\n            plt.show()\n    \n    dropdown = widgets.Dropdown(\n        options=dfs_dict.keys(),\n        description='Dataset:',\n        style={'description_width': 'initial'}\n    )\n    \n    dropdown.observe(on_select, names='value')\n    display(dropdown, output)\n    \n    first_key = next(iter(dfs_dict.keys()), None)\n    if first_key:\n        with output:\n            plot_func(dfs_dict[first_key], first_key)\n            plt.show()","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-04-28T08:04:02.921798Z","iopub.execute_input":"2025-04-28T08:04:02.922145Z","iopub.status.idle":"2025-04-28T08:04:02.929127Z","shell.execute_reply.started":"2025-04-28T08:04:02.922116Z","shell.execute_reply":"2025-04-28T08:04:02.927574Z"},"_kg_hide-input":true},"outputs":[],"execution_count":null},{"cell_type":"code","source":"def parallel_apply(series: pd.Series, func: Callable, desc: str, n_jobs=-1):\n    \"\"\"\n    Parallel apply with tqdm.\n    n_jobs: worker의 수(n개의 CPU 코어, 1개라면 그냥 직렬 처리와 다를 게 없다.)\n    \"\"\"\n    results = Parallel(n_jobs=n_jobs)(\n        delayed(func)(x) for x in tqdm(series, desc=desc)\n    )\n    return results","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-04-28T08:04:02.930352Z","iopub.execute_input":"2025-04-28T08:04:02.930689Z","iopub.status.idle":"2025-04-28T08:04:02.954776Z","shell.execute_reply.started":"2025-04-28T08:04:02.930662Z","shell.execute_reply":"2025-04-28T08:04:02.953691Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"# <div style=\"text-align:center; border-radius:10px; color:black; margin:0; font-size:100%; font-family:Carbon Plus; background-color:white; overflow:hidden\"><b>Loading the data</b></div>","metadata":{}},{"cell_type":"code","source":"class Config():\n    def __init__(self):\n        self.PATH = '/kaggle/input/stanford-rna-3d-folding'\n        self.CIF_PATH = '/kaggle/input/stanford-rna-3d-prediction-pdb-dataset'\n\nconfig = Config()\nos.makedirs('Figures', exist_ok=True)","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-04-28T08:04:02.955918Z","iopub.execute_input":"2025-04-28T08:04:02.956294Z","iopub.status.idle":"2025-04-28T08:04:02.977132Z","shell.execute_reply.started":"2025-04-28T08:04:02.956265Z","shell.execute_reply":"2025-04-28T08:04:02.975857Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"def load_data():\n    \"\"\"설정한 path에서 csv 파일을 가져오는 함수\"\"\"\n    train_seq = pd.read_csv(os.path.join(config.PATH, 'train_sequences.csv'))\n    valid_seq = pd.read_csv(os.path.join(config.PATH, 'validation_sequences.csv'))\n    test_seq = pd.read_csv(os.path.join(config.PATH, 'test_sequences.csv'))\n\n    train_labels = pd.read_csv(os.path.join(config.PATH, 'train_labels.csv'))\n    valid_labels = pd.read_csv(os.path.join(config.PATH, 'validation_labels.csv'))\n\n    return train_seq, valid_seq, test_seq, train_labels, valid_labels","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-04-28T08:04:02.981271Z","iopub.execute_input":"2025-04-28T08:04:02.981624Z","iopub.status.idle":"2025-04-28T08:04:03.000493Z","shell.execute_reply.started":"2025-04-28T08:04:02.981596Z","shell.execute_reply":"2025-04-28T08:04:02.998885Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"def summarize(df, desc='Summary'):\n    \"\"\"주어진 데이터프레임의 구조, null의 총개수, 컬럼 목록을 보여주는 함수\"\"\"\n\n    start = clr.S if 'Sequence' in desc else clr.L\n    \n    print(start+f'\\n----- {desc} -----\\n'+clr.E)\n    print(f'Shape: {df.shape}')\n    print(f'Missing: {df.isna().sum().sum()}')  # isna().sum()은 각 열의 isna()를 계산, 한 번 더 해야 df에 있는 전체 isna()를 반환\n    print(f'Columns: {df.columns.to_list()}\\n')\n    display_html(df.head(3))\n    print('\\n')","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-04-28T08:04:03.003450Z","iopub.execute_input":"2025-04-28T08:04:03.003832Z","iopub.status.idle":"2025-04-28T08:04:03.025635Z","shell.execute_reply.started":"2025-04-28T08:04:03.003801Z","shell.execute_reply":"2025-04-28T08:04:03.024222Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"train_seq, valid_seq, test_seq, train_labels, valid_labels = load_data()\n\nfor df, desc in zip([train_seq, valid_seq, test_seq, train_labels, valid_labels],\n                    ['Train Sequence', 'Validation Sequence','Test Sequence',\n                     'Train Label', 'Validation Label']):\n    summarize(df, desc)","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-04-28T08:04:03.026924Z","iopub.execute_input":"2025-04-28T08:04:03.027388Z","iopub.status.idle":"2025-04-28T08:04:03.578606Z","shell.execute_reply.started":"2025-04-28T08:04:03.027348Z","shell.execute_reply":"2025-04-28T08:04:03.577425Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"# <div style=\"text-align:center; border-radius:10px; color:black; margin:0; font-size:100%; font-family:Carbon Plus; background-color:white; overflow:hidden\"><b>Primary Structure</b></div>\n\nNucleic acids are macromolecules that exist as polymers called **polynucleotides**. As indicated by the name, each polynucleotide consists of monomers called **nucleotides**. A nucleotide, in general, is composed of three parts: \n+ A five-carbon sugar (a pentose)\n+ A nitrogen-containing (nitrogenous) base\n+ One to three phosphate groups\n\n![RNA Types](https://raw.githubusercontent.com/ammomahdavikia/asset-holding/main/rna_backbone.png)\n\nTo understand the structure of a single nucleotide, let’s first consider the **nitrogenous bases**. Each nitrogenous base has one or two rings that include nitrogen atoms. There are two families of nitrogenous bases: \n+ **Pyrimidines**: A pyrimidine has one six-membered ring of carbon and nitrogen atoms.\n    + Cytosine `C`\n    + Thymine `T`\n    + Uracil `U`\n+ **Purines**: Purine is larger, with a six-membered ring fused to a five-membered ring.\n    + Adenine `A`\n    + Guanine `G`\n\nAdenine, guanine, and cytosine are found in both DNA and RNA; thymine is found only in DNA, and uracil only in RNA. Let's get our hands dirty and explore the dataset a little bit:","metadata":{}},{"cell_type":"code","source":"print(clr.S+'\\n----- Nucleotide components -----\\n'+clr.E)\ndisplay_html(train_seq.sequence.apply(lambda x: set(x)).value_counts().reset_index())","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-04-28T08:04:03.579890Z","iopub.execute_input":"2025-04-28T08:04:03.580224Z","iopub.status.idle":"2025-04-28T08:04:03.604206Z","shell.execute_reply.started":"2025-04-28T08:04:03.580198Z","shell.execute_reply":"2025-04-28T08:04:03.603132Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"As you can see, `train_sequences.csv` contain 6 different types of characters based on [FASTA format](https://en.wikipedia.org/wiki/FASTA_format). But the organizers announced for `test_sequences.csv`, this is guaranteed to be a string of `A`, `C`, `G`, and `U`.\n\n| Symbol | Meaning                 |\n|--------|-------------------------|\n| A      | Adenine                 |\n| U      | Uracil                  |\n| C      | Cytosine                |\n| G      | Guanine                 |\n| X      | Any                     |\n| -      | Gap or missing base     |","metadata":{}},{"cell_type":"code","source":"def plot_nucleotide_frequency(df: pd.DataFrame, desc:str, ax=None):\n    \"\"\"시퀀스에서 각각의 뉴클레오타이드의 빈도 수를 표ㅎ현하는 막대 그래프를 그리는 함수\"\"\"\n    nucleotide_counts = Counter(''.join(df['sequence']))\n    c = {nt: [count, nt_clr[nt]] for nt, count in nucleotide_counts.items()}  # [개수, #ㅇㅇㅇㅇㅇㅇ(컬러코드)]\n    c = pd.DataFrame(c).T.reset_index().rename(columns={'index': 'nt', 0: 'count', 1: 'color'})  # 행/열 전환 -> 인덱스 없애고, 컬럼명 부여\n\n    if ax is None:\n        fig, ax = plt.subplots(figsize=(8, 5))\n\n    sns.barplot(data=c, x='nt', y='count', palette=c.color.to_list(), ax=ax)\n    ax.set_title('Nucleotide frequency')\n    ax.set_xlabel('')  # 현재 기본 값은 nt\n    if fig:\n        plt.savefig(f'Figures/{desc}_plot_nucleotide_frequency.png', dpi=300, bbox_inches='tight')  # tight = 여백 제거, dpi(dot per inch)","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-04-28T08:04:03.605468Z","iopub.execute_input":"2025-04-28T08:04:03.605794Z","iopub.status.idle":"2025-04-28T08:04:03.635503Z","shell.execute_reply.started":"2025-04-28T08:04:03.605768Z","shell.execute_reply":"2025-04-28T08:04:03.634056Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"plot_nucleotide_frequency(train_seq, 'Train')","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-04-28T08:04:03.636852Z","iopub.execute_input":"2025-04-28T08:04:03.637288Z","iopub.status.idle":"2025-04-28T08:04:04.278606Z","shell.execute_reply.started":"2025-04-28T08:04:03.637249Z","shell.execute_reply":"2025-04-28T08:04:04.277237Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"The higher the percent of **G:C base pairs** in the nucleic acids (and hence the lower the content of A:T base pairs), the higher is the melting point. How do we explain this behavior? \n\nG:C base pairs contribute more to the stability of nucleic acids than do A:T base pairs because of the greater number of hydrogen bonds for the former (three in a G:C base pair vs. two for A:T), but also, importantly, because the stacking interactions of G:C base pairs with adjacent base pairs are more favorable than the corresponding interactions of A:T base pairs with their neighboring base pairs.","metadata":{}},{"cell_type":"code","source":"def CG_ratio(seq: str):\n    \"\"\"전체 뉴클레오타이드 시퀀스에서 사이토신과 구아닌의 비율을 반환하는 함수\"\"\"\n    nucleotide_counts = Counter(''.join(seq))\n    return (nucleotide_counts['C'] + nucleotide_counts['G']) / len(seq)\n\ndef plot_CG_ratio(df, desc, ax=None):\n    \"\"\"전체 시퀀스에서 C와 G의 비율에 따른 시퀀스 길이 비율 그래프를 그려주는 함수\"\"\"\n    if ax is None:\n        fig, ax = plt.subplots(figsize=(8, 5))\n    sns.histplot(df.sequence.apply(CG_ratio), kde=True, ax=ax, color=colors[-1])  # 커널밀도추정(Kernel Density Estimation), 히스토그램을 스무딩한 곡선\n    ax.set_title('C+G Ratio frequency')\n    ax.set_xlabel('C+G Ratio')\n    ax.set_xlim(0, 1)  # (0,0)과 (0,1) 지점의 값을 무엇으로 둘 것인지 결정\n    if fig:\n        plt.savefig(f'Figures/{desc}_plot_CG_ratio.png', dpi=300, bbox_inches='tight')","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-04-28T08:04:04.279783Z","iopub.execute_input":"2025-04-28T08:04:04.280161Z","iopub.status.idle":"2025-04-28T08:04:04.287293Z","shell.execute_reply.started":"2025-04-28T08:04:04.280132Z","shell.execute_reply":"2025-04-28T08:04:04.286137Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"plot_CG_ratio(train_seq, 'Train')","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-04-28T08:04:04.288424Z","iopub.execute_input":"2025-04-28T08:04:04.288726Z","iopub.status.idle":"2025-04-28T08:04:05.063279Z","shell.execute_reply.started":"2025-04-28T08:04:04.288702Z","shell.execute_reply":"2025-04-28T08:04:05.062078Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"Now let’s add the sugar to which the nitrogenous base is attached. In DNA the sugar is **deoxyribose**; in RNA it is **ribose**. The only difference between these two sugars is that deoxyribose lacks an oxygen atom on the second carbon in the ring, hence the name deoxyribose.\n\nSo far, we have built a nucleoside (base plus sugar). To complete the construction of a nucleotide, we attach one to three **phosphate groups** to the 5' carbon of the sugar (the carbon numbers in the sugar include `'`, the prime symbol); With one phosphate, this is a nucleoside monophosphate, more often called a nucleotide.","metadata":{}},{"cell_type":"code","source":"def plot_backbone(pdb_id: str):\n    \"\"\"pdb_id의 3D 그림을 그리는 함수\"\"\"\n    cif_filename = f\"{config.CIF_PATH}/{pdb_id}.cif\"\n    \n    with open(cif_filename, 'r') as file:\n        cif_string = file.read()\n    \n    viewer = py3Dmol.view(width=600, height=400)\n    viewer.addModel(cif_string, 'cif')\n    \n    # viewer.setStyle({'cartoon': {'colorscheme': 'spectrum'}})\n    viewer.setStyle({'cartoon': {'colorscheme': 'greenCarbon'}})\n    viewer.setStyle({'atom': ['P']}, {'sphere': {'color': colors[1], 'radius': 0.8}})  # 인산의 색상과 형태 결정\n    viewer.setStyle({'atom': [\"C1'\", \"C2'\", \"C3'\", \"C4'\", \"O4'\"]}, {'stick': {'color': 'spectrum'}})  # 탄소와 산소 원자의 색상과 모양 결정\n    \n    viewer.zoomTo()\n    viewer.show()","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-04-28T08:04:05.064369Z","iopub.execute_input":"2025-04-28T08:04:05.064719Z","iopub.status.idle":"2025-04-28T08:04:05.071650Z","shell.execute_reply.started":"2025-04-28T08:04:05.064689Z","shell.execute_reply":"2025-04-28T08:04:05.070633Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"# Feel free to change the PDB ID\nsample_pdb_id = '1ATO'\nplot_backbone(sample_pdb_id)","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-04-28T08:04:05.072810Z","iopub.execute_input":"2025-04-28T08:04:05.073196Z","iopub.status.idle":"2025-04-28T08:04:05.139103Z","shell.execute_reply.started":"2025-04-28T08:04:05.073161Z","shell.execute_reply":"2025-04-28T08:04:05.137650Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"In this visualiztion, phosphate group and sugar are represented by red sphere and pentagon repectively. Now you might ask, Do we have to predict the location of each atom of the nucleotide? The short answer is **NO**, not in this competition! Here, our goal is to predict the location of each C'1 atom (first-red carbon sugar backbone) of each nucleotide in RNA sequence. \n\nWe will come back to the details of the conformation later in the notebook.\n\n**[Fact]:** I used the mmCIF (.cif) file format, representing each `target_id` of the notebook. You can access the same dataset [Here](https://www.kaggle.com/datasets/amirmmahdavikia/stanford-rna-3d-prediction-pdb-dataset)!","metadata":{}},{"cell_type":"code","source":"for df, desc in zip([train_seq, valid_seq, test_seq],\n                    ['Train Sequence', 'Validation Sequence','Test Sequence']):\n    \n    df['length'] = df.sequence.apply(lambda x: len(x))\n    print(clr.S+f'\\n----- {desc} Length -----\\n'+clr.E)\n    print(f'[MIN]: {df.length.min()}')\n    print(f'[MAX]: {df.length.max()}')\n    print(f'[AVG]: {df.length.mean():.2f}')\n    print(f'[STD]: {df.length.std():.2f}')\n    print(f'[Q1]: {df.length.quantile([0.25]).iloc[0]:.2f}')\n    print(f'[Q2]: {df.length.quantile([0.5]).iloc[0]:.2f}')\n    print(f'[Q3]: {df.length.quantile([0.75]).iloc[0]:.2f}')","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-04-28T08:04:05.140329Z","iopub.execute_input":"2025-04-28T08:04:05.140705Z","iopub.status.idle":"2025-04-28T08:04:05.170401Z","shell.execute_reply.started":"2025-04-28T08:04:05.140678Z","shell.execute_reply":"2025-04-28T08:04:05.169111Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"print(f'[Q2]: {df.length.quantile([0.5]).iloc[0]}')","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-04-28T08:04:05.171484Z","iopub.execute_input":"2025-04-28T08:04:05.171837Z","iopub.status.idle":"2025-04-28T08:04:05.179150Z","shell.execute_reply.started":"2025-04-28T08:04:05.171808Z","shell.execute_reply":"2025-04-28T08:04:05.178080Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"def plot_length_histogram(df, desc, ax=None):\n    \"\"\"RNA 시퀀스의 길이별 빈도를 확인하는 그래프를 그리는 함수\"\"\"\n    if ax is None:\n        fig, ax = plt.subplots(figsize=(8, 5))\n    sns.histplot(df.length, kde=True, ax=ax, color=colors[-1])\n    ax.set_title('Sequence Length Frequency')\n    ax.set_xlabel('Length')\n    if fig:\n        plt.savefig(f'Figures/{desc}_plot_length_histogram.png', dpi=300, bbox_inches='tight')","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-04-28T08:04:05.180317Z","iopub.execute_input":"2025-04-28T08:04:05.180736Z","iopub.status.idle":"2025-04-28T08:04:05.204146Z","shell.execute_reply.started":"2025-04-28T08:04:05.180698Z","shell.execute_reply":"2025-04-28T08:04:05.202767Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"plot_length_histogram(train_seq, 'Train')","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-04-28T08:04:05.205286Z","iopub.execute_input":"2025-04-28T08:04:05.205603Z","iopub.status.idle":"2025-04-28T08:04:06.434216Z","shell.execute_reply.started":"2025-04-28T08:04:05.205569Z","shell.execute_reply":"2025-04-28T08:04:06.432995Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"# <div style=\"text-align:center; border-radius:10px; color:black; margin:0; font-size:100%; font-family:Carbon Plus; background-color:white; overflow:hidden\"><b>Exploration</b></div>","metadata":{}},{"cell_type":"code","source":"for df, desc in zip([train_seq, valid_seq, test_seq],\n                    ['Train Sequence', 'Validation Sequence','Test Sequence']):\n    \n    df['temporal_cutoff'] = pd.to_datetime(df['temporal_cutoff'], errors='coerce') # coerce: 강요하다. datetime 형태로 변경할 수 없는 값을 만나면 Nan으로 변환한다.\n    print(clr.S+f'\\n----- {desc} Cutoff -----\\n'+clr.E)\n    print(f'[MIN]: {df.temporal_cutoff.min()}')\n    print(f'[MAX]: {df.temporal_cutoff.max()}')","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-04-28T08:04:06.435432Z","iopub.execute_input":"2025-04-28T08:04:06.435831Z","iopub.status.idle":"2025-04-28T08:04:06.453836Z","shell.execute_reply.started":"2025-04-28T08:04:06.435793Z","shell.execute_reply":"2025-04-28T08:04:06.452598Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"def plot_cutoff_histogram(df, desc, axs=None):\n    \"\"\"데이터별 cutoff와 그 유효성에 대한 그래프를 그리는 함수\"\"\"\n    if axs is None:\n        fig, axs = plt.subplots(nrows=1, ncols=2, figsize=(12, 6), gridspec_kw={'width_ratios': [3, 1]})\n    else:\n        fig = None\n\n    bars = sns.histplot(df.temporal_cutoff.dt.to_period('Y').astype('str'), kde=True, ax=axs[0], color=colors[-1])\n    axs[0].set_title('Temporal Cutoff Frequency')\n    axs[0].set_xlabel('Date')\n    \n    plt.setp(axs[0].xaxis.get_majorticklabels(), rotation=45)  # 특정 plot(axs[0]의) x축의 ticks만 45도 기울기로 표현\n\n    # 이건 왜 적용하는 거지?\n    for bar in bars.patches:\n        # get_x는 막대의 중간 포인트의 값을 반환\n        if (bar.get_x() + 1.5) >= 28.0:  # 바의 인덱스에 1.5를 더한 것이 28보다 크면 바의 색상을 변경\n            bar.set_facecolor(colors[1])\n\n    # 컷오프의 최대값보다 작은 값을 찾아야 하는 거 아닌가요?\n    valid_cutoff = (df.temporal_cutoff < valid_seq.temporal_cutoff.min()).value_counts().reset_index()\n    sizes = valid_cutoff['count'].to_list()\n    labels = valid_cutoff['temporal_cutoff'].to_list()\n    valid_map = {True: ['Valid', colors[3]], False: ['Invalid', colors[1]]}\n    validity = [f'{valid_map[label][0]} ({size})' for size, label in zip(sizes, labels)]\n    _colors = [valid_map[label][1] for label in labels]\n    axs[1].pie(sizes, labels=validity, colors=_colors)  # size는 [692, 152], labels는 [Valid (692), Invalid (152)]\n    axs[1].set_title('Cutoff Validity')\n\n    plt.xticks(rotation=45)\n\n    if fig:\n        plt.tight_layout()\n        plt.savefig(f'Figures/{desc}_plot_cutoff_histogram.png', dpi=300, bbox_inches='tight')","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-04-28T08:04:06.459983Z","iopub.execute_input":"2025-04-28T08:04:06.460463Z","iopub.status.idle":"2025-04-28T08:04:06.475971Z","shell.execute_reply.started":"2025-04-28T08:04:06.460419Z","shell.execute_reply":"2025-04-28T08:04:06.474566Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"plot_cutoff_histogram(train_seq, 'Train')","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-04-28T08:04:06.478079Z","iopub.execute_input":"2025-04-28T08:04:06.478423Z","iopub.status.idle":"2025-04-28T08:04:07.975433Z","shell.execute_reply.started":"2025-04-28T08:04:06.478395Z","shell.execute_reply":"2025-04-28T08:04:07.974247Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"for df, desc in zip([train_labels, valid_labels],\n                    ['Train Label', 'Validation Label']):\n    # 1SCL_A_1\n    df['chain'] = df.ID.apply(lambda x: x.split('_')[1])  # A\n    df['pdb_id'] = df.ID.apply(lambda x: x.split('_')[0]) # 1SCL\n    df['target_id'] = df.ID.apply(lambda x: re.sub(r'_(\\d+)$', '', x)) # 잔기를 제거하고, chain이 숫자라면 공백으로 변환\n    missing_count = df.groupby('target_id').x_1.apply(lambda x: x.isna().sum()) # targetID별 x_1 좌표가 null인 데이터의 개수를 가지고 있는 series\n\n    if 'Train' in desc:\n        o_missing_count = pd.merge(missing_count.reset_index(), train_seq, on='target_id', how='right')['x_1'] # train_seq에 맞추어서 정렬\n        train_seq['missing_label'] = o_missing_count\n        train_seq['missing_ratio'] = train_seq['missing_label'] / train_seq['length']\n    else:\n        o_missing_count = pd.merge(missing_count.reset_index(), valid_seq, on='target_id', how='right')['x_1']\n        valid_seq['missing_label'] = o_missing_count\n        valid_seq['missing_ratio'] = valid_seq['missing_label'] / valid_seq['length']\n\n    print(clr.L+f'\\n----- {desc} Missing -----\\n'+clr.E)\n    print(f'[COUNT]: {missing_count[missing_count>0].count()}')\n    print(f'[MIN]: {missing_count.min()}')\n    print(f'[MAX]: {missing_count.max()}')\n    print(f'[AVG]: {missing_count.mean():.2f}')","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-04-28T08:04:07.976746Z","iopub.execute_input":"2025-04-28T08:04:07.977092Z","iopub.status.idle":"2025-04-28T08:04:08.472108Z","shell.execute_reply.started":"2025-04-28T08:04:07.977065Z","shell.execute_reply":"2025-04-28T08:04:08.470564Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"def plot_missing_label(df, desc, axs=None):\n    \"\"\"missing_label과 missing_ratio를 확인할 수 있는 그래프를 그리는 함수\"\"\"\n    \n    if axs is None:\n        fig, axs = plt.subplots(nrows=1, ncols=2, figsize=(12, 6), gridspec_kw={'width_ratios': [1, 4]})\n    else:\n        fig = None\n\n    s=df['missing_label'] > 0\n    sns.countplot(\n        x=s,\n        ax=axs[0],\n        palette=[colors[4], colors[-1]]\n    )\n    axs[0].set_title('Any coordinate missing?')\n    axs[0].set_xlabel('')\n    ticks = axs[0].get_xticks()\n    _labels = ['No', 'Yes']\n    axs[0].set_xticks(ticks=ticks, labels=_labels[:len(ticks)])\n    axs[0].set_ylabel('Count')\n    ylim = axs[0].get_ylim()\n\n    sns.histplot(\n        df.loc[df.missing_label > 0, 'missing_ratio'], \n        ax=axs[1],\n        color=colors[-1],\n        kde=True\n    )\n\n    axs[1].set_title('Missing Ratio')\n    axs[1].set_xlabel('Ratio')\n    axs[1].set_ylabel('')\n    axs[1].set_ylim(ylim)  # y축의 값을 axs[0]의 것과 일치시킴\n\n    if fig:\n        plt.tight_layout()\n        plt.savefig(f'Figures/{desc}_plot_missing_label.png', dpi=300, bbox_inches='tight')","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-04-28T08:04:08.473724Z","iopub.execute_input":"2025-04-28T08:04:08.474044Z","iopub.status.idle":"2025-04-28T08:04:08.483535Z","shell.execute_reply.started":"2025-04-28T08:04:08.473993Z","shell.execute_reply":"2025-04-28T08:04:08.482218Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"plot_missing_label(train_seq, 'Train')","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-04-28T08:04:08.484663Z","iopub.execute_input":"2025-04-28T08:04:08.484966Z","iopub.status.idle":"2025-04-28T08:04:09.696835Z","shell.execute_reply.started":"2025-04-28T08:04:08.484941Z","shell.execute_reply":"2025-04-28T08:04:09.695677Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"def resample_array(arr, target_length=1024):\n    \"\"\"np 배열의 길이를 1024로 조정하여 정규화를 용이하게 하는 함수\"\"\"\n    return np.round(resample(arr, target_length)).astype(int)","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-04-28T08:04:09.698359Z","iopub.execute_input":"2025-04-28T08:04:09.698707Z","iopub.status.idle":"2025-04-28T08:04:09.704375Z","shell.execute_reply.started":"2025-04-28T08:04:09.698670Z","shell.execute_reply":"2025-04-28T08:04:09.702870Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"def binary_sequence(target_id):\n    \"\"\"각 시퀀스마다 null이면 0, null이 아니면 1을 부여하여 반환하는 함수\"\"\"\n    sequence = train_seq[train_seq.target_id == target_id].sequence\n    b_sequence = np.array(train_labels[train_labels.target_id == target_id].x_1.notna().astype(int)) # target_id가 같은 행을 필터링하고, 그것이 n/a가 아니라면 1, 맞다면 0으로 표시한 배열을 할당\n    resampled = resample_array(b_sequence)\n    return resampled","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-04-28T08:04:09.705358Z","iopub.execute_input":"2025-04-28T08:04:09.705690Z","iopub.status.idle":"2025-04-28T08:04:09.726536Z","shell.execute_reply.started":"2025-04-28T08:04:09.705663Z","shell.execute_reply":"2025-04-28T08:04:09.725289Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"def plot_missing_label_pattern(df, desc, ax=None):\n    if ax is None:\n        fig, ax = plt.subplots(figsize=(12, 6))\n\n    concat_matrix = []\n    selected_df = df.sort_values('missing_label', ascending=False).iloc[:30]  # missing_label이 많은 상위 30개의 요소를 가지는 df\n    for idx, rna in selected_df.iterrows():\n        b_seq = binary_sequence(rna.target_id)\n        concat_matrix.append(b_seq)\n\n    sns.heatmap(np.array(concat_matrix), cmap='viridis', ax=ax, cbar=False)\n    ax.set_title('Missing Label Pattern')\n    ax.set_yticklabels(labels=selected_df.target_id.to_list())\n    ax.xaxis.set_visible(False)\n    plt.setp(ax.yaxis.get_majorticklabels(), rotation=0)\n    if fig:\n        plt.savefig(f'Figures/{desc}_plot_missing_label_pattern.png', dpi=300, bbox_inches='tight')","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-04-28T08:04:09.727720Z","iopub.execute_input":"2025-04-28T08:04:09.728049Z","iopub.status.idle":"2025-04-28T08:04:09.748884Z","shell.execute_reply.started":"2025-04-28T08:04:09.727989Z","shell.execute_reply":"2025-04-28T08:04:09.747646Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"plot_missing_label_pattern(train_seq, 'Train')","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-04-28T08:04:09.750146Z","iopub.execute_input":"2025-04-28T08:04:09.750576Z","iopub.status.idle":"2025-04-28T08:04:11.640841Z","shell.execute_reply.started":"2025-04-28T08:04:09.750537Z","shell.execute_reply":"2025-04-28T08:04:11.639530Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"### Quick Takeaways:\n\n+ `[Train Dataset]`:\n    + 90% of sequences contain fewer than 160 nucleotides.\n    + Only 34 sequences exceed 1000 nucleotides.\n    + 152 sequences are **invalid for Phase 1 training** due to a temporal cutoff after `2022-05-27`.\n    + 238 sequences **lack at least one residue coordinate**.\n    + 46 sequences are **missing coordinate labels entirely**.\n    + No clear pattern is observed in the missing coordinate labels.","metadata":{}},{"cell_type":"markdown","source":"# <div style=\"text-align:center; border-radius:10px; color:black; margin:0; font-size:100%; font-family:Carbon Plus; background-color:white; overflow:hidden\"><b>Secondary Structure</b></div>","metadata":{}},{"cell_type":"markdown","source":"Despite being single-stranded, RNA molecules often exhibit a great deal of **double-helical** character. This is because RNA chains frequently fold back on themselves to form base-paired segments between short stretches of complementary sequences.\n\n![Secondary Structure](https://raw.githubusercontent.com/ammomahdavikia/asset-holding/main/rna_secondary_structure.png)\n\nIf the two stretches of complementary sequence are near each other, the RNA may adopt a **stem-loop structure** in which the intervening RNA is looped out from the end of the double-helical segment. \n\nStretches of double-helical RNA may also exhibit **internal loops** (unpaired nucleotides on either side of the stem), **bulges** (anunpaired nucleotide on one side of the bulge), or **junctions**.","metadata":{}},{"cell_type":"markdown","source":"### How Can We Predict RNA Secondary Structure?\n\nUltimately, a true secondary structure can be verifed when the 3D structure is determined. Crystallography, NMR, and now high-resolution cryo-electron microscopy (cryo-EM) represent the _crème de la crème_, with NMR and cryo-EM also having the ability to potentially detect conformational ensembles and dynamics.\n\n![Secondary Structure](https://raw.githubusercontent.com/ammomahdavikia/asset-holding/main/rna_secondary_structure_prediction.png)\n\nHowever, for large-scale analysis, computational prediction remains the most practical approach. For now, we will use automatic **thermodynamics-based** algorithm (**RNAfold**, part of the **ViennaRNA**) to predict the secondary structure for each sequence. ","metadata":{}},{"cell_type":"code","source":"def dot_bracket(seq):\n    \"\"\"\n    Predicts the minimum free energy (MFE) secondary structure of an RNA sequence in dot-bracket notation.\n    \"\"\"\n    fc = RNA.fold_compound(seq)  # 열역학적으로 가장 안정적인 2차원 RNA의 구조를 예측하는 메서드\n    return fc.mfe()","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-04-28T08:04:11.641921Z","iopub.execute_input":"2025-04-28T08:04:11.642474Z","iopub.status.idle":"2025-04-28T08:04:11.647072Z","shell.execute_reply.started":"2025-04-28T08:04:11.642443Z","shell.execute_reply":"2025-04-28T08:04:11.645901Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"def predict_secondary_structure(df, df_name='dataset'):\n    \"\"\"\"\"\"\n    print(clr.S+f\"\\n--- Dot-Bracket {df_name} ---\\n\"+clr.E)\n\n    ss_results = parallel_apply(df['sequence'], dot_bracket, desc=f'{df_name} - DB', n_jobs=-1)  # df['sequence']의 각 행의 값에 dot_bracket을 병렬로 적용\n    df[['secondary_structure', 'mfe']] = pd.DataFrame(ss_results, index=df.index)  # (Secondary Stucture, Minimum Free Energy)\n\n    return df","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-04-28T08:04:11.648232Z","iopub.execute_input":"2025-04-28T08:04:11.648550Z","iopub.status.idle":"2025-04-28T08:04:11.676635Z","shell.execute_reply.started":"2025-04-28T08:04:11.648526Z","shell.execute_reply":"2025-04-28T08:04:11.675465Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"for df, desc in zip([train_seq, valid_seq, test_seq],\n                    ['Train Sequence', 'Validation Sequence','Test Sequence']):\n    \n    df = predict_secondary_structure(df, desc)","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-04-28T08:04:11.677961Z","iopub.execute_input":"2025-04-28T08:04:11.678598Z","iopub.status.idle":"2025-04-28T08:09:19.573249Z","shell.execute_reply.started":"2025-04-28T08:04:11.678548Z","shell.execute_reply":"2025-04-28T08:09:19.571756Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"The **dot-bracket** notation is a simplified way to represent RNA secondary structures, where dots `.` indicate unpaired nucleotides, and matching parentheses `( )` represent base pairs (e.g., ”((…))” for a hairpin loop). ","metadata":{}},{"cell_type":"code","source":"train_seq[['target_id', 'sequence', 'secondary_structure']].head()","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-04-28T08:09:19.574651Z","iopub.execute_input":"2025-04-28T08:09:19.574990Z","iopub.status.idle":"2025-04-28T08:09:19.595598Z","shell.execute_reply.started":"2025-04-28T08:09:19.574961Z","shell.execute_reply":"2025-04-28T08:09:19.594539Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"It doesn't seem intuitive, right? Let's visualize better:","metadata":{}},{"cell_type":"code","source":"def plot_dot_bracket(seq: str, dot_bracket: str, desc=None, max_length=50):\n    \"\"\"dot bracket을 2차원으로 그리는 함수\"\"\"\n    fig, ax = plt.subplots(figsize=(4, 4))\n    bg = fgb.BulgeGraph.from_dotbracket(dot_bracket, seq)\n    fvm.plot_rna(bg, ax=ax, text_kwargs={'fontsize': 6, 'color': 'black'})\n    ax.set_title(f'(Potential) Secondary Structure - {desc}')\n    if fig:\n        plt.savefig(f'Figures/{desc}_plot_dot_bracket.png', dpi=300, bbox_inches='tight')","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-04-28T08:09:19.596789Z","iopub.execute_input":"2025-04-28T08:09:19.597141Z","iopub.status.idle":"2025-04-28T08:09:19.609761Z","shell.execute_reply.started":"2025-04-28T08:09:19.597113Z","shell.execute_reply":"2025-04-28T08:09:19.608618Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"sample_idx = np.random.randint(len(train_seq))  # train_seq의 길이(크기)보다 작은 음이 아닌 정수 하나를 무작위로 생성\nsample_seq = train_seq.sequence.iloc[sample_idx]\nsample_target_id = train_seq.target_id.iloc[sample_idx]\nsample_dot_bracket = train_seq.secondary_structure.iloc[sample_idx]\nplot_dot_bracket(sample_seq, sample_dot_bracket, desc=sample_target_id)","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-04-28T08:09:19.610893Z","iopub.execute_input":"2025-04-28T08:09:19.611238Z","iopub.status.idle":"2025-04-28T08:09:20.423250Z","shell.execute_reply.started":"2025-04-28T08:09:19.611210Z","shell.execute_reply":"2025-04-28T08:09:20.421818Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"### How can we use it for prediction?\n\nWe need better representation, **Base Pair Probability Matrix (BPPM)**. This matrix represents the likelihood of each nucleotide pairing with another in an RNA secondary structure, where each entry $(i, j)$ contains the probability that nucleotide $i$ forms a base pair with nucleotide $j$.","metadata":{}},{"cell_type":"code","source":"def base_pairing_probability(seq):\n    \"\"\"\n    Computes the base pairing probability matrix (BPPM) for a given RNA sequence.\n    \"\"\"\n    fc = RNA.fold_compound(seq) # {'sequence': \"...\", 'length: \"...\", 'strands(가닥 수)': \"...\"}\n    # RNA 구조의 모든 가능한 상태를 고려하여 각 상태의 확률을 계산하는 함수 -> 구조의 정확성 향상\n    fc.pf()  # partition function(분할 함수)\n    # 각 뉴클레오타이드가 염기쌍을 형성할 확률을 튜플로 반환\n    bpp = fc.bpp() # base paring probability(염기쌍 확률)\n    bppm = np.array([list(bp) for bp in bpp])\n    bppm = bppm[1:, 1:]  # Remove 0th dummy row and column\n    return bppm","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-04-28T08:09:20.424418Z","iopub.execute_input":"2025-04-28T08:09:20.424753Z","iopub.status.idle":"2025-04-28T08:09:20.430961Z","shell.execute_reply.started":"2025-04-28T08:09:20.424726Z","shell.execute_reply":"2025-04-28T08:09:20.429488Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"def plot_bppm(seq, dot_bracket, desc=None, max_length=50):\n    \"\"\"RNA 구조에 대한 bppm과 2차원 구조를 그리는 함수\"\"\"\n    bg = fgb.BulgeGraph.from_dotbracket(dot_bracket, seq)\n    if len(seq) > max_length:\n        seq = seq[:max_length]\n    \n    bppm = base_pairing_probability(seq)\n    fig, axs = plt.subplots(nrows=1, ncols=2, figsize=(12, 6), gridspec_kw={'width_ratios': [2, 1]})\n    \n    sns.heatmap(bppm, \n                cmap='viridis',\n                ax=axs[0],\n                annot=False  # 각 셀에 해당하는 값을 표기하지 않음\n               )\n    axs[0].set_xlabel('Position in sequence')\n    axs[0].set_ylabel('Position in sequence')\n    axs[0].set_title(f'Base Pairing Probability Matrix - {desc}')\n\n    \n    fvm.plot_rna(bg, ax=axs[1], text_kwargs={'fontsize': 8, 'color': 'black'})\n    axs[1].set_title(f'Secondary Structure - {desc}')\n    if fig:\n        plt.savefig(f'Figures/{desc}_plot_bppm.png', dpi=300, bbox_inches='tight')","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-04-28T08:09:20.432296Z","iopub.execute_input":"2025-04-28T08:09:20.432753Z","iopub.status.idle":"2025-04-28T08:09:20.453749Z","shell.execute_reply.started":"2025-04-28T08:09:20.432711Z","shell.execute_reply":"2025-04-28T08:09:20.452238Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"plot_bppm(sample_seq, sample_dot_bracket, desc=sample_target_id)","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-04-28T08:09:20.455665Z","iopub.execute_input":"2025-04-28T08:09:20.456063Z","iopub.status.idle":"2025-04-28T08:09:22.292028Z","shell.execute_reply.started":"2025-04-28T08:09:20.455996Z","shell.execute_reply":"2025-04-28T08:09:22.290684Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"### Pitfall?\n\nThe danger of only considering the lowest free-energy structure and showcasing it as **“The”** secondary structure is the major pitfall here. Thermodynamics-based algorithms consider all nucleotides as equally likely to be involved in secondary structure elements, which often leads to erroneous assumptions about base pairs.\n\nA feature of RNA that adds to its propensity to form double-helical structures is additional **non-Watson-Crick** base pairs. One such example is the G:U base pair. Non-Watson-Crick base pairs can be found in all\ncombinations in RNA (GA and GU are the most abundant in ribosomal RNA).\n\n![Base Pair Types](https://raw.githubusercontent.com/ammomahdavikia/asset-holding/main/base_pair_types.png)\n\nBecause such non-Watson-Crick base pairs can occur as well as the two conventional Watson-Crick base pairs, RNA chains have an enhanced capacity for self-complementarity.","metadata":{}},{"cell_type":"code","source":"def plot_mfe_length(df, desc, ax=None):\n    \"\"\"sequence의 mfe(minimum free energy)와 길이 사이의 관계를 확인하는 그래프\"\"\"\n    sns.jointplot(df, x='mfe', y='length', kind='reg', color=colors[-1])","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-04-28T08:09:22.293278Z","iopub.execute_input":"2025-04-28T08:09:22.293707Z","iopub.status.idle":"2025-04-28T08:09:22.299253Z","shell.execute_reply.started":"2025-04-28T08:09:22.293667Z","shell.execute_reply":"2025-04-28T08:09:22.297988Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"plot_mfe_length(train_seq, 'Train')","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-04-28T08:09:22.300344Z","iopub.execute_input":"2025-04-28T08:09:22.300687Z","iopub.status.idle":"2025-04-28T08:09:24.297039Z","shell.execute_reply.started":"2025-04-28T08:09:22.300658Z","shell.execute_reply":"2025-04-28T08:09:24.295954Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"train_seq.mfe.head()","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-04-28T08:09:24.297952Z","iopub.execute_input":"2025-04-28T08:09:24.298289Z","iopub.status.idle":"2025-04-28T08:09:24.306570Z","shell.execute_reply.started":"2025-04-28T08:09:24.298263Z","shell.execute_reply":"2025-04-28T08:09:24.305378Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"# <div style=\"text-align:center; border-radius:10px; color:black; margin:0; font-size:100%; font-family:Carbon Plus; background-color:white; overflow:hidden\"><b>🎯 Tertiary Structure 🎯</b></div>","metadata":{}},{"cell_type":"code","source":"def plot_xyz_frequency(df, desc, axs=None):\n    \"\"\"(x, y, z) 좌표의 값의 분포를 확인하는 그래프를 그리는 함수\"\"\"\n\n    if axs is None:\n        fig, axs = plt.subplots(nrows=1, ncols=3, figsize=(12, 4))\n    else:\n        fig = None\n\n    sns.histplot(df.x_1.replace(-1.000000e+18, np.nan), kde=True, ax=axs[0], color=colors[-1])\n    axs[0].set_title('X coordinate')\n    axs[0].set_xlabel('')\n    axs[0].set_ylabel('Count')\n    axs[0].set_xlim([-1000, 1000])\n    axs[0].set_ylim([0, 6000])\n\n    sns.histplot(df.y_1.replace(-1.000000e+18, np.nan), kde=True, ax=axs[1], color=colors[-1])\n    axs[1].set_title('Y coordinate')\n    axs[1].set_xlabel('')\n    axs[1].set_ylabel('Count')\n    axs[1].set_xlim([-1000, 1000])\n    axs[1].set_ylim([0, 6000])\n\n    sns.histplot(df.z_1.replace(-1.000000e+18, np.nan), kde=True, ax=axs[2], color=colors[-1])\n    axs[2].set_title('Z coordinate')\n    axs[2].set_xlabel('')\n    axs[2].set_ylabel('Count')\n    axs[2].set_xlim([-1000, 1000])\n    axs[2].set_ylim([0, 6000])\n    \n    if fig:\n        plt.tight_layout()\n        plt.savefig(f'Figures/{desc}_plot_xyz_frequency.png', dpi=300, bbox_inches='tight')","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-04-28T08:09:24.307628Z","iopub.execute_input":"2025-04-28T08:09:24.307930Z","iopub.status.idle":"2025-04-28T08:09:24.326663Z","shell.execute_reply.started":"2025-04-28T08:09:24.307895Z","shell.execute_reply":"2025-04-28T08:09:24.325521Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"plot_xyz_frequency(train_labels, 'Train')","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-04-28T08:09:24.327747Z","iopub.execute_input":"2025-04-28T08:09:24.328157Z","iopub.status.idle":"2025-04-28T08:09:28.939910Z","shell.execute_reply.started":"2025-04-28T08:09:24.328127Z","shell.execute_reply":"2025-04-28T08:09:28.938605Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"def plot_3d_strcutrue(pdb_id: str):\n    \"\"\"RNA의 3차원 구조를 그리는 함수\"\"\"\n    # pdb_id가 같은 것은 여러 개 존재할 수 있다.\n    rna = train_labels[train_labels.pdb_id == pdb_id]\n    chains = sorted(set(rna.chain))\n    num_chains = len(chains)\n\n    def interpolate_gray(index: int, total: int):\n        \"\"\"chain id에 따라서 label의 색상을 지정하는 함수\"\"\"\n        ratio = index / max(total - 1, 1)\n        gray_value = int(200 - (150 * ratio))\n        return f'rgb({gray_value},{gray_value},{gray_value})'\n    \n    chain_color_map = {chain: interpolate_gray(i, num_chains) for i, chain in enumerate(chains)}\n    \n    pdb_data = \"\"\n    for _, residue in rna.iterrows():  # 인덱스를 제외한 나머지 요소를 residue에 담음\n        #  레코드 타입 잔기ID 원자 이름 잔기 이름 체인ID 잔기ID 좌표x 좌표y 좌표z \n        pdb_data += f'ATOM  {residue.resid:5d}  P   {residue.resname} {residue.chain} {residue.resid:4d}    '\\\n                    f'{residue.x_1:8.3f}{residue.y_1:8.3f}{residue.z_1:8.3f}\\n'\n    view = py3Dmol.view(width=600, height=400)  # 너비 600, 높이 400의 그림 영역을 만드는 함수\n    view.addModel(pdb_data, 'pdb') # 그리고자 하는 data 추가\n    \n    for _, residue in rna.iterrows():\n        # c1 원자의 위치를 고려한 구체를 추가함\n        view.addSphere({\n            'center': {'x': residue.x_1, 'y': residue.y_1, 'z': residue.z_1},\n            'radius': 1.2,\n            'color': nt_clr.get(residue.resname, 'white')\n        })\n        # 구체 각각에 잔기 이름과 잔기 id를 결합한 label을 부착함\n        view.addLabel(f'{residue.resname}{residue.resid}', {\n            'position': {'x': residue.x_1, 'y': residue.y_1, 'z': residue.z_1},\n            'backgroundColor': chain_color_map[residue.chain],\n            'fontColor': 'white',\n            'fontSize': 10\n        })\n\n\n    cif_filename = f\"{config.CIF_PATH}/{pdb_id}.cif\"  # datasets의 pdb_id.cif\n    \n    parser = MMCIFParser(QUIET=True)\n    # get_structure는 모델 객체의 이터레이터를 가져 옴\n    structure = parser.get_structure(pdb_id, cif_filename)\n    \n    output_pdb = StringIO()\n    # 구조가 여러 모델을 포함할 수 있으나 첫 번째 모델을 선택하여 작업을 진행하고자 next() 사용\n    first_model = next(structure.get_models())\n    \n    for chain in first_model:\n        if chain.id in chains:  # Include only relevant chains\n            for residue in chain:\n                for atom in residue:\n                    # atom.name은 왼쪽 4자리 정렬\n                    # pdb_data에서는 atom_name을 p(인산)으로 고정하였는데 여기서는 다양한 원소 이름을  사용할 수 있으므로 변수로 지정\n                    output_pdb.write(f\"ATOM  {atom.serial_number:5d}  {atom.name:<4} {residue.resname} {chain.id} {residue.id[1]:4d}    \"\n                                     f\"{atom.coord[0]:8.3f}{atom.coord[1]:8.3f}{atom.coord[2]:8.3f}\\n\")\n                    \n    target_pdb_string = output_pdb.getvalue()\n\n    view.addModel(target_pdb_string, 'pdb') # 나선 구조를 추가함\n    view.setStyle({'cartoon': {'color': 'spectrum'}})  # 특정 색상을 지정할 수도 있음\n    view.zoomTo()  # (0, 0, 0) 좌표가 아니라 구조가 있는 것에 초점을 맞추어 보여 줌\n    view.show()","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-04-28T08:18:34.887736Z","iopub.execute_input":"2025-04-28T08:18:34.888138Z","iopub.status.idle":"2025-04-28T08:18:34.900684Z","shell.execute_reply.started":"2025-04-28T08:18:34.888097Z","shell.execute_reply":"2025-04-28T08:18:34.899167Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"plot_3d_strcutrue(sample_pdb_id)","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-04-28T08:18:35.297736Z","iopub.execute_input":"2025-04-28T08:18:35.298314Z","iopub.status.idle":"2025-04-28T08:18:35.636390Z","shell.execute_reply.started":"2025-04-28T08:18:35.298273Z","shell.execute_reply":"2025-04-28T08:18:35.635213Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"In this representation, the spheres correspond to XYZ coordinates of C'1 atom from `[]_labels.csv`, while the ribbon depicts the actual 3D structure obtained from the PDB.","metadata":{}},{"cell_type":"markdown","source":"# <div style=\"text-align:center; border-radius:10px; color:black; margin:0; font-size:100%; font-family:Carbon Plus; background-color:white; overflow:hidden\"><b>Multiple Sequence Alignment</b></div>","metadata":{}},{"cell_type":"markdown","source":"**Multiple Sequence Alignment (MSA)** is a technique to align structurally similar sequences in order to extract evolutionary information about the sequence structure and potentialy its function. These similarities can reveal evolutionary relationships, conserved functional domains, and structural motifs.\n\nRNAs with similar sequences often fold into similar structures, so aligning homologous sequences helps identify conserved residues and co-evolutionary patterns. These patterns are essential for methods like **AlphaFold** and **RoseTTAFold**.\n\nOne can extract MSAs for each sequence of dataset by retrieving the related datasets (NCBI, Rfam, ...) for similar patterns, But here we will use tha kaggle provided MSA dataset:","metadata":{}},{"cell_type":"code","source":"def msa_count(target_id: str):\n    \"\"\"target_id를 가진 MSA 파일에 포함된 시퀀스의 수를 계산하는 함수\"\"\"\n    msa_file = os.path.join(config.PATH, 'MSA', f'{target_id}.MSA.fasta')\n    try:\n        # fasta 파일의 형태로 파싱하여 각 시퀀스를 반환하는 이터레이터를 생성\n        return sum(1 for _ in SeqIO.parse(msa_file, 'fasta')) - 1  # 참조 시퀀스를 제외한(-1) 전체 시퀀스의 수 \n    except FileNotFoundError:\n        return np.nan\n\nfor df, desc in zip([train_seq, valid_seq, test_seq],\n                    ['Train Sequence', 'Validation Sequence','Test Sequence']):\n    \n    df['msa_count'] = df['target_id'].apply(msa_count)\n    df['msa_presence'] = pd.Series(df['msa_count']).gt(0) # msa_count가 0이라면 유전적, 기능적, 구조적 유사성을 가진 시퀀스가 존재하지 않는다는 미미\n    print(clr.S+f'\\n----- {desc} MSA -----\\n'+clr.E)\n    print(f'[HIT]: {df.msa_presence.sum()}')\n    print(f'[MIN]: {df.msa_count.min()}')\n    print(f'[MAX]: {df.msa_count.max()}')\n    print(f'[AVG]: {df.msa_count.mean():.2f}')\n    df['msa_presence'] = df['msa_presence'].map({True: 'Hit', False: 'No Hit'})\n    print(df.head())","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-04-28T08:53:20.355963Z","iopub.execute_input":"2025-04-28T08:53:20.356403Z","iopub.status.idle":"2025-04-28T08:53:30.640391Z","shell.execute_reply.started":"2025-04-28T08:53:20.356372Z","shell.execute_reply":"2025-04-28T08:53:30.639151Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"def plot_msa_analytics(df, desc, axs=None):\n    \"\"\"MSA 존재 여부와 존재하는 MSA_count의 통계를 그려주는 함수 \"\"\"\n    if axs is None:\n        fig, axs = plt.subplots(nrows=1, ncols=2, figsize=(12, 6), gridspec_kw={'width_ratios': [1, 4]})\n    else:\n        fig = None\n\n    sns.countplot(\n        data=df,\n        x='msa_presence',\n        ax=axs[0],\n        palette=[colors[3], colors[1]]\n    )\n    axs[0].set_title('MSA Presence')\n    axs[0].set_xlabel('')\n    axs[0].set_ylabel('Count')\n    ylim = axs[0].get_ylim()\n\n    sns.histplot(\n        df.loc[df.msa_count > 0, 'msa_count'].dropna(), \n        ax=axs[1],\n        color=colors[-1]\n    )\n\n    axs[1].set_title('Frequency of MSA Hits')\n    axs[1].set_xlabel('Hits')\n    axs[1].set_ylabel('')\n    axs[1].set_ylim(ylim)\n\n    if fig:\n        plt.tight_layout()\n        plt.savefig(f'Figures/{desc}_plot_msa_analytics.png', dpi=300, bbox_inches='tight')","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-04-28T08:09:47.758990Z","iopub.execute_input":"2025-04-28T08:09:47.759398Z","iopub.status.idle":"2025-04-28T08:09:47.767876Z","shell.execute_reply.started":"2025-04-28T08:09:47.759361Z","shell.execute_reply":"2025-04-28T08:09:47.766576Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"plot_msa_analytics(train_seq, 'Train')","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-04-28T08:09:47.769052Z","iopub.execute_input":"2025-04-28T08:09:47.769473Z","iopub.status.idle":"2025-04-28T08:09:49.083067Z","shell.execute_reply.started":"2025-04-28T08:09:47.769432Z","shell.execute_reply":"2025-04-28T08:09:49.081931Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"def plot_msa(target_id, max_sequence=20, max_position=40):\n    msa_file = os.path.join(config.PATH, 'MSA', f'{target_id}.MSA.fasta')\n    \n    sequences = []\n    try:\n        for record in SeqIO.parse(msa_file, 'fasta'):\n            sequences.append(list(str(record.seq)))\n    except FileNotFoundError:\n        print(np.nan)\n\n    # Q1.모든 Sequence의 길이가 동일한 이유?\n    msa_df = pd.DataFrame(sequences)\n    present_nucleotides = set(msa_df.values.flatten()) # 데이터프레임(2차원)을 1차원 set으로 변환\n    _nt_clr = {nt: nt_clr[nt] for nt in present_nucleotides} # e.g. {'U': '#310CA2'}\n    nt_to_num = {nt: idx for idx, nt in enumerate(_nt_clr.keys())}\n    # Q2. ListedColormap이란?\n    cmap = ListedColormap([_nt_clr[nt] for nt in _nt_clr.keys()]) # _nt_clr의 값을 리스트 형태로 바꾼 것\n    \n    msa_numeric = msa_df.replace(nt_to_num) # nt -> num\n    # Q3. max_sequence와 max_position을 정해서 서열을 자르는 이유는? 칸이 부족해서?\n    # 그리고 그 기준은 어떻게 구한 것인가?\n    if msa_numeric.shape[0] > max_sequence:\n        msa_numeric = msa_numeric[:max_sequence]\n    if msa_numeric.shape[1] > max_position:\n        msa_numeric = msa_numeric[:, :max_position]    \n\n    fig, ax = plt.subplots(figsize=(12, 6))\n    sns.heatmap(msa_numeric, \n                cmap=cmap, \n                cbar=False, # colorbar 유무(범례를 표현)\n                xticklabels=False,\n                yticklabels=False, \n                ax=ax,\n                annot=False  # 각 항목별 값을 표현\n               )\n    \n    for i in range(msa_numeric.shape[0]):\n        for j in range(msa_numeric.shape[1]):\n            nucleotide = msa_df.iloc[i, j]\n            ax.text(\n                j + 0.5, i + 0.5, # j(행), i(열)) 순으로 작성, 왼쪽 상단 꼭지점에 위치하므로 0.5씩 더한다.\n                nucleotide, \n                ha='center', # 수평 정렬(horizonal alignment)\n                va='center', # 수직 정렬(vertical alignment)\n                fontsize=8, \n                color='black', \n                fontweight='bold'\n                   )\n    \n    _labels = ['Original'] + [str(idx+1) for idx in range(msa_numeric.shape[0]-1)]\n    ax.set_yticks(np.arange(msa_numeric.shape[0]) + 0.5) # 이것도 상단에 정렬되어 있으므로 0.5씩 더하여 중간에 위치하도록 한다.\n    ax.set_yticklabels(_labels)\n    plt.xlabel('Position in sequence')\n    plt.title(f'MSA - {target_id}')\n\n    if fig:\n        plt.savefig(f'Figures/{target_id}_plot_msa.png', dpi=300, bbox_inches='tight')\n    plt.show()","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-04-28T10:47:04.557252Z","iopub.execute_input":"2025-04-28T10:47:04.557713Z","iopub.status.idle":"2025-04-28T10:47:04.572396Z","shell.execute_reply.started":"2025-04-28T10:47:04.557676Z","shell.execute_reply":"2025-04-28T10:47:04.571116Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"plot_msa(target_id='1EBQ_A') # Feel free to change","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-04-28T10:47:04.887504Z","iopub.execute_input":"2025-04-28T10:47:04.887875Z","iopub.status.idle":"2025-04-28T10:47:07.960826Z","shell.execute_reply.started":"2025-04-28T10:47:04.887848Z","shell.execute_reply":"2025-04-28T10:47:07.959270Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"# <div style=\"text-align:center; border-radius:10px; color:black; margin:0; font-size:100%; font-family:Carbon Plus; background-color:white; overflow:hidden\"><b>Baseline</b></div>\n\n### To Be Continued ...\n\n#### ✓ To Do:\n+ Pytorch Baseline implementation\n+ List Similar Datasets\n+ List Current Approaches","metadata":{}},{"cell_type":"markdown","source":"# <div style=\"text-align:center; border-radius:10px; color:black; margin:0; font-size:100%; font-family:Carbon Plus; background-color:white; overflow:hidden\"><b>Summary</b></div>\n\n![Base Pair Types](https://raw.githubusercontent.com/ammomahdavikia/asset-holding/main/stanford_rna_3d_folding_summary.png)","metadata":{}},{"cell_type":"markdown","source":"### Resources\n\n1.\tWatson, J. D., Baker, T. A., Bell, S. P., Gann, A., Levine, M., & Losick, R. (2013). Molecular Biology of the Gene (7th ed.). Pearson Education. ISBN: 978-0321762436.\n2. Nelson, D. L., & Cox, M. M. (2017). Lehninger Principles of Biochemistry (7th ed.). W.H. Freeman. ISBN: 978-1464126116.\n3. Berman, H. M., Westbrook, J., Feng, Z., Gilliland, G., Bhat, T. N., Weissig, H., … & Bourne, P. E. (2000). The Protein Data Bank. Nucleic Acids Research, 28(1), 235-242. DOI: 10.1093/nar/28.1.235.\n4. Urry, L. A., Cain, M. L., Wasserman, S. A., Minorsky, P. V., & Reece, J. B. (2020). Campbell Biology (12th ed.). Pearson. ISBN: 978-0135188743.\n5. Jumper, J., Evans, R., Pritzel, A., et al. (2021). Highly accurate protein structure prediction with AlphaFold. Nature, 596, 583–589. DOI: 10.1038/s41586-021-03819-2.\n6. Weeks, K. M. (2021). Thoughts on how to think (and talk) about RNA structure. Proceedings of the National Academy of Sciences, 118(37), e2112677119. DOI: 10.1073/pnas.2112677119.\n7. Wang, W., Feng, C., Han, R., Wang, Z., Ye, L., Du, Z., Wei, H., Zhang, F., Peng, Z., & Yang, J. (2023). trRosettaRNA: automated prediction of RNA 3D structure with transformer network. Nature Communications, 14, 7266. DOI: 10.1038/s41467-023-42528-4.","metadata":{}},{"cell_type":"markdown","source":"### <div style=\"text-align:center; border-radius:10px; color:black; margin:0; font-size:100%; font-family:Carbon Plus; background-color:white; overflow:hidden\"><b>🔼 Your Welcome 🔼</b></div>","metadata":{}}]}