{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"name":"python","version":"3.11.11","mimetype":"text/x-python","codemirror_mode":{"name":"ipython","version":3},"pygments_lexer":"ipython3","nbconvert_exporter":"python","file_extension":".py"},"kaggle":{"accelerator":"none","dataSources":[{"sourceId":87793,"databundleVersionId":12024591,"sourceType":"competition"}],"dockerImageVersionId":31012,"isInternetEnabled":true,"language":"python","sourceType":"notebook","isGpuEnabled":false}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"# We will cover :\n1. Goal of Comeptition\n2. About Data Set given\n3. Approach to DataSet and Guide to Train ML ( including import topics you must known )\n4. Data Visulization including 3D\n5. Discriptive and Infrential Statistiscs","metadata":{}},{"cell_type":"markdown","source":"# A. **Goal of the Competition ------------------>**","metadata":{}},{"cell_type":"markdown","source":"Predicting the 3D structure of RNA molecules from their sequence data. In short, **You are given RNA Sequence like ( GAACGA.. ) and you had to find out the position of Molecule in 3D Information**","metadata":{}},{"cell_type":"markdown","source":"Note : Unlike proteins, RNA structure prediction is less explored, and this competition aims to advance AI/ML techniques in this domain.","metadata":{}},{"cell_type":"markdown","source":"**Real World Application :**","metadata":{}},{"cell_type":"markdown","source":"\n* Drug Discovery: Many viruses (e.g., HIV, SARS-CoV-2) rely on RNA for replication. Knowing their 3D structure helps design antiviral drugs.\n* Disease Understanding: Mutations in RNA can lead to misfolding, causing diseases like ALS, muscular dystrophy, and cancer.\n* Synthetic Biology: Designing RNA-based therapeutics (e.g., mRNA vaccines like COVID-19 vaccines) requires understanding RNA folding.\n* CRISPR & Gene Editing: RNA guides (e.g., in CRISPR-Cas9) need correct folding to function properly.\n","metadata":{}},{"cell_type":"markdown","source":"## **Some Premelinary before Competition**","metadata":{}},{"cell_type":"markdown","source":"**What is RNA?**","metadata":{}},{"cell_type":"markdown","source":"![](https://media.geeksforgeeks.org/wp-content/uploads/20220711132035/image.png)","metadata":{}},{"cell_type":"markdown","source":"RNA (Ribonucleic Acid) is a single-stranded molecule made of nucleotides (A, U, C, G).\n\nIt plays crucial roles in protein synthesis (mRNA), gene regulation (miRNA, siRNA), and catalysis (ribozymes). Thus you could understand the Importance of this Competition in Real Life","metadata":{}},{"cell_type":"markdown","source":"**Why RNA Structure Matters**","metadata":{}},{"cell_type":"markdown","source":"Unlike DNA (double helix), RNA folds into complex 3D shapes that determine its function. RNA can switch structures (e.g., riboswitches) in response to cellular signals ( Don't worry as switching is not included in Competition ). \n\n\n**But This Competition about Predicting RNA is Important because Incorrect RNA folding can disrupt cellular processes, leading to diseases.** Thus, this Competition have significant Research and Human Benefit Competition","metadata":{}},{"cell_type":"markdown","source":"**And thus : If successful, winning models could revolutionize RNA-based drug design and unlock new therapies. This may gonna change Medicine Industry Forever**","metadata":{}},{"cell_type":"markdown","source":"The DataSet provided have 1 Folder and 8 files","metadata":{}},{"cell_type":"markdown","source":"# **B About the DataSet------------------>**","metadata":{}},{"cell_type":"markdown","source":"##    1. sample_submission.csv","metadata":{}},{"cell_type":"markdown","source":"This contains the DEMO file you had to submit the OUTPUT. It must have 18 Columns ! \n\n1st Column \"ID\" : The first contains the \"ID\" which is RNA ID from test_sequences.csv. This id will reflect the predicted value to actual ID  to be Tested from train_sequence.csv.\n\n2nd is Residue name \"resname\" : One of residue which link these 3 coordinates to the sequence\n\n\n3rd is \"resid\" : It's the reside id\n\n4,5,6 equals X,Y,Z coordinates......But wait they are from x1,x2,x3,y1,y2,y3....x5,y5,z5 in the Remaining Columns\n||  Basically you need to output 5 Possible Coordinates for DNA Sequence ","metadata":{}},{"cell_type":"markdown","source":"## 2. test_sequences.csv","metadata":{}},{"cell_type":"markdown","source":"Conatins 5 Columns for Testing... Here the \"sequence\" Column is actually for Predicting while the \"target_id\" is used for submiiting the recorded Predicted value to submission.csv value\nOther 3 columns wouldn't add any benefit in Training and is only there for Reference purpose","metadata":{}},{"cell_type":"markdown","source":"## **3. train_labels.csv and train_label.v2.csv**","metadata":{}},{"cell_type":"markdown","source":"These two DATA are Entirely same and is only diffrentiate for Training and Testing Model. The Data contains 6 Columns : 1 ID, 1 Residue Name, 1 Residue ID, and other 3 conatins the 3D Coordinates\n\nThe ID is reference for RNA Sequence in \"train_sequences.csv\" and \"train_sequenence.v2.csv\" as named \"target_id\" in that file. \n\nBasically the train_sequence contains the Sequence having ID, the id reference in this train_label file with 3D Coordinate which needed to be **TRAINED** ","metadata":{}},{"cell_type":"markdown","source":"Optional : You could Merge the 2 Train_label to one if you want. This may Increase Training Time but do helps in finding Correct ID of Sequence for Training","metadata":{}},{"cell_type":"code","source":"# Optional if want to Merge\n\nimport pandas as pd\n\n# Load the CSV file\ndata_label1 = pd.read_csv('/kaggle/input/stanford-rna-3d-folding/train_labels.csv')  \ndata_label2=pd.read_csv(\"/kaggle/input/stanford-rna-3d-folding/train_labels.v2.csv\")\n\ndata_label=pd.concat([data_label1,data_label2],ignore_index=True)\nprint(data_label.head())\nprint(\"Merge Successfull with Shape\")\ndata_label.describe()","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-05-08T07:24:26.425150Z","iopub.execute_input":"2025-05-08T07:24:26.425518Z","iopub.status.idle":"2025-05-08T07:24:31.702329Z","shell.execute_reply.started":"2025-05-08T07:24:26.425495Z","shell.execute_reply":"2025-05-08T07:24:31.701438Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"## **4. train_sequences.csv and train_sequences.v2.csv**","metadata":{}},{"cell_type":"markdown","source":"Conatins 5 Columns for Training... Here the \"sequence\" Column is actually for Training while the \"target_id\" is used for reference the Predicted 3D coordinate provided in train_label.csv\n\nOther 3 columns wouldn't add any benefit in Training and is only there for Reference purpose","metadata":{}},{"cell_type":"markdown","source":"## **5. validation_label.csv and validation_sequences.csv**","metadata":{}},{"cell_type":"markdown","source":"This is only for generalizes via validation and used to Predict structures for validation_sequences compare to validation_labels. This will help in Fine Tuning the Model. \nvalidation_sequence is same as train_sequence.\nvalidation_label have same as train_label except now it have 123 Columns. \n> 3 Coordinates x 40 each + 3 Columns of id, resname, resid = 123\n\nThese 120 Columns will help extra Tune the Model with approximate 40 real values.","metadata":{}},{"cell_type":"markdown","source":"## **6. MSA Folder**","metadata":{}},{"cell_type":"markdown","source":"It is critical inputs for predicting RNA 3D structures, similar to how they’re used in protein folding (e.g., AlphaFold). A FASTA-formatted file containing aligned sequences of homologous RNAs\n\n>>17RA_A (target sequence)\nAUG-CGU-AAU...\n>Homolog_1\nAUG-AGU-AAU...  # Notice the mutation (C→A) but conserved structure\n>Homolog_2\nGUG-CGU-ACU...  # Another variation\n\nGaps (-) indicate insertions/deletions, while aligned columns imply functional/structural conservation. ","metadata":{}},{"cell_type":"markdown","source":"**Note : This Folder is only for Tuning. Your model might miss long-range interactions, leading to poor 3D predictions ! But by using this you can leverage evolutionary data to outperform sequence-only approaches.","metadata":{}},{"cell_type":"markdown","source":"# **C. Approach to DataSet ---------------------->**","metadata":{}},{"cell_type":"markdown","source":"For WINNING this Competition, you need to first master Important Topics required to understand the Model Problem needed here, particular steps to be followed, and Fine Tuning Model using Ensemblem or Optimization.\n> We assume you known Fundamentals like Pandas, Supervised Learning, etc","metadata":{}},{"cell_type":"markdown","source":"## 1. Important Topics you should known","metadata":{}},{"cell_type":"markdown","source":"* Traditional methods (e.g., ViennaRNA, RosettaRNA)\n* Deep learning approaches (like AlphaFold for RNA)\n* Graph Neural Networks (GNNs) for modeling RNA as a 3D graph\n* Transformers for sequence-to-structure learning\n* Diffusion Models / VAEs for multi-conformation predictions\n\n* Extracting co-evolutionary signals (e.g., plmc, CCMpred)\n* Using MSA features in deep learning models\n* RMSD (Root Mean Square Deviation) for 3D structure accuracy\n\n* F1-score for base-pair prediction\n* Local Distance Difference Test (lDDT) for residue-level accuracy\n\n\n* Efficient cross-validation\n* Ensemble modeling\n* Hyperparameter tuning (Optuna, WandB)","metadata":{}},{"cell_type":"markdown","source":"With All this, You are good to go !","metadata":{}},{"cell_type":"markdown","source":"## 2. Steps to Solve ","metadata":{}},{"cell_type":"markdown","source":"###  **2.1. Install BioPython**","metadata":{}},{"cell_type":"markdown","source":"Biopython is a collection of Python libraries and tools designed for computational biology and bioinformatics. It provides modules for working with biological sequences, sequence annotations, and accessing various bioinformatics databases","metadata":{}},{"cell_type":"code","source":"!pip install biopython numpy pandas torch torch-geometric","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-05-08T07:24:31.703880Z","iopub.execute_input":"2025-05-08T07:24:31.704286Z","iopub.status.idle":"2025-05-08T07:26:07.449987Z","shell.execute_reply.started":"2025-05-08T07:24:31.704247Z","shell.execute_reply":"2025-05-08T07:26:07.448611Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"### **2.2. Load and Clean Data**","metadata":{}},{"cell_type":"markdown","source":"(A) Load FASTA Sequences from MSA Folders","metadata":{}},{"cell_type":"code","source":"from Bio import SeqIO\n\ndef load_fasta(file_path):\n    return {record.id: str(record.seq) for record in SeqIO.parse(file_path, \"fasta\")}\n\ntrain_seqs = load_fasta(\"/kaggle/input/stanford-rna-3d-folding/MSA/17RA_A.MSA.fasta\")\nprint(f\"First RNA: {list(train_seqs.items())[0]}\")","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-05-08T07:26:07.452063Z","iopub.execute_input":"2025-05-08T07:26:07.452545Z","iopub.status.idle":"2025-05-08T07:26:07.574852Z","shell.execute_reply.started":"2025-05-08T07:26:07.452509Z","shell.execute_reply":"2025-05-08T07:26:07.573940Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"(B) Load CSV Data ","metadata":{}},{"cell_type":"code","source":"import pandas as pd\n\n# Load sequence and structure data\ntrain_seqs = pd.read_csv(\"/kaggle/input/stanford-rna-3d-folding/train_sequences.csv\")\ntrain_labels = pd.read_csv(\"/kaggle/input/stanford-rna-3d-folding/train_labels.csv\")\n\nprint(\"Sequences:\\n\", train_seqs.head())\nprint(\"\\nLabels:\\n\", train_labels.head())","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-05-08T07:26:07.577479Z","iopub.execute_input":"2025-05-08T07:26:07.577720Z","iopub.status.idle":"2025-05-08T07:26:07.853056Z","shell.execute_reply.started":"2025-05-08T07:26:07.577703Z","shell.execute_reply":"2025-05-08T07:26:07.852071Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"## **2.3. Feature Engineering**","metadata":{}},{"cell_type":"markdown","source":"Convert RNA sequences to numerical features (one-hot encoding). Finally Machine learns through Numbers ","metadata":{}},{"cell_type":"code","source":"import numpy as np\n\ndef one_hot_encode(sequence):\n    mapping = {'A': [1,0,0,0], 'U': [0,1,0,0], \n               'C': [0,0,1,0], 'G': [0,0,0,1]}\n    return np.array([mapping.get(base, [0,0,0,0]) for base in sequence])\n\n# Example: One-hot encode the first sequence\nseq = train_seqs.iloc[0]['sequence']\nencoded_seq = one_hot_encode(seq)\nprint(f\"Encoded shape: {encoded_seq.shape}\")  ","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-05-08T07:26:07.854092Z","iopub.execute_input":"2025-05-08T07:26:07.854461Z","iopub.status.idle":"2025-05-08T07:26:07.861662Z","shell.execute_reply.started":"2025-05-08T07:26:07.854437Z","shell.execute_reply":"2025-05-08T07:26:07.860623Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"\nstructures = train_labels.groupby(['ID', 'resid'])[['x_1', 'y_1', 'z_1']].mean().reset_index()\nprint(structures.head())","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-05-08T07:26:07.862672Z","iopub.execute_input":"2025-05-08T07:26:07.862929Z","iopub.status.idle":"2025-05-08T07:26:08.069046Z","shell.execute_reply.started":"2025-05-08T07:26:07.862910Z","shell.execute_reply":"2025-05-08T07:26:08.068156Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"## **2.4. Training Model**","metadata":{}},{"cell_type":"markdown","source":"Now Go Ahead and Train the Model. Now we need a Real Data Scientist ! Good Luck ","metadata":{}},{"cell_type":"markdown","source":"## **2.5. Submit the CSV**","metadata":{}},{"cell_type":"code","source":"import pandas as pd\n\n\n#  submission = pd.DataFrame({\n#   'atom_id': range(len(pred_coords)),\n#    'resname': ['A'] * len(pred_coords),  # Placeholder\n#    'resid': range(1, len(pred_coords)+1),\n#    'x1': pred_coords[:, 0], 'y1': pred_coords[:, 1], 'z1': pred_coords[:, 2],\n              # ... Repeat for x2,y2,z2 to x5,y5,z5 \n#  })\n\n# submission.to_csv(\"submission.csv\", index=False)\nprint(\" Remove # above to Submit your Contest Files\")","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-05-08T07:26:08.070127Z","iopub.execute_input":"2025-05-08T07:26:08.070444Z","iopub.status.idle":"2025-05-08T07:26:08.076061Z","shell.execute_reply.started":"2025-05-08T07:26:08.070417Z","shell.execute_reply":"2025-05-08T07:26:08.074872Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"# D. **Visualizing Data -------------------------->** ","metadata":{}},{"cell_type":"markdown","source":"Visualizing data before machine learning is crucial because it helps identify patterns, trends, and anomalies in the data that might not be immediately obvious from a raw table. This visual understanding allows for better data cleaning, feature selection, and ultimately, the development of more effective and accurate machine learning models. ","metadata":{}},{"cell_type":"markdown","source":"## **1. Visualize Sequence Length Distribution**","metadata":{}},{"cell_type":"code","source":"import pandas as pd\nimport matplotlib.pyplot as plt\nimport seaborn as sns\n\ntrain_seqs = pd.read_csv(\"/kaggle/input/stanford-rna-3d-folding/test_sequences.csv\")\ntrain_seqs['seq_length'] = train_seqs['sequence'].apply(len)\n\n# Plot histogram\nplt.figure(figsize=(10, 5))\nsns.histplot(train_seqs['seq_length'], bins=30, kde=True)\nplt.title(\"Distribution of RNA Sequence Lengths\")\nplt.xlabel(\"Length\")\nplt.ylabel(\"Count\")\nplt.show()","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-05-08T07:26:08.077165Z","iopub.execute_input":"2025-05-08T07:26:08.077445Z","iopub.status.idle":"2025-05-08T07:26:08.383677Z","shell.execute_reply.started":"2025-05-08T07:26:08.077425Z","shell.execute_reply":"2025-05-08T07:26:08.382806Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"## **2. Nucleotide Frequency Bar Plot**","metadata":{}},{"cell_type":"code","source":"from collections import Counter\n\n# Count nucleotides in all sequences\nall_nucleotides = ''.join(train_seqs['sequence'])\nnucleotide_counts = Counter(all_nucleotides)\n\n# Plot\nplt.figure(figsize=(8, 4))\nsns.barplot(x=list(nucleotide_counts.keys()), y=list(nucleotide_counts.values()))\nplt.title(\"Nucleotide Frequency in Training Data\")\nplt.xlabel(\"Nucleotide\")\nplt.ylabel(\"Count\")\nplt.show()","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-05-08T07:26:08.384622Z","iopub.execute_input":"2025-05-08T07:26:08.384880Z","iopub.status.idle":"2025-05-08T07:26:08.551776Z","shell.execute_reply.started":"2025-05-08T07:26:08.384860Z","shell.execute_reply":"2025-05-08T07:26:08.550988Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"## **3. 3D Structure Visualization (Interactive)**","metadata":{}},{"cell_type":"code","source":"import plotly.express as px\n\ntrain_labels = pd.read_csv(\"/kaggle/input/stanford-rna-3d-folding/train_labels.csv\")\n# Extract first RNA's coordinates\nfirst_rna = train_labels[train_labels['ID'] == train_labels['ID'].iloc[0]]  # Change index as needed\n\n# 3D scatter plot\nfig = px.scatter_3d(first_rna, x='x_1', y='y_1', z='z_1', color='resid', \n                    title=\"3D RNA Structure (Plotly)\")\nfig.show()","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-05-08T07:26:08.554214Z","iopub.execute_input":"2025-05-08T07:26:08.554472Z","iopub.status.idle":"2025-05-08T07:26:08.805958Z","shell.execute_reply.started":"2025-05-08T07:26:08.554452Z","shell.execute_reply":"2025-05-08T07:26:08.805184Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"!pip install ViennaRNA","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-05-08T07:26:08.806848Z","iopub.execute_input":"2025-05-08T07:26:08.807105Z","iopub.status.idle":"2025-05-08T07:26:12.505944Z","shell.execute_reply.started":"2025-05-08T07:26:08.807087Z","shell.execute_reply":"2025-05-08T07:26:12.504944Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"import RNA  \n\n# Predict and plot secondary structure\nsequence = train_seqs['sequence'].iloc[0]  # First sequence\nstructure, _ = RNA.fold(sequence)\n\n# Draw structure\nprint(f\"Sequence: {sequence}\")\nprint(f\"Predicted Structure: {structure}\")\nRNA.svg_rna_plot(sequence, structure, \"secondary_structure.svg\")\n\n# Display in notebook\nfrom IPython.display import SVG\nSVG(\"secondary_structure.svg\")","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-05-08T07:26:12.507243Z","iopub.execute_input":"2025-05-08T07:26:12.507551Z","iopub.status.idle":"2025-05-08T07:26:12.528797Z","shell.execute_reply.started":"2025-05-08T07:26:12.507513Z","shell.execute_reply":"2025-05-08T07:26:12.527870Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"from scipy.spatial.distance import pdist, squareform\n\n# Calculate pairwise distances between residues\ncoords = first_rna[['x_1', 'y_1', 'z_1']].values\ndist_matrix = squareform(pdist(coords))\n\n# Plot heatmap\nplt.figure(figsize=(10, 8))\nsns.heatmap(dist_matrix, cmap=\"viridis\")\nplt.title(\"Residue-Residue Distance Matrix\")\nplt.xlabel(\"Residue ID\")\nplt.ylabel(\"Residue ID\")\nplt.show()","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-05-08T07:26:12.529807Z","iopub.execute_input":"2025-05-08T07:26:12.530483Z","iopub.status.idle":"2025-05-08T07:26:12.753454Z","shell.execute_reply.started":"2025-05-08T07:26:12.530445Z","shell.execute_reply":"2025-05-08T07:26:12.752518Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"import numpy as np\nimport matplotlib.pyplot as plt\nfrom sklearn.decomposition import PCA\n\n# 1. Generate sample data (10 residues, 3D)\ncoords = np.random.rand(10, 3)  # 10 residues, (x,y,z) coordinates\nresidue_ids = np.arange(10)     # Corresponding residue IDs (0-9)\n\n# 2. Verify data shape\nprint(\"Coordinates shape:\", coords.shape)  # Should be (10, 3)\nprint(\"Residue IDs shape:\", residue_ids.shape)  # Should be (10,)\n\n# 3. Perform PCA\npca = PCA(n_components=2)\ncoords_2d = pca.fit_transform(coords)\nprint(\"PCA result shape:\", coords_2d.shape)  # Should be (10, 2)\n\n# 4. Plot with consistent dimensions\nplt.figure(figsize=(8, 6))\nscatter = plt.scatter(coords_2d[:, 0], coords_2d[:, 1], \n                     c=residue_ids,        # Use matching residue IDs\n                     cmap=\"tab20\",\n                     s=100)               # Adjust point size\nplt.colorbar(scatter, label=\"Residue ID\")\nplt.title(\"PCA of RNA 3D Coordinates\")\nplt.xlabel(\"Principal Component 1\")\nplt.ylabel(\"Principal Component 2\")\nplt.grid(True)\nplt.show()","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-05-08T07:26:12.754723Z","iopub.execute_input":"2025-05-08T07:26:12.755092Z","iopub.status.idle":"2025-05-08T07:26:13.034345Z","shell.execute_reply.started":"2025-05-08T07:26:12.755062Z","shell.execute_reply":"2025-05-08T07:26:13.033309Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"# **E. Statistics----------------->**","metadata":{}},{"cell_type":"markdown","source":"## 1. Basic Sequence Statistics","metadata":{}},{"cell_type":"markdown","source":"GC-rich RNAs are more thermally stable (higher melting temperatures)","metadata":{}},{"cell_type":"code","source":"from collections import Counter\nimport numpy as np\n\nimport pandas as pd\n\ncounter =0\n\ndata_seq = pd.read_csv('/kaggle/input/stanford-rna-3d-folding/train_sequences.csv')  \ns = data_seq[\"sequence\"]\nfor seq in s:\n    print(seq)\n    counter+=1 \n    if(counter>8): # Only Print upto 8 Strand \n        break\n    nuc_counts = Counter(seq)\n    gc_content = (nuc_counts['G'] + nuc_counts['C']) / len(seq) * 100\n    print(f\"Nucleotide counts: {nuc_counts}\")\n    print(f\"GC content: {gc_content:.2f}%\")\n","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-05-08T07:26:13.035354Z","iopub.execute_input":"2025-05-08T07:26:13.035610Z","iopub.status.idle":"2025-05-08T07:26:13.071830Z","shell.execute_reply.started":"2025-05-08T07:26:13.035592Z","shell.execute_reply":"2025-05-08T07:26:13.071131Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"## 2. Pairwise Atomic Distances","metadata":{}},{"cell_type":"markdown","source":"Identifies compact vs. extended structure","metadata":{}},{"cell_type":"code","source":"from scipy.spatial.distance import pdist, squareform\n\n# coords = Nx3 array of atomic positions\ndist_matrix = squareform(pdist(coords))\nmean_distance = np.mean(dist_matrix)\nstd_distance = np.std(dist_matrix)\n\nprint(f\"Mean inter-atomic distance: {mean_distance:.2f} Å\")\nprint(f\"Standard deviation: {std_distance:.2f} Å\")","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-05-08T07:26:13.073310Z","iopub.execute_input":"2025-05-08T07:26:13.073944Z","iopub.status.idle":"2025-05-08T07:26:13.080407Z","shell.execute_reply.started":"2025-05-08T07:26:13.073920Z","shell.execute_reply":"2025-05-08T07:26:13.079421Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"## 3. Radius of Gyration (Compactness)","metadata":{}},{"cell_type":"code","source":"def radius_of_gyration(coords):\n    centroid = np.mean(coords, axis=0)\n    return np.sqrt(np.mean(np.sum((coords - centroid)**2, axis=1)))\n\nrgyr = radius_of_gyration(coords)\nprint(f\"Radius of gyration: {rgyr:.2f} Å\")","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-05-08T07:26:13.081323Z","iopub.execute_input":"2025-05-08T07:26:13.081665Z","iopub.status.idle":"2025-05-08T07:26:13.100041Z","shell.execute_reply.started":"2025-05-08T07:26:13.081646Z","shell.execute_reply":"2025-05-08T07:26:13.099253Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"## 4. Base Pairing Statistics","metadata":{}},{"cell_type":"markdown","source":"High pairing % suggests stable secondary structure","metadata":{}},{"cell_type":"code","source":"# Using ViennaRNA (install with !pip install ViennaRNA)\nimport RNA\n\ncounter =0\ndata = pd.read_csv(\"/kaggle/input/stanford-rna-3d-folding/train_sequences.csv\")\ntarget_col = data['sequence']\n\nfor s in target_col:\n    counter+=1\n    if(counter >8):  # Counter to Count Data \n        break\n    structure, _ = RNA.fold(s)\n    paired_bases = structure.count('(') + structure.count(')')\n    pairing_percentage = paired_bases / len(structure) * 100\n    \n    print(f\"Predicted structure: {structure}\")\n    print(f\"Paired bases: {pairing_percentage:.2f}%\")\n","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-05-08T07:26:13.101114Z","iopub.execute_input":"2025-05-08T07:26:13.101498Z","iopub.status.idle":"2025-05-08T07:26:13.148590Z","shell.execute_reply.started":"2025-05-08T07:26:13.101475Z","shell.execute_reply":"2025-05-08T07:26:13.147803Z"}},"outputs":[],"execution_count":null}]}