{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"name":"python","version":"3.10.12","mimetype":"text/x-python","codemirror_mode":{"name":"ipython","version":3},"pygments_lexer":"ipython3","nbconvert_exporter":"python","file_extension":".py"},"kaggle":{"accelerator":"none","dataSources":[{"sourceId":87793,"databundleVersionId":11553390,"sourceType":"competition"}],"dockerImageVersionId":30918,"isInternetEnabled":true,"language":"python","sourceType":"notebook","isGpuEnabled":false}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"# RNA Structure Analysis\n\nThis notebook focuses on analyzing RNA 3D structures using computational methods. RNA plays a critical role in various biological processes, and understanding its structure is essential for studying its function. \n\nThe workflow includes:\n- Loading and exploring RNA structural data.\n- Identifying and handling missing values.\n- Extracting key features such as residue positions and chain information.\n- Calculating structural statistics, including distances between residues and the radius of gyration.\n- Visualizing relationships between structural properties, such as sequence length and compactness.\n- Interactive 3D visualization of RNA structures.\n\nThis analysis provides insights into RNA structural properties and helps in understanding its biological significance.","metadata":{}},{"cell_type":"code","source":"# This Python 3 environment comes with many helpful analytics libraries installed\n# It is defined by the kaggle/python Docker image: https://github.com/kaggle/docker-python\n# For example, here's several helpful packages to load\n\nimport numpy as np # linear algebra\nimport pandas as pd # data processing, CSV file I/O (e.g. pd.read_csv)\n\n# Input data files are available in the read-only \"../input/\" directory\n# For example, running this (by clicking run or pressing Shift+Enter) will list all files under the input directory\n\nimport os\nfor dirname, _, filenames in os.walk('/kaggle/input'):\n    for filename in filenames:\n        print(os.path.join(dirname, filename))\n\n# You can write up to 20GB to the current directory (/kaggle/working/) that gets preserved as output when you create a version using \"Save & Run All\" \n# You can also write temporary files to /kaggle/temp/, but they won't be saved outside of the current session","metadata":{"_uuid":"8f2839f25d086af736a60e9eeb907d3b93b6e0e5","_cell_guid":"b1076dfc-b9ad-4769-8c92-a6c4dae69d19","trusted":true,"execution":{"iopub.status.busy":"2025-04-09T15:42:18.550345Z","iopub.execute_input":"2025-04-09T15:42:18.551398Z","iopub.status.idle":"2025-04-09T15:42:19.100406Z","shell.execute_reply.started":"2025-04-09T15:42:18.551360Z","shell.execute_reply":"2025-04-09T15:42:19.099306Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"import pandas as pd\nimport numpy as np\nfrom collections import defaultdict\nimport matplotlib.pyplot as plt\nimport seaborn as sns","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-04-09T15:42:19.101923Z","iopub.execute_input":"2025-04-09T15:42:19.102524Z","iopub.status.idle":"2025-04-09T15:42:19.107329Z","shell.execute_reply.started":"2025-04-09T15:42:19.102492Z","shell.execute_reply":"2025-04-09T15:42:19.106116Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"## Dataset Structure","metadata":{}},{"cell_type":"code","source":"# Load training labels\n\ntrain_labels=pd.read_csv(\"/kaggle/input/stanford-rna-3d-folding/train_labels.csv\")\n\ntrain_labels.shape\n","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-04-09T15:42:19.108365Z","iopub.execute_input":"2025-04-09T15:42:19.108679Z","iopub.status.idle":"2025-04-09T15:42:19.367030Z","shell.execute_reply.started":"2025-04-09T15:42:19.108654Z","shell.execute_reply":"2025-04-09T15:42:19.365894Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"train_labels.head() ","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-04-09T15:42:19.368821Z","iopub.execute_input":"2025-04-09T15:42:19.369149Z","iopub.status.idle":"2025-04-09T15:42:19.381336Z","shell.execute_reply.started":"2025-04-09T15:42:19.369121Z","shell.execute_reply":"2025-04-09T15:42:19.380281Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"The dataset contains RNA structural information with 137,095 entries and 6 columns:\n\n- **ID**: Unique identifier for each residue, combining PDB ID, chain, and residue number.\n- **resname**: Residue name (e.g., A, G, C, U).\n- **resid**: Residue number within the chain.\n- **x_1, y_1, z_1**: 3D coordinates of the residue's C1' atom.","metadata":{}},{"cell_type":"code","source":"train_labels.describe()","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-04-09T15:42:19.382573Z","iopub.execute_input":"2025-04-09T15:42:19.382894Z","iopub.status.idle":"2025-04-09T15:42:19.437486Z","shell.execute_reply.started":"2025-04-09T15:42:19.382859Z","shell.execute_reply":"2025-04-09T15:42:19.436437Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"\nThe statistical summary of the numerical columns in the `train_labels` DataFrame:\n\n1. **`resid`**:\n    - Represents the residue number within the chain.\n    - The `count` shows the total number of residues in the dataset.\n    - The `mean` indicates the average residue number.\n    - The `min` and `max` values represent the range of residue numbers.\n\n2. **`x_1`, `y_1`, `z_1`**:\n    - These columns represent the 3D coordinates of the residue's C1' atom.\n    - The `count` confirms the number of non-missing values for each coordinate.\n    - The `mean` provides the average position along each axis.\n    - The `std` (standard deviation) measures the spread of the coordinates.\n    - The `min` and `max` values define the range of positions along each axis.\n","metadata":{}},{"cell_type":"markdown","source":"## Define Two New Columns (`pdb_id` and `chain_id`)","metadata":{}},{"cell_type":"code","source":"# Define pdb_id and chain_id\n# Split the 'ID' column into 'pdb_id' and 'chain_id'\ntrain_labels['pdb_id'] = train_labels['ID'].apply(lambda x: x.split('_')[0])\n\ntrain_labels['chain'] = train_labels['ID'].apply(lambda x: x.split('_')[1])","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-04-09T15:42:19.438599Z","iopub.execute_input":"2025-04-09T15:42:19.439739Z","iopub.status.idle":"2025-04-09T15:42:19.549957Z","shell.execute_reply.started":"2025-04-09T15:42:19.439697Z","shell.execute_reply":"2025-04-09T15:42:19.548859Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"train_labels.head()","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-04-09T15:42:19.551117Z","iopub.execute_input":"2025-04-09T15:42:19.551427Z","iopub.status.idle":"2025-04-09T15:42:19.564311Z","shell.execute_reply.started":"2025-04-09T15:42:19.551401Z","shell.execute_reply":"2025-04-09T15:42:19.563303Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"## Handeling Missing Values\n\nHandling missing values is a critical step in data preprocessing. Common strategies to handle missing values include imputation, deletion, or using algorithms that can handle missing data directly.\n\nFirst, check the percentage of missing values in the dataset.","metadata":{}},{"cell_type":"code","source":"# Count missing values in each column\nmissing_per_column = train_labels.isna().sum()\n\nprint(missing_per_column)","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-04-09T15:42:19.565356Z","iopub.execute_input":"2025-04-09T15:42:19.565675Z","iopub.status.idle":"2025-04-09T15:42:19.613204Z","shell.execute_reply.started":"2025-04-09T15:42:19.565641Z","shell.execute_reply":"2025-04-09T15:42:19.612155Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"perc_missing_per_column = (missing_per_column/len(train_labels))* 100\nprint(f\"Percentage of the data is missing in each column:\\n{perc_missing_per_column.round(2)}\")","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-04-09T15:42:19.614306Z","iopub.execute_input":"2025-04-09T15:42:19.615061Z","iopub.status.idle":"2025-04-09T15:42:19.622560Z","shell.execute_reply.started":"2025-04-09T15:42:19.615019Z","shell.execute_reply":"2025-04-09T15:42:19.621481Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"The missing values are confined to the 3D coordinate columns (`x_1`, `y_1`, `z_1`), affecting **4.48%** of the dataset.","metadata":{}},{"cell_type":"code","source":"missing_pdb_ids = train_labels[train_labels[['x_1', 'y_1', 'z_1']].isna().any(axis=1)]['pdb_id'].unique()\nprint(f\"Missing pdb_ids: {missing_pdb_ids}\")\nprint(f\"Number of missing pdb_ids: {len(missing_pdb_ids)}\")","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-04-09T15:42:19.625850Z","iopub.execute_input":"2025-04-09T15:42:19.626205Z","iopub.status.idle":"2025-04-09T15:42:19.649193Z","shell.execute_reply.started":"2025-04-09T15:42:19.626178Z","shell.execute_reply":"2025-04-09T15:42:19.647975Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"**193** RNA structure in the total dataset have missing values.","metadata":{}},{"cell_type":"markdown","source":"### Visualize the number of missing values per residues","metadata":{}},{"cell_type":"code","source":"import pprint\n\nmissing_resids = train_labels[train_labels[['x_1', 'y_1', 'z_1']].isna().any(axis=1)][['pdb_id', 'resid']]\npprint.pprint(missing_resids)","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-04-09T15:42:19.650325Z","iopub.execute_input":"2025-04-09T15:42:19.651134Z","iopub.status.idle":"2025-04-09T15:42:19.666005Z","shell.execute_reply.started":"2025-04-09T15:42:19.651094Z","shell.execute_reply":"2025-04-09T15:42:19.664865Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"missing_dict = missing_resids.groupby('pdb_id')['resid'].apply(list).to_dict()\nfor pdb_id, resids in missing_dict.items():\n    pprint.pprint({pdb_id: resids})\n    break ","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-04-09T15:42:19.667150Z","iopub.execute_input":"2025-04-09T15:42:19.667552Z","iopub.status.idle":"2025-04-09T15:42:19.681440Z","shell.execute_reply.started":"2025-04-09T15:42:19.667514Z","shell.execute_reply":"2025-04-09T15:42:19.680541Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"# Group by resid and count the number of missing values\nmissing_resid_counts = missing_resids.groupby('resid').size().reset_index(name='count')\n\n# Filter out residues with more than 2 missing values to focus on significant ones\nfilter_missing_resid_counts = missing_resid_counts[missing_resid_counts['count']>2]\n\n\n# Plot the accumulation of missing values\n\nplt.figure(figsize=(12, 6))\n\nax = sns.barplot(\n    data=filter_missing_resid_counts,\n    x='resid',\n    y='count',\n    color='blue'\n)\n\n# Add title and labels\nplt.title('Accumulation of Missing Values by Residue', fontsize=16)\nplt.xlabel('Residues')\nplt.ylabel('Number of Missing Values')\n\n# Rotate x-axis labels for better readability\nplt.xticks(\n    ticks=range(0,len(filter_missing_resid_counts), 50),\n    rotation=45\n)\n\n# Adjust layout to prevent cutting off labels\nplt.tight_layout()\n\nplt.show()","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-04-09T15:42:19.682539Z","iopub.execute_input":"2025-04-09T15:42:19.682847Z","iopub.status.idle":"2025-04-09T15:42:21.105231Z","shell.execute_reply.started":"2025-04-09T15:42:19.682822Z","shell.execute_reply":"2025-04-09T15:42:21.104161Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"### Drop Missing Values\n\nIn this case, dropping rows with missing values is a suitable strategy because:\n- The missing values are confined to the 3D coordinate columns (`x_1`, `y_1`, `z_1`), which are essential for structural analysis.\n- Imputation may introduce inaccuracies, as the 3D coordinates are highly specific and cannot be reliably estimated.\n- The percentage of missing data (4.48%) is relatively small, so dropping these rows does not significantly impact the dataset's size or diversity.","metadata":{}},{"cell_type":"code","source":"# Drop missing values\ntrain_labels = train_labels.dropna(subset=['x_1', 'y_1', 'z_1'])\n\n# Check if there are any missing values left\nmissing_values_after = train_labels.isna().sum()\nprint(f\"Missing values after dropping: {missing_values_after}\")","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-04-09T15:42:21.106310Z","iopub.execute_input":"2025-04-09T15:42:21.106643Z","iopub.status.idle":"2025-04-09T15:42:21.163304Z","shell.execute_reply.started":"2025-04-09T15:42:21.106603Z","shell.execute_reply":"2025-04-09T15:42:21.162053Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"train_labels.shape","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-04-09T15:42:21.164548Z","iopub.execute_input":"2025-04-09T15:42:21.164880Z","iopub.status.idle":"2025-04-09T15:42:21.171342Z","shell.execute_reply.started":"2025-04-09T15:42:21.164852Z","shell.execute_reply":"2025-04-09T15:42:21.170436Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"- These missing values have been addressed by dropping the affected rows, leaving a clean dataset for further analysis.","metadata":{}},{"cell_type":"markdown","source":"## Group by `target_id` and `chain_id`","metadata":{}},{"cell_type":"code","source":"# Group by target_id and chain_id\n\nstructures_by_target = defaultdict(list)\nfor _, row in train_labels.iterrows():\n    full_id = row[\"ID\"]\n    target_chain_id = \"_\".join(full_id.split(\"_\")[0:2])\n    structures_by_target[target_chain_id].append(row)\n\nprint(f\"Number of unique structures: {len(structures_by_target)}\")","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-04-09T15:42:21.172246Z","iopub.execute_input":"2025-04-09T15:42:21.172500Z","iopub.status.idle":"2025-04-09T15:42:31.179459Z","shell.execute_reply.started":"2025-04-09T15:42:21.172477Z","shell.execute_reply":"2025-04-09T15:42:31.178505Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"# Display the first target ID and its associated data\n\ntarget_ids = list(structures_by_target.keys())\n\nfirst_target_id = target_ids[0]\nfirst_structures = structures_by_target[first_target_id]\n\nprint(f\"First target ID: {first_target_id}\")\nprint(f\"Number of structures for this target: {len(first_structures)}\")\nprint(\"First structure data:\")\nprint(first_structures[0])","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-04-09T15:42:31.180493Z","iopub.execute_input":"2025-04-09T15:42:31.180844Z","iopub.status.idle":"2025-04-09T15:42:31.188473Z","shell.execute_reply.started":"2025-04-09T15:42:31.180817Z","shell.execute_reply":"2025-04-09T15:42:31.187285Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"## Structural Statistics Calculation\n\nThis section of the code calculates key structural statistics for RNA 3D structures, including distances between consecutive residues and the radius of gyration. These metrics provide insights into the spatial arrangement and compactness of RNA structures.\n\n### Euclidean Distance\n- Measures the straight-line distance between two points in 3D space.\n- Provides insights into the spatial separation of consecutive residues.\n- The Euclidean distance between consecutive residues' C1' atoms is calculated using the Pythagorean formula:\n\n    \n    $$d = \\sqrt{(x_2 - x_1)^2 + (y_2 - y_1)^2 + (z_2 - z_1)^2}$$\n    \n\nwhere:\n- ${(\\mathbf{x}_1, \\mathbf{y}_1, \\mathbf{z}_1)}$ and ${(\\mathbf{x}_2, \\mathbf{y}_2, \\mathbf{z}_2)}$ are the 3D coordinates of two consecutive residues.\n\n### Radius of Gyration\n- Quantifies the spread of residues around the center of mass.\n- A smaller ${\\mathbf(R_g)}$ indicates a more compact structure, while a larger ${\\mathbf(R_g)}$ suggests a more extended structure.\n- The radius of gyration is calculated using the formula:\n\n\n$$R_g = \\sqrt{\\frac{1}{N} \\sum_{i=1}^{N} \\| \\mathbf{r}_i - \\mathbf{r}_{\\text{center}} \\|^2}$$\n\n\nwhere:\n- $N$ is the number of residues.\n- ${(\\mathbf{r}_i)}$ is the position vector of the $i$-th residue.\n- $\\mathbf{r}_{\\text{center}}$ is the center of mass of the structure.","metadata":{}},{"cell_type":"code","source":"structural_stats = {}\n\nfor target_id, residues in structures_by_target.items():\n    # Sort residues by resid\n    residues.sort(key=lambda x: int(x[\"resid\"]))\n\n    # Calculate the distance between consecutive C1' atoms\n    distances = []\n    for i in range(1, len(residues)):\n        prev = residues[i-1]\n        curr = residues[i]\n\n        dx = float(curr[\"x_1\"]) - float(prev[\"x_1\"])\n        dy = float(curr[\"y_1\"]) - float(prev[\"y_1\"])\n        dz = float(curr[\"z_1\"]) - float(prev[\"z_1\"])\n\n        # Calculates the Euclidean distance between consecutive C1' atoms using the Pythagorean formula\n        distance = np.sqrt(dx**2 + dy**2 + dz**2)\n        distances.append(distance)\n\n    # Calculate statistics \n    if distances:\n        avg_distance = np.mean(distances)\n        min_distance = np.min(distances)\n        max_distance = np.max(distances)\n    else:\n        avg_distance = min_distance = max_distance = 0\n\n    # Calculate radius of gyration (rough measure of compactness)\n    coords = np.array([[float(r['x_1']), float(r['y_1']), float(r['z_1'])] for r in residues])\n    center = np.mean(coords, axis=0)\n\n    # Calculate radius of gyration\n    rg = np.sqrt(np.mean(np.sum((coords - center) ** 2, axis=1)))\n\n    structural_stats[target_id] = {\n        'num_residues' : len(residues),\n        'avg_consecutive_distance' : avg_distance,\n        'min_consecutive_distance' : min_distance,\n        'max_consecutive_distance' : max_distance,\n        'radius_of_gyration' : rg\n    }\n","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-04-09T15:42:31.189546Z","iopub.execute_input":"2025-04-09T15:42:31.189905Z","iopub.status.idle":"2025-04-09T15:42:34.981375Z","shell.execute_reply.started":"2025-04-09T15:42:31.189872Z","shell.execute_reply":"2025-04-09T15:42:34.980475Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"1. **Distance Statistics**:\n    - If distances are available, the following statistics are computed:\n      - **Average Distance**: Mean of all consecutive distances.\n      - **Minimum Distance**: Smallest consecutive distance.\n      - **Maximum Distance**: Largest consecutive distance.\n    - If no distances are available, these values are set to `0`.\n\n2. **Radius of Gyration**:\n    - The radius of gyration is a measure of structural compactness.\n    - A smaller ${\\mathbf(R_g)}$ indicates a more compact structure, while a larger ${\\mathbf(R_g)}$ suggests a more extended structure.\n\n\n3. **Storing Results**:\n    - The computed statistics for each `target_id` are stored in the `structural_stats` dictionary with the following keys:\n      - `num_residues`: Total number of residues.\n      - `avg_consecutive_distance`: Average distance between consecutive residues.\n      - `min_consecutive_distance`: Minimum distance between consecutive residues.\n      - `max_consecutive_distance`: Maximum distance between consecutive residues.\n      - `radius_of_gyration`: Compactness measure of the structure.\n\n**This analysis provides a comprehensive understanding of RNA structural properties, aiding in the study of their biological functions.**","metadata":{}},{"cell_type":"code","source":"# Print statistics for a few examples\nprint('\\nStructural statistics for sample targets: ')\nfor target_id, stats in list(structural_stats.items())[:5]:\n    print(f\"\\nTarget ID: {target_id}\")\n    print(f\"Number of residues: {stats['num_residues']}\")\n    print(f\"Average consecutive C1' distance: {stats['avg_consecutive_distance']:.2f} Å\")\n    print(f\"Min/Max consecutive distances: {stats['min_consecutive_distance']:.2f}/{stats['max_consecutive_distance']:.2f} Å\")\n    print(f\"Radius of gyration: {stats['radius_of_gyration']:.2f} Å\")","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-04-09T15:42:34.982339Z","iopub.execute_input":"2025-04-09T15:42:34.982637Z","iopub.status.idle":"2025-04-09T15:42:34.989121Z","shell.execute_reply.started":"2025-04-09T15:42:34.982611Z","shell.execute_reply":"2025-04-09T15:42:34.988134Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"# Calculate overall statistics\nall_radii = [s['radius_of_gyration'] for s in structural_stats.values()]\navg_radius = np.mean(all_radii)\nmin_radius = np.min(all_radii)\nmax_radius = np.max(all_radii)\n\nprint(\"\\nOverall structural statistics: \")\nprint(f\"Radius of gyration (compactness measure): \")\nprint(f\"- Average: {avg_radius:.2f} Å\")\nprint(f\"- Minimum: {min_radius:.2f} Å\")\nprint(f\"- Maximum: {max_radius:.2f} Å\")","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-04-09T15:42:34.990230Z","iopub.execute_input":"2025-04-09T15:42:34.990495Z","iopub.status.idle":"2025-04-09T15:42:35.012143Z","shell.execute_reply.started":"2025-04-09T15:42:34.990472Z","shell.execute_reply":"2025-04-09T15:42:35.011206Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"### Visualize radius of gyration vs. sequence length","metadata":{}},{"cell_type":"code","source":"# Prepare data for the scatter plot\nsequence_lengths = [stats['num_residues'] for stats in structural_stats.values()]\nradii_of_gyration = [stats['radius_of_gyration'] for stats in structural_stats.values()]\n\n# Create the scatter plot\nplt.figure(figsize=(10, 6))\n#plt.style.use('dark_background')\n\nsns.scatterplot(\n    x=sequence_lengths,\n    y=radii_of_gyration,\n    alpha=0.6, \n    s=80,      \n    color='#1f77b4'  \n)\n\n# Add title and labels\nplt.title('Radius of Gyration vs. Sequence Length', fontsize=16)\nplt.xlabel('Sequence Length (number of residues)')\nplt.ylabel('Radius of Gyration (Å)')\n\n# Improve layout\nplt.tight_layout()\n\nplt.show()","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-04-09T15:42:35.013051Z","iopub.execute_input":"2025-04-09T15:42:35.013328Z","iopub.status.idle":"2025-04-09T15:42:35.265393Z","shell.execute_reply.started":"2025-04-09T15:42:35.013306Z","shell.execute_reply":"2025-04-09T15:42:35.264452Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"**To fit a trendline to the \"Radius of Gyration vs. Sequence Length\" plot, we can use a power-law model of the form:**\n\n$${R_g} = a \\cdot L^b$$\n\nWhere:\n - ${R_g}$ is the radius of gyration,\n - 𝐿 is the sequence length (number of residues),\n - 𝑎 and 𝑏  are fitted constants.\n","metadata":{}},{"cell_type":"code","source":"import plotly.express as px\nimport plotly.graph_objects as go\nfrom scipy.optimize import curve_fit\n\n# Define power law function\ndef power_law(x, a, b):\n    return a * x ** b\n# Fit the power law to the data\nparams, _ = curve_fit(power_law, sequence_lengths, radii_of_gyration, maxfev=10000)\na, b = params\n\n# Generate trendline\nx_fit = np.linspace(min(sequence_lengths), max(sequence_lengths), 500)\ny_fit = power_law(x_fit, a, b)\n\n# Create the scatter plot with trendline\nplt.figure(figsize=(10, 6))\n#plt.style.use('dark_background')\n\n# Create the scatter plot with Seaborn\nsns.scatterplot(\n    x=sequence_lengths,\n    y=radii_of_gyration,\n    alpha=0.6,  \n    s=80,       \n    color='#1f77b4', \n    label='RNA Structures'\n)\n\n# Add the power law trendline\nplt.plot(x_fit, y_fit, color='red', linewidth=2, \n         label=f'Power-Law Fit:\\n Rg = {a:.2f} * L^{b:.2f}')\n\nplt.title('Radius of Gyration vs. Sequence Length with Power Law Fit', fontsize=16)\nplt.xlabel('Sequence Length (number of residues)')\nplt.ylabel('Radius of Gyration (Å)')\n\nplt.legend()\n\nplt.tight_layout()\n\nplt.show()","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-04-09T15:42:35.266614Z","iopub.execute_input":"2025-04-09T15:42:35.267498Z","iopub.status.idle":"2025-04-09T15:42:35.560212Z","shell.execute_reply.started":"2025-04-09T15:42:35.267460Z","shell.execute_reply":"2025-04-09T15:42:35.559235Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"📊 **Interpretation of the Radius of Gyration vs. Sequence Length Plot (RNA Structures):**\n\n- Positive Correlation with Sublinear Growth\n\n    - There is a clear positive trend: as RNA sequence length increases, the radius of gyration (Rg) also increases.\n\n    - However, the power-law exponent of ~0.33 indicates sublinear growth, suggesting that larger RNAs fold more compactly rather than extending proportionally with length.\n\n    - This implies RNA structures in your dataset tend to become increasingly globular as they grow in size.\n\n- Clustered Region of Small-to-Medium RNAs\n\n    - Most data points are densely clustered under 500 nucleotides, with Rg values between ~10 and ~60 Å.\n\n    - This suggests that the majority of RNAs in your dataset are relatively short, likely including transfer RNAs, riboswitches, and small non-coding RNAs, all known for their compact tertiary structures.\n\n- Extended Range for Larger RNAs\n\n    - A few RNAs extend to ~4300 nucleotides, with Rg values approaching 100 Å.\n\n    - These likely correspond to long functional RNAs, such as ribosomal RNAs, long non-coding RNAs, or viral RNA genomes, which still follow the overall compact scaling trend.\n\n- Structural Diversity at Fixed Lengths\n\n    - Even at similar sequence lengths, Rg values can vary widely, showing heterogeneity in folding.\n\n    - This reflects the fact that some RNAs are tightly folded, while others are more branched or extended, depending on secondary/tertiary interactions, metal ion binding, and functional constraints.\n\n- Compactness of RNA Folding\n\n    - The exponent ~0.33 is lower than expected for random coils or partially unfolded RNAs, which typically show exponents of ~0.5–0.6.\n\n    - This implies that RNAs in this dataset are structurally compact, stabilized by strong intramolecular base pairing, non-canonical motifs, and higher-order architecture.","metadata":{}},{"cell_type":"markdown","source":"### Visualize distribution of consecutive C1' distances\n","metadata":{}},{"cell_type":"code","source":"\n# Prepare data for the histogram\nall_distances = [\n    stats['avg_consecutive_distance']\n    for stats in structural_stats.values()\n    if stats['num_residues'] > 1  # Only include structures with at least 2 residues\n]\n\nplt.figure(figsize=(10, 6))\n#plt.style.use('dark_background')\n\n# Create the histogram\nsns.histplot(\n    all_distances,\n    bins=50,\n    color='cornflowerblue',\n    kde=False\n)\n\nplt.title(\"Distribution of Average Consecutive C1' Distances\", fontsize=16)\nplt.xlabel(\"Distance (Å)\")\nplt.ylabel(\"Count\")\n\nplt.tight_layout()\n\nplt.show()","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-04-09T15:42:35.561654Z","iopub.execute_input":"2025-04-09T15:42:35.562061Z","iopub.status.idle":"2025-04-09T15:42:35.853418Z","shell.execute_reply.started":"2025-04-09T15:42:35.562025Z","shell.execute_reply":"2025-04-09T15:42:35.852455Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"📊 **Interpretation of the Plot:**\n- Distribution Characteristics\n\n    - Bell-Shaped Distribution: The data follows an approximately normal distribution centered around 6.0Å.\n\n    - Peak Range: The highest frequency occurs between 5.9-6.1Å, with approximately 90-95 structures.\n\n    - Range Boundaries: The main distribution spans from about 5.2Å to 6.7Å, with rare outliers extending to ~7.5Å and beyond.\n\n    - Sparse Outliers: Very few structures show distances beyond 7.5Å, with an isolated case at approximately 9Å.\n\n\nThis detailed analysis of C1' distances reveals the remarkable consistency of RNA backbone geometry. The predominant peak at approximately 6.0Å represents the standard spacing between consecutive C1' atoms in canonical RNA structures, with minor variations reflecting the diverse structural contexts found across the RNA structural landscape.","metadata":{}},{"cell_type":"markdown","source":"## Visualize RNA 3D Structure","metadata":{}},{"cell_type":"code","source":"! pip install py3Dmol\n! pip install Bio","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-04-09T15:42:35.854671Z","iopub.execute_input":"2025-04-09T15:42:35.855603Z","iopub.status.idle":"2025-04-09T15:42:43.857912Z","shell.execute_reply.started":"2025-04-09T15:42:35.855561Z","shell.execute_reply":"2025-04-09T15:42:43.856860Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"import py3Dmol\nfrom Bio import PDB\n\ndef fetch_and_visualize_rna_py3dmol(pdb_id):\n    \"\"\"\n    Fetches an RNA 3D structure from the RCSB PDB and visualizes it using py3Dmol.\n    \n    Parameters:\n    pdb_id (str): The PDB ID of the RNA structure to fetch (e.g, '1JWC').\n    \n    returns:\n    py3Dmol.view: An interactive 3D viewer for the RNA structure.\n    \"\"\"\n    pdb_parser = PDB.PDBList()\n    filename = pdb_parser.retrieve_pdb_file(pdb_id, file_format='pdb', pdir='.', overwrite=True)\n    \n    # Read the PDB file\n    with open(filename, 'r') as f:\n        pdb_data = f.read()\n    \n    # Create a py3Dmol view and add the structure\n    view = py3Dmol.view(width=800, height=600)\n    view.addModel(pdb_data, 'pdb')\n    \n    # Style the view\n    view.setStyle({'cartoon': {'color': 'spectrum'}})\n    view.addStyle({'hetflag': True}, {'stick': {}})\n    \n    # Configure the view\n    view.zoomTo()\n    view.setBackgroundColor('white')\n    \n    return view","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-04-09T15:42:43.859446Z","iopub.execute_input":"2025-04-09T15:42:43.860340Z","iopub.status.idle":"2025-04-09T15:42:43.867735Z","shell.execute_reply.started":"2025-04-09T15:42:43.860295Z","shell.execute_reply":"2025-04-09T15:42:43.866662Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"view = fetch_and_visualize_rna_py3dmol('1JWC')\nview.show()","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-04-09T15:42:43.868975Z","iopub.execute_input":"2025-04-09T15:42:43.869599Z","iopub.status.idle":"2025-04-09T15:42:44.311037Z","shell.execute_reply.started":"2025-04-09T15:42:43.869562Z","shell.execute_reply":"2025-04-09T15:42:44.309823Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"## Update The Datasets","metadata":{}},{"cell_type":"code","source":"train_labels.to_csv(\"/kaggle/working/train_label_pdb.csv\", index=False)","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-04-09T15:42:44.312362Z","iopub.execute_input":"2025-04-09T15:42:44.312828Z","iopub.status.idle":"2025-04-09T15:42:45.125865Z","shell.execute_reply.started":"2025-04-09T15:42:44.312762Z","shell.execute_reply":"2025-04-09T15:42:45.124901Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"train_sequences = pd.read_csv(\"/kaggle/input/stanford-rna-3d-folding/train_sequences.csv\")","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-04-09T15:42:45.131478Z","iopub.execute_input":"2025-04-09T15:42:45.131833Z","iopub.status.idle":"2025-04-09T15:42:45.170232Z","shell.execute_reply.started":"2025-04-09T15:42:45.131805Z","shell.execute_reply":"2025-04-09T15:42:45.169191Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"# Extract the relevant data into separate mapping dictionaries\navg_dist_map = {k: v['avg_consecutive_distance'] for k, v in structural_stats.items()}\nrg_map = {k: v['radius_of_gyration'] for k, v in structural_stats.items()}\n\n# Add new columns to train_sequences\ntrain_sequences['avg_consecutive_distance'] = train_sequences['target_id'].map(avg_dist_map)\ntrain_sequences['radius_of_gyration'] = train_sequences['target_id'].map(rg_map)","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-04-09T15:42:45.171243Z","iopub.execute_input":"2025-04-09T15:42:45.171530Z","iopub.status.idle":"2025-04-09T15:42:45.187371Z","shell.execute_reply.started":"2025-04-09T15:42:45.171505Z","shell.execute_reply":"2025-04-09T15:42:45.186183Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"train_sequences.head()","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-04-09T15:42:45.188483Z","iopub.execute_input":"2025-04-09T15:42:45.188859Z","iopub.status.idle":"2025-04-09T15:42:45.214422Z","shell.execute_reply.started":"2025-04-09T15:42:45.188821Z","shell.execute_reply":"2025-04-09T15:42:45.213330Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"train_sequences.to_csv(\"/kaggle/working/train_sequences_dist_rg.csv\", index=False)","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-04-09T15:42:45.215674Z","iopub.execute_input":"2025-04-09T15:42:45.216749Z","iopub.status.idle":"2025-04-09T15:42:45.302475Z","shell.execute_reply.started":"2025-04-09T15:42:45.216719Z","shell.execute_reply":"2025-04-09T15:42:45.301399Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"","metadata":{"trusted":true},"outputs":[],"execution_count":null}]}