{"metadata":{"kernelspec":{"name":"python3","display_name":"Python 3","language":"python"},"language_info":{"name":"python","version":"3.10.12","mimetype":"text/x-python","codemirror_mode":{"name":"ipython","version":3},"pygments_lexer":"ipython3","nbconvert_exporter":"python","file_extension":".py"},"kaggle":{"accelerator":"none","dataSources":[{"sourceId":87793,"databundleVersionId":11553390,"sourceType":"competition"}],"dockerImageVersionId":30918,"isInternetEnabled":true,"language":"python","sourceType":"notebook","isGpuEnabled":false}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"# Stanford RNA 3D Folding Exploratory Data Analysis","metadata":{}},{"cell_type":"markdown","source":"## Introduction\n\nRNA plays a crucial role in various biological processes, including gene expression regulation, enzymatic activity, and viral replication. Accurate 3D structural prediction of RNA sequences is essential for advancing research in molecular biology, drug discovery, and synthetic biology.\n\nThis exploratory data analysis (EDA) aims to provide a comprehensive overview of the dataset used in this project. The dataset consists of RNA sequences with their corresponding experimentally determined 3D structures, offering a rich resource for computational modeling and predictive analysis. By examining key aspects such as sequence length distribution, nucleotide composition, and structural diversity, we gain insights into the characteristics of the data and potential challenges for downstream machine learning applications.\n\nThe objectives of this data exploration are as follows:\n\n1. Dataset Overview – Understanding the size, structure, and temporal distribution of the dataset.\n\n2. Sequence Length Analysis – Identifying trends and variations in RNA sequence lengths.\n\n3. Nucleotide Composition – Evaluating the frequency and distribution of nucleotide bases (A, U, G, C) to assess sequence characteristics.\n\n4. Structural Labels and 3D Coordinates – Investigating the presence of 3D spatial coordinates and their implications for predictive modeling.\n\nBy analyzing these aspects, we aim to establish a foundational understanding of the dataset, highlighting its strengths, potential biases, and areas for further exploration. This analysis will also provide essential context for designing machine learning models capable of accurately predicting RNA 3D structures.","metadata":{}},{"cell_type":"markdown","source":"## Import Libraries","metadata":{}},{"cell_type":"code","source":"import pandas as pd\nimport numpy as np\nimport matplotlib.pyplot as plt\nimport plotly.express as px\nimport seaborn as sns\nfrom collections import Counter\nfrom datetime import datetime","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-04-03T12:33:44.395573Z","iopub.execute_input":"2025-04-03T12:33:44.395895Z","iopub.status.idle":"2025-04-03T12:33:46.187069Z","shell.execute_reply.started":"2025-04-03T12:33:44.395864Z","shell.execute_reply":"2025-04-03T12:33:46.185851Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"## Load training sequences and labels","metadata":{}},{"cell_type":"code","source":"train_sequences=pd.read_csv(\"/kaggle/input/stanford-rna-3d-folding/train_sequences.csv\")\ntrain_labels=pd.read_csv(\"/kaggle/input/stanford-rna-3d-folding/train_labels.csv\")","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-04-03T12:33:46.188021Z","iopub.execute_input":"2025-04-03T12:33:46.188494Z","iopub.status.idle":"2025-04-03T12:33:46.436107Z","shell.execute_reply.started":"2025-04-03T12:33:46.188465Z","shell.execute_reply":"2025-04-03T12:33:46.434862Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"## Basic Insight","metadata":{}},{"cell_type":"code","source":"print(f\"Number of training sequences: {len(train_sequences)}\")","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-04-03T12:33:46.437367Z","iopub.execute_input":"2025-04-03T12:33:46.437811Z","iopub.status.idle":"2025-04-03T12:33:46.444145Z","shell.execute_reply.started":"2025-04-03T12:33:46.437768Z","shell.execute_reply":"2025-04-03T12:33:46.442813Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"train_sequences.head()","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-04-03T12:33:46.445354Z","iopub.execute_input":"2025-04-03T12:33:46.445772Z","iopub.status.idle":"2025-04-03T12:33:46.477077Z","shell.execute_reply.started":"2025-04-03T12:33:46.445734Z","shell.execute_reply":"2025-04-03T12:33:46.476032Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"train_labels.head()","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-04-03T12:33:46.480187Z","iopub.execute_input":"2025-04-03T12:33:46.480560Z","iopub.status.idle":"2025-04-03T12:33:46.499063Z","shell.execute_reply.started":"2025-04-03T12:33:46.480500Z","shell.execute_reply":"2025-04-03T12:33:46.497976Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"## Analyze Sequence Lengths\n","metadata":{}},{"cell_type":"code","source":"train_sequences['length'] = train_sequences['sequence'].apply(len)\nmin_length = train_sequences['length'].min()\nmax_length = train_sequences['length'].max()\navg_length = train_sequences['length'].mean()\n","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-04-03T12:33:46.500729Z","iopub.execute_input":"2025-04-03T12:33:46.501139Z","iopub.status.idle":"2025-04-03T12:33:46.531740Z","shell.execute_reply.started":"2025-04-03T12:33:46.501101Z","shell.execute_reply":"2025-04-03T12:33:46.530352Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"print(f\"Sequence length statistics:\")\nprint(f\"- Minimum: {min_length}\")\nprint(f\"- Maximum: {max_length}\")\nprint(f\"- Average: {avg_length:.2f}\")","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-04-03T12:33:46.532754Z","iopub.execute_input":"2025-04-03T12:33:46.533364Z","iopub.status.idle":"2025-04-03T12:33:46.553648Z","shell.execute_reply.started":"2025-04-03T12:33:46.533308Z","shell.execute_reply":"2025-04-03T12:33:46.552506Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"### Visualize Sequence Length Distribution","metadata":{}},{"cell_type":"code","source":"fig = px.histogram(\n    train_sequences,\n    x='length',\n    nbins=30,\n    title='Distribution of RNA Sequence Lengths',\n    labels={'length': 'Sequence Length', 'count': 'Count'},\n    template='plotly_white'\n)\n\n# Customize the layout\nfig.update_layout(\n    width=900,\n    height=600,\n    title_font_size=20,\n    xaxis_title_font_size=16,\n    yaxis_title_font_size=16,\n    xaxis_title='Sequence Length',\n    yaxis_title='Count'\n)\n\n# Add grid lines for better readability\nfig.update_xaxes(showgrid=True, gridwidth=1, gridcolor='lightgray')\nfig.update_yaxes(showgrid=True, gridwidth=1, gridcolor='lightgray')\n\n# Customize histogram appearance\nfig.update_traces(\n    marker_color='royalblue',\n    marker_line_color='darkblue',\n    marker_line_width=1.5,\n    opacity=0.75\n)\n\n# Show the plot\nfig.show()","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-04-03T12:33:46.554798Z","iopub.execute_input":"2025-04-03T12:33:46.555233Z","iopub.status.idle":"2025-04-03T12:33:47.158451Z","shell.execute_reply.started":"2025-04-03T12:33:46.555192Z","shell.execute_reply":"2025-04-03T12:33:47.157247Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"Sequence lengths vary widely, which might pose challenges for machine learning models.","metadata":{}},{"cell_type":"markdown","source":"## Analyze Nucleotide Distribution","metadata":{}},{"cell_type":"code","source":"all_nucleotides = ''.join(train_sequences['sequence'].tolist())\nnucleotide_counts = Counter(all_nucleotides)\ntotal_nucleotides = sum(nucleotide_counts.values())","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-04-03T12:33:47.159431Z","iopub.execute_input":"2025-04-03T12:33:47.159754Z","iopub.status.idle":"2025-04-03T12:33:47.173048Z","shell.execute_reply.started":"2025-04-03T12:33:47.159723Z","shell.execute_reply":"2025-04-03T12:33:47.171914Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"print(\"\\nNucleotide distribution:\")\nfor nucleotide, count in nucleotide_counts.most_common():\n    percentage = (count / total_nucleotides) * 100\n    print(f\"{nucleotide}: {count} ({percentage:.2f}%)\")","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-04-03T12:33:47.174168Z","iopub.execute_input":"2025-04-03T12:33:47.174559Z","iopub.status.idle":"2025-04-03T12:33:47.199888Z","shell.execute_reply.started":"2025-04-03T12:33:47.174500Z","shell.execute_reply":"2025-04-03T12:33:47.198763Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"### Visualize Nucleotide Distribution","metadata":{}},{"cell_type":"code","source":"nucleotide_df = pd.DataFrame(nucleotide_counts.items(), columns=['Nucleotide', 'Count'])\nnucleotide_df['Percentage'] = (nucleotide_df['Count'] / total_nucleotides) * 100\n\n# Define colors for nucleotides (standard biological coloring)\nnucleotide_colors = {\n    'A': '#32CD32',  # Green\n    'U': '#FFD700',  # Gold\n    'G': '#4169E1',  # Blue\n    'C': '#FF6347',  # Red\n    '-': '#808080',  # Gray\n    'X': '#9370DB'   # Purple\n}\n\n# Create the bar plot\nfig_bar = px.bar(\n    nucleotide_df, \n    x='Nucleotide', \n    y='Count', \n    text='Percentage',\n #   labels={'count': 'Count', 'nucleotide': 'Nucleotide'},\n    title='Nucleotide Distribution',\n    color='Nucleotide',\n    color_discrete_map=nucleotide_colors,\n    template='plotly_white'\n)\n\n# Format the text to show percentages\nfig_bar.update_traces(\n    texttemplate='%{text:.2f}%', \n    textposition='outside'\n)\n\n# Set the size of the bar chart\nfig_bar.update_layout(\n    title_font_size=24,\n    width=800,\n    height=600,\n    font=dict(size=16)\n)\n\n\n\n# Create the pie chart (filter out very small values for better visualization)\nnucleotide_df_filtered = nucleotide_df[nucleotide_df['Percentage'] > 0.1]\n\nfig_pie = px.pie(\n    nucleotide_df_filtered, \n    values='Percentage', \n    names='Nucleotide',\n    title='Nucleotide Distribution (%)',\n    color='Nucleotide',\n    color_discrete_map=nucleotide_colors,\n    template='plotly_white'\n)\n\n# Update pie chart text format\nfig_pie.update_traces(\n    textinfo='label+percent', \n    textposition='inside',\n    insidetextorientation='radial'\n)\n\n# Increase the size of the pie chart\nfig_pie.update_layout(\n    width=800,    \n    height=600,  \n    title_font_size=24,  \n    font=dict(size=16)   \n)\n\n\n# Display plots one below the other\nfig_bar.show()\nfig_pie.show()","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-04-03T12:33:47.200721Z","iopub.execute_input":"2025-04-03T12:33:47.201071Z","iopub.status.idle":"2025-04-03T12:33:47.360510Z","shell.execute_reply.started":"2025-04-03T12:33:47.201040Z","shell.execute_reply":"2025-04-03T12:33:47.359607Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"The dataset has a GC-rich composition (G+C = 54.98%), which may indicate that many RNA structures in the dataset have strong secondary structures, such as hairpins and stems.","metadata":{}},{"cell_type":"markdown","source":"## Analyze Temporal Distribution","metadata":{}},{"cell_type":"code","source":"train_sequences['date'] = pd.to_datetime(train_sequences['temporal_cutoff'])\nmin_date = train_sequences['date'].min()\nmax_date = train_sequences['date'].max()\nprint(f\"\\nTemporal range: {min_date.date()} to {max_date.date()}\")","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-04-03T12:33:47.361508Z","iopub.execute_input":"2025-04-03T12:33:47.361826Z","iopub.status.idle":"2025-04-03T12:33:47.372633Z","shell.execute_reply.started":"2025-04-03T12:33:47.361797Z","shell.execute_reply":"2025-04-03T12:33:47.371484Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"### Visualize Temporal Distribution","metadata":{}},{"cell_type":"code","source":"# Calculate year counts\nyear_counts = train_sequences['date'].dt.year.value_counts().sort_index()\n\n# Convert to DataFrame for Plotly\nyear_df = pd.DataFrame({'Year': year_counts.index, 'Count': year_counts.values})\n\n# Create bar chart using Plotly Express\nfig = px.bar(\n    year_df,\n    x='Year',\n    y='Count',\n    title='Number of RNA Structures by Year',\n    labels={'Year': 'Year', 'Count': 'Count'},\n    template='plotly_white'\n)\n\n# Improve layout\nfig.update_layout(\n    width=900,\n    height=600,\n    title_font_size=20,\n    xaxis_title_font_size=16,\n    yaxis_title_font_size=16\n)\n\n# Format x-axis to ensure years display as integers without decimal points\nfig.update_xaxes(\n    type='category',  # Treat years as categories to maintain order\n    tickmode='linear'  # Show all years\n)\n\n# Add grid lines for y-axis only\nfig.update_yaxes(showgrid=True, gridwidth=1, gridcolor='lightgray')\n\n# Customize bar appearance\nfig.update_traces(\n    marker_color='royalblue',\n    marker_line_color='darkblue',\n    marker_line_width=1.5\n)\n\n# Show the plot\nfig.show()","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-04-03T12:33:47.373821Z","iopub.execute_input":"2025-04-03T12:33:47.374180Z","iopub.status.idle":"2025-04-03T12:33:47.455980Z","shell.execute_reply.started":"2025-04-03T12:33:47.374149Z","shell.execute_reply":"2025-04-03T12:33:47.454806Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"## Count Unique PDB IDs","metadata":{}},{"cell_type":"code","source":"pdb_ids = set([target_id.split('_')[0] for target_id in train_sequences['target_id']])\nprint(f\"Number of unique PDB IDs: {len(pdb_ids)}\")","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-04-03T12:33:47.457422Z","iopub.execute_input":"2025-04-03T12:33:47.457817Z","iopub.status.idle":"2025-04-03T12:33:47.463788Z","shell.execute_reply.started":"2025-04-03T12:33:47.457788Z","shell.execute_reply":"2025-04-03T12:33:47.462819Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"","metadata":{"trusted":true},"outputs":[],"execution_count":null}]}