{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"name":"python","version":"3.10.12","mimetype":"text/x-python","codemirror_mode":{"name":"ipython","version":3},"pygments_lexer":"ipython3","nbconvert_exporter":"python","file_extension":".py"},"kaggle":{"accelerator":"none","dataSources":[{"sourceId":87793,"databundleVersionId":11403143,"sourceType":"competition"}],"dockerImageVersionId":30918,"isInternetEnabled":true,"language":"python","sourceType":"notebook","isGpuEnabled":false}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"# Introduction\n\n#### In this notebook, we will explore two datasets: train_sequences and train_labels (containing structural information for each nucleotide). The goal is to understand the patterns in the sequences and their 3D structures. We will also analyze sequence lengths, nucleotide composition, and k-mer frequencies.","metadata":{}},{"cell_type":"markdown","source":"# Section 1: Importing Libraries\n## In this section, we import the necessary Python libraries for data analysis and visualization.\n\n\n**pandas:**  For data manipulation and analysis.\n\n**matplotlib and seaborn:** For data visualization.\n","metadata":{}},{"cell_type":"code","source":"import pandas as pd\nimport matplotlib.pyplot as plt\nimport seaborn as sns","metadata":{"_uuid":"8f2839f25d086af736a60e9eeb907d3b93b6e0e5","_cell_guid":"b1076dfc-b9ad-4769-8c92-a6c4dae69d19","trusted":true,"execution":{"iopub.status.busy":"2025-03-16T07:09:23.618277Z","iopub.execute_input":"2025-03-16T07:09:23.618561Z","iopub.status.idle":"2025-03-16T07:09:24.955684Z","shell.execute_reply.started":"2025-03-16T07:09:23.618523Z","shell.execute_reply":"2025-03-16T07:09:24.954497Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"# Section 2: Loading the Datasets\n## Here, we load the training datasets (train_sequences and train_labels) and the validation datasets (validation_sequences and validation_labels).\n\n**train_sequences**: Contains RNA sequences and metadata.\n\n**train_labels:** Contains structural information for each nucleotide.\n\n**validation_sequences and validation_labels:** Similar to the training datasets but used for validation","metadata":{}},{"cell_type":"code","source":"# Load the datasets\ntrain_sequences = pd.read_csv('/kaggle/input/stanford-rna-3d-folding/train_sequences.csv')\ntrain_labels = pd.read_csv('/kaggle/input/stanford-rna-3d-folding/train_labels.csv')","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-03-16T07:09:24.956872Z","iopub.execute_input":"2025-03-16T07:09:24.957549Z","iopub.status.idle":"2025-03-16T07:09:25.385031Z","shell.execute_reply.started":"2025-03-16T07:09:24.957509Z","shell.execute_reply":"2025-03-16T07:09:25.384008Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"# Load the datasets\nvalidation_sequences = pd.read_csv('/kaggle/input/stanford-rna-3d-folding/validation_sequences.csv')\nvalidation_labels = pd.read_csv('/kaggle/input/stanford-rna-3d-folding/validation_labels.csv')","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-03-16T07:09:25.386036Z","iopub.execute_input":"2025-03-16T07:09:25.386403Z","iopub.status.idle":"2025-03-16T07:09:25.480012Z","shell.execute_reply.started":"2025-03-16T07:09:25.386367Z","shell.execute_reply":"2025-03-16T07:09:25.479075Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"# Section 3: Data Inspection\n## In this section, we inspect the first few rows of the datasets to understand their structure and content.\n\n**train_sequences:** 844 rows and 5 columns.\n\n**train_labels:** 137,095 rows and 6 columns.\n\n**validation_sequences:** 12 rows and 5 columns.\n\n**validation_labels:** 2,515 rows and 123 columns.","metadata":{}},{"cell_type":"code","source":"train_sequences.head(2)","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-03-16T07:09:25.480872Z","iopub.execute_input":"2025-03-16T07:09:25.481185Z","iopub.status.idle":"2025-03-16T07:09:25.506366Z","shell.execute_reply.started":"2025-03-16T07:09:25.481159Z","shell.execute_reply":"2025-03-16T07:09:25.505171Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"train_labels.head(2)","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-03-16T07:09:25.507372Z","iopub.execute_input":"2025-03-16T07:09:25.507738Z","iopub.status.idle":"2025-03-16T07:09:25.519583Z","shell.execute_reply.started":"2025-03-16T07:09:25.507703Z","shell.execute_reply":"2025-03-16T07:09:25.518610Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"train_sequences.shape","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-03-16T07:09:25.522822Z","iopub.execute_input":"2025-03-16T07:09:25.523144Z","iopub.status.idle":"2025-03-16T07:09:25.539232Z","shell.execute_reply.started":"2025-03-16T07:09:25.523119Z","shell.execute_reply":"2025-03-16T07:09:25.538204Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"train_labels.shape","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-03-16T07:09:25.541029Z","iopub.execute_input":"2025-03-16T07:09:25.541397Z","iopub.status.idle":"2025-03-16T07:09:25.556477Z","shell.execute_reply.started":"2025-03-16T07:09:25.541367Z","shell.execute_reply":"2025-03-16T07:09:25.555485Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"validation_sequences.shape","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-03-16T07:09:25.557410Z","iopub.execute_input":"2025-03-16T07:09:25.557690Z","iopub.status.idle":"2025-03-16T07:09:25.574482Z","shell.execute_reply.started":"2025-03-16T07:09:25.557667Z","shell.execute_reply":"2025-03-16T07:09:25.573329Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"validation_labels.shape","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-03-16T07:09:25.575469Z","iopub.execute_input":"2025-03-16T07:09:25.575886Z","iopub.status.idle":"2025-03-16T07:09:25.593981Z","shell.execute_reply.started":"2025-03-16T07:09:25.575847Z","shell.execute_reply":"2025-03-16T07:09:25.593059Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"## Basic Dataset Information\n### In this section, we display basic information about the datasets, including data types and missing values\n\n**train_sequences:** Contains 5 columns, all of type object. There are 5 missing values in the all_sequences column.\n\n**train_labels:** Contains 6 columns, with some missing values in the x_1, y_1, and z_1 columns.","metadata":{}},{"cell_type":"code","source":"# Display basic info about the datasets\nprint(\"===== Train Sequences Info =====\")\nprint(train_sequences.info())","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-03-16T07:09:25.594972Z","iopub.execute_input":"2025-03-16T07:09:25.595275Z","iopub.status.idle":"2025-03-16T07:09:25.635386Z","shell.execute_reply.started":"2025-03-16T07:09:25.595247Z","shell.execute_reply":"2025-03-16T07:09:25.634204Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"print(\"\\n===== Train Labels Info =====\")\nprint(train_labels.info())","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-03-16T07:09:25.636522Z","iopub.execute_input":"2025-03-16T07:09:25.636885Z","iopub.status.idle":"2025-03-16T07:09:25.680252Z","shell.execute_reply.started":"2025-03-16T07:09:25.636850Z","shell.execute_reply":"2025-03-16T07:09:25.679259Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"# Basic statistics for train_sequences\nprint(\"\\n===== Train Sequences Description =====\")\ntrain_sequences.describe(include='all')","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-03-16T07:09:25.681193Z","iopub.execute_input":"2025-03-16T07:09:25.681560Z","iopub.status.idle":"2025-03-16T07:09:25.703596Z","shell.execute_reply.started":"2025-03-16T07:09:25.681525Z","shell.execute_reply":"2025-03-16T07:09:25.702681Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"# Check for missing values\nprint(\"\\nMissing values in train_sequences:\")\nprint(train_sequences.isnull().sum())","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-03-16T07:09:25.704606Z","iopub.execute_input":"2025-03-16T07:09:25.705078Z","iopub.status.idle":"2025-03-16T07:09:25.722932Z","shell.execute_reply.started":"2025-03-16T07:09:25.705033Z","shell.execute_reply":"2025-03-16T07:09:25.721918Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"# Check for missing values\nprint(\"\\nMissing values in train_labels:\")\nprint(train_labels.isnull().sum())","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-03-16T07:09:25.723888Z","iopub.execute_input":"2025-03-16T07:09:25.724261Z","iopub.status.idle":"2025-03-16T07:09:25.761376Z","shell.execute_reply.started":"2025-03-16T07:09:25.724225Z","shell.execute_reply":"2025-03-16T07:09:25.760068Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"# Check unique values in each column\nprint(\"\\nUnique values in each column:\")\nprint(train_sequences.nunique())","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-03-16T07:09:25.762469Z","iopub.execute_input":"2025-03-16T07:09:25.762813Z","iopub.status.idle":"2025-03-16T07:09:25.771431Z","shell.execute_reply.started":"2025-03-16T07:09:25.762770Z","shell.execute_reply":"2025-03-16T07:09:25.770487Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"**sequence:** 784 unique, 60 repeated.\n\n**temporal_cutoff:** 476 unique, 368 repeated.\n\n**description:** 716 unique, 128 repeated.\n\n**all_sequences:** 732 unique, 107 repeated.","metadata":{}},{"cell_type":"code","source":"# Check for non-unique values in the 'sequence' column\nnon_unique_sequences = train_sequences['sequence'].value_counts()\nprint(\"Non-unique sequences and their counts:\")\nprint(non_unique_sequences[non_unique_sequences > 1])","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-03-16T07:09:25.772339Z","iopub.execute_input":"2025-03-16T07:09:25.772687Z","iopub.status.idle":"2025-03-16T07:09:25.789467Z","shell.execute_reply.started":"2025-03-16T07:09:25.772653Z","shell.execute_reply":"2025-03-16T07:09:25.788427Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"# Basic statistics for train_labels\nprint(\"\\n===== Train Labels Description =====\")\ntrain_labels.describe(include='all')","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-03-16T07:09:25.790601Z","iopub.execute_input":"2025-03-16T07:09:25.791010Z","iopub.status.idle":"2025-03-16T07:09:25.969893Z","shell.execute_reply.started":"2025-03-16T07:09:25.790960Z","shell.execute_reply":"2025-03-16T07:09:25.968832Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"# Section 4: EDA\n\n## Exploratory Data Analysis (EDA) on RNA 3D Folding Datasets\n## Introduction to EDA\nExploratory Data Analysis (EDA) is a critical step in any data science or machine learning project. It involves analyzing and summarizing datasets to understand their structure, patterns, and relationships. EDA helps us uncover insights, detect anomalies, and prepare the data for further modeling or analysis.\n\nIn this notebook, we will perform EDA on two datasets related to RNA 3D folding:\n\n* **train_sequences.csv:** Contains RNA sequences along with metadata such as sequence IDs, descriptions, and temporal cutoffs.\n\n* **train_labels.csv:** Contains structural information for each nucleotide in the sequences, including coordinates (x, y, z) and residue information.\n\nGoals of EDA\nUnderstand the Data: Explore the structure, content, and characteristics of the datasets.\n\nIdentify Patterns: Analyze sequence lengths, nucleotide composition, and temporal trends.","metadata":{}},{"cell_type":"markdown","source":"The sequence_length analysis shows that the RNA sequences in train_sequences vary widely in length, ranging from 3 to 4298 nucleotides, with an average length of 162 nucleotides. The high standard deviation (515.03) indicates significant variability in sequence lengths.\n\n3: The shortest sequence in the dataset has 3 nucleotides.\n\n4298: The longest sequence in the dataset has 4298 nucleotides.","metadata":{}},{"cell_type":"code","source":"# Analyze the 'sequence' column in train_sequences\nprint(\"\\n===== Sequence Length Analysis =====\")\ntrain_sequences['sequence_length'] = train_sequences['sequence'].apply(len)\ntrain_sequences['sequence_length'].describe()","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-03-16T07:09:25.970917Z","iopub.execute_input":"2025-03-16T07:09:25.971337Z","iopub.status.idle":"2025-03-16T07:09:25.983675Z","shell.execute_reply.started":"2025-03-16T07:09:25.971300Z","shell.execute_reply":"2025-03-16T07:09:25.982625Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"# Plot the distribution of sequence lengths\nplt.figure(figsize=(10, 6))\nsns.histplot(train_sequences['sequence_length'], bins=50, kde=True)\nplt.title('Distribution of Sequence Lengths')\nplt.xlabel('Sequence Length')\nplt.ylabel('Frequency')\nplt.show()","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-03-16T07:09:25.984612Z","iopub.execute_input":"2025-03-16T07:09:25.984952Z","iopub.status.idle":"2025-03-16T07:09:26.382051Z","shell.execute_reply.started":"2025-03-16T07:09:25.984903Z","shell.execute_reply":"2025-03-16T07:09:26.380735Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"from collections import Counter\n\n# Count nucleotide frequencies\nnucleotide_counts = Counter(\"\".join(train_sequences[\"sequence\"]))\nnucleotide_counts","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-03-16T07:09:26.387101Z","iopub.execute_input":"2025-03-16T07:09:26.387433Z","iopub.status.idle":"2025-03-16T07:09:26.403287Z","shell.execute_reply.started":"2025-03-16T07:09:26.387406Z","shell.execute_reply":"2025-03-16T07:09:26.402302Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"The nucleotide composition analysis shows that G (Guanine) is the most frequent nucleotide (41,450 occurrences), followed by C (Cytosine), A (Adenine), and U (Uracil). Rare characters like '-' and 'X' (likely placeholders or errors) appear only 4 and 2 times, respectively.","metadata":{}},{"cell_type":"code","source":"# Plot nucleotide composition\nplt.figure(figsize=(8, 6))\nsns.barplot(x=list(nucleotide_counts.keys()), y=list(nucleotide_counts.values()))\nplt.title(\"Nucleotide Composition\")\nplt.xlabel(\"Nucleotide\")\nplt.ylabel(\"Frequency\")\nplt.show()","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-03-16T07:09:26.404922Z","iopub.execute_input":"2025-03-16T07:09:26.405282Z","iopub.status.idle":"2025-03-16T07:09:26.619706Z","shell.execute_reply.started":"2025-03-16T07:09:26.405253Z","shell.execute_reply":"2025-03-16T07:09:26.618726Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"The temporal_cutoff analysis shows that the RNA sequences in the dataset span from 1995-01-26 to 2024-12-18, with the median date around 2014-07-09. This indicates the dataset covers nearly 30 years of RNA sequence data.","metadata":{}},{"cell_type":"code","source":"# Analyze the 'temporal_cutoff' column in train_sequences\nprint(\"\\n===== Temporal Cutoff Analysis =====\")\ntrain_sequences['temporal_cutoff'] = pd.to_datetime(train_sequences['temporal_cutoff'])\ntrain_sequences['temporal_cutoff'].describe()","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-03-16T07:09:26.620699Z","iopub.execute_input":"2025-03-16T07:09:26.621007Z","iopub.status.idle":"2025-03-16T07:09:26.637077Z","shell.execute_reply.started":"2025-03-16T07:09:26.620982Z","shell.execute_reply":"2025-03-16T07:09:26.636250Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"# Find the starting and ending dates\nstart_date = train_sequences['temporal_cutoff'].min()\nend_date = train_sequences['temporal_cutoff'].max()\n\n# Print the results\nprint(\"Starting Date:\", start_date)\nprint(\"Ending Date:\", end_date)","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-03-16T07:09:26.638141Z","iopub.execute_input":"2025-03-16T07:09:26.638517Z","iopub.status.idle":"2025-03-16T07:09:26.656114Z","shell.execute_reply.started":"2025-03-16T07:09:26.638478Z","shell.execute_reply":"2025-03-16T07:09:26.654975Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"# Plot the distribution of temporal cutoff dates\nplt.figure(figsize=(10, 6))\nsns.histplot(train_sequences['temporal_cutoff'], bins=50, kde=True)\nplt.title('Distribution of Temporal Cutoff Dates')\nplt.xlabel('Temporal Cutoff')\nplt.ylabel('Frequency')\nplt.show()","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-03-16T07:09:26.656965Z","iopub.execute_input":"2025-03-16T07:09:26.657300Z","iopub.status.idle":"2025-03-16T07:09:27.034468Z","shell.execute_reply.started":"2025-03-16T07:09:26.657271Z","shell.execute_reply":"2025-03-16T07:09:27.033444Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"The resid column represents the residue ID (position of nucleotides in the RNA sequences). The values range from 1 to 4298, with a mean of 897.26, indicating the sequences vary significantly in length. The high standard deviation (1014.32) confirms this variability.\n","metadata":{}},{"cell_type":"code","source":"# Analyze the 'resid' column in train_labels\nprint(\"\\n===== Residue ID Analysis =====\")\ntrain_labels['resid'].describe()","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-03-16T07:09:27.035546Z","iopub.execute_input":"2025-03-16T07:09:27.035927Z","iopub.status.idle":"2025-03-16T07:09:27.048626Z","shell.execute_reply.started":"2025-03-16T07:09:27.035889Z","shell.execute_reply":"2025-03-16T07:09:27.047617Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"# Plot the distribution of residue IDs\nplt.figure(figsize=(10, 6))\nsns.histplot(train_labels['resid'], bins=50, kde=True)\nplt.title('Distribution of Residue IDs')\nplt.xlabel('Residue ID')\nplt.ylabel('Frequency')\nplt.show()","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-03-16T07:09:27.049567Z","iopub.execute_input":"2025-03-16T07:09:27.049876Z","iopub.status.idle":"2025-03-16T07:09:27.927211Z","shell.execute_reply.started":"2025-03-16T07:09:27.049850Z","shell.execute_reply":"2025-03-16T07:09:27.925858Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"The 3D coordinates (x_1, y_1, z_1) represent the spatial positions of nucleotides in the RNA structures. The values range widely (e.g., x_1 from -821.09 to 849.89), indicating diverse structural conformations, with a mean around 80-99 units for each axis.","metadata":{}},{"cell_type":"code","source":"# Analyze the 3D coordinates in train_labels\nprint(\"\\n===== 3D Coordinates Analysis =====\")\ncoordinates = train_labels[['x_1', 'y_1', 'z_1']].describe()\nprint(coordinates)","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-03-16T07:09:27.928315Z","iopub.execute_input":"2025-03-16T07:09:27.928623Z","iopub.status.idle":"2025-03-16T07:09:27.964888Z","shell.execute_reply.started":"2025-03-16T07:09:27.928596Z","shell.execute_reply":"2025-03-16T07:09:27.963791Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"# Plot the distribution of x, y, z coordinates\nplt.figure(figsize=(15, 5))\nplt.subplot(1, 3, 1)\nsns.histplot(train_labels['x_1'], bins=50, kde=True)\nplt.title('Distribution of x-coordinates')\n\nplt.subplot(1, 3, 2)\nsns.histplot(train_labels['y_1'], bins=50, kde=True)\nplt.title('Distribution of y-coordinates')\n\nplt.subplot(1, 3, 3)\nsns.histplot(train_labels['z_1'], bins=50, kde=True)\nplt.title('Distribution of z-coordinates')\n\nplt.tight_layout()\nplt.show()","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-03-16T07:09:27.966147Z","iopub.execute_input":"2025-03-16T07:09:27.966544Z","iopub.status.idle":"2025-03-16T07:09:30.817270Z","shell.execute_reply.started":"2025-03-16T07:09:27.966503Z","shell.execute_reply":"2025-03-16T07:09:30.816095Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"The correlation plot shows the relationships between the 3D coordinates (x_1, y_1, z_1) of the RNA structures. The values indicate the strength and direction of the correlation, ranging from -1 to 1. For example, x_1 and y_1 have a correlation of 0.66, suggesting a moderate positive relationship, while other pairs show weaker correlations.","metadata":{}},{"cell_type":"code","source":"# Compute correlations\ncorrelation_matrix = train_labels[[\"x_1\", \"y_1\", \"z_1\"]].corr()\n\n# Plot correlation matrix\nplt.figure(figsize=(8, 6))\nsns.heatmap(correlation_matrix, annot=True, cmap=\"coolwarm\", vmin=-1, vmax=1)\nplt.title(\"Correlation Matrix for x, y, z Coordinates\")\nplt.show()","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-03-16T07:09:30.818705Z","iopub.execute_input":"2025-03-16T07:09:30.819061Z","iopub.status.idle":"2025-03-16T07:09:31.062460Z","shell.execute_reply.started":"2025-03-16T07:09:30.819024Z","shell.execute_reply":"2025-03-16T07:09:31.061374Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"The scatter plot visualizes the relationship between sequence length and the number of sequences in the dataset. It shows how sequences are distributed across different lengths.","metadata":{}},{"cell_type":"code","source":"# Group by sequence length and analyze coordinates\nsequence_length_groups = train_sequences.groupby(\"sequence_length\").size().reset_index(name=\"count\")\n\n# Plot sequence length vs. coordinate statistics\nplt.figure(figsize=(10, 6))\nsns.scatterplot(data=sequence_length_groups, x=\"sequence_length\", y=\"count\")\nplt.title(\"Sequence Length vs. Number of Sequences\")\nplt.xlabel(\"Sequence Length\")\nplt.ylabel(\"Number of Sequences\")\nplt.show()","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-03-16T07:09:31.063405Z","iopub.execute_input":"2025-03-16T07:09:31.063757Z","iopub.status.idle":"2025-03-16T07:09:31.287268Z","shell.execute_reply.started":"2025-03-16T07:09:31.063730Z","shell.execute_reply":"2025-03-16T07:09:31.286186Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"The boxplot visualizes the distribution of the 3D coordinates (x_1, y_1, z_1) for the RNA structures. It highlights the median, quartiles, and potential outliers, showing the spread and variability of the spatial positions.","metadata":{}},{"cell_type":"code","source":"# Boxplot for x, y, z coordinates\nplt.figure(figsize=(10, 6))\nsns.boxplot(data=train_labels[[\"x_1\", \"y_1\", \"z_1\"]])\nplt.title(\"Boxplot of x, y, z Coordinates\")\nplt.show()","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-03-16T07:09:31.288208Z","iopub.execute_input":"2025-03-16T07:09:31.288541Z","iopub.status.idle":"2025-03-16T07:09:31.497264Z","shell.execute_reply.started":"2025-03-16T07:09:31.288506Z","shell.execute_reply":"2025-03-16T07:09:31.496231Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"# Section 5: Visualizing RNA 3D /2D Structures\n\n## Introduction\nIn this section, we will visualize the 3D structure of RNA sequences using the train_labels dataset. The dataset contains structural information for each nucleotide in the RNA sequences, including their coordinates (x_1, y_1, z_1) and residue names (resname).","metadata":{}},{"cell_type":"code","source":"import pandas as pd\nimport plotly.graph_objects as go","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-03-16T07:09:31.498285Z","iopub.execute_input":"2025-03-16T07:09:31.498599Z","iopub.status.idle":"2025-03-16T07:09:31.515863Z","shell.execute_reply.started":"2025-03-16T07:09:31.498573Z","shell.execute_reply":"2025-03-16T07:09:31.514907Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"## Function plot_structure:\n\n1. Filters the dataset for a specific RNA sequence (sequence_id).\n\n2. Uses Plotly to create an interactive 3D scatter plot, with nucleotides (A, G, C, U) as colored markers and the RNA backbone as a gray line.\n\n\n## Nucleotide Count and Sequence:\n\n\n1. Counts occurrences of each nucleotide in the sequence.\n   \n2. Reconstructs and prints the nucleotide sequence in one line.\n","metadata":{}},{"cell_type":"code","source":"def plot_structure(df: pd.DataFrame, sequence_id: str) -> None:\n    sequence_df = df[df[\"sequence_id\"] == sequence_id]\n    sequence_points = sequence_df[[\"x_1\", \"y_1\", \"z_1\", \"resname\"]]\n    \n    colors = {\"A\": \"red\", \"G\": \"blue\", \"C\": \"green\", \"U\": \"yellow\"}\n    fig = go.Figure()\n    \n    for resname, color in colors.items():\n        subset = sequence_df[sequence_df[\"resname\"] == resname]\n        fig.add_trace(go.Scatter3d(\n            x=subset[\"x_1\"], y=subset[\"y_1\"], z=subset[\"z_1\"],\n            mode='markers',\n            marker=dict(size=5, color=color),\n            name=resname,\n        ))\n    \n    fig.add_trace(go.Scatter3d(\n        x=sequence_df[\"x_1\"], y=sequence_df[\"y_1\"], z=sequence_df[\"z_1\"],\n        mode='lines',\n        line=dict(color='gray', width=2),\n        name='RNA Backbone'\n    ))\n    \n    fig.update_layout(\n            scene=dict(xaxis_title='X', yaxis_title='Y', zaxis_title='Z'),\n            title=f'3D RNA Structure of sequence {sequence_id}',\n        )\n            \n    fig.show(renderer=\"iframe\")","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-03-16T07:09:31.517277Z","iopub.execute_input":"2025-03-16T07:09:31.517670Z","iopub.status.idle":"2025-03-16T07:09:31.525570Z","shell.execute_reply.started":"2025-03-16T07:09:31.517630Z","shell.execute_reply":"2025-03-16T07:09:31.524420Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"train_labels[\"sequence_id\"] = train_labels[\"ID\"].apply(lambda x: \"_\".join(x.split(\"_\")[:-1]))","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-03-16T07:09:31.526588Z","iopub.execute_input":"2025-03-16T07:09:31.526917Z","iopub.status.idle":"2025-03-16T07:09:31.626123Z","shell.execute_reply.started":"2025-03-16T07:09:31.526890Z","shell.execute_reply":"2025-03-16T07:09:31.625148Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"train_labels.head(20)","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-03-16T07:09:31.627077Z","iopub.execute_input":"2025-03-16T07:09:31.627471Z","iopub.status.idle":"2025-03-16T07:09:31.642730Z","shell.execute_reply.started":"2025-03-16T07:09:31.627434Z","shell.execute_reply":"2025-03-16T07:09:31.641817Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"plot_structure(train_labels, \"1SCL_A\")","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-03-16T07:09:31.643755Z","iopub.execute_input":"2025-03-16T07:09:31.644113Z","iopub.status.idle":"2025-03-16T07:09:32.281309Z","shell.execute_reply.started":"2025-03-16T07:09:31.644077Z","shell.execute_reply":"2025-03-16T07:09:32.280069Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"## Shows how many A, U, G, C nucleotides appear in 1SCL_A.","metadata":{}},{"cell_type":"code","source":"# Select only rows where Group == '1SCL_A'\nfiltered_df = train_labels[train_labels[\"sequence_id\"] == \"1SCL_A\"]\n# Count occurrences of each nucleotide\nnucleotide_counts = filtered_df[\"resname\"].value_counts()\n# Print the counts\nprint(nucleotide_counts)","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-03-16T07:09:32.282369Z","iopub.execute_input":"2025-03-16T07:09:32.282729Z","iopub.status.idle":"2025-03-16T07:09:32.302014Z","shell.execute_reply.started":"2025-03-16T07:09:32.282700Z","shell.execute_reply":"2025-03-16T07:09:32.300728Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"# Filter nucleotides where Group is \"1SCL_A\"\nnucleotide_sequence = \"\".join(train_labels[train_labels[\"sequence_id\"] == \"1SCL_A\"][\"resname\"])\n# Print the nucleotide sequence in one line\nprint(nucleotide_sequence)","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-03-16T07:09:32.303156Z","iopub.execute_input":"2025-03-16T07:09:32.303533Z","iopub.status.idle":"2025-03-16T07:09:32.334403Z","shell.execute_reply.started":"2025-03-16T07:09:32.303498Z","shell.execute_reply":"2025-03-16T07:09:32.333183Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"plot_structure(train_labels, \"1RHT_A\")","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-03-16T07:09:32.335488Z","iopub.execute_input":"2025-03-16T07:09:32.336065Z","iopub.status.idle":"2025-03-16T07:09:32.427224Z","shell.execute_reply.started":"2025-03-16T07:09:32.336027Z","shell.execute_reply":"2025-03-16T07:09:32.426105Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"# only in X-Y plane\nplt.figure(figsize=(8, 6))\nsns.scatterplot(x=train_labels[\"x_1\"], y=train_labels[\"y_1\"])\nplt.title(\"Scatterplot of RNA C1' Atoms (X-Y Plane)\")\nplt.xlabel(\"X Coordinate\")\nplt.ylabel(\"Y Coordinate\")\nplt.show()","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-03-16T07:09:32.428088Z","iopub.execute_input":"2025-03-16T07:09:32.428380Z","iopub.status.idle":"2025-03-16T07:09:32.930519Z","shell.execute_reply.started":"2025-03-16T07:09:32.428354Z","shell.execute_reply":"2025-03-16T07:09:32.929473Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"# Section 6: Analyzing k-mers in RNA Sequences\n\n## Introduction\nIn this section, we will analyze k-mers (subsequences of length k) in RNA sequences. K-mers are essential for understanding patterns and motifs in RNA sequences, which can provide insights into their structure and function.\n\nA k-mer is a small sequence of k nucleotides (A, C, G, U) from a longer RNA or DNA sequence.","metadata":{}},{"cell_type":"markdown","source":"It shows all possible 4-mers (subsequences of 4 nucleotides) from each sequence in  dataset and counts how often each 4-mer appears. It then prints the 20 most common 4-mers and their frequencies, such as GGGG appearing 1,232 times.","metadata":{}},{"cell_type":"code","source":"# K-mer- Function, k=4\ndef get_kmers(sequence, k=3):\n    return [sequence[i:i+k] for i in range(len(sequence) - k + 1)]\n\nkmer_counts = Counter()\nfor seq in train_sequences[\"sequence\"]:\n    kmer_counts.update(get_kmers(seq, k=4))\nprint(kmer_counts.most_common(20))","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-03-16T07:09:32.931835Z","iopub.execute_input":"2025-03-16T07:09:32.932373Z","iopub.status.idle":"2025-03-16T07:09:32.973853Z","shell.execute_reply.started":"2025-03-16T07:09:32.932344Z","shell.execute_reply":"2025-03-16T07:09:32.972838Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"import matplotlib.pyplot as plt\nimport seaborn as sns\n\n# Get the top 10 most common k-mers\ntop_kmers = kmer_counts.most_common(10)\nkmer_names = [kmer for kmer, count in top_kmers]\nkmer_counts = [count for kmer, count in top_kmers]\n\n# Plot\nplt.figure(figsize=(10, 6))\nsns.barplot(x=kmer_names, y=kmer_counts)\nplt.title(\"Top 10 Most Frequent 4-mers\")\nplt.xlabel(\"4-mer\")\nplt.ylabel(\"Frequency\")\nplt.show()","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-03-16T07:09:32.974797Z","iopub.execute_input":"2025-03-16T07:09:32.975237Z","iopub.status.idle":"2025-03-16T07:09:33.228492Z","shell.execute_reply.started":"2025-03-16T07:09:32.975207Z","shell.execute_reply":"2025-03-16T07:09:33.227445Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"It shows the most frequent k-mers for different values of k (2, 3, 4) in  RNA sequences. For example, the top 2-mer is GG (appearing 13,170 times), the top 3-mer is GGG (appearing 4,021 times), and the top 4-mer is GGGG (appearing 1,232 times).","metadata":{}},{"cell_type":"code","source":"def generate_kmer_counts(sequences, k):\n    kmer_counts = Counter()\n    for seq in sequences:\n        kmer_counts.update(get_kmers(seq, k=k))\n    return kmer_counts\n\n# Generate k-mer counts for k=2, 3, 4\nk2_counts = generate_kmer_counts(train_sequences[\"sequence\"], k=2)\nk3_counts = generate_kmer_counts(train_sequences[\"sequence\"], k=3)\nk4_counts = generate_kmer_counts(train_sequences[\"sequence\"], k=4)\n\n# Print top 5 for each k\nprint(\"Top 5 2-mers:\", k2_counts.most_common(5))\nprint(\"Top 5 3-mers:\", k3_counts.most_common(5))\nprint(\"Top 5 4-mers:\", k4_counts.most_common(5))","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-03-16T07:09:33.229531Z","iopub.execute_input":"2025-03-16T07:09:33.229978Z","iopub.status.idle":"2025-03-16T07:09:33.339099Z","shell.execute_reply.started":"2025-03-16T07:09:33.229918Z","shell.execute_reply":"2025-03-16T07:09:33.338053Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"It finds all the starting positions of a specific k-mer (e.g., AUGG) in a list of sequences. It returns a list of positions where the k-mer appears.","metadata":{}},{"cell_type":"code","source":"def find_kmer_positions(sequences, kmer):\n    positions = []\n    for seq in sequences:\n        start_indices = [i for i in range(len(seq) - len(kmer) + 1) if seq[i:i+len(kmer)] == kmer]\n        positions.extend(start_indices)\n    return positions\n\n# Example: Find positions of the k-mer \"AUGG\"\nkmer = \"AUGG\"\npositions = find_kmer_positions(train_sequences[\"sequence\"], kmer)\nprint(f\"Positions of '{kmer}': {positions}\")","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-03-16T07:09:33.340092Z","iopub.execute_input":"2025-03-16T07:09:33.340406Z","iopub.status.idle":"2025-03-16T07:09:33.376763Z","shell.execute_reply.started":"2025-03-16T07:09:33.340371Z","shell.execute_reply":"2025-03-16T07:09:33.375649Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"A 4-mer is a subsequence of 4 nucleotides (A, C, G, U) from an RNA sequence.\n\n**For example,** in the sequence AUGGCAU:\n\n**The 4-mers are:** AUGG, UGGC, GGCA, GCAU","metadata":{}},{"cell_type":"code","source":"from collections import Counter\nfrom wordcloud import WordCloud\nimport matplotlib.pyplot as plt\n\n# Function to generate k-mers\ndef get_kmers(sequence, k=3):\n    return [sequence[i:i+k] for i in range(len(sequence) - k + 1)]\n\n# Initialize a Counter to store k-mer frequencies\nkmer_counts = Counter()\n\n# Loop through each sequence in the dataset\nfor seq in train_sequences[\"sequence\"]:\n    kmer_counts.update(get_kmers(seq, k=4))\n\nkmer_dict = dict(kmer_counts.most_common(30))\n\n# Generate word cloud\nwordcloud = WordCloud(width=800, height=400, background_color=\"white\").generate_from_frequencies(kmer_dict)\n\n# Plot\nplt.figure(figsize=(5, 5))\nplt.imshow(wordcloud, interpolation=\"bilinear\")\nplt.axis(\"off\")\nplt.title(\"4-mer Word Cloud\")\nplt.show()","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-03-16T07:09:33.377920Z","iopub.execute_input":"2025-03-16T07:09:33.378344Z","iopub.status.idle":"2025-03-16T07:09:34.137237Z","shell.execute_reply.started":"2025-03-16T07:09:33.378309Z","shell.execute_reply":"2025-03-16T07:09:34.136042Z"}},"outputs":[],"execution_count":null}]}