{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"name":"python","version":"3.10.12","mimetype":"text/x-python","codemirror_mode":{"name":"ipython","version":3},"pygments_lexer":"ipython3","nbconvert_exporter":"python","file_extension":".py"},"kaggle":{"accelerator":"none","dataSources":[{"sourceId":87793,"databundleVersionId":11228175,"sourceType":"competition"}],"dockerImageVersionId":30918,"isInternetEnabled":false,"language":"python","sourceType":"notebook","isGpuEnabled":false}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"### **Introduction:**\n\nThis notebook performs an exploratory data analysis (EDA) on the Stanford RNA 3D Folding dataset. The datasets include training, validation, and test RNA sequences along with corresponding per-residue label data and a sample submission template.\n\n**Key Points:**\n- **Data Overview:**  \n  - The `train_sequences.csv` file contains 844 entries with basic information about each RNA target.  \n  - The `train_labels.csv` file provides per-residue coordinate data (with some missing values) for 137,095 rows corresponding to 735 unique targets.  \n  - The validation and test sequence files each have 12 entries.  \n  - The sample submission file is provided as a template with coordinate columns set to zero.\n\n- **Sequence Analysis:**  \n  - RNA sequence lengths vary widely (from as few as 3 nucleotides up to 4298 nucleotides).  \n  - Base composition analysis reveals that the sequences are predominantly composed of G, C, A, and U. Note that extra characters such as '-' and 'X' appear in the training sequences.\n\n- **Label Data Analysis:**  \n  - Coordinate columns in the label files have meaningful ranges and summary statistics.  \n  - Correlation heatmaps of coordinate columns indicate expected relationships.\n\n- **Data Integration & Quality:**  \n  - Merging of sequence and label data shows discrepancies in the training set (i.e. target IDs do not match perfectly).  \n  - Missing values exist, particularly in the training label coordinates.\n\n**Conclusion Preview:**  \nThe EDA reveals essential characteristics of the dataset, including wide variation in RNA sequence lengths, predominant base composition, and areas for data cleaning (missing values and merge key adjustments). These insights provide a strong foundation for further analysis and modeling.\n","metadata":{}},{"cell_type":"markdown","source":"### Importing required libraries","metadata":{}},{"cell_type":"code","source":"import pandas as pd\nimport numpy as np\nimport seaborn as sns\nimport matplotlib.pyplot as plt\nimport missingno as msno   # for missing data visualization (pip install missingno)\nfrom pathlib import Path\n\n# Configure plots\nsns.set(style=\"whitegrid\", context=\"notebook\")\nplt.rcParams[\"figure.figsize\"] = (10, 6)","metadata":{"_uuid":"8f2839f25d086af736a60e9eeb907d3b93b6e0e5","_cell_guid":"b1076dfc-b9ad-4769-8c92-a6c4dae69d19","trusted":true,"execution":{"iopub.status.busy":"2025-03-13T07:45:43.808383Z","iopub.execute_input":"2025-03-13T07:45:43.808751Z","iopub.status.idle":"2025-03-13T07:45:46.931848Z","shell.execute_reply.started":"2025-03-13T07:45:43.808719Z","shell.execute_reply":"2025-03-13T07:45:46.930387Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"### Load the dataset files","metadata":{}},{"cell_type":"code","source":"data_dir = Path(\"/kaggle/input/stanford-rna-3d-folding\")\n\ntrain_sequences_path = data_dir / \"train_sequences.csv\"\ntrain_labels_path    = data_dir / \"train_labels.csv\"\nvalidation_sequences_path = data_dir / \"validation_sequences.csv\"\nvalidation_labels_path    = data_dir / \"validation_labels.csv\"\ntest_sequences_path       = data_dir / \"test_sequences.csv\"\nsample_submission_path    = data_dir / \"sample_submission.csv\"\n\n# Read CSVs\ndf_train_seq = pd.read_csv(train_sequences_path)\ndf_train_lbl = pd.read_csv(train_labels_path)\ndf_val_seq   = pd.read_csv(validation_sequences_path)\ndf_val_lbl   = pd.read_csv(validation_labels_path)\ndf_test_seq  = pd.read_csv(test_sequences_path)\ndf_sub_sample= pd.read_csv(sample_submission_path)\n","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-03-13T07:46:23.843501Z","iopub.execute_input":"2025-03-13T07:46:23.843989Z","iopub.status.idle":"2025-03-13T07:46:24.423728Z","shell.execute_reply.started":"2025-03-13T07:46:23.843955Z","shell.execute_reply":"2025-03-13T07:46:24.422608Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"### Quick overview of each DataFrame","metadata":{}},{"cell_type":"code","source":"print(\"== train_sequences.csv ==\")\ndisplay(df_train_seq.head(5))\nprint(df_train_seq.info())\n\nprint(\"\\n== train_labels.csv ==\")\ndisplay(df_train_lbl.head(5))\nprint(df_train_lbl.info())\n\nprint(\"\\n== validation_sequences.csv ==\")\ndisplay(df_val_seq.head(5))\nprint(df_val_seq.info())\n\nprint(\"\\n== validation_labels.csv ==\")\ndisplay(df_val_lbl.head(5))\nprint(df_val_lbl.info())\n\nprint(\"\\n== test_sequences.csv ==\")\ndisplay(df_test_seq.head(5))\nprint(df_test_seq.info())\n\nprint(\"\\n== sample_submission.csv ==\")\ndisplay(df_sub_sample.head(5))\nprint(df_sub_sample.info())\n","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-03-13T07:46:56.224700Z","iopub.execute_input":"2025-03-13T07:46:56.225165Z","iopub.status.idle":"2025-03-13T07:46:56.414527Z","shell.execute_reply.started":"2025-03-13T07:46:56.225136Z","shell.execute_reply":"2025-03-13T07:46:56.413290Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"**Inference:**\n- `train_sequences.csv` shows 844 entries with 5 columns.\n- `train_labels.csv` contains 137,095 rows with 6 columns (note that some coordinate columns have missing values).\n- The validation and test sequences each have 12 entries, while the sample submission has 2515 rows.\nThis provides a quick snapshot of the dataset structure and reveals any immediate issues with missing values or incorrect data types.","metadata":{}},{"cell_type":"markdown","source":"### Basic descriptive statistics for numeric columns in label files","metadata":{}},{"cell_type":"code","source":"print(\"== train_labels.csv numeric summary ==\")\ndisplay(df_train_lbl.describe(include=[np.number]))\n\nprint(\"\\n== validation_labels.csv numeric summary ==\")\ndisplay(df_val_lbl.describe(include=[np.number]))\n\nprint(\"\\n== sample_submission.csv numeric summary ==\")\ndisplay(df_sub_sample.describe(include=[np.number]))","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-03-13T07:47:51.189978Z","iopub.execute_input":"2025-03-13T07:47:51.190516Z","iopub.status.idle":"2025-03-13T07:47:51.498378Z","shell.execute_reply.started":"2025-03-13T07:47:51.190475Z","shell.execute_reply":"2025-03-13T07:47:51.497302Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"**Inference:**\n- For `train_labels.csv`, the median residue index is 481, and coordinate columns (x_1, y_1, z_1) show medians around 62–70 with wide ranges.\n- These statistics give an idea of the data scale and variability, which is essential for understanding the spatial distribution in RNA structures.","metadata":{}},{"cell_type":"markdown","source":"### Checking for missing values in all dataframes","metadata":{}},{"cell_type":"code","source":"print(\"== Missing values: train_sequences ==\")\ndisplay(df_train_seq.isnull().sum())\n\nprint(\"\\n== Missing values: train_labels ==\")\ndisplay(df_train_lbl.isnull().sum())\n\nprint(\"\\n== Missing values: validation_sequences ==\")\ndisplay(df_val_seq.isnull().sum())\n\nprint(\"\\n== Missing values: validation_labels ==\")\ndisplay(df_val_lbl.isnull().sum())\n\nprint(\"\\n== Missing values: test_sequences ==\")\ndisplay(df_test_seq.isnull().sum())\n\nprint(\"\\n== Missing values: sample_submission ==\")\ndisplay(df_sub_sample.isnull().sum())","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-03-13T07:48:27.627972Z","iopub.execute_input":"2025-03-13T07:48:27.628464Z","iopub.status.idle":"2025-03-13T07:48:27.692709Z","shell.execute_reply.started":"2025-03-13T07:48:27.628415Z","shell.execute_reply":"2025-03-13T07:48:27.691818Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"**Inference:**\n- In `train_sequences.csv`, only the `all_sequences` column has 5 missing values.\n- In `train_labels.csv`, the coordinate columns (`x_1`, `y_1`, `z_1`) have 6145 missing entries each.\n- All other datasets are complete.\nThis highlights that missing data is primarily an issue in the train labels, which may require imputation or careful handling in downstream analysis.","metadata":{}},{"cell_type":"markdown","source":"### Visualizing missing data patterns","metadata":{}},{"cell_type":"code","source":"msno.bar(df_train_seq, sort=\"descending\", figsize=(8,4), color='blue')\nplt.title(\"Missing values in train_sequences.csv\")\nplt.show()\n\nmsno.bar(df_train_lbl, sort=\"descending\", figsize=(8,4), color='green')\nplt.title(\"Missing values in train_labels.csv\")\nplt.show()\n","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-03-13T07:49:14.435376Z","iopub.execute_input":"2025-03-13T07:49:14.435872Z","iopub.status.idle":"2025-03-13T07:49:15.807913Z","shell.execute_reply.started":"2025-03-13T07:49:14.435837Z","shell.execute_reply":"2025-03-13T07:49:15.806625Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"**Inference:**\nThe visualizations confirm the numerical missing values and indicate that while sequence files are largely complete, the label files have notable gaps that could affect model training or evaluation.","metadata":{}},{"cell_type":"markdown","source":"### Distribution : A look at 'sequence' columns in train_sequences/validation_sequences/test_sequences","metadata":{}},{"cell_type":"code","source":"def get_rna_length_stats(df, seq_col=\"sequence\"):\n    \"\"\"Compute length stats from a DataFrame that has an RNA sequence column.\"\"\"\n    df[\"seq_length\"] = df[seq_col].apply(len)\n    return df[\"seq_length\"].describe()\n\nprint(\"== Distribution of RNA sequence lengths in train_sequences ==\")\ndisplay(get_rna_length_stats(df_train_seq.copy(), \"sequence\"))\n\nprint(\"\\n== Distribution of RNA sequence lengths in validation_sequences ==\")\ndisplay(get_rna_length_stats(df_val_seq.copy(), \"sequence\"))\n\nprint(\"\\n== Distribution of RNA sequence lengths in test_sequences ==\")\ndisplay(get_rna_length_stats(df_test_seq.copy(), \"sequence\"))","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-03-13T07:50:24.759282Z","iopub.execute_input":"2025-03-13T07:50:24.759869Z","iopub.status.idle":"2025-03-13T07:50:24.792165Z","shell.execute_reply.started":"2025-03-13T07:50:24.759826Z","shell.execute_reply":"2025-03-13T07:50:24.791142Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"**Inference:**\n- The average sequence length in the training set is approximately 162 nucleotides, but the range spans from 3 to 4298.\n- The validation and test sets show similar distributions.\nThis indicates a broad variation in sequence lengths, which might influence modeling approaches and require normalization or stratification.","metadata":{}},{"cell_type":"markdown","source":"### Plotting histograms of RNA sequence lengths","metadata":{}},{"cell_type":"code","source":"df_train_seq[\"seq_length\"] = df_train_seq[\"sequence\"].apply(len)\ndf_val_seq[\"seq_length\"]   = df_val_seq[\"sequence\"].apply(len)\ndf_test_seq[\"seq_length\"]  = df_test_seq[\"sequence\"].apply(len)\n\nsns.histplot(df_train_seq[\"seq_length\"], bins=30, color='blue', kde=True)\nplt.title(\"Distribution of sequence lengths (train)\")\nplt.show()\n\nsns.histplot(df_val_seq[\"seq_length\"], bins=30, color='green', kde=True)\nplt.title(\"Distribution of sequence lengths (validation)\")\nplt.show()\n\nsns.histplot(df_test_seq[\"seq_length\"], bins=30, color='orange', kde=True)\nplt.title(\"Distribution of sequence lengths (test)\")\nplt.show()","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-03-13T07:51:04.935041Z","iopub.execute_input":"2025-03-13T07:51:04.935554Z","iopub.status.idle":"2025-03-13T07:51:06.206765Z","shell.execute_reply.started":"2025-03-13T07:51:04.935510Z","shell.execute_reply":"2025-03-13T07:51:06.205585Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"**Inference:**\n- The training set shows a wide spread in sequence lengths.\n- The validation and test sets display similar patterns.\nThese plots help visualize the distribution and identify any potential outliers or multimodal distributions in the RNA lengths.","metadata":{}},{"cell_type":"markdown","source":"### Analyze base composition in train_sequences, e.g. fraction of A, C, G, U (and others)","metadata":{}},{"cell_type":"code","source":"def base_composition(seq):\n    \"\"\"Return counts/fractions of each base in an RNA sequence.\"\"\"\n    from collections import Counter\n    c = Counter(seq)\n    length = len(seq)\n    return {base: c[base]/length for base in c}\n\ntrain_comps = df_train_seq[\"sequence\"].apply(base_composition)\n\n# Flatten out into a DataFrame\ndf_bases_train = pd.DataFrame(train_comps.tolist()).fillna(0)  # fill missing with 0 if a base not present\nprint(\"Base composition columns in train_sequences:\")\ndisplay(df_bases_train.describe())\n\n# Plot average fraction of each base\nmean_bases = df_bases_train.mean().sort_values(ascending=False)\nsns.barplot(x=mean_bases.index, y=mean_bases.values, palette=\"Blues_d\")\nplt.title(\"Average base composition in train_sequences\")\nplt.xlabel(\"Base\")\nplt.ylabel(\"Mean fraction\")\nplt.show()","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-03-13T07:51:55.209190Z","iopub.execute_input":"2025-03-13T07:51:55.209795Z","iopub.status.idle":"2025-03-13T07:51:55.506504Z","shell.execute_reply.started":"2025-03-13T07:51:55.209756Z","shell.execute_reply":"2025-03-13T07:51:55.505167Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"**Inference:**\n- The RNA sequences are primarily composed of G (29.9%), C (25.4%), A (23.1%), and U (21.7%).\n- There are minor fractions of '-' and 'X', which may indicate gaps or unknown bases.\nThis analysis informs us about the overall nucleotide composition, which could be relevant for downstream RNA structure prediction.","metadata":{}},{"cell_type":"markdown","source":"### Distribution : Inspect train_labels","metadata":{}},{"cell_type":"code","source":"coords_cols = [col for col in df_train_lbl.columns if col.startswith(\"x_\") or col.startswith(\"y_\") or col.startswith(\"z_\")]\nprint(\"Coordinate columns found in train_labels:\", coords_cols)\n\n# Summaries\ndisplay(df_train_lbl[coords_cols].describe())\n\n# A quick correlation heatmap among x_1, y_1, z_1, x_2, y_2, z_2, etc.\ncorr = df_train_lbl[coords_cols].corr()\nsns.heatmap(corr, cmap=\"vlag\", center=0)\nplt.title(\"Correlation among coordinate columns in train_labels\")\nplt.show()","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-03-13T07:53:13.302297Z","iopub.execute_input":"2025-03-13T07:53:13.302956Z","iopub.status.idle":"2025-03-13T07:53:13.634110Z","shell.execute_reply.started":"2025-03-13T07:53:13.302918Z","shell.execute_reply":"2025-03-13T07:53:13.632882Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"**Inference:**\n- The coordinate columns show a wide range (with medians around 62 for x, 67 for y, and 73 for z).\n- The heatmap displays correlations that suggest certain linear relationships or consistent spatial patterns.\nThese results are important for understanding the 3D spatial distribution of RNA residues.","metadata":{}},{"cell_type":"markdown","source":"### Similarly for validation_labels","metadata":{}},{"cell_type":"code","source":"coords_cols_val = [col for col in df_val_lbl.columns if col.startswith(\"x_\") or col.startswith(\"y_\") or col.startswith(\"z_\")]\nprint(\"Coordinate columns found in validation_labels:\", coords_cols_val)\n\n# Summaries\ndisplay(df_val_lbl[coords_cols_val].describe())\n\n# A quick correlation heatmap among these\ncorr_val = df_val_lbl[coords_cols_val].corr()\nsns.heatmap(corr_val, cmap=\"vlag\", center=0)\nplt.title(\"Correlation among coordinate columns in validation_labels\")\nplt.show()","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-03-13T07:53:44.673700Z","iopub.execute_input":"2025-03-13T07:53:44.674189Z","iopub.status.idle":"2025-03-13T07:53:45.774582Z","shell.execute_reply.started":"2025-03-13T07:53:44.674156Z","shell.execute_reply":"2025-03-13T07:53:45.773392Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"**Inference:**\n- The validation labels data are complete (with no missing values) and display correlation patterns similar to the training labels.\nThis confirms data consistency and offers a comparison point for further analyses.","metadata":{}},{"cell_type":"markdown","source":"### Merging train_sequences and train_labels for deeper analysis","metadata":{}},{"cell_type":"code","source":"# The 'ID' in labels is of the form \"target_id_resnum\", so let's parse out the 'target_id' from ID if needed\n# Alternatively, if train_labels has a 'target_id' column or we can do a direct merge on \"ID\".\n\ndf_train_lbl['target_id'] = df_train_lbl['ID'].apply(lambda x: x.split('_')[0])\n\n# Now let's see if we can join with df_train_seq on 'target_id' if that is consistent:\nmerged_train = pd.merge(df_train_lbl, df_train_seq, on=\"target_id\", how=\"inner\")\nprint(\"Merged shape:\", merged_train.shape)\ndisplay(merged_train.head(5))\n\n# We can similarly do for validation if 'ID' format is consistent with 'target_id'\ndf_val_lbl['target_id'] = df_val_lbl['ID'].apply(lambda x: x.split('_')[0])\nmerged_val = pd.merge(df_val_lbl, df_val_seq, on=\"target_id\", how=\"inner\")\nprint(\"Merged shape (validation):\", merged_val.shape)\ndisplay(merged_val.head(5))","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-03-13T07:55:08.562402Z","iopub.execute_input":"2025-03-13T07:55:08.562962Z","iopub.status.idle":"2025-03-13T07:55:08.697483Z","shell.execute_reply.started":"2025-03-13T07:55:08.562930Z","shell.execute_reply":"2025-03-13T07:55:08.696460Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"### Further EDA","metadata":{}},{"cell_type":"code","source":"if \"resid\" in merged_train.columns:\n    # Check how resid compares to the length of the sequence\n    # For each target_id, the maximum 'resid' might match the length of the sequence if it's 1-based indexing\n    summary_df = merged_train.groupby(\"target_id\").agg({\n        \"resid\": \"max\",\n        \"sequence\": lambda s: len(s.iloc[0])  # length of the first sequence in that group\n    }).rename(columns={\"resid\": \"max_resid\", \"sequence\": \"seq_length\"})\n    summary_df[\"diff\"] = summary_df[\"seq_length\"] - summary_df[\"max_resid\"]\n    print(summary_df.head(10))\n    sns.histplot(summary_df[\"diff\"], kde=True)\n    plt.title(\"Difference between sequence length and max residue index\")\n    plt.show()","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-03-13T07:56:08.610088Z","iopub.execute_input":"2025-03-13T07:56:08.610553Z","iopub.status.idle":"2025-03-13T07:56:08.869708Z","shell.execute_reply.started":"2025-03-13T07:56:08.610518Z","shell.execute_reply":"2025-03-13T07:56:08.868511Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"**Inference:**\n- The empty merge result for training data indicates that the `target_id` values in `train_sequences.csv` and `train_labels.csv` do not match perfectly.\n- This discrepancy suggests that the merge key may require additional cleaning or a different approach.\nResolving this is important for integrated analyses combining sequence and label information.","metadata":{}},{"cell_type":"markdown","source":"### Observing target distributions","metadata":{}},{"cell_type":"code","source":"# Count how many distinct 'target_id' are in each set:\nprint(\"Unique target_id in train_sequences:\", df_train_seq[\"target_id\"].nunique())\nprint(\"Unique target_id in train_labels:\", df_train_lbl[\"target_id\"].nunique())\nprint(\"Unique target_id in validation_sequences:\", df_val_seq[\"target_id\"].nunique())\nprint(\"Unique target_id in validation_labels:\", df_val_lbl[\"target_id\"].nunique())\nprint(\"Unique target_id in test_sequences:\", df_test_seq[\"target_id\"].nunique())","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-03-13T07:57:21.318204Z","iopub.execute_input":"2025-03-13T07:57:21.318694Z","iopub.status.idle":"2025-03-13T07:57:21.339866Z","shell.execute_reply.started":"2025-03-13T07:57:21.318652Z","shell.execute_reply":"2025-03-13T07:57:21.338773Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"**Inference:**\n- The difference between train_sequences and train_labels suggests that some RNA targets in the training sequences do not have associated label data.\n- Validation and test sets appear consistent.\nThis observation is crucial for understanding data coverage across the different files.","metadata":{}},{"cell_type":"markdown","source":"### Additional domain-specific EDA","metadata":{}},{"cell_type":"code","source":"\n# How many unique nucleotides appear in 'sequence' columns (like A, C, G, U, or others)?\n\ndef unique_bases(df, seq_col=\"sequence\"):\n    all_bases = set()\n    for s in df[seq_col]:\n        all_bases.update(list(s))\n    return all_bases\n\nprint(\"train_sequences unique bases:\", unique_bases(df_train_seq))\nprint(\"validation_sequences unique bases:\", unique_bases(df_val_seq))\nprint(\"test_sequences unique bases:\", unique_bases(df_test_seq))\n\n# If there are 'N' or 'T' or 'X' in some sequences, that might require special handling.","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-03-13T07:58:22.414750Z","iopub.execute_input":"2025-03-13T07:58:22.415287Z","iopub.status.idle":"2025-03-13T07:58:22.429038Z","shell.execute_reply.started":"2025-03-13T07:58:22.415251Z","shell.execute_reply":"2025-03-13T07:58:22.427593Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"**Inference:**\n- The presence of '-' and 'X' in the train sequences might indicate gaps, unknown bases, or special annotations.\n- This may require additional preprocessing before modeling.","metadata":{}},{"cell_type":"markdown","source":"### **Conclusion:**\n\n**Summary of Findings:**\n- **Data Coverage:** The training sequences and labels differ in granularity (844 vs. 735 unique targets), indicating that some targets lack label data.\n- **Sequence Analysis:** RNA sequences show wide variability in length (3 to 4298 nucleotides) and have a predominant base composition (G, C, A, U) with some extra characters in the training set.\n- **Coordinate Information:** The label files provide detailed 3D spatial coordinates with descriptive statistics and expected correlations, though the training labels include missing data.\n- **Data Integration:** An attempted merge of training sequences with labels revealed a mismatch in `target_id` values, suggesting a need for additional cleaning.\n- **Submission Template:** The sample submission file is correctly formatted as a placeholder with all zeros.\n\n**Overall Conclusion:**\nThe EDA provides a solid foundation for understanding the dataset's structure, quality, and key characteristics. It highlights areas that require further attention (such as handling missing values and resolving merge key discrepancies) before proceeding with modeling. These insights will guide the development of robust RNA 3D structure prediction methods.\n","metadata":{}},{"cell_type":"code","source":"# Create submission.csv directly from the sample_submission DataFrame\ndf_sub_sample.to_csv(\"submission.csv\", index=False)","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-03-13T08:20:51.604848Z","iopub.execute_input":"2025-03-13T08:20:51.605402Z","iopub.status.idle":"2025-03-13T08:20:51.663931Z","shell.execute_reply.started":"2025-03-13T08:20:51.605364Z","shell.execute_reply":"2025-03-13T08:20:51.662141Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"","metadata":{"trusted":true},"outputs":[],"execution_count":null}]}