{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"name":"python","version":"3.10.12","mimetype":"text/x-python","codemirror_mode":{"name":"ipython","version":3},"pygments_lexer":"ipython3","nbconvert_exporter":"python","file_extension":".py"},"kaggle":{"accelerator":"nvidiaTeslaT4","dataSources":[{"sourceId":87793,"databundleVersionId":12276181,"sourceType":"competition"}],"dockerImageVersionId":30918,"isInternetEnabled":false,"language":"python","sourceType":"notebook","isGpuEnabled":true}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"code","source":"# This Python 3 environment comes with many helpful analytics libraries installed\n# It is defined by the kaggle/python Docker image: https://github.com/kaggle/docker-python\n# For example, here's several helpful packages to load\n\nimport numpy as np # linear algebra\nimport pandas as pd # data processing, CSV file I/O (e.g. pd.read_csv)\n\n# Input data files are available in the read-only \"../input/\" directory\n# For example, running this (by clicking run or pressing Shift+Enter) will list all files under the input directory\n\nimport os\n#for dirname, _, filenames in os.walk('/kaggle/input'):\n#    for filename in filenames:\n#        print(os.path.join(dirname, filename))\n\n# You can write up to 20GB to the current directory (/kaggle/working/) that gets preserved as output when you create a version using \"Save & Run All\" \n# You can also write temporary files to /kaggle/temp/, but they won't be saved outside of the current session","metadata":{"_uuid":"8f2839f25d086af736a60e9eeb907d3b93b6e0e5","_cell_guid":"b1076dfc-b9ad-4769-8c92-a6c4dae69d19","trusted":true,"execution":{"iopub.status.busy":"2025-10-18T04:27:53.303895Z","iopub.execute_input":"2025-10-18T04:27:53.304079Z","iopub.status.idle":"2025-10-18T04:27:55.122006Z","shell.execute_reply.started":"2025-10-18T04:27:53.304060Z","shell.execute_reply":"2025-10-18T04:27:55.121295Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"**In biology, Rnd proteins and DNA play distinct but interconnected roles in cellular processes**\nRnd Proteins:\n\n- Rnd1: Stimulates neurite outgrowth through inhibition of RhoA signaling in PC12 cells in culture.\n- Rnd2: Promotes migration by regulating the multipolar-to-bipolar transition in the intermediate zone, partially through RhoA inhibition in cortical neurons in vivo.\n- Rnd3: Controls interkinetic nuclear migration, orientation of cleavage plane, and maintenance of adherens junctions by inhibiting actin polymerization in neural stem cells (NSCs) in the embryonic cerebral cortex in vivo.\nDNA (Deoxyribonucleic Acid):\n\n- Structure: DNA is a double-stranded molecule composed of two polynucleotide chains that coil around each other to form a double helix. Each strand is made up of nucleotides, which consist of a phosphate group, a deoxyribose sugar, and one of four nitrogenous bases: adenine (A), thymine (T), cytosine (C), or guanine (G).\n- Function: DNA stores and transfers genetic information for the cell. It contains the instructions for the synthesis of other molecules, such as proteins, and is essential for the development, functioning, growth, and reproduction of all known organisms and many viruses.\n\nInteractions and Regulation:\n\n- Transcriptional Regulation: Rnd proteins are regulated at the transcriptional level by various transcription factors. For example, the proneural transcription factor Neurogenin2 (Neurog2) promotes the migration of nascent cortical neurons by inducing the expression of Rnd2, while another proneural factor, Ascl1, promotes neuronal migration in the cortex through direct regulation of Rnd3.\nDNA and RNA: DNA serves as the template for RNA synthesis during transcription. Messenger RNA (mRNA) carries the genetic information from - DNA to the ribosomes, where proteins are synthesized. Transfer RNA (tRNA) brings amino acids to the ribosomes to build proteins according to the mRNA code.\n\nIn summary, while Rnd proteins are involved in regulating cellular processes such as migration and cytoskeletal dynamics, DNA is the primary molecule for storing and transmitting genetic information. The expression and function of Rnd proteins are often controlled by the genetic information encoded in DNA.","metadata":{}},{"cell_type":"markdown","source":"### Respect learning from\nhttps://www.kaggle.com/code/mpwolke/pyrococcus-furiosus-fireball-eternafold and discusion group","metadata":{}},{"cell_type":"markdown","source":"## Load the Data\nFirst, let's load the dataset and inspect it.","metadata":{}},{"cell_type":"code","source":"import pandas as pd\n\npath = '/kaggle/input/stanford-rna-3d-folding/'\n# Load the training sequences dataset\ntrain_sequences_df = pd.read_csv(path + 'train_sequences.csv')\n\n# Load the training labels dataset\ntrain_labels_df = pd.read_csv(path + 'train_labels.csv')\n\n# Load the test sequences dataset\ntest_sequences_df = pd.read_csv(path + 'test_sequences.csv')\n\n# Load the validation sequences dataset\nvalidation_sequences_df = pd.read_csv(path + 'validation_sequences.csv')\n\n# Load the validation labels dataset\nvalidation_labels_df = pd.read_csv(path + 'validation_labels.csv')\n\n# Load the sample submission dataset\nsample_submission_df = pd.read_csv(path + 'sample_submission.csv')","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-10-18T04:27:55.122685Z","iopub.execute_input":"2025-10-18T04:27:55.123115Z","iopub.status.idle":"2025-10-18T04:27:55.579371Z","shell.execute_reply.started":"2025-10-18T04:27:55.123090Z","shell.execute_reply":"2025-10-18T04:27:55.578470Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"train_sequences_df.head(5)","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-10-18T04:27:55.580192Z","iopub.execute_input":"2025-10-18T04:27:55.580460Z","iopub.status.idle":"2025-10-18T04:27:55.606310Z","shell.execute_reply.started":"2025-10-18T04:27:55.580439Z","shell.execute_reply":"2025-10-18T04:27:55.605641Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"train_labels_df.head(5)","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-10-18T04:27:55.607028Z","iopub.execute_input":"2025-10-18T04:27:55.607269Z","iopub.status.idle":"2025-10-18T04:27:55.616605Z","shell.execute_reply.started":"2025-10-18T04:27:55.607248Z","shell.execute_reply":"2025-10-18T04:27:55.615967Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"test_sequences_df.head(5)","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-10-18T04:27:55.617486Z","iopub.execute_input":"2025-10-18T04:27:55.617805Z","iopub.status.idle":"2025-10-18T04:27:55.635956Z","shell.execute_reply.started":"2025-10-18T04:27:55.617776Z","shell.execute_reply":"2025-10-18T04:27:55.635272Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"validation_sequences_df.head(5)","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-10-18T04:27:55.638052Z","iopub.execute_input":"2025-10-18T04:27:55.638321Z","iopub.status.idle":"2025-10-18T04:27:55.656974Z","shell.execute_reply.started":"2025-10-18T04:27:55.638302Z","shell.execute_reply":"2025-10-18T04:27:55.656043Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"validation_labels_df.head(5)","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-10-18T04:27:55.659099Z","iopub.execute_input":"2025-10-18T04:27:55.659395Z","iopub.status.idle":"2025-10-18T04:27:55.708892Z","shell.execute_reply.started":"2025-10-18T04:27:55.659365Z","shell.execute_reply":"2025-10-18T04:27:55.707937Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"sample_submission_df.head(5)","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-10-18T04:27:55.709909Z","iopub.execute_input":"2025-10-18T04:27:55.710228Z","iopub.status.idle":"2025-10-18T04:27:55.732476Z","shell.execute_reply.started":"2025-10-18T04:27:55.710169Z","shell.execute_reply":"2025-10-18T04:27:55.731762Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"## Exploratory Data Analysis (EDA)","metadata":{}},{"cell_type":"markdown","source":"Let's perform some basic exploratory data analysis to understand the dataset better.\n\nSummary Statistics:\n\nCount of unique target_id values.\n\nDistribution of resname values.\n\nSummary statistics for coordinates.","metadata":{}},{"cell_type":"code","source":"import matplotlib.pyplot as plt\nfrom mpl_toolkits.mplot3d import Axes3D\n\n# 3D plot of coordinates\nfig = plt.figure(figsize=(10, 8))\nax = fig.add_subplot(111, projection='3d')\n\n# Plot the coordinates\nax.scatter(train_labels_df['x_1'], train_labels_df['y_1'], train_labels_df['z_1'], c='b', marker='o')\n\n# Set labels\nax.set_xlabel('X Coordinate')\nax.set_ylabel('Y Coordinate')\nax.set_zlabel('Z Coordinate')\nax.set_title('3D Coordinates of Residues')\n\nplt.show()","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-10-18T04:27:55.733359Z","iopub.execute_input":"2025-10-18T04:27:55.733577Z","iopub.status.idle":"2025-10-18T04:27:58.635483Z","shell.execute_reply.started":"2025-10-18T04:27:55.733548Z","shell.execute_reply":"2025-10-18T04:27:58.634388Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"We can visualize the 3D coordinates to better understand the spatial distribution of the residues. For simplicity, let's visualize the first structure (x_1, y_1, z_1).","metadata":{}},{"cell_type":"code","source":"import matplotlib.pyplot as plt\nfrom mpl_toolkits.mplot3d import Axes3D\n\n# 3D plot of coordinates for the first structure\nfig = plt.figure(figsize=(10, 8))\nax = fig.add_subplot(111, projection='3d')\n\n# Plot the coordinates\nax.scatter(validation_labels_df['x_1'], validation_labels_df['y_1'], validation_labels_df['z_1'], c='b', marker='o')\n\n# Set labels\nax.set_xlabel('X Coordinate')\nax.set_ylabel('Y Coordinate')\nax.set_zlabel('Z Coordinate')\nax.set_title('3D Coordinates of Residues (First Structure)')\n\nplt.show()","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-10-18T04:27:58.636488Z","iopub.execute_input":"2025-10-18T04:27:58.636807Z","iopub.status.idle":"2025-10-18T04:27:58.881019Z","shell.execute_reply.started":"2025-10-18T04:27:58.636777Z","shell.execute_reply":"2025-10-18T04:27:58.880086Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"## Evaluation","metadata":{}},{"cell_type":"markdown","source":"$$\n\\text{TM-score} = \\max \\left( \\frac{1}{L_{\\text{ref}}} \\sum_{i=1}^{L_{\\text{align}}} \\frac{1}{1 + \\left(\\frac{d_i}{d_0}\\right)^2} \\right)\n$$","metadata":{}},{"cell_type":"markdown","source":"where:\n\n\n- $d_i$ is the distance between the \\(i\\)-th residue in the predicted structure and the corresponding residue in the reference structure.\n\n- $L_align$ is the number of aligned residues.\n\n- $d_i$ is the distance between the ith pair of aligned residues, in Angstroms.\n\n- $d_0$ is a distance scaling factor in Angstroms, defined as:","metadata":{}},{"cell_type":"markdown","source":"$$\nd_0 = 0.6 (L_{\\text{ref}} - 0.5)^{1/2} - 2.5\n$$","metadata":{}},{"cell_type":"markdown","source":"$$\nd_0 = \n\\begin{cases} \n0.3 & \\text{if } L_{\\text{ref}} < 12 \\\\\n0.4 & \\text{if } 12 \\leq L_{\\text{ref}} < 16 \\\\\n0.5 & \\text{if } 16 \\leq L_{\\text{ref}} < 20 \\\\\n0.6 & \\text{if } 20 \\leq L_{\\text{ref}} < 24 \\\\\n0.7 & \\text{if } 24 \\leq L_{\\text{ref}} < 29 \\\\\n0.6 (L_{\\text{ref}} - 0.5)^{1/2} - 2.5 & \\text{if } L_{\\text{ref}} \\geq 30\n\\end{cases}\n$$","metadata":{}},{"cell_type":"markdown","source":">  Summarize the Data","metadata":{}},{"cell_type":"code","source":"def summarize_sequences(df):\n    \"\"\"Summarize the train_sequences.csv data.\"\"\"\n    summary = {\n        'total_sequences': len(df),\n        'sequence_lengths': df['sequence'].apply(len).describe(),\n        'temporal_cutoffs': df['temporal_cutoff'].describe(),\n        'descriptions': df['description'].unique()\n    }\n    return summary\n\ndef summarize_labels(df):\n    \"\"\"Summarize the train_labels.csv data.\"\"\"\n    summary = {\n        'total_residues': len(df),\n        'residue_names': df['resname'].value_counts(),\n        'coordinate_ranges': {\n            'x_1': (df['x_1'].min(), df['x_1'].max()),\n            'y_1': (df['y_1'].min(), df['y_1'].max()),\n            'z_1': (df['z_1'].min(), df['z_1'].max())\n        }\n    }\n    return summary\n\ndef summarize_sample_submission(df):\n    \"\"\"Summarize the sample_submission.csv data.\"\"\"\n    summary = {\n        'total_residues': len(df),\n        'residue_names': df['resname'].value_counts(),\n        'coordinate_ranges': {\n            'x_1': (df['x_1'].min(), df['x_1'].max()),\n            'y_1': (df['y_1'].min(), df['y_1'].max()),\n            'z_1': (df['z_1'].min(), df['z_1'].max())\n        }\n    }\n    return summary","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-10-18T04:27:58.881931Z","iopub.execute_input":"2025-10-18T04:27:58.882250Z","iopub.status.idle":"2025-10-18T04:27:58.888814Z","shell.execute_reply.started":"2025-10-18T04:27:58.882216Z","shell.execute_reply":"2025-10-18T04:27:58.887844Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"> Calculate TM-score","metadata":{}},{"cell_type":"code","source":"import numpy as np\n\ndef calculate_d0(L_ref):\n    \"\"\"Calculate the distance scaling factor d0.\"\"\"\n    if L_ref >= 30:\n        d0 = 0.6 * (L_ref - 0.5) ** 0.5 - 2.5\n    else:\n        if L_ref < 12:\n            d0 = 0.3\n        elif 12 <= L_ref < 15:\n            d0 = 0.4\n        elif 15 <= L_ref < 19:\n            d0 = 0.5\n        elif 19 <= L_ref < 23:\n            d0 = 0.6\n        else:\n            d0 = 0.7\n    return d0\n\ndef calculate_tm_score(L_ref, L_align, distances):\n    \"\"\"Calculate the TM-score.\"\"\"\n    d0 = calculate_d0(L_ref)\n    tm_score = np.max(1 / L_ref * np.sum(1 / (1 + (distances / d0) ** 2)))\n    return tm_score","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-10-18T04:27:58.889726Z","iopub.execute_input":"2025-10-18T04:27:58.889908Z","iopub.status.idle":"2025-10-18T04:27:58.913876Z","shell.execute_reply.started":"2025-10-18T04:27:58.889891Z","shell.execute_reply":"2025-10-18T04:27:58.913048Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"> Usage","metadata":{}},{"cell_type":"code","source":"# Read the data\ntrain_sequences = train_sequences_df\ntrain_labels = train_labels_df\nsample_submission = sample_submission_df\n\n# Summarize the data\nsequences_summary = summarize_sequences(train_sequences)\nlabels_summary = summarize_labels(train_labels)\nsample_submission_summary = summarize_sample_submission(sample_submission)\n\n# Print summaries\nprint(\"Train Sequences Summary:\")\nprint(sequences_summary)\n\nprint(\"\\nTrain Labels Summary:\")\nprint(labels_summary)\n\nprint(\"\\nSample Submission Summary:\")\nprint(sample_submission_summary)\n\n# Example TM-score calculation\n# Assume we have predictions for the first residue of the first sequence\nL_ref = 27  # Length of the reference sequence\nL_align = 27  # Number of aligned residues\ndistances = np.array([0.1, 0.2, 0.3, 0.4, 0.5, 0.6, 0.7, 0.8, 0.9, 1.0, 1.1, 1.2, 1.3, 1.4, 1.5, 1.6, 1.7, 1.8, 1.9, 2.0, 2.1, 2.2, 2.3, 2.4, 2.5, 2.6, 2.7])\n\n# Calculate TM-score\ntm_score = calculate_tm_score(L_ref, L_align, distances)\nprint(f\"\\nTM-score: {tm_score}\")","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-10-18T04:27:58.914807Z","iopub.execute_input":"2025-10-18T04:27:58.915065Z","iopub.status.idle":"2025-10-18T04:27:58.963692Z","shell.execute_reply.started":"2025-10-18T04:27:58.915038Z","shell.execute_reply":"2025-10-18T04:27:58.962904Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"## Predicion and Submssion","metadata":{}},{"cell_type":"code","source":"# copy datasets to new name to match\ntest_seq = test_sequences_df\nsample_sub = sample_submission_df","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-10-18T04:27:58.964928Z","iopub.execute_input":"2025-10-18T04:27:58.965319Z","iopub.status.idle":"2025-10-18T04:27:58.969038Z","shell.execute_reply.started":"2025-10-18T04:27:58.965286Z","shell.execute_reply":"2025-10-18T04:27:58.968412Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"1. Function to Generate Fixed RNA Positions:","metadata":{}},{"cell_type":"code","source":"def generate_positions(sequence_length, shape='linear'):\n    \"\"\"\n    Generate a basic RNA shape with the specified number of residues\n    \"\"\"\n    if shape == 'linear':\n        # Create a simple straight line with evenly spaced residues\n        positions = np.zeros((sequence_length, 3))\n        for i in range(sequence_length):\n            positions[i] = [i * 5.0, 0.0, 0.0]  # 5Å spacing\n     \n    elif shape == 'circle':\n        # Create a circle\n        positions = np.zeros((sequence_length, 3))\n        radius = sequence_length / (2 * np.pi)  # Adjust radius based on sequence length\n        for i in range(sequence_length):\n            angle = 2 * np.pi * i / sequence_length\n            positions[i] = [radius * np.cos(angle), radius * np.sin(angle), 0.0]\n     \n    elif shape == 'helix':\n        # Create a helix (like A-form RNA)\n        positions = np.zeros((sequence_length, 3))\n        radius = 10.0  # Radius of helix\n        rise_per_residue = 2.8  # Å rise per residue\n        residues_per_turn = 11  # ~11 residues per turn for A-form RNA\n        \n        for i in range(sequence_length):\n            angle = 2 * np.pi * i / residues_per_turn\n            positions[i] = [\n                radius * np.cos(angle), \n                radius * np.sin(angle), \n                i * rise_per_residue\n            ]\n     \n    return positions","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-10-18T04:27:58.969864Z","iopub.execute_input":"2025-10-18T04:27:58.970193Z","iopub.status.idle":"2025-10-18T04:27:58.986300Z","shell.execute_reply.started":"2025-10-18T04:27:58.970150Z","shell.execute_reply":"2025-10-18T04:27:58.985531Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"2. Purpose: Generate fixed 3D positions for RNA sequences based on the specified shape (linear, circle, or helix).\n   \nFunction to Generate Diverse Shapes:","metadata":{}},{"cell_type":"code","source":"def generate_diverse_shapes(sequence, n_conformations=5):\n    \"\"\"Generate diverse RNA shapes for the 5 required conformations\"\"\"\n    sequence_length = len(sequence)\n    conformations = []\n    \n    # Basic shapes to use\n    shapes = ['linear', 'circle', 'helix', 'helix', 'circle']\n    \n    for i in range(n_conformations):\n        shape = shapes[i % len(shapes)]\n        \n        # Generate basic shape\n        coords = generate_positions(sequence_length, shape)\n        \n        # Apply transformations for additional diversity\n        if i > 0:\n            # Add some rotation\n            angle = np.radians(i * 72)  # 72 degrees = 360/5\n            c, s = np.cos(angle), np.sin(angle)\n            \n            # Rotation matrix\n            if i % 3 == 1:\n                R = np.array([[c, 0, s], [0, 1, 0], [-s, 0, c]])  # Y-axis\n            elif i % 3 == 2:\n                R = np.array([[1, 0, 0], [0, c, -s], [0, s, c]])  # X-axis\n            else:\n                R = np.array([[c, -s, 0], [s, c, 0], [0, 0, 1]])  # Z-axis\n            \n            # Center, rotate, and translate back\n            center = np.mean(coords, axis=0)\n            coords = coords - center\n            coords = np.dot(coords, R)\n            coords = coords + center\n            \n            # Add some translation\n            coords = coords + np.random.normal(0, 5, 3)\n        \n        conformations.append(coords)\n    \n    return conformations","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-10-18T04:27:58.987214Z","iopub.execute_input":"2025-10-18T04:27:58.987480Z","iopub.status.idle":"2025-10-18T04:27:59.007062Z","shell.execute_reply.started":"2025-10-18T04:27:58.987454Z","shell.execute_reply":"2025-10-18T04:27:59.006357Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"Purpose: Generate multiple diverse conformations for a given RNA sequence by combining different basic shapes and applying transformations (rotation and translation).\n3. Process Test Sequences and Create Submission:","metadata":{}},{"cell_type":"code","source":"print(\"\\n=== Generating Predictions for Submission ===\")\nsubmission = sample_sub.copy()\n\nfor idx, row in test_seq.iterrows():\n    target_id = row['target_id']\n    sequence = row['sequence']\n    \n    print(f\"Processing {target_id} (length: {len(sequence)})\")\n    \n    # Generate 5 diverse conformations\n    conformations = generate_diverse_shapes(sequence)\n    \n    # Fill submission with predictions\n    for i, conformation in enumerate(conformations):\n        # Get rows for this RNA\n        mask = submission['ID'].str.startswith(target_id)\n        \n        # Get sorted indices by residue ID\n        sorted_indices = submission.loc[mask].sort_values('resid').index\n        \n        # Fill coordinates for each residue\n        for j, idx in enumerate(sorted_indices):\n            if j < len(conformation):\n                submission.loc[idx, f'x_{i+1}'] = float(conformation[j][0])\n                submission.loc[idx, f'y_{i+1}'] = float(conformation[j][1])\n                submission.loc[idx, f'z_{i+1}'] = float(conformation[j][2])\n            else:\n                # Just in case we have a mismatch - shouldn't happen\n                submission.loc[idx, f'x_{i+1}'] = float(j * 5.0)\n                submission.loc[idx, f'y_{i+1}'] = 0.0\n                submission.loc[idx, f'z_{i+1}'] = 0.0\n\n# Check for any NaN values and fill them\nif submission.isna().any().any():\n    print(\"Warning: NaN values detected in submission. Filling with zeros.\")\n    submission = submission.fillna(0.0)\n\n# Save submission file\nsubmission.to_csv('submission.csv', index=False)\nprint(\"\\nSaved submission file: submission.csv\")\n\n# Display sample of submission\nprint(\"\\nSubmission preview:\")\nprint(submission.head())","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-10-18T04:27:59.007864Z","iopub.execute_input":"2025-10-18T04:27:59.008160Z","iopub.status.idle":"2025-10-18T04:28:05.424721Z","shell.execute_reply.started":"2025-10-18T04:27:59.008131Z","shell.execute_reply":"2025-10-18T04:28:05.423937Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"Explanation\n\n1- Initialization:","metadata":{}},{"cell_type":"code","source":"print(\"\\n=== Generating Predictions for Submission ===\")\nsubmission = sample_sub.copy()","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-10-18T04:28:05.425515Z","iopub.execute_input":"2025-10-18T04:28:05.425738Z","iopub.status.idle":"2025-10-18T04:28:05.430475Z","shell.execute_reply.started":"2025-10-18T04:28:05.425718Z","shell.execute_reply":"2025-10-18T04:28:05.429574Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"Purpose: Initialize the submission DataFrame by copying the sample submission.\n\n2- Processing Each Test Sequence:","metadata":{}},{"cell_type":"code","source":"for idx, row in test_seq.iterrows():\n    target_id = row['target_id']\n    sequence = row['sequence']\n    \n    print(f\"Processing {target_id} (length: {len(sequence)})\")\n    \n    # Generate 5 diverse conformations\n    conformations = generate_diverse_shapes(sequence)","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-10-18T04:28:05.431293Z","iopub.execute_input":"2025-10-18T04:28:05.431604Z","iopub.status.idle":"2025-10-18T04:28:05.535520Z","shell.execute_reply.started":"2025-10-18T04:28:05.431572Z","shell.execute_reply":"2025-10-18T04:28:05.534866Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"Purpose: Iterate over each test sequence, generate 5 diverse conformations, and print a progress message.\n\n3- Filling the Submission DataFrame:","metadata":{}},{"cell_type":"code","source":"for i, conformation in enumerate(conformations):\n    # Get rows for this RNA\n    mask = submission['ID'].str.startswith(target_id)\n    \n    # Get sorted indices by residue ID\n    sorted_indices = submission.loc[mask].sort_values('resid').index\n    \n    # Fill coordinates for each residue\n    for j, idx in enumerate(sorted_indices):\n        if j < len(conformation):\n            submission.loc[idx, f'x_{i+1}'] = float(conformation[j][0])\n            submission.loc[idx, f'y_{i+1}'] = float(conformation[j][1])\n            submission.loc[idx, f'z_{i+1}'] = float(conformation[j][2])\n        else:\n            # Just in case we have a mismatch - shouldn't happen\n            submission.loc[idx, f'x_{i+1}'] = float(j * 5.0)\n            submission.loc[idx, f'y_{i+1}'] = 0.0\n            submission.loc[idx, f'z_{i+1}'] = 0.0","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-10-18T04:28:05.538182Z","iopub.execute_input":"2025-10-18T04:28:05.538392Z","iopub.status.idle":"2025-10-18T04:28:05.872462Z","shell.execute_reply.started":"2025-10-18T04:28:05.538375Z","shell.execute_reply":"2025-10-18T04:28:05.871758Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"Purpose: For each conformation, fill the corresponding coordinates in the submission DataFrame. If there is a mismatch in the number of residues, fill with default values.\n\n4- Handling NaN Values:","metadata":{}},{"cell_type":"code","source":"if submission.isna().any().any():\n    print(\"Warning: NaN values detected in submission. Filling with zeros.\")\n    submission = submission.fillna(0.0)","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-10-18T04:28:05.873671Z","iopub.execute_input":"2025-10-18T04:28:05.874006Z","iopub.status.idle":"2025-10-18T04:28:05.878988Z","shell.execute_reply.started":"2025-10-18T04:28:05.873954Z","shell.execute_reply":"2025-10-18T04:28:05.878095Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"Purpose: Check for any NaN values in the submission DataFrame and fill them with zeros if found.\n\n5- Saving the Submission File:","metadata":{}},{"cell_type":"code","source":"submission.to_csv('submission.csv', index=False)\nprint(\"\\nSaved submission file: submission.csv\")","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-10-18T04:28:05.879804Z","iopub.execute_input":"2025-10-18T04:28:05.880025Z","iopub.status.idle":"2025-10-18T04:28:05.929044Z","shell.execute_reply.started":"2025-10-18T04:28:05.879997Z","shell.execute_reply":"2025-10-18T04:28:05.928079Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":" 6. Displaying a Sample of the Submission","metadata":{}},{"cell_type":"code","source":"# Display a sample of the submission\nprint(\"\\nSubmission preview:\")\nprint(submission.head())","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-10-18T04:28:05.929912Z","iopub.execute_input":"2025-10-18T04:28:05.930221Z","iopub.status.idle":"2025-10-18T04:28:05.947748Z","shell.execute_reply.started":"2025-10-18T04:28:05.930174Z","shell.execute_reply":"2025-10-18T04:28:05.946723Z"},"_kg_hide-output":true},"outputs":[],"execution_count":null}]}