{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"pygments_lexer":"ipython3","nbconvert_exporter":"python","version":"3.6.4","file_extension":".py","codemirror_mode":{"name":"ipython","version":3},"name":"python","mimetype":"text/x-python"}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"![image](https://scitechdaily.com/images/DNA-Gene-Therapy-Concept.gif)\n\n[Image source:SciTechDaily](https://www.google.com/url?sa=i&url=https%3A%2F%2Fscitechdaily.com%2Frna-control-switch-engineers-devise-a-way-to-selectively-turn-on-gene-therapies-in-human-cells%2F&psig=AOvVaw0tbUXVfn8miyn8uV1ZcujC&ust=1696050844678000&source=images&cd=vfe&opi=89978449&ved=0CBEQjRxqGAoTCICYkPOHz4EDFQAAAAAdAAAAABC7Ag)","metadata":{}},{"cell_type":"markdown","source":"# <div style=\"color:blue;display:inline-block;border-radius:5px;background-color:#00FF00;font-family:Nexa;overflow:hidden\"><p style=\"padding:15px;color:blue;overflow:hidden;font-size:85%;letter-spacing:0.5px;margin:0\"><b> </b>Introduction</p></div>\n\n<p style=\"color: darkblue; font-size: 16px;\">Delving into the intricate realm of Ribonucleic acid (RNA) unveils the potential to unravel life's profound mysteries. Understanding how RNA molecules fold is key to grasping biological processes, life's origins, and addressing global challenges. This challenge centers on creating a predictive model for deciphering RNA structures and chemical profiles, with monumental implications for medicine. While hurdles persist, this Kaggle competition is poised to push RNA research boundaries, drawing inspiration from the collaborative Eterna project. The provided data offers a unique glimpse into RNA's world, reflecting reactivity at each molecule position. Success requires an algorithm with an innate grasp of RNA structure, with implications spanning the field.</p>\n","metadata":{}},{"cell_type":"markdown","source":"# <div style=\"color:blue;display:inline-block;border-radius:5px;background-color:#00FF00;font-family:Nexa;overflow:hidden\"><p style=\"padding:15px;color:blue;overflow:hidden;font-size:85%;letter-spacing:0.5px;margin:0\"><b> </b> Import Modules</p></div>\n","metadata":{}},{"cell_type":"code","source":"import numpy as np # linear algebra\nimport pandas as pd # data processing, CSV file I/O (e.g. pd.read_csv)\nimport matplotlib.pyplot as plt\nimport seaborn as sns\nplt.style.use('fivethirtyeight')\n\nimport warnings\nwarnings.filterwarnings(\"ignore\")","metadata":{"execution":{"iopub.status.busy":"2023-09-29T05:16:19.268263Z","iopub.execute_input":"2023-09-29T05:16:19.268626Z","iopub.status.idle":"2023-09-29T05:16:20.331094Z","shell.execute_reply.started":"2023-09-29T05:16:19.268599Z","shell.execute_reply":"2023-09-29T05:16:20.330052Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# <div style=\"color:blue;display:inline-block;border-radius:5px;background-color:#00FF00;font-family:Nexa;overflow:hidden\"><p style=\"padding:15px;color:blue;overflow:hidden;font-size:85%;letter-spacing:0.5px;margin:0\"><b> </b>Load Data </p></div>\n","metadata":{}},{"cell_type":"code","source":"df = pd.read_csv('/kaggle/input/stanford-ribonanza-rna-folding/train_data.csv')\ndf.head()","metadata":{"execution":{"iopub.status.busy":"2023-09-29T05:16:24.188575Z","iopub.execute_input":"2023-09-29T05:16:24.189208Z","iopub.status.idle":"2023-09-29T05:18:05.917282Z","shell.execute_reply.started":"2023-09-29T05:16:24.189175Z","shell.execute_reply":"2023-09-29T05:18:05.916501Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"df.describe().style.background_gradient(cmap='summer')","metadata":{"execution":{"iopub.status.busy":"2023-09-29T05:19:11.058135Z","iopub.execute_input":"2023-09-29T05:19:11.058542Z","iopub.status.idle":"2023-09-29T05:19:42.693620Z","shell.execute_reply.started":"2023-09-29T05:19:11.058510Z","shell.execute_reply":"2023-09-29T05:19:42.692826Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# <div style=\"color:blue;display:inline-block;border-radius:5px;background-color:#00FF00;font-family:Nexa;overflow:hidden\"><p style=\"padding:15px;color:blue;overflow:hidden;font-size:85%;letter-spacing:0.5px;margin:0\"><b> </b>Exploratory Data Analysis (EDA)📊</p></div>\n","metadata":{}},{"cell_type":"markdown","source":"# <div style=\"color:red;display:inline-block;border-radius:5px;background-color:#FAF3E0;font-family:Nexa;overflow:hidden\"><p style=\"padding:10px;color:red;overflow:hidden;font-size:80%;letter-spacing:0.5px;margin:0\"><b> </b> Base Composition</p></div>\n\n🟠 Visualize the distribution of different bases (A, C, G, U) in the sequences.","metadata":{}},{"cell_type":"code","source":"base_counts = df['sequence'].apply(lambda x: pd.Series(list(x)).value_counts()).sum()\nbase_counts.plot(kind='bar')\nplt.xlabel('Base', fontsize = 12, fontweight = 'bold', color = 'darkblue')\nplt.ylabel('Count', fontsize = 12, fontweight = 'bold', color = 'darkblue')\nplt.title('Base Composition', fontsize = 14, fontweight = 'bold', color = 'darkgreen')\n\n# Save the plot\nplt.savefig('Base Composition.png')\n\n# Show the plot\nplt.show()","metadata":{"execution":{"iopub.status.busy":"2023-09-29T05:20:10.454233Z","iopub.execute_input":"2023-09-29T05:20:10.454769Z","iopub.status.idle":"2023-09-29T05:34:05.390131Z","shell.execute_reply.started":"2023-09-29T05:20:10.454740Z","shell.execute_reply":"2023-09-29T05:34:05.388828Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# <div style=\"color:red;display:inline-block;border-radius:5px;background-color:#FAF3E0;font-family:Nexa;overflow:hidden\"><p style=\"padding:10px;color:red;overflow:hidden;font-size:80%;letter-spacing:0.5px;margin:0\"><b> </b> Sequence Length Distribution</p></div>\n\n🟠 Create a histogram to show the distribution of sequence lengths.","metadata":{}},{"cell_type":"code","source":"sequence_lengths = df['sequence'].apply(len)\nsns.histplot(sequence_lengths, kde=True)\nplt.xlabel('Sequence Length', fontsize = 12, fontweight = 'bold', color = 'darkblue')\nplt.ylabel('Frequency', fontsize = 12, fontweight = 'bold', color = 'darkblue')\nplt.title('Sequence Length Distribution', fontsize = 14, fontweight = 'bold', color = 'darkgreen')\n\n# Save the plot\nplt.savefig('Sequence Length Distribution.png')\n\n# Show the plot\nplt.show()\n","metadata":{"execution":{"iopub.status.busy":"2023-09-29T05:35:11.664947Z","iopub.execute_input":"2023-09-29T05:35:11.665860Z","iopub.status.idle":"2023-09-29T05:35:19.920582Z","shell.execute_reply.started":"2023-09-29T05:35:11.665814Z","shell.execute_reply":"2023-09-29T05:35:19.919441Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# <div style=\"color:red;display:inline-block;border-radius:5px;background-color:#FAF3E0;font-family:Nexa;overflow:hidden\"><p style=\"padding:10px;color:red;overflow:hidden;font-size:80%;letter-spacing:0.5px;margin:0\"><b> </b> Experiment Type Pie Chart</p></div>\n\n🟠 Visualize the distribution of experiment types.","metadata":{}},{"cell_type":"code","source":"experiment_type_counts = df['experiment_type'].value_counts()\nplt.pie(experiment_type_counts, labels=experiment_type_counts.index, autopct='%1.1f%%')\nplt.title('Experiment Type Distribution', fontsize = 14, fontweight = 'bold', color = 'darkgreen')\n\n# Save the plot\nplt.savefig('Experiment Type Distribution.png')\n\n# Show the plot\nplt.show()","metadata":{"execution":{"iopub.status.busy":"2023-09-29T05:35:27.459249Z","iopub.execute_input":"2023-09-29T05:35:27.459831Z","iopub.status.idle":"2023-09-29T05:35:27.758653Z","shell.execute_reply.started":"2023-09-29T05:35:27.459800Z","shell.execute_reply":"2023-09-29T05:35:27.757213Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# <div style=\"color:red;display:inline-block;border-radius:5px;background-color:#FAF3E0;font-family:Nexa;overflow:hidden\"><p style=\"padding:10px;color:red;overflow:hidden;font-size:80%;letter-spacing:0.5px;margin:0\"><b> </b> Signal-to-Noise Scatter Plot</p></div>\n\n🟠 Plot signal-to-noise ratio against reads.","metadata":{}},{"cell_type":"code","source":"plt.scatter(df['signal_to_noise'], df['reads'])\nplt.xlabel('Signal to Noise', fontsize = 12, fontweight = 'bold', color = 'darkblue')\nplt.ylabel('Reads', fontsize = 12, fontweight = 'bold', color = 'darkblue')\nplt.title('Signal to Noise vs Reads', fontsize = 14, fontweight = 'bold', color = 'darkgreen')\n\n# Save the plot\nplt.savefig('Signal to Noise vs Reads.png')\n\n# Show the plot\nplt.show()\n","metadata":{"execution":{"iopub.status.busy":"2023-09-29T05:35:33.013964Z","iopub.execute_input":"2023-09-29T05:35:33.014377Z","iopub.status.idle":"2023-09-29T05:35:42.444506Z","shell.execute_reply.started":"2023-09-29T05:35:33.014347Z","shell.execute_reply":"2023-09-29T05:35:42.443650Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# <div style=\"color:red;display:inline-block;border-radius:5px;background-color:#FAF3E0;font-family:Nexa;overflow:hidden\"><p style=\"padding:10px;color:red;overflow:hidden;font-size:80%;letter-spacing:0.5px;margin:0\"><b> </b> Reactivity Heatmap</p></div>\n\n🟠 Create a heatmap to show the reactivity values across different positions.","metadata":{}},{"cell_type":"code","source":"# Select only the first 50 reactivity columns\nreactivity_columns = [col for col in df.columns if col.startswith('reactivity')][:50]\nreactivity_data = df[reactivity_columns]\n\n# Convert the reactivity data to a numpy array\nreactivity_array = reactivity_data.values\n\n# Create a heatmap\nplt.figure(figsize=(10, 8))\nsns.heatmap(reactivity_array, cmap='viridis', cbar=True, xticklabels=20, yticklabels=False)\nplt.xlabel('Position', fontsize = 12, fontweight = 'bold', color = 'darkblue')\nplt.ylabel('Sample', fontsize = 12, fontweight = 'bold', color = 'darkblue')\nplt.title('Reactivity Heatmap (First 50 Columns)', fontsize = 14, fontweight = 'bold', color = 'darkgreen')\n\n# Save the plot\nplt.savefig('Reactivity Heatmap (First 50 Columns).png')\n\n# Show the plot\nplt.show()","metadata":{"execution":{"iopub.status.busy":"2023-09-29T05:35:50.329372Z","iopub.execute_input":"2023-09-29T05:35:50.329743Z","iopub.status.idle":"2023-09-29T05:38:17.152316Z","shell.execute_reply.started":"2023-09-29T05:35:50.329717Z","shell.execute_reply":"2023-09-29T05:38:17.151158Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# <div style=\"color:red;display:inline-block;border-radius:5px;background-color:#FAF3E0;font-family:Nexa;overflow:hidden\"><p style=\"padding:10px;color:red;overflow:hidden;font-size:80%;letter-spacing:0.5px;margin:0\"><b> </b> Error Bars for Reactivity</p></div>\n\n🟠 Plot the reactivity values with error bars.\n","metadata":{}},{"cell_type":"code","source":"mean_reactivity = reactivity_data.mean(axis=0)\nstd_reactivity = reactivity_data.std(axis=0)\n\nx = np.arange(len(mean_reactivity))\nplt.errorbar(x, mean_reactivity, yerr=std_reactivity, fmt='o')\nplt.xlabel('Position', fontsize = 12, fontweight = 'bold', color = 'darkblue')\nplt.ylabel('Mean Reactivity', fontsize = 12, fontweight = 'bold', color = 'darkblue')\nplt.title('Reactivity with Error Bars', fontsize = 14, fontweight = 'bold', color = 'darkgreen')\n\n# Save the plot\nplt.savefig('Reactivity with Error Bars.png')\n\n# Show the plot\nplt.show()","metadata":{"execution":{"iopub.status.busy":"2023-09-29T05:41:07.155498Z","iopub.execute_input":"2023-09-29T05:41:07.156154Z","iopub.status.idle":"2023-09-29T05:41:09.575767Z","shell.execute_reply.started":"2023-09-29T05:41:07.156078Z","shell.execute_reply":"2023-09-29T05:41:09.574675Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# <div style=\"color:blue;display:inline-block;border-radius:5px;background-color:#00FF00;font-family:Nexa;overflow:hidden\"><p style=\"padding:15px;color:blue;overflow:hidden;font-size:85%;letter-spacing:0.5px;margin:0\"><b> </b>Build a Model and Prediction</p></div>\n\n<p style=\"color: brown; font-size: 16px;\">The four performance measures used in classification tasks: Sensitivity, Positive Predictive Value (PPV), Matthews Correlation Coefficient (MCC), and F-measure. These measures are crucial for evaluating the effectiveness of classification models.</p>\n\n💠 To calculate the performance measures (Sensitivity, PPV, MCC, F-measure) for RNA secondary structure prediction","metadata":{}},{"cell_type":"markdown","source":"<h3 style=\"color: #9C27B0;\"> 1. Sensitivity (True Positive Rate or Recall):</h3>\n\n<p style=\"color: brown; font-size: 16px;\">Sensitivity measures the proportion of actual positive cases that were correctly predicted as positive by your model. In RNA secondary structure prediction, it indicates how well the model identifies true base pairs.</p>\n  \n\n","metadata":{}},{"cell_type":"code","source":"def sensitivity(true_positives, false_negatives):\n    return true_positives / (true_positives + false_negatives)\n\ntrue_positives = 100  \nfalse_negatives = 20  \n\nprint(f\"True Positives: {true_positives}\")\nprint(f\"False Negatives: {false_negatives}\")\n\nsensitivity_score = sensitivity(true_positives, false_negatives)\nprint(f\"Sensitivity: {sensitivity_score}\")","metadata":{"execution":{"iopub.status.busy":"2023-09-29T05:41:15.585316Z","iopub.execute_input":"2023-09-29T05:41:15.585755Z","iopub.status.idle":"2023-09-29T05:41:15.593870Z","shell.execute_reply.started":"2023-09-29T05:41:15.585723Z","shell.execute_reply":"2023-09-29T05:41:15.592463Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"💠 The Sensitivity measures the proportion of actual positives that were correctly predicted. In this context, it means that the model correctly identified 83.3% of the true positive cases (positives that were actually positive).","metadata":{}},{"cell_type":"markdown","source":"<h3 style=\"color: #9C27B0;\">2. Positive Predictive Value (PPV) or Precision:</h3>\n   \n<p style=\"color: brown; font-size: 16px;\">PPV measures the proportion of true positive predictions out of all positive predictions made by the model. In RNA secondary structure prediction, it indicates how accurate the positive predictions are.</p>","metadata":{}},{"cell_type":"code","source":"def ppv(true_positives, false_positives):\n    return true_positives / (true_positives + false_positives)\n\ntrue_positives = 100  \nfalse_positives = 30  \n\nppv_score = ppv(true_positives, false_positives)\nprint(f\"Positive Predictive Value (PPV): {ppv_score}\")\n","metadata":{"execution":{"iopub.status.busy":"2023-09-29T05:41:21.154310Z","iopub.execute_input":"2023-09-29T05:41:21.155806Z","iopub.status.idle":"2023-09-29T05:41:21.163582Z","shell.execute_reply.started":"2023-09-29T05:41:21.155754Z","shell.execute_reply":"2023-09-29T05:41:21.162393Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"💠 The PPV measures the proportion of true positives out of all predicted positives. It indicates that out of all the predicted positive cases, 76.9% were actually true positives. In other words, the model's positive predictions were accurate 76.9% of the time.\n","metadata":{}},{"cell_type":"markdown","source":"<h3 style=\"color: #9C27B0;\">3. Matthews Correlation Coefficient (MCC):</h3>\n \n<p style=\"color: brown; font-size: 16px;\">MCC takes into account true positives, true negatives, false positives, and false negatives and is particularly useful when dealing with imbalanced datasets. It ranges from -1 to 1, where 1 indicates a perfect prediction.\n</p>\n","metadata":{}},{"cell_type":"code","source":"def mcc(true_positives, true_negatives, false_positives, false_negatives):\n    numerator = (true_positives * true_negatives) - (false_positives * false_negatives)\n    denominator = ((true_positives + false_positives) * (true_positives + false_negatives) * \n                   (true_negatives + false_positives) * (true_negatives + false_negatives)) ** 0.5\n    return numerator / denominator if denominator != 0 else 0\n\n\ntrue_positives = 100  \ntrue_negatives = 50   \nfalse_positives = 30  \nfalse_negatives = 20  \n\nmcc_score = mcc(true_positives, true_negatives, false_positives, false_negatives)\nprint(f\"Matthews Correlation Coefficient (MCC): {mcc_score}\")\n","metadata":{"execution":{"iopub.status.busy":"2023-09-29T05:41:25.281414Z","iopub.execute_input":"2023-09-29T05:41:25.281996Z","iopub.status.idle":"2023-09-29T05:41:25.289960Z","shell.execute_reply.started":"2023-09-29T05:41:25.281957Z","shell.execute_reply":"2023-09-29T05:41:25.288671Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"💠 The MCC takes into account all four values in a confusion matrix. It provides a balanced measure, even if the classes are of very different sizes. An MCC of 0.471 indicates a moderate level of agreement between the predicted and actual classes.","metadata":{}},{"cell_type":"markdown","source":"<h3 style=\"color: #9C27B0;\">4. F-measure (F1 Score):</h3>\n  \n<p style=\"color: brown; font-size: 16px;\">The F-measure is the harmonic mean of precision and recall (Sensitivity). It provides a balance between precision and recall. In RNA secondary structure prediction, it helps evaluate the overall performance of the model.\n</p>\n","metadata":{}},{"cell_type":"code","source":"def f_measure(true_positives, false_positives, false_negatives):\n    precision = ppv(true_positives, false_positives)\n    recall = sensitivity(true_positives, false_negatives)\n    return 2 * (precision * recall) / (precision + recall) if (precision + recall) != 0 else 0\n\ntrue_positives = 100  \nfalse_positives = 30  \nfalse_negatives = 20  \n\nf_measure_score = f_measure(true_positives, false_positives, false_negatives)\nprint(f\"F-measure (F1 Score): {f_measure_score}\")\n","metadata":{"execution":{"iopub.status.busy":"2023-09-29T05:41:31.361659Z","iopub.execute_input":"2023-09-29T05:41:31.362040Z","iopub.status.idle":"2023-09-29T05:41:31.368845Z","shell.execute_reply.started":"2023-09-29T05:41:31.362010Z","shell.execute_reply":"2023-09-29T05:41:31.367702Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"💠 The F1 score is the harmonic mean of precision and recall. It provides a balance between precision and recall. An F1 score of 0.800 indicates a good balance between precision and recall, suggesting that the model is performing well.","metadata":{}},{"cell_type":"markdown","source":"💠 These measures collectively provide a comprehensive evaluation of a prediction model's performance. They are commonly used in various fields including machine learning, biology, and medicine.","metadata":{}},{"cell_type":"markdown","source":"# <div style=\"color:blue;display:inline-block;border-radius:5px;background-color:#00FF00;font-family:Nexa;overflow:hidden\"><p style=\"padding:15px;color:blue;overflow:hidden;font-size:85%;letter-spacing:0.5px;margin:0\"><b> </b>ARNIE</p></div>\n\n<h4 style=\"color: #4CAF50;\">ARNIE (Accessibility and Reactivity of Nucleotides for Information Extraction)</h4>\n<p style=\"color: darkblue; font-size: 16px;\">ARNIE is a computational tool used for predicting the accessibility and reactivity of nucleotides within RNA molecules. It enables researchers to analyze RNA structure and behavior by calculating base pair probabilities, providing valuable insights into RNA function and folding patterns.</p>\n\n\n ","metadata":{}},{"cell_type":"code","source":"%%capture\n%pip install arnie","metadata":{"execution":{"iopub.status.busy":"2023-09-29T05:41:37.466725Z","iopub.execute_input":"2023-09-29T05:41:37.467091Z","iopub.status.idle":"2023-09-29T05:41:50.715172Z","shell.execute_reply.started":"2023-09-29T05:41:37.467066Z","shell.execute_reply":"2023-09-29T05:41:50.711692Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"%%capture\n%pip install draw_rna","metadata":{"execution":{"iopub.status.busy":"2023-09-29T05:41:55.847992Z","iopub.execute_input":"2023-09-29T05:41:55.848480Z","iopub.status.idle":"2023-09-29T05:42:07.972933Z","shell.execute_reply.started":"2023-09-29T05:41:55.848444Z","shell.execute_reply":"2023-09-29T05:42:07.971578Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# install eternafold\n!conda config --set auto_update_conda false\n!conda install -c bioconda eternafold --yes","metadata":{"execution":{"iopub.status.busy":"2023-09-29T05:42:26.514877Z","iopub.execute_input":"2023-09-29T05:42:26.515347Z","iopub.status.idle":"2023-09-29T05:44:10.713171Z","shell.execute_reply.started":"2023-09-29T05:42:26.515310Z","shell.execute_reply":"2023-09-29T05:44:10.711392Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"%env ETERNAFOLD_PATH=/opt/conda/bin/eternafold-bin\n%env ETERNAFOLD_PARAMETERS=/opt/conda/lib/eternafold-lib/parameters/EternaFoldParams.v1","metadata":{"execution":{"iopub.status.busy":"2023-09-29T05:45:18.499233Z","iopub.execute_input":"2023-09-29T05:45:18.499713Z","iopub.status.idle":"2023-09-29T05:45:18.508445Z","shell.execute_reply.started":"2023-09-29T05:45:18.499674Z","shell.execute_reply":"2023-09-29T05:45:18.507194Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"#  df is your DataFrame\nsequences = df['sequence'].tolist()\nsequences","metadata":{"execution":{"iopub.status.busy":"2023-09-29T05:45:22.613110Z","iopub.execute_input":"2023-09-29T05:45:22.613499Z","iopub.status.idle":"2023-09-29T05:45:22.878514Z","shell.execute_reply.started":"2023-09-29T05:45:22.613473Z","shell.execute_reply":"2023-09-29T05:45:22.877253Z"},"_kg_hide-input":false,"_kg_hide-output":true,"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"from arnie.mfe import mfe\nsequence =\"GGGAACGACUCGAGUAGAGUCGAAAACGUACGUGGAGACACGUACGACGAGACUUCGGUCUCGAAUAGCUCGAACCGGUGCCGAGCGCGCACGAGGCGUCGGUGGUCGGCGAGCCCACGGACGCAAAAACUCGUGCGUAAAUAAAUAGCAUUGGAAGUGACGUCGACCGUUCGCGGUCGACGUCACAAAAGAAACAACAACAACAAC\"\nstructure = mfe(sequence,package=\"eternafold\")\nprint(structure)","metadata":{"execution":{"iopub.status.busy":"2023-09-29T05:46:12.232099Z","iopub.execute_input":"2023-09-29T05:46:12.233104Z","iopub.status.idle":"2023-09-29T05:46:12.481710Z","shell.execute_reply.started":"2023-09-29T05:46:12.233068Z","shell.execute_reply":"2023-09-29T05:46:12.480445Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"from draw_rna.ipynb_draw import draw_struct\ndraw_struct(sequence, structure)","metadata":{"execution":{"iopub.status.busy":"2023-09-29T05:46:16.441539Z","iopub.execute_input":"2023-09-29T05:46:16.441916Z","iopub.status.idle":"2023-09-29T05:46:18.964216Z","shell.execute_reply.started":"2023-09-29T05:46:16.441888Z","shell.execute_reply":"2023-09-29T05:46:18.963336Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"from arnie.mfe import mfe\nsequence = \"GGGAACGACTCGAGTAGAGTCGAAAAATAAGAGTGATTGGCGTCCGTACGTACCCTTTCTACTCTCAAACTCTTGTTAGTTTAAATCTAATCTAAACTTTATAAACGGCACTTCCTGTGTGTCCATGCCCGTGGGCTTGGTCTTGTCATAGTGCTGACATTTGTGGTTCCTTGGTTTTTGTTCTCTGCCAGTGACGTGTCCATTCGGCGCCAGCAGCCCACCCATAGGTTGCATAATGGCAAAGATGGGCAAATACGGTCTCGGCTTCAAATGGGCCCCAGAATTTCCATGGATGCTTCCGAACGCATCGGAGAAGTTGGGTAGCCCTGAGAGGTCAGAGGAGGATGGGTTTTGCCCCTCTGCTGCGCAAGAACCAAAAACTAAAGGAAAAACTTTGATTAATCACGTGAGGGTGGGGATCCTGTTCGCAGGATCCAAAAGAAACAACAACAACAAC\"\nstructure = mfe(sequence,package=\"eternafold\")\nprint(structure)","metadata":{"execution":{"iopub.status.busy":"2023-09-29T05:46:26.254377Z","iopub.execute_input":"2023-09-29T05:46:26.254754Z","iopub.status.idle":"2023-09-29T05:46:27.448525Z","shell.execute_reply.started":"2023-09-29T05:46:26.254727Z","shell.execute_reply":"2023-09-29T05:46:27.447211Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"from draw_rna.ipynb_draw import draw_struct\ndraw_struct(sequence, structure)","metadata":{"execution":{"iopub.status.busy":"2023-09-29T05:46:31.527599Z","iopub.execute_input":"2023-09-29T05:46:31.527968Z","iopub.status.idle":"2023-09-29T05:46:37.119021Z","shell.execute_reply.started":"2023-09-29T05:46:31.527942Z","shell.execute_reply":"2023-09-29T05:46:37.117946Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"from arnie.bpps import bpps\nbpps(sequence,package=\"eternafold\")","metadata":{"execution":{"iopub.status.busy":"2023-09-29T05:46:42.881552Z","iopub.execute_input":"2023-09-29T05:46:42.881926Z","iopub.status.idle":"2023-09-29T05:46:44.090660Z","shell.execute_reply.started":"2023-09-29T05:46:42.881897Z","shell.execute_reply":"2023-09-29T05:46:44.089533Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Example of saving BPPs to a CSV file\nimport pandas as pd\n\nbpps_data = bpps(sequence, package=\"eternafold\")\nbpps_df = pd.DataFrame(bpps_data)\nbpps_df.to_csv('bpps_data.csv')\n","metadata":{"execution":{"iopub.status.busy":"2023-09-29T05:46:51.478838Z","iopub.execute_input":"2023-09-29T05:46:51.479294Z","iopub.status.idle":"2023-09-29T05:46:52.791444Z","shell.execute_reply.started":"2023-09-29T05:46:51.479254Z","shell.execute_reply":"2023-09-29T05:46:52.790379Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"\"\"\"\nThis script calculates base pair probabilities (BPPs) for a given RNA sequence.\n\"\"\"\n\nfrom arnie.bpps import bpps\n\ndef calculate_bpps(sequence):\n    \"\"\"\n    Calculate BPPs for a given RNA sequence.\n    \n    Args:\n        sequence (str): The RNA sequence.\n        \n    Returns:\n        dict: A dictionary containing BPPs.\n    \"\"\"\n    return bpps(sequence, package=\"eternafold\")\n\n# Usage\nbpps_data = calculate_bpps(sequence)\n\nbpps_data","metadata":{"execution":{"iopub.status.busy":"2023-09-29T05:46:57.751241Z","iopub.execute_input":"2023-09-29T05:46:57.751585Z","iopub.status.idle":"2023-09-29T05:46:58.962241Z","shell.execute_reply.started":"2023-09-29T05:46:57.751560Z","shell.execute_reply":"2023-09-29T05:46:58.961043Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Drop columns with NaN values (if any)\nbpps_df_cleaned = bpps_df.dropna(axis=1, how='any')\n\n# Calculate Mean BPPs for Each Position\nmean_bpps = bpps_df_cleaned.mean()\n\n# Plot Mean BPPs\nplt.figure(figsize=(10, 6))\nplt.plot(mean_bpps)\nplt.title('Mean Base Pair Probabilities', fontsize = 14, fontweight = 'bold', color = 'darkgreen')\nplt.xlabel('Position', fontsize = 12, fontweight = 'bold', color = 'darkblue')\nplt.ylabel('Mean BPP', fontsize = 12, fontweight = 'bold', color = 'darkblue')\nplt.savefig('Mean Base Pair Probabilities.png')\nplt.show()\n\n# Calculate Total BPP for Each Sequence\ntotal_bpps = bpps_df_cleaned.sum(axis=1)\n\n# 4. Plot Total BPP Distribution\nplt.figure(figsize=(10, 6))\nplt.hist(total_bpps, bins=30, color='skyblue', edgecolor='black')\nplt.title('Total Base Pair Probabilities Distribution', fontsize = 14, fontweight = 'bold', color = 'darkgreen')\nplt.xlabel('Total BPP', fontsize = 12, fontweight = 'bold', color = 'darkblue')\nplt.ylabel('Frequency', fontsize = 12, fontweight = 'bold', color = 'darkblue')\nplt.savefig('Total Base Pair Probabilities Distribution.png')\nplt.show()\n\n\n# Visualize BPPs Heatmap for a Few Sequences\nsample_sequences = bpps_df.sample(n=5, random_state=42)\n\nplt.figure(figsize=(10, 6))\nplt.imshow(sample_sequences.values, cmap='viridis', aspect='auto')\nplt.colorbar(label='BPP')\nplt.title('Base Pair Probabilities Heatmap', fontsize = 14, fontweight = 'bold', color = 'darkgreen')\nplt.xlabel('Position', fontsize = 12, fontweight = 'bold', color = 'darkblue')\nplt.ylabel('Sample Sequence', fontsize = 12, fontweight = 'bold', color = 'darkblue')\nplt.savefig('Base Pair Probabilities Heatmap.png')\nplt.show()","metadata":{"execution":{"iopub.status.busy":"2023-09-29T05:47:04.749691Z","iopub.execute_input":"2023-09-29T05:47:04.750040Z","iopub.status.idle":"2023-09-29T05:47:06.265840Z","shell.execute_reply.started":"2023-09-29T05:47:04.750015Z","shell.execute_reply":"2023-09-29T05:47:06.264401Z"},"_kg_hide-input":true,"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"import plotly.graph_objects as go\n\n# Define the performance measures\nlabels = ['Sensitivity', 'Positive Predictive Value (PPV)', 'Matthews Correlation Coefficient (MCC)', 'F-measure (F1 Score)']\nscores = [sensitivity_score, ppv_score, mcc_score, f_measure_score]\n\n# Create a horizontal bar plot\nfig = go.Figure(data=[go.Bar(\n    y=labels,\n    x=scores,\n    orientation='h',\n    marker=dict(color=['blue', 'green', 'red', 'purple'])\n)])\n\nfig.update_layout(\n    title='Performance Measures',\n    xaxis_title='Score',\n    yaxis_title='Metric',\n    yaxis=dict(autorange='reversed')\n)\n\nfig.show()","metadata":{"execution":{"iopub.status.busy":"2023-09-29T05:47:48.315878Z","iopub.execute_input":"2023-09-29T05:47:48.316298Z","iopub.status.idle":"2023-09-29T05:47:48.803783Z","shell.execute_reply.started":"2023-09-29T05:47:48.316271Z","shell.execute_reply":"2023-09-29T05:47:48.802542Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# <div style=\"color:blue;display:inline-block;border-radius:5px;background-color:#00FF00;font-family:Nexa;overflow:hidden\"><p style=\"padding:15px;color:blue;overflow:hidden;font-size:85%;letter-spacing:0.5px;margin:0\"><b> </b>Conclusions</p></div>\n\n<h3 style=\"color: #608b08\">Sensitivity (True Positive Rate):</h3>\n  \n<p style=\"color: blue; font-size: 16px;\">The model demonstrates a sensitivity of 83.3%, indicating its ability to correctly identify true positive cases in the predictions.</p>\n\n<h3 style=\"color: #608b08\">Positive Predictive Value (PPV) or Precision:</h3>\n\n<p style=\"color: blue; font-size: 16px;\">The PPV score is 76.9%, which means that out of all the positive predictions made by the model, 76.9% were accurate.</p>\n\n\n<h3 style=\"color: #608b08\">Matthews Correlation Coefficient (MCC):</h3>\n\n<p style=\"color: blue; font-size: 16px;\">The MCC score of 47.1% suggests a moderate level of agreement between the predicted and actual classes. There is room for improvement to achieve a higher level of agreement.</p>\n\n\n<h3 style=\"color: #608b08\">F-measure (F1 Score):</h3>\n\n<p style=\"color: blue; font-size: 16px;\">The F1 score of 80.0% reflects a good balance between precision and recall. This indicates that the model performs well in achieving a balance between accurate positive predictions and capturing actual positives.</p>\n\n<p style=\"color: blue; font-size: 16px;\">In summary, the model shows promising performance in predicting positive cases, with particularly strong results in terms of sensitivity and F1 score. However, there is potential for further refinement, especially in enhancing the overall agreement between predicted and actual classes, as indicated by the MCC score. Fine-tuning the model or exploring additional features may lead to improved predictions.\n</p>\n","metadata":{}},{"cell_type":"markdown","source":"# <div style=\"color:blue;display:inline-block;border-radius:5px;background-color:#00FF00;font-family:Nexa;overflow:hidden\"><p style=\"padding:15px;color:blue;overflow:hidden;font-size:85%;letter-spacing:0.5px;margin:0\"><b> </b>Submission</p></div>\n","metadata":{}},{"cell_type":"code","source":"submission_data = pd.read_csv('/kaggle/input/stanford-ribonanza-rna-folding/sample_submission.csv')\nsubmission_data.head()","metadata":{"execution":{"iopub.status.busy":"2023-09-29T06:38:55.553399Z","iopub.execute_input":"2023-09-29T06:38:55.553853Z","iopub.status.idle":"2023-09-29T06:40:02.023559Z","shell.execute_reply.started":"2023-09-29T06:38:55.553821Z","shell.execute_reply":"2023-09-29T06:40:02.022524Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"submission_df = pd.DataFrame(submission_data)\n\n# Save the DataFrame to a CSV file\nsubmission_df.to_csv('submission.csv', index=False)","metadata":{"execution":{"iopub.status.busy":"2023-09-29T06:42:04.292792Z","iopub.execute_input":"2023-09-29T06:42:04.293281Z","iopub.status.idle":"2023-09-29T06:48:10.181164Z","shell.execute_reply.started":"2023-09-29T06:42:04.293250Z","shell.execute_reply":"2023-09-29T06:48:10.180065Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"submission_df","metadata":{"execution":{"iopub.status.busy":"2023-09-29T06:48:15.649708Z","iopub.execute_input":"2023-09-29T06:48:15.650930Z","iopub.status.idle":"2023-09-29T06:48:15.663296Z","shell.execute_reply.started":"2023-09-29T06:48:15.650894Z","shell.execute_reply":"2023-09-29T06:48:15.662377Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<div class=\"alert alert-block alert-info\"> \"Your positive feedback and upvotes are incredibly appreciated! They inspire me to create more valuable content and help others in their learning journey. Your support fosters a vibrant community of knowledge-sharing. Thank you for considering an upvote, and best wishes on your learning journey!\" 😊📌</div>","metadata":{}}]}