{"cells":[{"cell_type":"markdown","id":"766ef24e","metadata":{},"source":"# Ribonanza TM-Score: Basic Analysis\n\nThis notebook performs a basic analysis of the TM-scores provided in the dataset, likely related to the [Stanford Ribonanza RNA Folding competition](https://www.kaggle.com/competitions/stanford-ribonanza-rna-folding).\n\n**Dataset:** [Ribonanza RNA Folding TM-Score Results](https://www.kaggle.com/datasets/dynamo14324/ribonanza-tm-score)\n\n**Goal:** Load the TM-score data and visualize its distribution.\n\n**Sections:**\n1.  Setup and Data Loading\n2.  Score Distribution Analysis\n3.  Conclusion"},{"cell_type":"markdown","id":"0157c0fc","metadata":{},"source":"## 1. Setup and Data Loading\n\nImport libraries and load the TM-score data. We need to determine the file name within the dataset."},{"cell_type":"code","execution_count":null,"id":"d79e4c44","metadata":{},"outputs":[],"source":"import pandas as pd\nimport numpy as np\nimport matplotlib.pyplot as plt\nimport seaborn as sns\nimport os\n\n# Plotting style\nsns.set_style(\"whitegrid\")\nplt.style.use(\"fivethirtyeight\")\n\nprint(\"Libraries imported.\")\n\n# --- Determine the file name --- \n# Option 1: Assume a common name like 'scores.csv' or 'tm_scores.csv'\n# Option 2: List files in the directory (works when running locally, not directly in Kaggle notebook unless data is downloaded first)\nDATA_DIR = \"/kaggle/input/dynamo14324/ribonanza-tm-score\"\n\n# Let's try listing files if possible, otherwise guess common names\nscore_file = None\ntry:\n    files = os.listdir(DATA_DIR)\n    print(f\"Files in dataset directory: {{files}}\")\n    # Heuristic: find the first csv or txt file\n    for f in files:\n        if f.endswith(\".csv\") or f.endswith(\".txt\") or f.endswith(\".py\"): # Added .py based on previous upload\n            score_file = os.path.join(DATA_DIR, f)\n            print(f\"Found potential score file: {{score_file}}\")\n            break\nexcept FileNotFoundError:\n    print(f\"Directory not found: {{DATA_DIR}}. This is expected if run outside Kaggle environment.\")\n    # Fallback guesses if directory listing fails\n    potential_files = [\"tm_scores.csv\", \"scores.csv\", \"results.csv\", \"metric.py\"] # Added metric.py\n    for f_guess in potential_files:\n        if os.path.exists(os.path.join(DATA_DIR, f_guess)): # This check might fail outside Kaggle\n             score_file = os.path.join(DATA_DIR, f_guess)\n             print(f\"Assuming score file: {{score_file}}\")\n             break\n\n# If still not found, assign a default guess for the code to proceed\nif score_file is None:\n    score_file = os.path.join(DATA_DIR, \"metric.py\") # Defaulting to metric.py based on upload log\n    print(f\"Warning: Could not definitively find score file. Assuming: {{score_file}}\")\n\n# Load data (adjust based on actual file format)\ndf_scores = None\nif score_file and os.path.exists(score_file):\n    try:\n        if score_file.endswith(\".csv\"):\n            df_scores = pd.read_csv(score_file)\n            print(f\"Score dataset loaded successfully: {{df_scores.shape[0]}} rows, {{df_scores.shape[1]}} columns\")\n            print(\"First 5 rows:\")\n            display(df_scores.head())\n        elif score_file.endswith(\".py\"):\n             print(f\"Score file is a Python script ({{score_file}}). Cannot load as DataFrame for simple analysis. Analysis might need to be adapted based on script content.\")\n             # Optionally, read the script content\n             # with open(score_file, 'r') as f:\n             #     script_content = f.read()\n             # print(\"\nScript Content:\n\", script_content[:500], \"...\")\n        else:\n            # Add handling for other formats if necessary (e.g., txt, json)\n            print(f\"Loading logic for file type {{score_file.split('.')[-1]}} not implemented.\")\n            \n    except Exception as e:\n        print(f\"Error loading score file {{score_file}}: {{e}}\")\n\nelif score_file:\n     print(f\"Error: Score file specified ({{score_file}}) but not found or accessible in this environment.\")\nelse:\n    print(f\"Error: Score file could not be identified.\")"},{"cell_type":"markdown","id":"32853f3b","metadata":{},"source":"## 2. Score Distribution Analysis\n\nIf the scores were loaded into a DataFrame, let's visualize their distribution."},{"cell_type":"code","execution_count":null,"id":"2288c3e1","metadata":{},"outputs":[],"source":"if df_scores is not None:\n    # --- Identify the score column --- \n    # This requires knowing the column name. Common names: 'score', 'tm_score', 'TM_score', 'value'\n    score_col = None\n    potential_score_cols = [\"tm_score\", \"TM_score\", \"score\", \"Score\", \"value\"]\n    for col in potential_score_cols:\n        if col in df_scores.columns:\n            score_col = col\n            print(f\"Identified score column: {{score_col}}\")\n            break\n            \n    if score_col:\n        print(f\"\nSummary Statistics for {{score_col}}:\")\n        display(df_scores[score_col].describe())\n        \n        # Plot distribution\n        plt.figure(figsize=(10, 6))\n        sns.histplot(df_scores[score_col], kde=True, bins=30)\n        plt.title(f\"Distribution of {{score_col}}\")\n        plt.xlabel(\"TM-Score\")\n        plt.ylabel(\"Frequency\")\n        plt.show()\n    else:\n        print(\"Error: Could not identify the score column in the DataFrame.\")\n        print(f\"Available columns: {{df_scores.columns.tolist()}}\")\nelif score_file and score_file.endswith(\".py\"):\n    print(\"Skipping distribution analysis as data is in a Python script, not a DataFrame.\")\nelse:\n    print(\"Score data not loaded, skipping distribution analysis.\")"},{"cell_type":"markdown","id":"416b0ff5","metadata":{},"source":"## 3. Conclusion\n\nThis notebook attempted to load and analyze the TM-scores from the dataset.\n\n**Observations:**\n*   [Summarize findings, e.g., Data was loaded successfully/unsuccessfully. If successful, describe the score distribution (mean, median, range, shape)].\n*   If the data was a script: \"The dataset appears to contain a Python script (metric.py), likely used for calculating the TM-score itself, rather than a list of pre-calculated scores.\"\n\n**Further Steps:**\n*   If scores were loaded: Compare this distribution to public leaderboards or results from the Ribonanza competition.\n*   If it was a script: Understand how the metric.py script works and potentially use it to score predictions.\n*   Investigate the source of the scores/script for more context."}],"metadata":{},"nbformat":4,"nbformat_minor":5}