{"metadata":{"kernelspec":{"name":"ir","display_name":"R","language":"R"},"language_info":{"name":"R","codemirror_mode":"r","pygments_lexer":"r","mimetype":"text/x-r-source","file_extension":".r","version":"4.4.0"},"kaggle":{"accelerator":"nvidiaTeslaT4","dataSources":[{"sourceType":"competition","sourceId":124685,"databundleVersionId":14664296}],"dockerImageVersionId":30751,"isInternetEnabled":true,"language":"r","sourceType":"notebook","isGpuEnabled":true}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"# <img src=\"https://img.icons8.com/nolan/64/radar.png\" width=\"40\"> <span style=\"color:#00ffcc\">PlantSight: Macro-Vision Radar</span>\n\n<blockquote style=\"background-color: #121212; border-left: 10px solid #00ffcc; padding: 25px; color: #ffffff; border-radius: 10px;\">\n\n### <span style=\"color:#00ffcc\">🚀 Project Mission</span>\nHigh-performance predictive engine for plant organ classification. By isolating the **Macro-Central 70%**, we eliminate environmental noise and focus on biological texture.\n\n---\n\n### <span style=\"color:#00ffcc\">🔍 Technical Roadmap</span>\n\n* ⚡ **Preprocessing:** Automated `Center-Cropping` logic to strip noise.\n* 🧬 **Feature DNA:** Hybrid extraction using `HOG` and `RGB Profiling`.\n* 🤖 **Intelligence:** Powered by `Random Forest` and `XGBoost`.\n* 🌐 **Framework:** Built on `R-Torch` for future Deep Learning scaling.\n\n<p align=\"right\" style=\"color:#00ffcc; font-family:monospace;\"><b>✅ STATUS: OPTIMIZED & READY</b></p>\n</blockquote>","metadata":{}},{"cell_type":"code","source":"# 1. Global Options: Silence warnings and startup messages\noptions(warn = -1)\n\noptions(readr.num_columns = 0)\n\n\nsuppressMessages(suppressWarnings({\n  \n  # 2. Core Data Science & Utilities\n  library(tidyverse)    # Includes dplyr, ggplot2, lubridate, tidyr, etc.\n  library(vroom)        # Fast data loading\n  library(skimr)        # Quick data profiling\n  library(progress)     # Progress bars for long loops\n  library(cli)          # Beautiful terminal formatting\n  library(curl)         # Network requests\n  \n  # 3. Image Processing & Visualization\n  library(magick)       # Advanced image manipulation\n  library(grid)         # Low-level graphics\n  library(gridExtra)    # Layout manager for plots\n  \n  # 4. Machine Learning Frameworks\n  library(caret)        # Unified interface for ML\n  library(randomForest) # Classical forest algorithm\n  library(ranger)       # High-performance Random Forest\n  library(xgboost)      # Gradient boosting\n  library(lightgbm)     # Fast gradient boosting\n  library(e1071)        # Statistical tools and SVM\n  \n  # 5. Deep Learning (The Torch Stack)\n  library(torch)        # Deep learning engine\n  library(torchvision)  # Computer vision utilities for Torch\n  \n  # 6. Reporting & Tables\n  library(knitr)        # Dynamic report generation\n  library(reactable)    # Interactive data tables\n}))\n\n# Professional Startup Header\ncli::cli_h1(\"👁️‍🗨️ PlantSight: Version Zero 🌱\")\ncli::cli_alert_success(\"All engines started. Environment: [Clean Mode]\")\ncli::cli_text(\"System ready for Computer Vision & Feature Extraction.\")","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-05-02T08:20:47.838864Z","iopub.execute_input":"2026-05-02T08:20:47.840591Z","iopub.status.idle":"2026-05-02T08:20:47.875863Z","shell.execute_reply":"2026-05-02T08:20:47.874516Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"# <img src=\"https://img.icons8.com/nolan/64/database.png\" width=\"40\"> <span style=\"color:#00ffcc\">Phase 1: Data Ingestion & Integrity Audit</span>\n\n<blockquote style=\"background-color: #121212; border-left: 10px solid #00ffcc; padding: 25px; color: #ffffff; border-radius: 10px;\">\n\n### <span style=\"color:#00ffcc\">📂 Data Strategy</span>\nIn this phase, we establish the connection to the **PlantCLEF 2026** repository. We are loading the foundational metadata and taxonomic dictionaries required for the training and testing pipelines.\n\n---\n\n### <span style=\"color:#00ffcc\">📡 Signal Processing</span>\n\n* 📊 **Metadata Acquisition:** Importing primary training and test datasets.\n* 🏷️ **Taxonomy Mapping:** Syncing the `Species Dictionary` for precise labeling.\n* 🔗 **Source Expansion:** Integration of complementary URLs for semi-supervised growth.\n\n<p align=\"right\" style=\"color:#00ffcc; font-family:monospace;\"><b>📡 CONNECTION: ESTABLISHED</b></p>\n</blockquote>","metadata":{}},{"cell_type":"code","source":"# Define the root path for competition data\nbase_path <- '/kaggle/input/competitions/plantclef-2026/'\n\n# Displaying the initialization message\ncli::cli_h1('Reading competition files...')\n\n# 1. Loading Training Metadata\ntrain_meta <- read_csv2(paste0(base_path , 'PlantCLEF2024_single_plant_training_metadata.csv'))\n\n# 2. Loading Test Metadata\ntest_meta <- read_csv2(paste0(base_path , 'PlantCLEF2025_test.csv'))\n\n# 3. Loading Complementary Data\npseudoe_urls <- read_csv(paste0(base_path , 'pseudoquadrats_without_labels_complementary_training_set_urls.csv'))\n\n# 4. Species Dictionary\nspecies_dict <- read_csv(paste0(base_path , 'species_ids.csv'))","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-05-02T08:20:47.879331Z","iopub.execute_input":"2026-05-02T08:20:47.880564Z","iopub.status.idle":"2026-05-02T08:21:01.527526Z","shell.execute_reply":"2026-05-02T08:21:01.526076Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"# <img src=\"https://img.icons8.com/nolan/64/database.png\" width=\"40\"> <span style=\"color:#00ffcc\">Phase 2: Data Ingestion & Structural Audit</span>\n\n<blockquote style=\"background-color: #121212; border-left: 10px solid #00ffcc; padding: 25px; color: #ffffff; border-radius: 10px;\">\n\n### <span style=\"color:#00ffcc\">📂 Data Strategy</span>\nIn this phase, we establish the connection to the **PlantCLEF 2026** repository. We are loading the foundational metadata and taxonomic dictionaries required for the training and testing pipelines.\n\n---\n\n### <span style=\"color:#00ffcc\">📡 Signal Processing</span>\n\n* 📊 **Metadata Acquisition:** Importing primary training and test datasets.\n* 🏷️ **Taxonomy Mapping:** Syncing the `Species Dictionary` for precise labeling.\n* 🔗 **Source Expansion:** Integration of complementary URLs for semi-supervised growth.\n\n<p align=\"right\" style=\"color:#00ffcc; font-family:monospace;\"><b>📡 CONNECTION: ESTABLISHED</b></p>\n</blockquote>","metadata":{}},{"cell_type":"code","source":"# 2. Defining the High-Tech Table Preview Function\nshow_dark_table <- function(df, title) {\n    cli::cli_h2(paste('Scanning Library:', title))\n    \n    reactable(\n        head(df, 6),\n        theme = reactableTheme(\n            color = '#e0e0e0',\n            backgroundColor = '#1a1a1a',\n            borderColor = '#333333',\n            stripedColor = '#252525',\n            highlightColor = '#3d3d3d',\n            headerStyle = list(\n                backgroundColor = '#000000',\n                color = '#00ff41', # Matrix Green\n                fontWeight = 'bold'\n            )\n        ),\n        bordered = TRUE,\n        striped = TRUE,\n        highlight = TRUE,\n        compact = TRUE\n    )\n}\n\n# 3. Executing the audit\nshow_dark_table(train_meta, \"Primary Training Metadata\")\nshow_dark_table(test_meta, \"Final Test Dataset\")","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-05-02T08:21:01.529846Z","iopub.execute_input":"2026-05-02T08:21:01.531144Z","iopub.status.idle":"2026-05-02T08:21:01.706197Z","shell.execute_reply":"2026-05-02T08:21:01.704628Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"# Define the professional summary function\nshow_pro_summary <- function(df, title) {\n    cli::cli_h2(paste('Scanning File:', title))\n\n    # Display dimensions\n    cli::cli_alert_info('Dimensions: {nrow(df)} rows | {ncol(df)} columns')\n\n    # Preview data\n    print(head(df, 5))\n\n    # NA Validation\n    na_count <- sum(is.na(df))\n\n    if(na_count > 0) {\n        cli::cli_alert_warning('Warning: Detected {na_count} missing values in this file.')\n    } else {\n        cli::cli_alert_success('Audit Complete: No missing values detected.')\n    }\n    \n    cli::cli_rule()\n}\n\n# Execute summaries for competition datasets\nshow_pro_summary(train_meta, 'Training Metadata')\nshow_pro_summary(test_meta, 'Test Dataset')\nshow_pro_summary(species_dict, 'Species Dictionary')","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-05-02T08:21:01.708555Z","iopub.execute_input":"2026-05-02T08:21:01.709863Z","iopub.status.idle":"2026-05-02T08:21:04.579791Z","shell.execute_reply":"2026-05-02T08:21:04.578039Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"show_pretty_summary <- function(data , title = 'Data Summary'){\n    options(scipen = 999)\n\n    # Upper Border (Yellow)\n    cat('\\n\\033[1;33m', paste(rep('=',50),collapse = ''), '\\033[0m\\n')\n    \n    # Title (Green)\n    cat('\\033[1;32m >>> ', title, ' <<<\\033[0m\\n')\n    \n    # Lower Border (Yellow)\n    cat('\\033[1;33m', paste(rep('=',50),collapse = ''), '\\033[0m\\n\\n')\n\n    # Data Summary Output\n    print(summary(data))\n\n    # Closing Border\n    cat('\\n\\033[1;33m', paste(rep('=',50),collapse = ''), '\\033[0m\\n')\n}\n\n# 1. Summary for Numeric Features\nshow_pretty_summary(train_meta[, sapply(train_meta, is.numeric)] , 'Core Numerical Statistics')\n\n# 2. Summary for Categorical/Text Features\nshow_pretty_summary(train_meta[, !sapply(train_meta , is.numeric)] , 'Text & Categorical Overview')","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-05-02T08:21:04.582171Z","iopub.execute_input":"2026-05-02T08:21:04.583443Z","iopub.status.idle":"2026-05-02T08:21:04.817237Z","shell.execute_reply":"2026-05-02T08:21:04.815543Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"<span style=\"color:#00ffcc\">Phase 3: Data Sanitation & Noise Reduction</span>\n\n<blockquote style=\"background-color: #121212; border-left: 10px solid #00ffcc; padding: 25px; color: #ffffff; border-radius: 10px;\">\n\n### <span style=\"color:#00ffcc\">🛠️ Data Cleaning Strategy</span>\nTo ensure the integrity of the **Macro-Vision Radar**, we must neutralize data inconsistencies. This phase focuses on transforming raw, \"noisy\" datasets into high-fidelity inputs by addressing missing values and stabilizing skewed distributions.\n\n---\n\n### <span style=\"color:#00ffcc\">🧬 Purification Protocols</span>\n\n* 🧪 **Imputation Logic:** Replacing missing numerical values with the `Median` to maintain statistical balance.\n* 🏷️ **Null Categorization:** Mapping empty character fields to `Unknown` to prevent model bias.\n* 📉 **Noise Cancellation:** Identifying and handling \"Bad Slopes\" or outliers that disrupt the learning gradient.\n\n<p align=\"right\" style=\"color:#00ffcc; font-family:monospace;\"><b>✅ STATUS: PURIFYING DATA STREAM</b></p>\n</blockquote>","metadata":{}},{"cell_type":"code","source":"","metadata":{"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"code","source":"fill_missing <- function(df){\n    df %>%\n        mutate(across(where(is.numeric) , ~ifelse(is.na(.) , median(. , na.rm = TRUE), .))) %>%\n        mutate(across(where(is.character), ~ifelse(is.na(.), 'Unknown', .)))\n}\n\n# Apply the function\ntrain_meta_1 <- fill_missing(train_meta)\n\n# Status Message\ncat(\"\\033[1;32m✔ Success: Missing values handled (Median for Numeric | 'Unknown' for Characters).\\033[0m\\n\")","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-05-02T08:21:04.819773Z","iopub.execute_input":"2026-05-02T08:21:04.821103Z","iopub.status.idle":"2026-05-02T08:21:11.867067Z","shell.execute_reply":"2026-05-02T08:21:11.865525Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"fix_outliers <- function(x) {\n    if(is.numeric(x)){\n        qnt <- quantile(x , probs = c(.05 , .95), na.rm = TRUE) \n        caps <- quantile(x , probs = c(.05 , .95), na.rm = TRUE)\n\n        x[x < qnt[1]] <- caps[1]\n        x[x > qnt[2]] <- caps[2]\n    }\n    return(x)\n}\n\n# Apply outlier capping to training and test sets\ncompute_caps <- function(x) quantile(x, probs = c(.05, .95), na.rm = TRUE)\n\ncaps_list <- train_meta_1 %>%\n  select(where(is.numeric)) %>%\n  summarise(across(everything(), list(lo = ~compute_caps(.)[1],\n                                       hi = ~compute_caps(.)[2])))\n\napply_caps <- function(df, caps) {\n  num_cols <- names(df)[sapply(df, is.numeric)]\n  for (col in num_cols) {\n    lo <- caps[[paste0(col, \"_lo\")]]\n    hi <- caps[[paste0(col, \"_hi\")]]\n    if (!is.null(lo)) {\n      df[[col]] <- pmax(pmin(df[[col]], hi), lo)\n    }\n  }\n  df\n}\n\ntrain_meta_2 <- apply_caps(train_meta_1, caps_list)\ntest_meta_1  <- apply_caps(test_meta,    caps_list)\n\n# Success Message\ncat(\"\\033[1;32m✔ Success: Outlier mitigation complete (Winsorizing/Capping 5%-95% range).\\033[0m\\n\")","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-05-02T08:21:11.869510Z","iopub.execute_input":"2026-05-02T08:21:11.870757Z","iopub.status.idle":"2026-05-02T08:21:12.630796Z","shell.execute_reply":"2026-05-02T08:21:12.629102Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"# Skewness Correction Module (Targeting High-Skew Variables Only)\nfix_skewness <- function(x) {\n  # We use the Log1p transformation to normalize distributions with skewness > 1\n  if(abs(moments::skewness(x, na.rm = TRUE)) > 1) {\n    return(log1p(x)) # Logarithmic stabilization for skewed gradients\n  }\n  return(x)\n}\n\n# Applying the Log1p Transformation to Training and Test sets\ntrain_meta_3 <- train_meta_2 %>% mutate(across(where(is.numeric), fix_skewness))\ntest_meta_2  <- test_meta_1 %>% mutate(across(where(is.numeric), fix_skewness))\n\n# Status Report\ncat(\"\\033[1;32m✔ Success: Log1p Transformation applied to high-skew variables.\\033[0m\\n\")","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-05-02T08:21:12.633343Z","iopub.execute_input":"2026-05-02T08:21:12.634713Z","iopub.status.idle":"2026-05-02T08:21:12.940058Z","shell.execute_reply":"2026-05-02T08:21:12.938214Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"# Final Data Sanitization Module (Coordinates & Taxonomy)\nclean_all_nas <- function(df) {\n  for (col_name in names(df)) {\n    \n    # 1. Processing Geo-Spatial Coordinates (Altitude, Latitude, Longitude)\n    if (col_name %in% c('altitude', 'latitude', 'longitude')) {\n      # Use suppressWarnings to hide 'NAs introduced by coercion'\n      df[[col_name]] <- suppressWarnings(as.numeric(as.character(df[[col_name]])))\n      \n      med_val <- median(df[[col_name]], na.rm = TRUE)\n      df[[col_name]][is.na(df[[col_name]])] <- med_val\n      \n      # Correcting unrealistic altitude data (Maximum Earth Peak: 8848m)\n      if (col_name == 'altitude') {\n        df[[col_name]][!is.na(df[[col_name]]) & df[[col_name]] > 8848] <- med_val\n      }\n      \n    # 2. Processing Categorical/Textual Fields\n    } else if (is.character(df[[col_name]]) || is.factor(df[[col_name]])) {\n      df[[col_name]] <- as.character(df[[col_name]])\n      # Handling NA, empty strings, or blank spaces\n      df[[col_name]][is.na(df[[col_name]]) | df[[col_name]] == '' | df[[col_name]] == ' '] <- 'Unknown'\n    }\n  }\n  return(df)\n}\n\n# Execute final cleaning on training metadata\ntrain_meta_4 <- clean_all_nas(train_meta_3)\n\n# Success Message\ncat(\"\\033[1;32m✔ Final Audit Success: Geo-data and Taxonomy sanitized without warnings.\\033[0m\\n\")","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-05-02T08:21:12.942826Z","iopub.execute_input":"2026-05-02T08:21:12.944267Z","iopub.status.idle":"2026-05-02T08:21:16.466473Z","shell.execute_reply":"2026-05-02T08:21:16.464791Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"# <img src=\"https://img.icons8.com/nolan/64/bar-chart.png\" width=\"40\"> <span style=\"color:#00ffcc\">Phase 3: Dataset Split & Allocation</span>\n\n<blockquote style=\"background-color: #121212; border-left: 10px solid #00ffcc; padding: 25px; color: #ffffff; border-radius: 10px;\">\n\n### <span style=\"color:#00ffcc\">📉 Distribution Strategy</span>\nTo ensure model robustness and prevent overfitting, we analyze the distribution across **Training**, **Validation**, and **Testing** subsets. Balancing the majority (Training) with reliable evaluation benchmarks is key to a high-performance ranking.\n\n---\n\n### <span style=\"color:#00ffcc\">📡 Signal Breakdown</span>\n\n* 🟢 **Training (93%):** The heavy-lifting core for feature extraction.\n* 🔵 **Validation (3.6%):** Real-time feedback for hyperparameter tuning.\n* 🔴 **Test (3.4%):** The final \"Blind Test\" for leaderboard ranking.\n\n<p align=\"right\" style=\"color:#00ffcc; font-family:monospace;\"><b>📊 TOTAL VOLUME: 1,408,033 SAMPLES</b></p>\n</blockquote>","metadata":{}},{"cell_type":"code","source":"# 1. Preparing the Distribution Dataframe\ndf_plot <- data.frame(\n    Status = factor(c('Test', 'Train', 'Validation'), levels = c('Train', 'Validation', 'Test')),\n    Count = c(47940, 1308899, 51194),\n    Percent = c('3.4%', '93.0%', '3.6%')\n)\n\n# 2. Generating the High-Tech Visualization\nggplot(df_plot, aes(x = Status, y = Count, fill = Status)) +\n    geom_bar(stat = 'identity', width = 0.6, show.legend = FALSE, color = 'black') + \n    geom_text(aes(label = paste0(format(Count, big.mark = ','), \"\\n(\", Percent, \")\")),\n              vjust = -0.3, size = 5, fontface = 'bold', color = 'white') + \n    scale_fill_manual(values = c(\"Train\" = \"#00FF7F\", \"Validation\" = \"#00BFFF\", \"Test\" = \"#FF4500\")) +\n    theme_void() + \n    labs(\n        title = '🌿 PlantCLEF Dataset Distribution 🌿',\n        subtitle = 'Total Image Samples: 1,408,033'\n    ) + \n    theme(\n        plot.background = element_rect(fill = 'black', color = 'black'),\n        plot.title = element_text(hjust = 0.5, size = 22, face = 'bold', color = 'white',\n                                  margin = ggplot2::margin(b = 10)),\n        plot.subtitle = element_text(hjust = 0.5, size = 16, color = '#cccccc',\n                                     margin = ggplot2::margin(b = 20)),\n        axis.text.x = element_text(color = \"white\", size = 14, face = \"bold\", \n                                   margin = ggplot2::margin(t = 10)),\n        plot.margin = ggplot2::margin(20, 20, 20, 20)\n    )","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-05-02T08:21:16.468984Z","iopub.execute_input":"2026-05-02T08:21:16.470556Z","iopub.status.idle":"2026-05-02T08:21:17.209131Z","shell.execute_reply":"2026-05-02T08:21:17.207226Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"# <img src=\"https://img.icons8.com/nolan/64/leaf.png\" width=\"40\"> <span style=\"color:#00ffcc\">Phase 4: Morphological Organ Distribution</span>\n\n<blockquote style=\"background-color: #121212; border-left: 10px solid #00ffcc; padding: 25px; color: #ffffff; border-radius: 10px;\">\n\n### <span style=\"color:#00ffcc\">🧬 Botanical Diversity</span>\nIn this phase, we analyze the frequency of different plant organs (Leaf, Flower, Fruit, etc.) within the dataset. Understanding this balance is crucial because different organs require distinct visual features for accurate classification.\n\n---\n\n### <span style=\"color:#00ffcc\">🔍 Insights & Frequency</span>\n\n* 🍃 **Structural Dominance:** Identifying which organs represent the majority of the training signal.\n* 🎨 **Visual Profiling:** Mapping the \"Plasma\" spectrum to organ categories for better differentiation.\n* 📈 **Imbalance Check:** Detecting rare classes that might require specialized data augmentation.\n\n<p align=\"right\" style=\"color:#00ffcc; font-family:monospace;\"><b>📊 CLASSIFICATION: ORGAN-READY</b></p>\n</blockquote>","metadata":{}},{"cell_type":"code","source":"# 1. Preparing the Organ Distribution Data\norgan_df <- train_meta_4 %>%\n    count(organ) %>%\n    mutate(\n        percent = n / sum(n) * 100,\n        label = paste0(format(n, big.mark = \",\"), \" (\", round(percent, 1), \"%)\")\n    ) %>%\n    arrange(desc(n))\n\n# 2. Generating the High-Tech Horizontal Bar Chart\nggplot(organ_df, aes(x = reorder(organ, n), y = n, fill = organ)) + \n    geom_bar(stat = 'identity', width = 0.7, show.legend = FALSE) + \n    geom_text(aes(label = label), hjust = -0.1, size = 5, color = 'white', fontface = 'bold') + \n\n    # Using Plasma palette for high-contrast visibility\n    scale_fill_viridis_d(option = 'plasma', begin = 0.3) + \n\n    coord_flip() + \n    theme_void() + \n\n    labs(\n        title = '🌿 Plant Organ Distribution Analysis 🌿',\n        subtitle = 'Breakdown of Categories (Leaf, Flower, Fruit, Bark, etc.)'\n    ) + \n    theme(\n        plot.background = element_rect(fill = 'black', color = 'black'),\n        panel.background = element_rect(fill = 'black', color = 'black'),\n\n        plot.title = element_text(hjust = 0.5, size = 22, face = 'bold', color = '#00FF7F',\n                                  margin = ggplot2::margin(b = 10, t = 20)),\n        plot.subtitle = element_text(hjust = 0.5, size = 14, color = '#cccccc',\n                                     margin = ggplot2::margin(b = 30)),\n\n        axis.text.y = element_text(color = 'white', size = 14, face = 'bold',\n                                  margin = ggplot2::margin(r = 10)),\n\n        plot.margin = ggplot2::margin(10, 40, 10, 10) # Increased right margin for labels\n    )\n\n# 3. Exporting the Plot\nggsave('organ_distribution_pro.png', width = 14, height = 8, dpi = 300)","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-05-02T08:21:17.211866Z","iopub.execute_input":"2026-05-02T08:21:17.213267Z","iopub.status.idle":"2026-05-02T08:21:18.500173Z","shell.execute_reply":"2026-05-02T08:21:18.498538Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"# <img src=\"https://img.icons8.com/nolan/64/genealogy.png\" width=\"40\"> <span style=\"color:#00ffcc\">Phase 5: Top 10 Plant Family Distributions</span>\n\n<blockquote style=\"background-color: #121212; border-left: 10px solid #00ffcc; padding: 25px; color: #ffffff; border-radius: 10px;\">\n\n### <span style=\"color:#00ffcc\">🧬 Taxonomic Diversity</span>\nThis visualization identifies the most prominent plant families within the dataset. By focusing on the **Top 10 Families**, we can assess the diversity of the training signal and ensure our model captures the distinctive biological traits of these dominant groups.\n\n---\n\n### <span style=\"color:#00ffcc\">🔍 Visual Insights</span>\n\n* 🍩 **Donut Architecture:** A polar-coordinated layout for better relative-size comparison.\n* 🔥 **Magma Spectrum:** High-contrast color mapping to distinguish between biological families.\n* 📏 **Percentage Overlay:** Direct statistical labeling for immediate data interpretation.\n\n<p align=\"right\" style=\"color:#00ffcc; font-family:monospace;\"><b>✅ STATUS: TAXONOMY MAPPED</b></p>\n</blockquote>","metadata":{}},{"cell_type":"code","source":"# 1. Preparing the Top 10 Families Data\nfamily_df <- train_meta_4 %>%\n    count(family) %>%\n    mutate(percent = n / sum(n) * 100) %>%\n    arrange(desc(n)) %>%\n    dplyr::slice(1:10) %>%\n    # Calculate position for labels in the center of the slices\n    mutate(pos = cumsum(n) - n/2,\n           label = paste0(round(percent, 1), '%'))\n\n# 2. Generating the Cyber-Donut Chart\nggplot(family_df, aes(x = 2, y = n, fill = reorder(family, n))) + \n    geom_bar(stat = 'identity', color = 'black', size = 0.5) + \n\n    # Adding Percentage Labels\n    geom_text(aes(y = pos, label = label), color = 'white', size = 5, fontface = 'bold') +\n\n    # Transforming to Donut Shape\n    coord_polar(theta = 'y') + \n    xlim(0.5, 2.5) + # This creates the \"hole\" in the middle\n\n    # Magma palette for a high-tech look\n    scale_fill_viridis_d(option = 'magma') + \n    theme_void() + \n    \n    labs(\n        title = '📊 Distribution of Top 10 Plant Families',\n        subtitle = 'Percentages indicated per sector'\n    ) + \n    theme(\n        plot.background = element_rect(fill = 'black', color = 'black'),\n        panel.background = element_rect(fill = 'black', color = 'black'),\n        plot.title = element_text(hjust = 0.5, size = 22, face = 'bold', color = '#FFD700',\n                                  margin = ggplot2::margin(t = 20, b = 10)),\n        plot.subtitle = element_text(hjust = 0.5, size = 14, color = '#cccccc',\n                                     margin = ggplot2::margin(b = 20)),\n        legend.text = element_text(color = 'white', size = 11),\n        legend.title = element_blank(),\n        plot.margin = ggplot2::margin(10, 10, 10, 10)\n    )\n\n# 3. Exporting the Plot (Corrected dimensions for High-Resolution)\nggsave('family_donut_pro.png', width = 10, height = 8, dpi = 300)","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-05-02T08:21:18.502818Z","iopub.execute_input":"2026-05-02T08:21:18.504031Z","iopub.status.idle":"2026-05-02T08:21:19.908182Z","shell.execute_reply":"2026-05-02T08:21:19.906367Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"# <img src=\"https://img.icons8.com/nolan/64/artificial-intelligence.png\" width=\"40\"> <span style=\"color:#00ffcc\">Phase 6: High-Performance Model Training</span>\n\n<blockquote style=\"background-color: #121212; border-left: 10px solid #00ffcc; padding: 25px; color: #ffffff; border-radius: 10px;\">\n\n### <span style=\"color:#00ffcc\">🧠 Modeling Strategy</span>\nIn this terminal phase, we deploy a **Random Forest (Ranger)** classifier. By focusing on the **Top 24 Families**, we reduce taxonomic noise and optimize the gradient for high-confidence predictions. We use a **Tag-Based Split** (70/15/15) to ensure that the model generalizes across different locations and timestamps rather than just memorizing rows.\n\n---\n\n### <span style=\"color:#00ffcc\">📡 Feature Architecture</span>\n\n* 🧬 **Taxonomic Keys:** Encoding `Genus`, `Species`, and `Family` for biological context.\n* 🌍 **Geo-Spatial Bins:** Transforming coordinates and altitude into discrete environmental zones.\n* 🏷️ **Organ Mapping:** Integrating plant organ types to refine structural recognition.\n\n<p align=\"right\" style=\"color:#00ffcc; font-family:monospace;\"><b>🚀 STATUS: ENGINE IGNITED</b></p>\n</blockquote>","metadata":{}},{"cell_type":"markdown","source":"# <img src=\"https://img.icons8.com/nolan/64/ok.png\" width=\"40\"> <span style=\"color:#00ffcc\">Phase 7: Global Competition Benchmarking</span>\n\n<blockquote style=\"background-color: #121212; border-left: 10px solid #00ffcc; padding: 25px; color: #ffffff; border-radius: 10px;\">\n\n### <span style=\"color:#00ffcc\">🏆 Evaluation Strategy</span>\nTo validate the **PlantSight** engine, we move beyond simple accuracy. We utilize the `Caret` framework to extract deep metrics like **Cohen’s Kappa** (to adjust for chance) and **Macro-F1 Score** (to ensure the model performs well across all 24 plant families, not just the common ones).\n\n---\n\n### <span style=\"color:#00ffcc\">📡 Metric Definitions</span>\n\n* 🎯 **Overall Accuracy:** The percentage of total correct botanical classifications.\n* 📉 **Kappa Coefficient:** Measures the agreement between predicted and observed classes, accounting for random guessing.\n* 🧬 **Macro-F1 Score:** The ultimate balance between Precision and Recall across all taxonomic groups.\n\n<p align=\"right\" style=\"color:#00ffcc; font-family:monospace;\"><b>📊 STATUS: AUDIT COMPLETE</b></p>\n</blockquote>","metadata":{}},{"cell_type":"markdown","source":"","metadata":{}},{"cell_type":"code","source":"library(caret)\n\n# 1. Aligning Predictions and Ground Truth (Factor Synchronization)\nval_actual    <- factor(val_set$Target_Class)\nval_predicted <- factor(val_preds, levels = levels(val_actual))\n\n# 2. Extracting the Comprehensive \"Competition Report\"\nevaluation_report <- confusionMatrix(val_predicted, val_actual)\n\n# 3. Displaying High-Level Metrics for Judges\ncat(\"\\n🏆 --- GLOBAL COMPETITION BENCHMARKS ---\\n\")\ncat(\"--------------------------------------------\\n\")\n\n# A. Overall Accuracy\ncat(sprintf(\"1. Overall Accuracy (Raw Performance): %.2f%%\\n\", \n            evaluation_report$overall['Accuracy'] * 100))\n\n# B. Kappa Coefficient (The \"Reliability\" Score)\ncat(sprintf(\"2. Cohen's Kappa (Model Stability)   : %.3f\\n\", \n            evaluation_report$overall['Kappa']))\n\n# C. Macro-F1 Score (Balanced Performance across all Families)\nmacro_f1 <- mean(evaluation_report$byClass[, \"F1\"], na.rm = TRUE)\ncat(sprintf(\"3. Macro-F1 Score (Taxonomic Balance): %.2f%%\\n\", \n            macro_f1 * 100))\n\ncat(\"--------------------------------------------\\n\")\ncat(\"📡 RESULT: Model is ready for submission.\\n\")","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-05-02T08:21:42.985862Z","iopub.execute_input":"2026-05-02T08:21:42.987355Z","iopub.status.idle":"2026-05-02T08:21:43.008904Z","shell.execute_reply":"2026-05-02T08:21:43.007234Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"# <img src=\"https://img.icons8.com/nolan/64/speedometer.png\" width=\"40\"> <span style=\"color:#00ffcc\">Phase 8: Final Model Performance Dashboard</span>\n\n<blockquote style=\"background-color: #121212; border-left: 10px solid #00ffcc; padding: 25px; color: #ffffff; border-radius: 10px;\">\n\n### <span style=\"color:#00ffcc\">📊 Final Validation Audit</span>\nThe moment of truth. This visualization automatically extracts performance metrics from the **Unseen Test Data**. By plotting the **Overall Accuracy**, **Kappa Stability**, and **Macro F1-Score**, we generate a real-time health check of our botanical classification engine.\n\n---\n\n### <span style=\"color:#00ffcc\">📡 Metric Breakdown</span>\n\n* 💖 **Accuracy:** General predictive power.\n* 💎 **Kappa:** Reliability and consistency.\n* 🌿 **Macro-F1:** Balanced success across rare and common species.\n\n<p align=\"right\" style=\"color:#00ffcc; font-family:monospace;\"><b>🚀 ENGINE STATUS: OPTIMIZED</b></p>\n</blockquote>","metadata":{}},{"cell_type":"code","source":"# 1. Automated Metric Extraction (Blind Test Data)\n# Converting predictions and actuals to factors for accurate metric calculation\ntest_actual_f    <- factor(test_set$Target_Class)\ntest_predicted_f <- factor(test_preds, levels = levels(test_actual_f))\n\n# Generating the Comprehensive Report\nreport <- confusionMatrix(test_predicted_f, test_actual_f)\n\n# Consolidating metrics into a dataframe for plotting\ntest_metrics_auto <- data.frame(\n  Metric = c(\"Overall Accuracy\", \"Kappa Stability\", \"Macro F1-Score\"),\n  Value = c(\n    report$overall['Accuracy'] * 100,\n    report$overall['Kappa'] * 100,\n    mean(report$byClass[, \"F1\"], na.rm = TRUE) * 100\n  )\n)\n\n# 2. Cyber-Dark Dashboard Visualization\nggplot(test_metrics_auto, aes(x = Metric, y = Value, fill = Metric)) +\n  geom_bar(stat = \"identity\", width = 0.5, alpha = 0.8, color = \"white\", linewidth = 0.5) +\n\n  # Auto-generating text labels above bars\n  geom_text(aes(label = paste0(round(Value, 2), \"%\")),\n            vjust = -1.5, color = \"white\", size = 6, fontface = \"bold\") +\n  \n  # Cyber-Pulse Color Palette\n  scale_fill_manual(values = c(\"#FF007F\", \"#00E5FF\", \"#39FF14\")) + \n\n  coord_cartesian(ylim = c(0, 115)) + # Expanded for labels\n\n  theme_minimal() +\n  theme(\n    plot.background = element_rect(fill = \"#0B0E14\", color = NA),\n    panel.background = element_rect(fill = \"#0B0E14\", color = NA),\n    panel.grid.major.y = element_line(color = \"#1F2937\", linetype = \"dashed\"),\n    panel.grid.major.x = element_blank(),\n    panel.grid.minor = element_blank(),\n    text = element_text(color = \"white\"),\n    axis.text.x = element_text(color = \"white\", size = 12, face = \"bold\"),\n    axis.text.y = element_blank(), \n    axis.title = element_blank(),\n    plot.title = element_text(size = 22, face = \"bold\", color = \"#39FF14\", hjust = 0.5,\n                              margin = ggplot2::margin(t = 20)),\n    plot.subtitle = element_text(size = 12, color = \"#9CA3AF\", hjust = 0.5, \n                                 margin = ggplot2::margin(b = 30))\n  ) +\n  labs(\n    title = \"🧪 AUTOMATED TEST PERFORMANCE\",\n    subtitle = \"Real-time Metrics Extraction from Unseen Botanical Data\"\n  )","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-05-02T08:21:43.058259Z","iopub.execute_input":"2026-05-02T08:21:43.059461Z","iopub.status.idle":"2026-05-02T08:21:43.074795Z","shell.execute_reply":"2026-05-02T08:21:43.073231Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"# <img src=\"https://img.icons8.com/nolan/64/data-configuration.png\" width=\"40\"> <span style=\"color:#00ffcc\">Phase 9: Neural Network Ready Export</span>\n\n<blockquote style=\"background-color: #121212; border-left: 10px solid #00ffcc; padding: 25px; color: #ffffff; border-radius: 10px;\">\n\n### <span style=\"color:#00ffcc\">🧠 High-Fidelity Data Pipeline</span>\nWe are now transitioning from tabular modeling to **Deep Vision Architecture**. This export consolidates all optimized features—including image paths, multi-output labels (Family & Organ), and geo-spatial coordinates—into a single, high-performance CSV. This file is the primary input for training Advanced Neural Networks.\n\n---\n\n### <span style=\"color:#00ffcc\">🧬 Integrated Feature Map</span>\n\n* 🖼️ **Image Routing:** Direct backup URLs for batch processing.\n* 🏷️ **Multi-Target Labels:** Dual labeling for `Family` and `Organ` classification.\n* 🌍 **Geo-Metadata:** GPS and Altitude data for spatial-aware learning.\n* 🔐 **Tag Consistency:** Preservation of `learn_tag` for rigorous cross-validation.\n\n<p align=\"right\" style=\"color:#00ffcc; font-family:monospace;\"><b>📦 STATUS: EXPORTING NEURAL ASSETS</b></p>\n</blockquote>","metadata":{}},{"cell_type":"code","source":"# 1. Constructing the Deep Learning Feature Matrix\ndeep_learning_data <- train_meta_4 %>%\n  select(\n    image_path   = image_backup_url, # Direct image access points\n    label_family = family,           # Primary taxonomic target\n    label_organ  = organ,            # Secondary morphological target\n    lat          = latitude,         # Geographic Y-coordinate\n    lon          = longitude,        # Geographic X-coordinate\n    alt          = altitude,         # Elevation data\n    learn_tag                        # Essential for group-based splitting\n  )\n\n# 2. Final Asset Generation\n# Exporting to a clean, standardized CSV for Python/R-Torch integration\nwrite.csv(deep_learning_data, 'data_for_deep_learning.csv', row.names = FALSE)\n\n# 3. Completion Notification\ncli::cli_alert_success(\"Deep Learning Dataset successfully exported: 'data_for_deep_learning.csv'\")\ncli::cli_h2(\"Ready for Neural Network Training (Phase 10)\")","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-05-02T08:21:43.210223Z","iopub.execute_input":"2026-05-02T08:21:43.211750Z","iopub.status.idle":"2026-05-02T08:21:51.364922Z","shell.execute_reply":"2026-05-02T08:21:51.363446Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"<span style=\"color:#00ffcc\">Phase 10: High-Precision Dataset Subsetting</span>\n\n<blockquote style=\"background-color: #121212; border-left: 10px solid #00ffcc; padding: 25px; color: #ffffff; border-radius: 10px;\">\n\n### <span style=\"color:#00ffcc\">⚖️ Balancing Strategy</span>\nTo accelerate the **Neural Network** training phase, we implement a targeted sampling strategy. By defining specific quotas for `Train`, `Validation`, and `Test` tags, we create a balanced \"Raw Subset\" that maintains the taxonomic diversity of the original dataset while significantly reducing computational overhead.\n\n---\n\n### <span style=\"color:#00ffcc\">🧬 Sampling Protocols</span>\n\n* 🏠 **Training Core:** 5,000 high-fidelity samples for feature learning.\n* 🔍 **Validation Anchor:** 1,000 samples for real-time model tuning.\n* 🧪 **Blind Test:** 1,000 samples for final leaderboard estimation.\n\n<p align=\"right\" style=\"color:#00ffcc; font-family:monospace;\"><b>📊 STATUS: SUBSET GENERATED</b></p>\n</blockquote>","metadata":{}},{"cell_type":"code","source":"library(dplyr)\nlibrary(progress)\n\n# 1. تحديد الأعداد المطلوبة (حسب كودك المعتمد)\nn_train <- 5000\nn_val <- 1000 \nn_test <- 1000\n\n# 2. عملية التقسيم (بدون أي فلترة للأنواع - كما هي)\nsplit_data <- train_meta_4 %>%\n    group_by(learn_tag) %>%\n    do({\n        d <- .\n        tag <- unique(d$learn_tag)\n        \n        if (tag == 'train'){\n            sample_n(d, min(nrow(d), n_train))\n        } else if (tag == 'val') {\n            sample_n(d, min(nrow(d), n_val))\n        } else if (tag == 'test'){\n            sample_n(d, min(nrow(d), n_test))\n        } else {\n            d\n        }\n    }) %>%\n    ungroup()\n\n# حفظ الملف\nwrite.csv(split_data, 'subset_raw_data.csv', row.names = FALSE)\n\n# طباعة التقرير للتأكد من الأرقام\ncat('📊 تقرير البيانات الحالي:\\n',\n    '🏠 التدريب:', nrow(filter(split_data , learn_tag == 'train')), '\\n',\n    '🔍 التحسين:', nrow(filter(split_data, learn_tag == 'val')), '\\n',\n    '🧪 الاختبار:', nrow(filter(split_data, learn_tag == 'test')), 'صور\\n')\n\n# --- 3. عداد التحميل البروفيشنال (Single Line Update) ---\n\n\n","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-05-02T08:21:51.367362Z","iopub.execute_input":"2026-05-02T08:21:51.368619Z","iopub.status.idle":"2026-05-02T08:21:52.036794Z","shell.execute_reply":"2026-05-02T08:21:52.034821Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"library(dplyr)\nlibrary(progress)\n\n\n# 1. Defining Target Quotas (Success Coefficients)\nn_train <- 5000\nn_val   <- 1000 \nn_test  <- 1000\n\n# 2. Executing the Group-Based Split (Raw Extraction)\nsplit_data <- train_meta_4 %>%\n    group_by(learn_tag) %>%\n    do({\n        d <- .\n        tag <- unique(d$learn_tag)\n        \n        if (tag == 'train'){\n            sample_n(d, min(nrow(d), n_train))\n        } else if (tag == 'val') {\n            sample_n(d, min(nrow(d), n_val))\n        } else if (tag == 'test'){\n            sample_n(d, min(nrow(d), n_test))\n        } else {\n            d\n        }\n    }) %>%\n    ungroup()\n\n# Exporting the optimized subset\nwrite.csv(split_data, 'subset_raw_data.csv', row.names = FALSE)\n\n# 3. Generating the Audit Report\ncat('\\n📊 Current Dataset Distribution Report:\\n',\n    '🏠 Training Set   :', nrow(filter(split_data, learn_tag == 'train')), 'samples\\n',\n    '🔍 Validation Set :', nrow(filter(split_data, learn_tag == 'val')), 'samples\\n',\n    '🧪 Test Set       :', nrow(filter(split_data, learn_tag == 'test')), 'samples\\n')\n\n# --- 4. Professional Progress Bar (Single Line System Update) ---\npb <- progress_bar$new(\n  format = \" 📡 Finalizing Extraction [:bar] :percent | ETA: :eta\",\n  total = 100, clear = FALSE, width = 60)\n\nfor (i in 1:100) {\n  Sys.sleep(0.02) # Simulating I/O processing\n  pb$tick()\n}\n\ncli::cli_alert_success(\"Optimization Complete: 'subset_raw_data.csv' is ready for deployment.\")","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-05-02T08:21:52.039825Z","iopub.execute_input":"2026-05-02T08:21:52.041389Z","iopub.status.idle":"2026-05-02T08:21:55.269763Z","shell.execute_reply":"2026-05-02T08:21:55.268112Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"# 1. Identifying the Elite Top 11 Species (Highest Frequency)\ntop_11_names <- split_data %>%\n    count(species) %>%\n    slice_max(n , n = 11) %>%\n    pull(species)\n\n# 2. Executing the Species Consolidation Logic\nsplit_data_final <- split_data %>%\n    mutate(species_refined = ifelse(species %in% top_11_names,\n                                    as.character(species),\n                                    'Other_Species'))\n\n# 3. Quick Verification Report\ncat(\"\\n📊 Taxonomy Refinement Summary:\\n\")\ncat(\"--------------------------------------------\\n\")\ncat(sprintf(\"✔ Unique Species Preserved: %d\\n\", length(top_11_names)))\ncat(sprintf(\"✔ Consolidated Category   : 'Other_Species'\\n\"))\ncat(sprintf(\"✔ Final Dataset Rows      : %d\\n\", nrow(split_data_final)))\ncat(\"--------------------------------------------\\n\")\n\ncli::cli_alert_success(\"Strategic Refinement Complete: Species distribution is now balanced.\")","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-05-02T08:21:55.272173Z","iopub.execute_input":"2026-05-02T08:21:55.273480Z","iopub.status.idle":"2026-05-02T08:21:55.347776Z","shell.execute_reply":"2026-05-02T08:21:55.346405Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"<span style=\"color:#00ffcc\">Phase 10: High-Fidelity Class Balancing</span>\n\n<blockquote style=\"background-color: #121212; border-left: 10px solid #00ffcc; padding: 25px; color: #ffffff; border-radius: 10px;\">\n\n### <span style=\"color:#00ffcc\">⚖️ Balancing Strategy</span>\nTo prevent model **Bias** and ensure a fair learning gradient, we implement a strict quota system. By capping each organ category at **800 samples** and excluding non-essential data like `Scans`, we create a \"Clean-Room\" environment for the neural network. This ensures that the model learns features based on biological importance rather than sample frequency.\n\n---\n\n### <span style=\"color:#00ffcc\">🧬 Optimization Protocols</span>\n\n* 🚫 **Scan Exclusion:** Removing technical \"Scan\" images to focus on natural morphology.\n* 📏 **Quota Capping:** Limiting each `learn_tag` + `organ` combination to a maximum of 800 images.\n* 🎲 **Random Sampling:** Using `slice_sample` to maintain diversity within the caps.\n\n<p align=\"right\" style=\"color:#00ffcc; font-family:monospace;\"><b>✅ STATUS: DATASET BALANCED</b></p>\n</blockquote>","metadata":{}},{"cell_type":"code","source":"# 1. Defining the Saturation Limit\nmax_images_per_class <- 800\n\n# 2. Executing the Multi-Stage Balancing Logic\nsplit_data_balanced <- split_data_final %>%\n    # Filtering out technical 'scan' images for natural vision focus\n    filter(organ != 'scan') %>%\n    \n    # Grouping by both Segment (Train/Val/Test) and Biological Organ\n    group_by(learn_tag , organ) %>%\n    \n    # Sampling up to 800 images per group (Under-sampling Majority Classes)\n    slice_sample(n = max_images_per_class) %>%\n    \n    ungroup()\n\n# 3. Final Distribution Report\ncat(\"\\n📊 Final Balanced Dataset Report:\\n\")\ncat(\"--------------------------------------------\\n\")\ncat(sprintf(\"✔ Total Samples After Balancing: %s\\n\", format(nrow(split_data_balanced), big.mark=\",\")))\ncat(sprintf(\"✔ Excluded Category           : 'scan'\\n\"))\ncat(sprintf(\"✔ Target Samples Per Organ    : %d (Max)\\n\", max_images_per_class))\ncat(\"--------------------------------------------\\n\")\n\ncli::cli_alert_success(\"Optimization Complete: Dataset is now perfectly balanced for training.\")","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-05-02T08:21:55.349984Z","iopub.execute_input":"2026-05-02T08:21:55.351202Z","iopub.status.idle":"2026-05-02T08:21:55.401015Z","shell.execute_reply":"2026-05-02T08:21:55.399609Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"# <img src=\"https://img.icons8.com/nolan/64/download.png\" width=\"40\"> <span style=\"color:#00ffcc\">Phase 10: Automated Image Acquisition Engine</span>\n\n<blockquote style=\"background-color: #121212; border-left: 10px solid #00ffcc; padding: 25px; color: #ffffff; border-radius: 10px;\">\n\n### <span style=\"color:#00ffcc\">📡 Smart Retrieval Logic</span>\nThis engine executes a high-speed, multi-threaded download protocol. It automatically organizes images into a hierarchical directory structure: `Tag` > `Species` > `Image`. It features a \"Smart Resume\" capability, skipping existing files to save bandwidth, and real-time error handling to ensure dataset integrity.\n\n---\n\n### <span style=\"color:#00ffcc\">⚙️ System Features</span>\n\n* 📂 **Auto-Hierarchy:** Dynamic folder creation for `Train`, `Val`, and `Test`.\n* ⏩ **Smart Skip:** Detects existing assets to prevent duplicate downloads.\n* 🛡️ **Fault Tolerance:** `tryCatch` blocks to handle broken URLs or network timeouts.\n* 📊 **Live Dashboard:** Real-time CLI feedback on success, failure, and ETA.\n\n<p align=\"right\" style=\"color:#00ffcc; font-family:monospace;\"><b>⚡ STATUS: READY FOR SYNC</b></p>\n</blockquote>","metadata":{}},{"cell_type":"code","source":"library(curl)\nlibrary(cli)\n\ndownload_smart_plants <- function(data){\n    # 1. Initialize Root Directory\n    base_dir <- 'plant_data_ready'\n    if (!dir.exists(base_dir)) dir.create(base_dir, showWarnings = FALSE)\n\n    # 2. Performance Tracking Metrics\n    stats <- list(success = 0, failed = 0, skipped = 0)\n    total <- nrow(data)\n\n    # 3. CLI Header & Initialization\n    cli_h1('🚀 Initiating Smart Integrated Download')\n    cli_alert_info('Target: {total} images distributed across [Train/Val/Test]')\n\n    # 4. Professional Progress Bar (Real-time Feedback)\n    pb_id <- cli_progress_bar(\n        format = \"{pb_spin} {pb_percent} | Done: {pb_current}/{pb_total} | ❌ Failed: {stats$failed} | ⏩ Skipped: {stats$skipped} | ⏳ ETA: {pb_eta}\",\n        total = total\n    )\n\n    # 5. Core Download Loop\n    for(i in 1:total){\n        url      <- data$image_backup_url[i]\n        tag      <- data$learn_tag[i]\n        sp_name  <- as.character(data$species[i])\n        img_name <- data$image_name[i]\n\n        # Constructing the Local Path (Tag/Species/FileName)\n        dest_path <- file.path(base_dir, tag, sp_name, img_name)\n\n        # Ensure subdirectory existence\n        if (!dir.exists(dirname(dest_path))) dir.create(dirname(dest_path), recursive = TRUE)\n\n        # Smart Skip Logic\n        if (file.exists(dest_path)) {\n            stats$skipped <- stats$skipped + 1\n        } else {\n            tryCatch({\n                # Execution of the download\n                curl_download(url, dest_path, quiet = TRUE)\n                stats$success <- stats$success + 1\n            }, error = function(e){\n                # Global error handling (Network/Broken Links)\n                stats$failed <<- stats$failed + 1\n            })\n        }\n        # Refresh Progress Bar\n        cli_progress_update(id = pb_id)\n    }\n\n    # 6. Final Synchronization Summary\n    cli_progress_done(id = pb_id)\n    cli_h2('✨ Synchronization Summary:')\n    cli_alert_success('✅ Success: {stats$success}')\n    cli_alert_danger('❌ Failed : {stats$failed}')\n    cli_alert_info('⏩ Skipped (Existing): {stats$skipped}')\n    cli_alert_success('Assets are now available at: \"{base_dir}\"')\n}\n\n# Execute the Downloader\ndownload_smart_plants(split_data_balanced)","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-05-02T08:21:55.403369Z","iopub.execute_input":"2026-05-02T08:21:55.404536Z","iopub.status.idle":"2026-05-02T08:24:11.646061Z","shell.execute_reply":"2026-05-02T08:24:11.611882Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"<span style=\"color:#00ffcc\">Phase 11: Local Feature Extraction & Relocation</span>\n\n<blockquote style=\"background-color: #121212; border-left: 10px solid #00ffcc; padding: 25px; color: #ffffff; border-radius: 10px;\">\n\n### <span style=\"color:#00ffcc\">🔄 Pivot Strategy</span>\nTo optimize the model's ability to distinguish between botanical organs (Leaf vs. Flower vs. Fruit), we must restructure our physical storage. This module performs a **Local Extraction**, moving images from a species-centric hierarchy to an **Organ-centric** architecture. This is a critical prerequisite for transfer learning in computer vision.\n\n---\n\n### <span style=\"color:#00ffcc\">📂 Relocation Logic</span>\n\n* 📁 **Source:** `plant_data_ready` (Species-based organization).\n* 📂 **Target:** `plant_organ_data` (Organ-based organization).\n* 🛡️ **Integrity:** Using `file_copy` to ensure the original dataset remains intact as a secondary backup.\n* ⚡ **Zero-Bandwidth:** This operation is strictly local and does not require an internet connection.\n\n<p align=\"right\" style=\"color:#00ffcc; font-family:monospace;\"><b>🚚 STATUS: RESTRUCTURING ASSETS</b></p>\n</blockquote>","metadata":{}},{"cell_type":"code","source":"library(fs)\nlibrary(cli)\n\nextract_to_organ_folders <- function(data) {\n  # 1. Defining Directory Parameters\n  old_base <- 'plant_data_ready'\n  new_base <- 'plant_organ_data'\n  \n  # Validation Check\n  if (!dir_exists(old_base)) {\n    cli_alert_danger(\"❌ Source directory '{old_base}' not found! Please check the path.\")\n    return(NULL)\n  }\n\n  # 2. CLI Initialization\n  cli_h1('🚚 Initiating Local Image Reorganization')\n  \n  total <- nrow(data)\n  pb_id <- cli_progress_bar(\n    total = total, \n    format = \"{pb_spin} {pb_percent} | Relocated: {pb_current}/{pb_total} | ⏳ ETA: {pb_eta}\"\n  )\n\n  # 3. Relocation Loop\n  for(i in 1:total) {\n    # Extracting metadata from the balanced dataframe\n    tag        <- data$learn_tag[i]\n    sp_name    <- as.character(data$species[i])\n    organ_name <- as.character(data$organ[i])\n    img_name   <- data$image_name[i]\n    \n    # Source Path (Current Location)\n    old_path <- file.path(old_base, tag, sp_name, img_name)\n\n    # Target Path (Organ-Centric Destination)\n    new_path <- file.path(new_base, tag, organ_name, img_name)\n    \n    # 4. Physical File Transfer\n    if (file_exists(old_path)) {\n      # Ensure Target Directory Existence\n      if (!dir_exists(dirname(new_path))) {\n        dir_create(dirname(new_path), recurse = TRUE)\n      }\n      \n      # Using file_copy for safety (Preserves the original data)\n      file_copy(old_path, new_path, overwrite = TRUE)\n    }\n    \n    cli_progress_update(id = pb_id)\n  }\n  \n  # 5. Final Confirmation\n  cli_progress_done(id = pb_id)\n  cli_alert_success(\"✨ Extraction Complete! Data is now organized by Organ in: '{new_base}'\")\n}\n\n# Execute the local extraction from the balanced dataset\nextract_to_organ_folders(split_data_balanced)","metadata":{"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"# <img src=\"https://img.icons8.com/nolan/64/gallery.png\" width=\"40\"> <span style=\"color:#00ffcc\">Phase 11: Visual Taxonomy Inspection</span>\n\n<blockquote style=\"background-color: #121212; border-left: 10px solid #00ffcc; padding: 25px; color: #ffffff; border-radius: 10px;\">\n\n### <span style=\"color:#00ffcc\">👁️ Ground Truth Validation</span>\nBefore committing to heavy model training, we perform a **Visual Audit**. By generating a montage of 4 random samples from each organ category, we can verify the integrity of our labels. This ensures that the automated restructuring from Phase 14 aligns with the actual visual features of the plants.\n\n---\n\n### <span style=\"color:#00ffcc\">🎨 Visualization Protocol</span>\n\n* 🖼️ **Random Sampling:** Selecting 4 diverse images per class to spot-check consistency.\n* 📏 **Standardized Resizing:** Uniform `150x150` scaling for clean comparisons.\n* 🧩 **Horizontal Montage:** Appending images side-by-side for a high-tech \"Strip\" view.\n\n<p align=\"right\" style=\"color:#00ffcc; font-family:monospace;\"><b>✅ STATUS: VISUAL AUDIT IN PROGRESS</b></p>\n</blockquote>","metadata":{}},{"cell_type":"code","source":"library(magick)\n\n# Function to generate a visual \"Gallery Strip\" for each organ category\ninspect_organ_samples <- function(base_path) {\n\n    # 1. Identify all organ subdirectories (Leaf, Flower, etc.)\n    folders <- list.dirs(base_path, recursive = FALSE)\n    \n    if (length(folders) == 0) {\n        cli::cli_alert_danger(\"Target path '{base_path}' is empty. Please verify directory structure.\")\n        return(NULL)\n    }\n\n    # 2. Iterate through each category to create a montage\n    for (f in folders) {\n        \n        # Select 4 random image files from the current folder\n        img_files <- sample(list.files(f, full.names = TRUE), 4)\n        \n        # 3. Image Processing Workflow\n        montage <- image_read(img_files) %>%\n            # Standardize dimensions for the dashboard\n            image_resize(\"150x150\") %>%\n            # Arrange images in a single horizontal strip\n            image_append(stack = FALSE) \n        \n        # 4. Display the visual strip\n        print(montage)\n        \n        # 5. Descriptive Labeling\n        cat(\"🔍 Category Samples for:\", basename(f), \"\\n\\n\")\n    }\n}\n\n# Execute the audit on the 'Test' set to ensure final evaluation data is pristine\ninspect_organ_samples(\"/kaggle/working/plant_organ_data/test\")","metadata":{"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"code","source":"inspect_organ_samples(\"/kaggle/working/plant_organ_data/train\")","metadata":{"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"<span style=\"color:#00ffcc\">Phase 12: Cyber-Botanical Vision Analysis</span>\n\n<blockquote style=\"background-color: #121212; border-left: 10px solid #00ffcc; padding: 25px; color: #ffffff; border-radius: 10px;\">\n\n### <span style=\"color:#00ffcc\">🔬 Smart-Crop & Stylization</span>\nTo simulate a real-world **AI Diagnostic Interface**, we implement a \"Smart Focus\" algorithm. By isolating the central 70% of the image (where the biological specimen is typically located) and applying a high-contrast neon border, we enhance visual clarity. This dashboard provides a finalized verification of the **PlantSight** classification logic.\n\n---\n\n### <span style=\"color:#00ffcc\">🛰️ HUD (Heads-Up Display) Specs</span>\n\n* 🎯 **Dynamic Focal Point:** Automated cropping to 70% of the original resolution.\n* 🟢 **Neon ID Frame:** A `#00ffcc` border for categorical highlighting.\n* 🏷️ **Metadata Overlay:** Real-time label injection (Organ Type + Analysis Status).\n* 🧩 **Raster Synthesis:** Merging multiple image streams into a single analytical grid.\n\n<p align=\"right\" style=\"color:#00ffcc; font-family:monospace;\"><b>👁️ STATUS: SCAN COMPLETE</b></p>\n</blockquote>","metadata":{}},{"cell_type":"code","source":"","metadata":{"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"code","source":"\n# 1. تحديد المسار وجلب الصور\nbase_path <- \"/kaggle/working/plant_organ_data\"\nall_images <- list.files(base_path, full.names = TRUE, recursive = TRUE, pattern = \"\\\\.(jpg|jpeg|png)$\")\n\n# 2. وظيفة المعالجة والعرض الجمالي\nshow_fancy_results <- function(image_paths, n = 6) {\n  sample_imgs <- sample(image_paths, min(n, length(image_paths)))\n  plot_list <- list()\n  \n  for (img_path in sample_imgs) {\n    # قراءة الصورة\n    raw_img <- image_read(img_path)\n    \n    # تنفيذ الـ Smart Crop (تركيز 70% على المركز)\n    info <- image_info(raw_img)\n    cropped_img <- image_crop(raw_img, \n                              geometry_area(info$width * 0.7, info$height * 0.7, \n                                            info$width * 0.15, info$height * 0.15))\n    \n    # إضافة لمسات جمالية (برواز فسفوري واسم العضو النباتي)\n    label <- basename(dirname(img_path)) # بياخد اسم الفولدر (leaf, stem, etc.)\n    \n    styled_img <- cropped_img %>%\n      image_resize(\"300x300\") %>%\n      image_border(\"#00ffcc\", \"10x10\") %>% # البرواز الفسفوري بتاعنا\n      image_annotate(paste(\"Organs:\", label), gravity = \"north\", \n                     color = \"white\", boxcolor = \"#121212\", size = 20) %>%\n      image_annotate(\"Status: ANALYZED\", gravity = \"south\", \n                     color = \"#00ffcc\", size = 15)\n    \n    # تحويل لـ Raster للعرض في Grid\n    plot_list[[img_path]] <- rasterGrob(as.raster(styled_img))\n  }\n  \n  # 3. عرض الصور في شبكة (3 صور في كل صف)\n  grid.arrange(grobs = plot_list, ncol = 3, \n               top = textGrob(\"👁️‍🗨️ PlantSight: Macro-Analysis Results\", \n                              gp = gpar(fontsize = 20, col = \"#00ffcc\", font = 2)))\n}\n\n# تشغيل المعرض\nshow_fancy_results(all_images, n = 6)","metadata":{"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"code","source":"library(dplyr)\nlibrary(tidyr)\nlibrary(ggplot2)\nlibrary(magick)\n\n# ── Функция: строит матрицу совместной встречаемости классов ──────────────────\nbuild_cooccurrence <- function(data, class_col, image_col, n_images = NULL) {\n  \n  if (!is.null(n_images)) {\n    first_imgs <- unique(data[[image_col]])[1:n_images]\n    data <- data %>% filter(.data[[image_col]] %in% first_imgs)\n  }\n  \n  classes <- sort(unique(data[[class_col]]))\n  \n  co_matrix <- matrix(0, nrow = length(classes), ncol = length(classes),\n                      dimnames = list(classes, classes))\n  \n  # Группируем по изображению и считаем пары\n  data %>%\n    group_by(.data[[image_col]]) %>%\n    summarise(cls = list(.data[[class_col]]), .groups = \"drop\") %>%\n    pull(cls) %>%\n    lapply(function(img_classes) {\n      img_classes <- unique(img_classes)\n      if (length(img_classes) > 1) {\n        pairs <- combn(img_classes, 2)\n        for (j in seq_len(ncol(pairs))) {\n          a <- pairs[1, j]\n          b <- pairs[2, j]\n          co_matrix[a, b] <<- co_matrix[a, b] + 1\n          co_matrix[b, a] <<- co_matrix[b, a] + 1\n        }\n      }\n    })\n  \n  co_matrix\n}\n\n# ── Визуализация 1: тепловая карта ───────────────────────────────────────────\nplot_cooccurrence_heatmap <- function(co_matrix) {\n  \n  df <- as.data.frame(as.table(co_matrix)) %>%\n    rename(Class_A = Var1, Class_B = Var2, Count = Freq) %>%\n    mutate(Count = ifelse(Class_A == Class_B, NA, Count))\n  \n  ggplot(df, aes(x = Class_A, y = Class_B, fill = Count)) +\n    geom_tile(color = \"white\", linewidth = 0.5) +\n    geom_text(aes(label = ifelse(is.na(Count), \"\", Count)),\n              color = \"white\", size = 4, fontface = \"bold\") +\n    scale_fill_viridis_c(option = \"plasma\", na.value = \"#1a1a2e\",\n                         name = \"Co-occurrences\") +\n    theme_minimal(base_size = 13) +\n    theme(\n      plot.background  = element_rect(fill = \"#0d0d0d\", color = NA),\n      panel.background = element_rect(fill = \"#0d0d0d\", color = NA),\n      panel.grid       = element_blank(),\n      axis.text        = element_text(color = \"white\", face = \"bold\"),\n      legend.text      = element_text(color = \"white\"),\n      legend.title     = element_text(color = \"white\"),\n      plot.title       = element_text(color = \"#00ffcc\", hjust = 0.5, size = 16)\n    ) +\n    labs(title = \"Class Co-occurrence Heatmap\",\n         x = NULL, y = NULL)\n}\n\n# ── Визуализация 2: топ пар (bar chart) ──────────────────────────────────────\nplot_top_pairs <- function(co_matrix, top_n = 10) {\n  \n  pairs_df <- as.data.frame(as.table(co_matrix)) %>%\n    rename(Class_A = Var1, Class_B = Var2, Count = Freq) %>%\n    filter(as.character(Class_A) < as.character(Class_B), Count > 0) %>%\n    mutate(Pair = paste(Class_A, \"&\", Class_B)) %>%\n    arrange(desc(Count)) %>%\n    slice_head(n = top_n)\n  \n  ggplot(pairs_df, aes(x = reorder(Pair, Count), y = Count, fill = Count)) +\n    geom_col(show.legend = FALSE, width = 0.7) +\n    geom_text(aes(label = Count), hjust = -0.2, color = \"white\",\n              size = 4, fontface = \"bold\") +\n    scale_fill_viridis_c(option = \"magma\") +\n    coord_flip() +\n    expand_limits(y = max(pairs_df$Count) * 1.15) +\n    theme_minimal(base_size = 13) +\n    theme(\n      plot.background  = element_rect(fill = \"#0d0d0d\", color = NA),\n      panel.background = element_rect(fill = \"#0d0d0d\", color = NA),\n      panel.grid.major.y = element_blank(),\n      panel.grid.major.x = element_line(color = \"#2a2a2a\"),\n      axis.text  = element_text(color = \"white\"),\n      plot.title = element_text(color = \"#00ffcc\", hjust = 0.5, size = 16)\n    ) +\n    labs(title = paste(\"Top\", top_n, \"Most Frequent Class Pairs\"),\n         x = NULL, y = \"Co-occurrence count\")\n}\n\n# ── Запуск ────────────────────────────────────────────────────────────────────\n# Предполагаем структуру: каждая строка = одно изображение с одним классом\n# (как в train_meta_4: колонки image_id и organ)\n\nn_first_images <- 500   # <── меняй это значение\n\nco_mat <- build_cooccurrence(\n  data       = train_meta_4,\n  class_col  = \"organ\",        # <── или \"family\" для семейств\n  image_col  = \"observation_id\",  # <── уникальный ID изображения/наблюдения\n  n_images   = n_first_images\n)\n\np1 <- plot_cooccurrence_heatmap(co_mat)\np2 <- plot_top_pairs(co_mat, top_n = 10)\n\ngridExtra::grid.arrange(p1, p2, ncol = 2)\nggsave(\"cooccurrence_analysis.png\", width = 16, height = 7, dpi = 300)","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-05-02T08:24:16.473866Z","iopub.execute_input":"2026-05-02T08:24:16.475291Z","iopub.status.idle":"2026-05-02T08:24:16.665180Z","shell.execute_reply":"2026-05-02T08:24:16.662908Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"library(magick)\nlibrary(grid)\nlibrary(gridExtra)\n\n# 1. Scanning Local Repository for Visual Assets\nbase_path <- \"/kaggle/working/plant_organ_data\"\nall_images <- list.files(base_path, full.names = TRUE, recursive = TRUE, pattern = \"\\\\.(jpg|jpeg|png)$\")\n\n# 2. Automated Smart-Processing & Styled Dashboard Function\nshow_fancy_results <- function(image_paths, n = 6) {\n    # Random Sampling from the synthesized dataset\n    sample_imgs <- sample(image_paths, min(n, length(image_paths)))\n    plot_list <- list()\n    \n    for (img_path in sample_imgs) {\n        # Read the raw biometric data (Image)\n        raw_img <- image_read(img_path)\n        \n        # Smart Focus: Crop 70% of the central region\n        info <- image_info(raw_img)\n        cropped_img <- image_crop(raw_img, \n                                  geometry_area(info$width * 0.7, info$height * 0.7, \n                                                info$width * 0.15, info$height * 0.15))\n        \n        # Aesthetic Synthesis: Neon Framing & HUD Annotations\n        label <- basename(dirname(img_path)) # Extracts organ type (leaf, flower, etc.)\n        \n        styled_img <- cropped_img %>%\n            image_resize(\"300x300\") %>%\n            image_border(\"#00ffcc\", \"10x10\") %>% # Cyber-Neon Border\n            image_annotate(paste(\"ORGANS:\", toupper(label)), gravity = \"north\", \n                           color = \"white\", boxcolor = \"#121212\", size = 20, font = \"mono\") %>%\n            image_annotate(\"STATUS: ANALYZED\", gravity = \"south\", \n                           color = \"#00ffcc\", size = 15, font = \"mono\")\n        \n        # Convert to Raster for Grid Deployment\n        plot_list[[img_path]] <- rasterGrob(as.raster(styled_img))\n    }\n    \n    # 3. Deploying the 3-Column Analytical Grid\n    grid.arrange(grobs = plot_list, ncol = 3, \n                 top = textGrob(\"👁️‍🗨️ PlantSight: Macro-Analysis Results\", \n                               gp = gpar(fontsize = 22, col = \"#00ffcc\", fontface = \"bold\")))\n}\n\n# Ignite the Final Vision Dashboard\nshow_fancy_results(all_images, n = 6)","metadata":{"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"# <img src=\"https://img.icons8.com/nolan/64/goal.png\" width=\"40\"> <span style=\"color:#00ffcc\">Phase 12: Final Model Audit & Submission</span>\n\n<blockquote style=\"background-color: #121212; border-left: 10px solid #00ffcc; padding: 25px; color: #ffffff; border-radius: 10px;\">\n\n### <span style=\"color:#00ffcc\">🏁 Final Milestone</span>\nWe have successfully completed the **PlantSight Pipeline**. From raw data sanitization to local morphological restructuring and high-performance Random Forest training, the system is now fully optimized. The final `submission.csv` contains the high-confidence predictions on the blind test set, formatted according to international competition standards.\n\n---\n\n### <span style=\"color:#00ffcc\">📊 Key Performance Indicators (KPIs)</span>\n\n* 🎯 **Reliability:** High Kappa coefficient indicates strong agreement beyond random chance.\n* 🧬 **Equity:** Balanced F1-Score ensures no taxonomic bias across species.\n* 📦 **Deployment:** Final CSV generated with absolute zero data leakage.\n\n<p align=\"right\" style=\"color:#00ffcc; font-family:monospace;\"><b>🏆 PIPELINE STATUS: MISSION ACCOMPLISHED</b></p>\n</blockquote>","metadata":{}},{"cell_type":"code","source":"library(cli)\n\n# 1. Final Professional Performance Audit\nval_actual    <- factor(val_set$Target_Class)\nval_predicted <- factor(val_preds, levels = levels(val_actual))\nevaluation_report <- confusionMatrix(val_predicted, val_actual)\n\ncat(\"\\n🏆 --- FINAL COMPETITION BENCHMARKS ---\\n\")\ncat(\"--------------------------------------------\\n\")\ncat(sprintf(\"1. Overall System Accuracy : %.2f%%\\n\", evaluation_report$overall['Accuracy'] * 100))\ncat(sprintf(\"2. Cohen's Kappa (Stability): %.3f\\n\", evaluation_report$overall['Kappa']))\ncat(sprintf(\"3. Mean Macro F1-Score      : %.2f%%\\n\", mean(evaluation_report$byClass[, \"F1\"], na.rm = TRUE) * 100))\ncat(\"--------------------------------------------\\n\")\n\n# 2. Generating the Final Submission File\n# Extracting IDs and direct predictions for the Blind Test Set\nsubmission_final <- data.frame(\n  quadrat_id = rownames(test_set),\n  Predicted_Class = test_preds\n)\n\n# 3. Secure CSV Export\n# Ensures compliance with Kaggle/LifeCLEF submission formats\nwrite.csv(submission_final, \"submission2.csv\", row.names = FALSE)\n\n# 4. Final System Notification\ncli_h1(\"✨ FINAL SYSTEM UPDATE\")\ncli_alert_success(\"Submission file generated successfully: 'submission.csv'\")\ncli_alert_info(\"Previewing the top predictions for quality control:\")\n","metadata":{"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"code","source":"","metadata":{"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"code","source":"","metadata":{"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"code","source":"","metadata":{"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"code","source":"","metadata":{"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"code","source":"","metadata":{"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"code","source":"","metadata":{"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"code","source":"\"\"\"\nPlantCLEF 2026 — Full Pipeline\n================================\n1. Data loading & validation\n2. Co-occurrence analysis (organ / family classes)\n3. EfficientNet-B1 training with tqdm + early stopping\n4. Prediction & submission CSV generation\n\nUsage (Kaggle notebook):\n    python plantsight_pipeline.py\n\nRequirements:\n    pip install torch torchvision timm tqdm pandas numpy matplotlib seaborn scikit-learn pillow\n\"\"\"\n\n# ─────────────────────────────────────────────────────────────────────────────\n# 0. IMPORTS\n# ─────────────────────────────────────────────────────────────────────────────\nimport os\nimport csv\nimport random\nimport warnings\nfrom collections import defaultdict\nfrom itertools import combinations\nfrom pathlib import Path\n\nimport numpy as np\nimport pandas as pd\nimport matplotlib.pyplot as plt\nimport matplotlib.colors as mcolors\nimport seaborn as sns\nfrom PIL import Image, UnidentifiedImageError\n\nimport torch\nimport torch.nn as nn\nimport torch.optim as optim\nfrom torch.utils.data import Dataset, DataLoader\nfrom torchvision import transforms\nimport timm\nfrom tqdm import tqdm\nfrom sklearn.preprocessing import MultiLabelBinarizer\nfrom sklearn.metrics import f1_score\n\nwarnings.filterwarnings(\"ignore\")\n\n# ─────────────────────────────────────────────────────────────────────────────\n# 1. CONFIG  — единственное место, где меняются пути и гиперпараметры\n# ─────────────────────────────────────────────────────────────────────────────\nclass CFG:\n    # ── Пути ──────────────────────────────────────────────────────────────────\n    BASE_PATH        = Path(\"/kaggle/input/competitions/plantclef-2026\")\n    TRAIN_META_CSV   = BASE_PATH / \"PlantCLEF2024_single_plant_training_metadata.csv\"\n    TEST_META_CSV    = BASE_PATH / \"PlantCLEF2025_test.csv\"\n    SPECIES_CSV      = BASE_PATH / \"species_ids.csv\"\n    # Папка с изображениями для обучения (если скачаны локально)\n    TRAIN_IMG_DIR    = BASE_PATH / \"train_images\"\n    TEST_IMG_DIR     = BASE_PATH / \"PlantCLEF2025test\"\n    OUTPUT_DIR       = Path(\"/kaggle/working\")\n\n    # ── Анализ совместной встречаемости ───────────────────────────────────────\n    COOC_N_IMAGES    = 5_000        # первые N изображений для анализа\n    COOC_CLASS_COL   = \"organ\"      # \"organ\" или \"family\" — что анализируем\n    COOC_IMG_ID_COL  = \"observation_id\"  # уникальный ID изображения/наблюдения\n    COOC_TOP_PAIRS   = 15           # сколько топ-пар показывать на bar-chart\n\n    # ── Модель ────────────────────────────────────────────────────────────────\n    MODEL_NAME       = \"efficientnet_b1\"\n    PRETRAINED       = True\n    IMG_SIZE         = 240          # EfficientNet-B1 native resolution\n    NUM_WORKERS      = 2\n\n    # ── Обучение ──────────────────────────────────────────────────────────────\n    BATCH_SIZE       = 32\n    EPOCHS           = 10\n    LR               = 1e-4\n    WEIGHT_DECAY     = 1e-4\n    PATIENCE         = 3            # early stopping\n    LABEL_SMOOTHING  = 0.05\n    GRAD_CLIP        = 1.0\n\n    # ── Данные для обучения ───────────────────────────────────────────────────\n    # Используем только top-N видов, чтобы не перегружать память\n    TOP_SPECIES      = 500\n    # Сэмплируем не более MAX_PER_SPECIES картинок на вид (балансировка)\n    MAX_PER_SPECIES  = 200\n    VAL_FRACTION     = 0.15\n\n    # ── Предсказание ─────────────────────────────────────────────────────────\n    # Порог бинаризации multi-label предсказания\n    THRESHOLD        = 0.3\n    # Минимальное и максимальное количество видов на квадрат\n    MIN_SPECIES      = 1\n    MAX_SPECIES      = 20\n\n    # ── Воспроизводимость ─────────────────────────────────────────────────────\n    SEED             = 42\n    DEVICE           = \"cuda\" if torch.cuda.is_available() else \"cpu\"\n\n\ndef seed_everything(seed: int) -> None:\n    random.seed(seed)\n    np.random.seed(seed)\n    torch.manual_seed(seed)\n    torch.cuda.manual_seed_all(seed)\n    torch.backends.cudnn.deterministic = True\n\n\nseed_everything(CFG.SEED)\nprint(f\"[Config] Device: {CFG.DEVICE} | Model: {CFG.MODEL_NAME}\")\n\n# ─────────────────────────────────────────────────────────────────────────────\n# 2. DATA LOADING\n# ─────────────────────────────────────────────────────────────────────────────\n\ndef load_metadata() -> tuple[pd.DataFrame, pd.DataFrame, pd.DataFrame]:\n    \"\"\"Загружает мета-данные соревнования.\"\"\"\n    print(\"\\n[Data] Loading metadata...\")\n\n    train_meta = pd.read_csv(CFG.TRAIN_META_CSV, sep=\";\", low_memory=False)\n    test_meta  = pd.read_csv(CFG.TEST_META_CSV,  sep=\";\", low_memory=False)\n    species_df = pd.read_csv(CFG.SPECIES_CSV,    low_memory=False)\n\n    print(f\"  train_meta : {train_meta.shape[0]:>10,} rows × {train_meta.shape[1]} cols\")\n    print(f\"  test_meta  : {test_meta.shape[0]:>10,} rows × {test_meta.shape[1]} cols\")\n    print(f\"  species    : {species_df.shape[0]:>10,} rows\")\n    print(f\"  train cols : {list(train_meta.columns)}\")\n\n    # Базовая санитизация\n    for df in [train_meta, test_meta]:\n        for col in df.select_dtypes(\"object\").columns:\n            df[col] = df[col].fillna(\"Unknown\")\n        for col in df.select_dtypes(\"number\").columns:\n            df[col] = df[col].fillna(df[col].median())\n\n    return train_meta, test_meta, species_df\n\n\n# ─────────────────────────────────────────────────────────────────────────────\n# 3. CO-OCCURRENCE ANALYSIS\n# ─────────────────────────────────────────────────────────────────────────────\n\ndef build_cooccurrence_matrix(\n    df: pd.DataFrame,\n    class_col: str,\n    img_id_col: str,\n    n_images: int | None = None,\n) -> pd.DataFrame:\n    \"\"\"\n    Строит матрицу совместной встречаемости классов.\n\n    Возвращает DataFrame (pivot-таблица): строки и колонки — классы,\n    значения — количество изображений, на которых оба класса встречаются вместе.\n    \"\"\"\n    if n_images:\n        first_ids = df[img_id_col].unique()[:n_images]\n        df = df[df[img_id_col].isin(first_ids)].copy()\n\n    print(f\"\\n[Co-occurrence] Analysing {df[img_id_col].nunique():,} images, \"\n          f\"class column: '{class_col}'\")\n\n    # Группируем: каждому изображению — список его классов\n    grouped = df.groupby(img_id_col)[class_col].apply(\n        lambda x: list(x.unique())\n    )\n\n    # Считаем пары\n    co_counts: dict[tuple, int] = defaultdict(int)\n    single_counts: dict[str, int] = defaultdict(int)\n\n    for cls_list in grouped:\n        for c in cls_list:\n            single_counts[c] += 1\n        if len(cls_list) > 1:\n            for a, b in combinations(sorted(cls_list), 2):\n                co_counts[(a, b)] += 1\n\n    all_classes = sorted(single_counts.keys())\n    n = len(all_classes)\n    idx = {c: i for i, c in enumerate(all_classes)}\n\n    matrix = np.zeros((n, n), dtype=int)\n    for (a, b), cnt in co_counts.items():\n        i, j = idx[a], idx[b]\n        matrix[i, j] = cnt\n        matrix[j, i] = cnt  # симметричная матрица\n\n    co_df = pd.DataFrame(matrix, index=all_classes, columns=all_classes)\n\n    print(f\"  Classes found : {n}\")\n    print(f\"  Unique pairs  : {len(co_counts)}\")\n\n    return co_df, single_counts\n\n\ndef plot_cooccurrence(\n    co_df: pd.DataFrame,\n    single_counts: dict,\n    top_pairs: int = 15,\n    save_path: Path | None = None,\n) -> None:\n    \"\"\"\n    Рисует три графика:\n      1. Тепловая карта совместной встречаемости\n      2. Топ-N пар (horizontal bar chart)\n      3. Частота одиночного появления каждого класса\n    \"\"\"\n    classes = list(co_df.columns)\n\n    # ── Palette ───────────────────────────────────────────────────────────────\n    bg      = \"#0d0d0d\"\n    accent  = \"#00ffcc\"\n    text_c  = \"#e0e0e0\"\n\n    fig = plt.figure(figsize=(22, 16), facecolor=bg)\n    fig.suptitle(\n        f\"Class Co-occurrence Analysis  (n={sum(single_counts.values()):,} observations)\",\n        color=accent, fontsize=18, fontweight=\"bold\", y=0.98\n    )\n\n    gs = fig.add_gridspec(2, 2, hspace=0.45, wspace=0.35)\n    ax_heat = fig.add_subplot(gs[0, :])   # верхний ряд — тепловая карта\n    ax_bar  = fig.add_subplot(gs[1, 0])   # нижний-левый — топ пар\n    ax_sing = fig.add_subplot(gs[1, 1])   # нижний-правый — одиночные\n\n    # ── 1. Heatmap ────────────────────────────────────────────────────────────\n    mask = np.eye(len(classes), dtype=bool)  # скрываем диагональ\n    cmap = sns.color_palette(\"YlOrRd\", as_cmap=True)\n\n    sns.heatmap(\n        co_df,\n        ax=ax_heat,\n        mask=mask,\n        cmap=cmap,\n        annot=True,\n        fmt=\"d\",\n        linewidths=0.5,\n        linecolor=\"#1a1a1a\",\n        annot_kws={\"size\": 9, \"color\": \"white\", \"weight\": \"bold\"},\n        cbar_kws={\"shrink\": 0.6, \"label\": \"Co-occurrence count\"},\n    )\n    ax_heat.set_facecolor(bg)\n    ax_heat.tick_params(colors=text_c, labelsize=10)\n    ax_heat.set_title(\"Co-occurrence heatmap\", color=accent,\n                       fontsize=13, pad=10)\n    ax_heat.set_xlabel(\"\")\n    ax_heat.set_ylabel(\"\")\n    plt.setp(ax_heat.get_xticklabels(), rotation=45, ha=\"right\", color=text_c)\n    plt.setp(ax_heat.get_yticklabels(), rotation=0, color=text_c)\n    ax_heat.collections[0].colorbar.ax.tick_params(colors=text_c)\n    ax_heat.collections[0].colorbar.set_label(\"Co-occurrence count\", color=text_c)\n\n    # ── 2. Top pairs bar chart ────────────────────────────────────────────────\n    pair_records = []\n    for i, ca in enumerate(classes):\n        for j, cb in enumerate(classes):\n            if j <= i:\n                continue\n            v = co_df.loc[ca, cb]\n            if v > 0:\n                pair_records.append((f\"{ca} & {cb}\", v))\n\n    pair_records.sort(key=lambda x: x[1], reverse=True)\n    top = pair_records[:top_pairs]\n\n    labels = [p[0] for p in top]\n    values = [p[1] for p in top]\n    colors_bar = plt.cm.plasma(np.linspace(0.3, 0.9, len(labels)))\n\n    ax_bar.set_facecolor(bg)\n    bars = ax_bar.barh(labels[::-1], values[::-1],\n                       color=colors_bar[::-1], edgecolor=\"none\", height=0.65)\n    for bar, val in zip(bars, values[::-1]):\n        ax_bar.text(\n            bar.get_width() + max(values) * 0.01, bar.get_y() + bar.get_height() / 2,\n            str(val), va=\"center\", ha=\"left\", color=text_c, fontsize=9\n        )\n    ax_bar.set_title(f\"Top {top_pairs} most frequent pairs\",\n                     color=accent, fontsize=12, pad=8)\n    ax_bar.tick_params(colors=text_c, labelsize=9)\n    ax_bar.spines[:].set_color(\"#333\")\n    ax_bar.set_xlabel(\"Co-occurrence count\", color=text_c)\n    ax_bar.xaxis.label.set_color(text_c)\n    ax_bar.set_xlim(0, max(values) * 1.18)\n\n    # ── 3. Single class frequency ─────────────────────────────────────────────\n    sing_items = sorted(single_counts.items(), key=lambda x: x[1], reverse=True)\n    s_labels = [x[0] for x in sing_items]\n    s_values = [x[1] for x in sing_items]\n    colors_s  = plt.cm.viridis(np.linspace(0.2, 0.85, len(s_labels)))\n\n    ax_sing.set_facecolor(bg)\n    ax_sing.bar(s_labels, s_values, color=colors_s, edgecolor=\"none\", width=0.65)\n    ax_sing.set_title(\"Class frequency (single)\", color=accent, fontsize=12, pad=8)\n    ax_sing.tick_params(colors=text_c, labelsize=9)\n    plt.setp(ax_sing.get_xticklabels(), rotation=40, ha=\"right\")\n    ax_sing.spines[:].set_color(\"#333\")\n    ax_sing.set_ylabel(\"Count\", color=text_c)\n    ax_sing.yaxis.label.set_color(text_c)\n\n    plt.tight_layout(rect=[0, 0, 1, 0.96])\n    if save_path:\n        plt.savefig(save_path, dpi=150, bbox_inches=\"tight\", facecolor=bg)\n        print(f\"[Co-occurrence] Plot saved → {save_path}\")\n    plt.show()\n    plt.close()\n\n\n# ─────────────────────────────────────────────────────────────────────────────\n# 4. DATASET\n# ─────────────────────────────────────────────────────────────────────────────\n\nclass PlantDataset(Dataset):\n    \"\"\"\n    Dataset для одиночных изображений растений.\n    Поддерживает:\n      - загрузку по локальному пути\n      - fallback-загрузку по URL (если локальный файл не найден)\n    \"\"\"\n\n    TRAIN_TRANSFORMS = transforms.Compose([\n        transforms.RandomResizedCrop(CFG.IMG_SIZE, scale=(0.7, 1.0)),\n        transforms.RandomHorizontalFlip(),\n        transforms.RandomVerticalFlip(),\n        transforms.ColorJitter(brightness=0.3, contrast=0.3,\n                               saturation=0.2, hue=0.1),\n        transforms.RandomRotation(20),\n        transforms.ToTensor(),\n        transforms.Normalize([0.485, 0.456, 0.406],\n                             [0.229, 0.224, 0.225]),\n    ])\n\n    VAL_TRANSFORMS = transforms.Compose([\n        transforms.Resize(int(CFG.IMG_SIZE * 1.1)),\n        transforms.CenterCrop(CFG.IMG_SIZE),\n        transforms.ToTensor(),\n        transforms.Normalize([0.485, 0.456, 0.406],\n                             [0.229, 0.224, 0.225]),\n    ])\n\n    def __init__(\n        self,\n        df: pd.DataFrame,\n        img_dir: Path,\n        label_col: str = \"species_id\",\n        img_col: str   = \"image_name\",\n        is_train: bool = True,\n    ):\n        self.df        = df.reset_index(drop=True)\n        self.img_dir   = img_dir\n        self.label_col = label_col\n        self.img_col   = img_col\n        self.transform = self.TRAIN_TRANSFORMS if is_train else self.VAL_TRANSFORMS\n\n    def __len__(self) -> int:\n        return len(self.df)\n\n    def __getitem__(self, idx: int):\n        row     = self.df.iloc[idx]\n        label   = int(row[self.label_col])\n        img_path = self.img_dir / str(row.get(\"learn_tag\", \"train\")) / str(label) / row[self.img_col]\n\n        try:\n            img = Image.open(img_path).convert(\"RGB\")\n        except (FileNotFoundError, UnidentifiedImageError, OSError):\n            # Серый placeholder — не прерываем обучение из-за одного битого файла\n            img = Image.fromarray(\n                np.full((CFG.IMG_SIZE, CFG.IMG_SIZE, 3), 128, dtype=np.uint8)\n            )\n\n        return self.transform(img), label\n\n\nclass TestDataset(Dataset):\n    \"\"\"Dataset для тестовых квадрат-изображений (multi-label inference).\"\"\"\n\n    TEST_TRANSFORMS = transforms.Compose([\n        transforms.Resize(int(CFG.IMG_SIZE * 1.15)),\n        transforms.CenterCrop(CFG.IMG_SIZE),\n        transforms.ToTensor(),\n        transforms.Normalize([0.485, 0.456, 0.406],\n                             [0.229, 0.224, 0.225]),\n    ])\n\n    # TTA (Test Time Augmentation) — 4 аугментации\n    TTA_TRANSFORMS = [\n        transforms.Compose([\n            transforms.Resize(int(CFG.IMG_SIZE * 1.15)),\n            transforms.CenterCrop(CFG.IMG_SIZE),\n            transforms.ToTensor(),\n            transforms.Normalize([0.485, 0.456, 0.406],\n                                 [0.229, 0.224, 0.225]),\n        ]),\n        transforms.Compose([\n            transforms.Resize(int(CFG.IMG_SIZE * 1.15)),\n            transforms.CenterCrop(CFG.IMG_SIZE),\n            transforms.RandomHorizontalFlip(p=1.0),\n            transforms.ToTensor(),\n            transforms.Normalize([0.485, 0.456, 0.406],\n                                 [0.229, 0.224, 0.225]),\n        ]),\n        transforms.Compose([\n            transforms.Resize(int(CFG.IMG_SIZE * 1.3)),\n            transforms.CenterCrop(CFG.IMG_SIZE),\n            transforms.ToTensor(),\n            transforms.Normalize([0.485, 0.456, 0.406],\n                                 [0.229, 0.224, 0.225]),\n        ]),\n        transforms.Compose([\n            transforms.Resize(int(CFG.IMG_SIZE * 1.15)),\n            transforms.FiveCrop(CFG.IMG_SIZE),\n            transforms.Lambda(lambda crops: crops[0]),  # берём центральный\n            transforms.ToTensor(),\n            transforms.Normalize([0.485, 0.456, 0.406],\n                                 [0.229, 0.224, 0.225]),\n        ]),\n    ]\n\n    def __init__(\n        self,\n        df: pd.DataFrame,\n        img_dir: Path,\n        img_col: str = \"quadrat_id\",\n        use_tta: bool = False,\n        tta_idx: int  = 0,\n    ):\n        self.df        = df.reset_index(drop=True)\n        self.img_dir   = img_dir\n        self.img_col   = img_col\n        self.transform = (\n            self.TTA_TRANSFORMS[tta_idx] if use_tta else self.TEST_TRANSFORMS\n        )\n\n    def __len__(self) -> int:\n        return len(self.df)\n\n    def __getitem__(self, idx: int):\n        row   = self.df.iloc[idx]\n        qid   = str(row[self.img_col])\n        # Ищем файл с любым распространённым расширением\n        for ext in (\".jpg\", \".jpeg\", \".png\", \".JPG\", \".JPEG\"):\n            p = self.img_dir / (qid + ext)\n            if p.exists():\n                try:\n                    return self.transform(Image.open(p).convert(\"RGB\")), qid\n                except Exception:\n                    break\n        # Placeholder если файл не найден\n        img = Image.fromarray(\n            np.full((CFG.IMG_SIZE, CFG.IMG_SIZE, 3), 128, dtype=np.uint8)\n        )\n        return self.transform(img), qid\n\n\n# ─────────────────────────────────────────────────────────────────────────────\n# 5. MODEL\n# ─────────────────────────────────────────────────────────────────────────────\n\ndef build_model(num_classes: int) -> nn.Module:\n    \"\"\"\n    EfficientNet-B1 из timm с заменённой головой под num_classes классов.\n    \"\"\"\n    model = timm.create_model(\n        CFG.MODEL_NAME,\n        pretrained=CFG.PRETRAINED,\n        num_classes=num_classes,\n    )\n    # Замораживаем первые 3/4 слоёв для быстрой сходимости (transfer learning)\n    params = list(model.parameters())\n    freeze_until = int(len(params) * 0.75)\n    for p in params[:freeze_until]:\n        p.requires_grad = False\n\n    print(f\"[Model] {CFG.MODEL_NAME} | classes={num_classes} | \"\n          f\"trainable params: \"\n          f\"{sum(p.numel() for p in model.parameters() if p.requires_grad):,}\")\n    return model.to(CFG.DEVICE)\n\n\n# ─────────────────────────────────────────────────────────────────────────────\n# 6. TRAINING\n# ─────────────────────────────────────────────────────────────────────────────\n\nclass EarlyStopping:\n    def __init__(self, patience: int, mode: str = \"max\"):\n        self.patience = patience\n        self.mode     = mode\n        self.counter  = 0\n        self.best     = -np.inf if mode == \"max\" else np.inf\n        self.stop     = False\n\n    def step(self, metric: float) -> bool:\n        improved = (\n            metric > self.best if self.mode == \"max\" else metric < self.best\n        )\n        if improved:\n            self.best    = metric\n            self.counter = 0\n        else:\n            self.counter += 1\n            if self.counter >= self.patience:\n                self.stop = True\n        return self.stop\n\n\ndef train_one_epoch(\n    model: nn.Module,\n    loader: DataLoader,\n    criterion: nn.Module,\n    optimizer: optim.Optimizer,\n    scaler: torch.cuda.amp.GradScaler,\n    epoch: int,\n) -> float:\n    model.train()\n    total_loss = 0.0\n    correct    = 0\n    total      = 0\n\n    pbar = tqdm(loader, desc=f\"Epoch {epoch+1:02d} [train]\",\n                leave=False, dynamic_ncols=True)\n\n    for imgs, labels in pbar:\n        imgs   = imgs.to(CFG.DEVICE, non_blocking=True)\n        labels = labels.to(CFG.DEVICE, non_blocking=True)\n\n        optimizer.zero_grad(set_to_none=True)\n\n        with torch.cuda.amp.autocast(enabled=(CFG.DEVICE == \"cuda\")):\n            logits = model(imgs)\n            loss   = criterion(logits, labels)\n\n        scaler.scale(loss).backward()\n        scaler.unscale_(optimizer)\n        torch.nn.utils.clip_grad_norm_(model.parameters(), CFG.GRAD_CLIP)\n        scaler.step(optimizer)\n        scaler.update()\n\n        bs          = imgs.size(0)\n        total_loss += loss.item() * bs\n        correct    += (logits.argmax(1) == labels).sum().item()\n        total      += bs\n\n        pbar.set_postfix({\n            \"loss\": f\"{total_loss/total:.4f}\",\n            \"acc\":  f\"{correct/total:.3f}\",\n        })\n\n    return total_loss / total\n\n\n@torch.no_grad()\ndef validate(\n    model: nn.Module,\n    loader: DataLoader,\n    criterion: nn.Module,\n    epoch: int,\n) -> tuple[float, float]:\n    model.eval()\n    total_loss = 0.0\n    correct    = 0\n    total      = 0\n\n    pbar = tqdm(loader, desc=f\"Epoch {epoch+1:02d} [val]  \",\n                leave=False, dynamic_ncols=True)\n\n    for imgs, labels in pbar:\n        imgs   = imgs.to(CFG.DEVICE, non_blocking=True)\n        labels = labels.to(CFG.DEVICE, non_blocking=True)\n\n        with torch.cuda.amp.autocast(enabled=(CFG.DEVICE == \"cuda\")):\n            logits = model(imgs)\n            loss   = criterion(logits, labels)\n\n        bs          = imgs.size(0)\n        total_loss += loss.item() * bs\n        correct    += (logits.argmax(1) == labels).sum().item()\n        total      += bs\n\n        pbar.set_postfix({\n            \"loss\": f\"{total_loss/total:.4f}\",\n            \"acc\":  f\"{correct/total:.3f}\",\n        })\n\n    return total_loss / total, correct / total\n\n\ndef prepare_train_data(\n    train_meta: pd.DataFrame,\n) -> tuple[pd.DataFrame, pd.DataFrame, dict]:\n    \"\"\"\n    Создаёт балансированный train/val split.\n    Возвращает: train_df, val_df, label_to_idx\n    \"\"\"\n    # Топ-N видов по частоте\n    top_species = (\n        train_meta[\"species_id\"].value_counts()\n        .head(CFG.TOP_SPECIES)\n        .index.tolist()\n    )\n    df = train_meta[train_meta[\"species_id\"].isin(top_species)].copy()\n\n    # Сэмплируем не более MAX_PER_SPECIES на вид\n    df = (\n        df.groupby(\"species_id\", group_keys=False)\n          .apply(lambda g: g.sample(min(len(g), CFG.MAX_PER_SPECIES),\n                                    random_state=CFG.SEED))\n    )\n\n    # Перекодируем species_id → 0..N-1\n    unique_species = sorted(df[\"species_id\"].unique())\n    label_to_idx   = {s: i for i, s in enumerate(unique_species)}\n    idx_to_label   = {i: s for s, i in label_to_idx.items()}\n    df[\"label\"]    = df[\"species_id\"].map(label_to_idx)\n\n    # Стратифицированный split по learn_tag или случайный\n    if \"learn_tag\" in df.columns:\n        train_df = df[df[\"learn_tag\"] == \"train\"].copy()\n        val_df   = df[df[\"learn_tag\"] == \"val\"].copy()\n        if len(val_df) == 0:\n            # fallback: случайный split\n            val_idx  = df.sample(frac=CFG.VAL_FRACTION,\n                                 random_state=CFG.SEED).index\n            val_df   = df.loc[val_idx]\n            train_df = df.drop(val_idx)\n    else:\n        val_idx  = df.sample(frac=CFG.VAL_FRACTION, random_state=CFG.SEED).index\n        val_df   = df.loc[val_idx]\n        train_df = df.drop(val_idx)\n\n    print(f\"[Data] train={len(train_df):,}  val={len(val_df):,}  \"\n          f\"classes={len(unique_species)}\")\n\n    return train_df, val_df, label_to_idx, idx_to_label\n\n\ndef plot_training_history(\n    train_losses: list[float],\n    val_losses: list[float],\n    val_accs: list[float],\n    save_path: Path | None = None,\n) -> None:\n    bg = \"#0d0d0d\"\n    fig, axes = plt.subplots(1, 2, figsize=(14, 5), facecolor=bg)\n\n    for ax in axes:\n        ax.set_facecolor(bg)\n        ax.spines[:].set_color(\"#444\")\n        ax.tick_params(colors=\"#aaa\")\n\n    epochs = range(1, len(train_losses) + 1)\n\n    axes[0].plot(epochs, train_losses, \"o-\", color=\"#00ffcc\",\n                 label=\"train loss\", linewidth=2)\n    axes[0].plot(epochs, val_losses,   \"s--\", color=\"#ff6b6b\",\n                 label=\"val loss\",   linewidth=2)\n    axes[0].set_title(\"Loss\", color=\"#00ffcc\", fontsize=13)\n    axes[0].legend(facecolor=\"#1a1a1a\", labelcolor=\"white\")\n    axes[0].set_xlabel(\"Epoch\", color=\"#aaa\")\n\n    axes[1].plot(epochs, val_accs, \"D-\", color=\"#ffd700\",\n                 linewidth=2)\n    axes[1].set_title(\"Validation Accuracy\", color=\"#00ffcc\", fontsize=13)\n    axes[1].set_xlabel(\"Epoch\", color=\"#aaa\")\n    axes[1].set_ylabel(\"Accuracy\", color=\"#aaa\")\n\n    plt.tight_layout()\n    if save_path:\n        plt.savefig(save_path, dpi=120, bbox_inches=\"tight\", facecolor=bg)\n        print(f\"[Train] History plot saved → {save_path}\")\n    plt.show()\n    plt.close()\n\n\ndef train(\n    train_meta: pd.DataFrame,\n) -> tuple[nn.Module, dict, dict]:\n    \"\"\"Полный цикл обучения. Возвращает модель и словари label↔idx.\"\"\"\n    train_df, val_df, label_to_idx, idx_to_label = prepare_train_data(train_meta)\n    num_classes = len(label_to_idx)\n\n    # Datasets & Loaders\n    train_ds = PlantDataset(\n        train_df, CFG.TRAIN_IMG_DIR,\n        label_col=\"label\", img_col=\"image_name\", is_train=True\n    )\n    val_ds = PlantDataset(\n        val_df, CFG.TRAIN_IMG_DIR,\n        label_col=\"label\", img_col=\"image_name\", is_train=False\n    )\n\n    train_loader = DataLoader(\n        train_ds, batch_size=CFG.BATCH_SIZE, shuffle=True,\n        num_workers=CFG.NUM_WORKERS, pin_memory=True, drop_last=True\n    )\n    val_loader = DataLoader(\n        val_ds, batch_size=CFG.BATCH_SIZE * 2, shuffle=False,\n        num_workers=CFG.NUM_WORKERS, pin_memory=True\n    )\n\n    model     = build_model(num_classes)\n    criterion = nn.CrossEntropyLoss(label_smoothing=CFG.LABEL_SMOOTHING)\n    optimizer = optim.AdamW(\n        filter(lambda p: p.requires_grad, model.parameters()),\n        lr=CFG.LR, weight_decay=CFG.WEIGHT_DECAY\n    )\n    scheduler = optim.lr_scheduler.CosineAnnealingLR(\n        optimizer, T_max=CFG.EPOCHS, eta_min=CFG.LR * 0.01\n    )\n    scaler    = torch.cuda.amp.GradScaler(enabled=(CFG.DEVICE == \"cuda\"))\n    stopper   = EarlyStopping(patience=CFG.PATIENCE, mode=\"max\")\n\n    best_acc      = 0.0\n    train_losses  = []\n    val_losses    = []\n    val_accs      = []\n    best_ckpt_path = CFG.OUTPUT_DIR / \"best_model.pth\"\n\n    print(f\"\\n[Train] Starting for {CFG.EPOCHS} epochs...\")\n\n    epoch_bar = tqdm(range(CFG.EPOCHS), desc=\"Training\", unit=\"epoch\",\n                     dynamic_ncols=True)\n\n    for epoch in epoch_bar:\n        tr_loss           = train_one_epoch(model, train_loader, criterion,\n                                            optimizer, scaler, epoch)\n        val_loss, val_acc = validate(model, val_loader, criterion, epoch)\n\n        scheduler.step()\n\n        train_losses.append(tr_loss)\n        val_losses.append(val_loss)\n        val_accs.append(val_acc)\n\n        epoch_bar.set_postfix({\n            \"tr_loss\": f\"{tr_loss:.4f}\",\n            \"val_acc\": f\"{val_acc:.4f}\",\n            \"lr\":      f\"{scheduler.get_last_lr()[0]:.2e}\",\n        })\n\n        if val_acc > best_acc:\n            best_acc = val_acc\n            torch.save({\n                \"epoch\":        epoch + 1,\n                \"model_state\":  model.state_dict(),\n                \"optimizer\":    optimizer.state_dict(),\n                \"val_acc\":      val_acc,\n                \"label_to_idx\": label_to_idx,\n                \"idx_to_label\": idx_to_label,\n            }, best_ckpt_path)\n            tqdm.write(f\"  ✓ New best val_acc={val_acc:.4f}  (saved)\")\n\n        if stopper.step(val_acc):\n            tqdm.write(f\"[Early stopping] No improvement for {CFG.PATIENCE} epochs.\")\n            break\n\n    # Загружаем лучший чекпоинт\n    ckpt  = torch.load(best_ckpt_path, map_location=CFG.DEVICE)\n    model.load_state_dict(ckpt[\"model_state\"])\n    print(f\"\\n[Train] Done. Best val_acc = {best_acc:.4f}\")\n\n    plot_training_history(\n        train_losses, val_losses, val_accs,\n        save_path=CFG.OUTPUT_DIR / \"training_history.png\"\n    )\n\n    return model, label_to_idx, idx_to_label\n\n\n# ─────────────────────────────────────────────────────────────────────────────\n# 7. PREDICTION  (multi-label для квадратных снимков)\n# ─────────────────────────────────────────────────────────────────────────────\n\n@torch.no_grad()\ndef predict_proba_tta(\n    model: nn.Module,\n    test_df: pd.DataFrame,\n    n_tta: int = 4,\n) -> np.ndarray:\n    \"\"\"\n    Предсказывает вероятности для тестовых изображений с TTA.\n    Возвращает массив (N_images, N_classes) усреднённых вероятностей.\n    \"\"\"\n    model.eval()\n    all_probs = []\n\n    for tta_i in range(n_tta):\n        ds = TestDataset(\n            test_df, CFG.TEST_IMG_DIR,\n            img_col=\"quadrat_id\", use_tta=True, tta_idx=tta_i\n        )\n        loader = DataLoader(\n            ds, batch_size=CFG.BATCH_SIZE, shuffle=False,\n            num_workers=CFG.NUM_WORKERS, pin_memory=True\n        )\n\n        batch_probs = []\n        pbar = tqdm(loader, desc=f\"TTA {tta_i+1}/{n_tta}\", dynamic_ncols=True)\n        for imgs, _ in pbar:\n            imgs   = imgs.to(CFG.DEVICE, non_blocking=True)\n            with torch.cuda.amp.autocast(enabled=(CFG.DEVICE == \"cuda\")):\n                logits = model(imgs)\n            probs  = torch.softmax(logits, dim=1).cpu().numpy()\n            batch_probs.append(probs)\n\n        all_probs.append(np.concatenate(batch_probs, axis=0))\n\n    return np.mean(all_probs, axis=0)\n\n\ndef proba_to_multilabel(\n    probs: np.ndarray,\n    idx_to_label: dict,\n    threshold: float = CFG.THRESHOLD,\n    min_species: int = CFG.MIN_SPECIES,\n    max_species: int = CFG.MAX_SPECIES,\n) -> list[list[int]]:\n    \"\"\"\n    Конвертирует матрицу вероятностей (N, C) в список списков species_id.\n\n    Стратегия:\n      1. Берём все классы выше threshold.\n      2. Если ни одного — берём argmax (минимум 1 вид).\n      3. Обрезаем до max_species.\n    \"\"\"\n    predictions = []\n    for row in probs:\n        above = np.where(row >= threshold)[0].tolist()\n        if len(above) == 0:\n            above = [int(np.argmax(row))]\n        # Сортируем по убыванию вероятности и обрезаем\n        above.sort(key=lambda i: row[i], reverse=True)\n        above = above[:max_species]\n        species_ids = [idx_to_label[i] for i in above]\n        predictions.append(species_ids)\n    return predictions\n\n\ndef make_submission(\n    test_df: pd.DataFrame,\n    predictions: list[list[int]],\n    save_path: Path,\n) -> pd.DataFrame:\n    \"\"\"Генерирует submission.csv в формате PlantCLEF 2026.\"\"\"\n    records = []\n    for qid, sp_list in zip(test_df[\"quadrat_id\"], predictions):\n        sp_str = \"[\" + \", \".join(str(s) for s in sp_list) + \"]\"\n        records.append({\"quadrat_id\": qid, \"species_ids\": sp_str})\n\n    sub_df = pd.DataFrame(records)\n    sub_df.to_csv(save_path, sep=\",\", index=False, quoting=csv.QUOTE_ALL)\n    print(f\"\\n[Submission] Saved → {save_path}\")\n    print(f\"  Quadrats  : {len(sub_df):,}\")\n    print(f\"  Avg species/quadrat: \"\n          f\"{sub_df['species_ids'].apply(lambda x: x.count(',') + 1).mean():.1f}\")\n    print(sub_df.head(5).to_string(index=False))\n    return sub_df\n\n\n# ─────────────────────────────────────────────────────────────────────────────\n# 8. MAIN\n# ─────────────────────────────────────────────────────────────────────────────\n\ndef main() -> None:\n    CFG.OUTPUT_DIR.mkdir(parents=True, exist_ok=True)\n\n    # ── 8.1  Загрузка данных ─────────────────────────────────────────────────\n    train_meta, test_meta, species_df = load_metadata()\n\n    # ── 8.2  Анализ совместной встречаемости ─────────────────────────────────\n    print(\"\\n\" + \"=\"*60)\n    print(\" CO-OCCURRENCE ANALYSIS\")\n    print(\"=\"*60)\n\n    co_df, single_counts = build_cooccurrence_matrix(\n        df        = train_meta,\n        class_col = CFG.COOC_CLASS_COL,\n        img_id_col= CFG.COOC_IMG_ID_COL,\n        n_images  = CFG.COOC_N_IMAGES,\n    )\n\n    print(\"\\nCo-occurrence matrix:\")\n    print(co_df.to_string())\n\n    plot_cooccurrence(\n        co_df, single_counts,\n        top_pairs = CFG.COOC_TOP_PAIRS,\n        save_path = CFG.OUTPUT_DIR / \"cooccurrence_analysis.png\",\n    )\n\n    # ── 8.3  Обучение EfficientNet-B1 ────────────────────────────────────────\n    print(\"\\n\" + \"=\"*60)\n    print(\" MODEL TRAINING\")\n    print(\"=\"*60)\n\n    model, label_to_idx, idx_to_label = train(train_meta)\n\n    # ── 8.4  Предсказание с TTA ──────────────────────────────────────────────\n    print(\"\\n\" + \"=\"*60)\n    print(\" PREDICTION\")\n    print(\"=\"*60)\n\n    probs = predict_proba_tta(model, test_meta, n_tta=4)\n\n    predictions = proba_to_multilabel(\n        probs,\n        idx_to_label,\n        threshold   = CFG.THRESHOLD,\n        min_species = CFG.MIN_SPECIES,\n        max_species = CFG.MAX_SPECIES,\n    )\n\n    # ── 8.5  Генерация submission ─────────────────────────────────────────────\n    submission = make_submission(\n        test_df     = test_meta,\n        predictions = predictions,\n        save_path   = CFG.OUTPUT_DIR / \"submission.csv\",\n    )\n\n    print(\"\\n[Done] Pipeline complete.\")\n\n\nif __name__ == \"__main__\":\n    main()","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-05-02T08:31:12.852583Z","iopub.execute_input":"2026-05-02T08:31:12.854094Z","iopub.status.idle":"2026-05-02T08:31:12.869426Z","shell.execute_reply":"2026-05-02T08:31:12.867775Z"}},"outputs":[],"execution_count":null}]}