{"metadata":{"kernelspec":{"display_name":"Python (notebook-env)","language":"python","name":"notebook-env"},"language_info":{"codemirror_mode":{"name":"ipython","version":3},"file_extension":".py","mimetype":"text/x-python","name":"python","nbconvert_exporter":"python","pygments_lexer":"ipython3","version":"3.14.6"},"kaggle":{"accelerator":"none","dataSources":[],"dockerImageVersionId":28755,"isGpuEnabled":false,"isInternetEnabled":false,"language":"python","sourceType":"notebook"}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"# 📊 Exploratory Data Analysis (EDA) & Data Quality Diagnostics\n**Team:** MedVision - Group 4 - SIC\n\n**Project:** Diabetic Retinopathy Detection and Grading Using\nRetinal Images  \n\n**Stage:** Phase 1 — Data Auditing, Artifact Identification & Feature Engineering\n\n---\n\n## 📌 Executive Summary\n\nBefore training deep learning models for Diabetic Retinopathy (DR) detection, a comprehensive Exploratory Data Analysis (EDA) was performed across the dataset. Due to multi-clinic data acquisition, fundus imagery exhibits significant hardware-induced variance, spatial inconsistency, and extreme class imbalance. \n\nThis notebook establishes data integrity, quantifies imaging artifacts, and engineers persistence columns in the primary DataFrame to justify downstream preprocessing decisions.\n\n---\n\n## 📑 Diagnostic Checklist & Key Findings\n\n1. **📁 Data Integrity & Corruption Scan**\n   * **Objective:** Scan $100\\%$ of image paths for missing files, zero-byte allocations, unreadable headers, or pitch-black frames.\n   * **Finding:** Pipeline verified; clean dataset confirmed prior to memory-intensive preprocessing.\n<br></br>\n\n2. **⚖️ Target Class Distribution & Imbalance Audit**\n   * **Objective:** Quantify sample representation across DR severity grades ($0$ to $4$).\n   * **Finding:** Identified severe class imbalance (dominant non-DR vs. minority severe DR cases), dictating the need for loss-function reweighting or sampling strategies.\n<br></br>\n\n3. **📐 Spatial Geometry Analysis (Dimensions & Aspect Ratios)**\n   * **Objective:** Measure width, height, and aspect ratio variations across raw images.\n   * **Finding:** Significant hardware resolution discrepancies; establishes requirement for uniform letterbox padding and fixed resizing (e.g., $512 \\times 512$).\n<br></br>\n\n4. **💡 Overall Image Brightness & Illumination Bias Validation**\n   * **Objective:** Calculate global image mean intensity (`mean_intensity`), visualize extreme exposure variance, and evaluate bias across severity grades.\n   * **Finding:** Lighting variance is an independent artifact (not correlated with disease severity). Visual examples confirm the necessity of dynamic contrast normalization (CLAHE / Ben Graham method).\n<br></br>\n\n5. **🎨 RGB Channel Spectrum & Signal-to-Noise Profile**\n   * **Objective:** Quantify separate channel distributions (`red_mean`, `green_mean`, `blue_mean`).\n   * **Finding:** Confirms Red saturation dominance and Blue signal deficiency. Proves the **Green channel** is the primary signal carrier for retinal micro-lesion contrast.\n<br></br>\n\n6. **👁️ Tissue Coverage Ratio & Dead Space Audit**\n   * **Objective:** Calculate ocular frame coverage (`eye_tissue_ratio`) to evaluate empty black border margins.\n   * **Finding:** Highlights camera cropping inconsistency across clinics, justifying automated eye-globe contour cropping prior to model training.\n\n---","metadata":{}},{"cell_type":"markdown","source":"### i. Library Initialization & Environment Setup\n\nThis section imports all foundational dependencies required for the Exploratory Data Analysis (EDA) pipeline:\n\n* **`cv2` (OpenCV) & `PIL`**: Core computer vision engines used for loading, color-space transformations, grayscale conversion, and channel decomposition.\n* **`pandas` & `numpy`**: Vectorized data structures and fast matrix math for managing metadata and pixel array calculations.\n* **`matplotlib.pyplot` & `seaborn`**: Statistical visualization libraries used to plot KDE distributions, histogram counts, and diagnostic image grids.\n* **`tqdm`**: Provides lightweight progress bar monitoring across iterative image audits without impacting I/O performance.","metadata":{}},{"cell_type":"code","source":"# System and File Handling\nimport os\n\n# Data Manipulation, Numerical Processing & Statistical Testing\nimport numpy as np\nimport pandas as pd\nfrom scipy import stats\n\n# Computer Vision & Image Processing\nimport cv2\nfrom PIL import Image\n\n# Data Visualization\nimport matplotlib.pyplot as plt\nimport seaborn as sns\n\n# Progress Monitoring\nfrom tqdm import tqdm","metadata":{"execution":{"iopub.execute_input":"2026-09-19T19:36:14.717921Z","iopub.status.busy":"2026-09-19T19:36:14.717532Z","iopub.status.idle":"2026-09-19T19:36:14.72439Z","shell.execute_reply":"2026-09-19T19:36:14.723276Z","shell.execute_reply.started":"2026-09-19T19:36:14.717889Z"},"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"### ii. Loading Dataset & Creating File Paths\n\nIn this step, we:\n* Load `train.csv` to get the image IDs (`id_code`) and retinopathy grades (`diagnosis`).\n* Create a new column, `file_path`, containing the full location of each `.png` image so we can easily load them in upcoming steps.","metadata":{}},{"cell_type":"code","source":"# Load dataset\ndf = pd.read_csv('/kaggle/input/competitions/aptos2019-blindness-detection/train.csv')\n\nimage_dir = '/kaggle/input/competitions/aptos2019-blindness-detection/train_images'\n\n# Add a column with the direct file path for each image\ndf['file_path'] = df['id_code'].apply(lambda x: os.path.join(image_dir, f\"{x}.png\"))","metadata":{"execution":{"iopub.execute_input":"2026-09-19T18:46:07.31709Z","iopub.status.busy":"2026-09-19T18:46:07.316591Z","iopub.status.idle":"2026-09-19T18:46:07.351096Z","shell.execute_reply":"2026-09-19T18:46:07.350154Z","shell.execute_reply.started":"2026-09-19T18:46:07.317057Z"},"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"### 1. Data Integrity & Corruption Scan\n\nThis step verifies the physical integrity of every image in the dataset before running heavy computer vision tasks:\n\n* **Objective:** Scan $100\\%$ of image paths for missing files, zero-byte allocations, unreadable headers, or pitch-black frames.\n* **Audit Criteria:**\n  1. File existence on disk.\n  2. Non-zero file size allocation.\n  3. Image decode validation using OpenCV (`cv2.imread`).\n  4. Pitch-black frame detection (mean pixel intensity $< 5$).\n* **Finding:** Pipeline verified; clean dataset confirmed prior to memory-intensive preprocessing.","metadata":{}},{"cell_type":"code","source":"# Function to scan all image paths for file or image-level corruption\ndef scan_for_corrupted_images(df, min_brightness_threshold=5):\n    corrupted_files = []\n    \n    print(f\"Scanning {len(df)} images for corruption and errors...\\n\")\n    \n    for idx, row in tqdm(df.iterrows(), total=len(df)):\n        path = row['file_path']\n        \n        # 1. Check if the file actually exists on disk\n        if not os.path.exists(path):\n            corrupted_files.append({\n                'index': idx, 'file_path': path, 'reason': 'File Not Found'\n            })\n            continue\n            \n        # 2. Check for zero-byte / empty files\n        if os.path.getsize(path) == 0:\n            corrupted_files.append({\n                'index': idx, 'file_path': path, 'reason': 'Zero-Byte File (Empty)'\n            })\n            continue\n            \n        # 3. Attempt to read the image via OpenCV\n        img = cv2.imread(path)\n        if img is None:\n            corrupted_files.append({\n                'index': idx, 'file_path': path, 'reason': 'Unreadable / Corrupted Image'\n            })\n            continue\n            \n        # 4. Check for extreme low-intensity \"blackout\" images\n        mean_val = img.mean()\n        if mean_val < min_brightness_threshold:\n            corrupted_files.append({\n                'index': idx, 'file_path': path, 'reason': f'Severe Blackout Image (Mean Pixel = {mean_val:.2f})'\n            })\n\n    # Summary Output\n    corrupted_df = pd.DataFrame(corrupted_files)\n    \n    if len(corrupted_df) == 0:\n        print(\"\\nSuccess! All images are valid, readable, and non-corrupted.\")\n    else:\n        print(f\"\\nWarning: Found {len(corrupted_df)} problematic files!\")\n        \n    return corrupted_df\n\n# Execute scan and output findings report\ncorrupted_report = scan_for_corrupted_images(df)\n\nif not corrupted_report.empty:\n    display(corrupted_report)","metadata":{"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"### 2. Target Class Distribution & Imbalance Audit\n\nThis section evaluates the frequency distribution of Diabetic Retinopathy severity grades ($0–4$) across the training set:\n\n* **Objective:** Quantify sample representation across DR severity grades to identify class imbalance.\n* **Severity Mapping:** `0` (No DR), `1` (Mild), `2` (Moderate), `3` (Severe), `4` (Proliferative DR).\n* **Finding:** Significant class imbalance exists—majority non-DR (Grade 0) cases dominate, while severe cases (Grades 3 & 4) represent small minority classes. This justifies downstream loss-function reweighting, targeted augmentation, or class-balanced sampling.","metadata":{}},{"cell_type":"code","source":"# Plot Class Distribution (Brighter for highest counts, darker for lower counts)\nplt.figure(figsize=(8, 4))\nax = sns.countplot(x='diagnosis', data=df)\n\nplt.title('Class Imbalance: DR Severity Grades (0–4)', fontsize=12)\nplt.xlabel('Diagnosis Grade (0: No DR, 1: Mild, 2: Moderate, 3: Severe, 4: Proliferative)', fontsize=10)\nplt.ylabel('Image Count', fontsize=10)\n\ntotal = len(df)\nfor p in ax.patches:\n    height = p.get_height()\n    percentage = (height / total) * 100\n    ax.annotate(\n        f'{int(height)}\\n({percentage:.1f}%)', \n        (p.get_x() + p.get_width() / 2., height),\n        ha='center', va='bottom', \n        xytext=(0, 3), textcoords='offset points',\n        fontsize=9\n    )\n\nplt.ylim(0, max([p.get_height() for p in ax.patches]) * 1.15)\nplt.tight_layout()\nplt.show()\n\n# Print summary metrics and compute loss weights\nprint(f\"Total Dataset Size: {len(df)} images\\n\")\nprint(\"Class Counts & Percentages:\")\nbreakdown = pd.DataFrame({\n    'Count': df['diagnosis'].value_counts(),\n    'Percentage (%)': np.round(df['diagnosis'].value_counts(normalize=True) * 100, 2)\n}).sort_index()\ndisplay(breakdown)","metadata":{"execution":{"iopub.execute_input":"2026-09-19T18:47:27.752965Z","iopub.status.busy":"2026-09-19T18:47:27.752591Z","iopub.status.idle":"2026-09-19T18:47:27.989344Z","shell.execute_reply":"2026-09-19T18:47:27.988232Z","shell.execute_reply.started":"2026-09-19T18:47:27.752934Z"},"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"### 3. Spatial Geometry Analysis (Dimensions & Aspect Ratios)\n\nThis section evaluates the raw spatial dimensions and geometric variance across the fundus camera acquisitions:\n\n* **Objective:** Measure width, height, and aspect ratio variations across raw images to identify spatial inconsistencies.\n* **Audit Scope:** Extract pixel dimensions $(W \\times H)$ and compute aspect ratios $(W / H)$ directly from image headers.\n* **Finding:** Significant hardware resolution discrepancies exist across clinical capture sites. This confirms the requirement for standardized aspect-ratio letterboxing and fixed target resizing (e.g., $512 \\times 512$) during preprocessing to avoid structural distortion of retinal lesions.","metadata":{}},{"cell_type":"code","source":"widths, heights, aspect_ratios = [], [], []\n\n# Extract image dimensions from headers\nfor path in df['file_path']:\n    with Image.open(path) as img:\n        w, h = img.size\n        widths.append(w)\n        heights.append(h)\n        aspect_ratios.append(w / h)\n\n# Plot Spatial Geometry Distributions\nfig, ax = plt.subplots(2, 2, figsize=(14, 8))\n\n# 1. Width Distribution\nsns.histplot(widths, ax=ax[0, 0], color='teal', kde=True, bins=20)\nax[0, 0].set_title('Raw Image Widths Distribution', fontsize=11)\nax[0, 0].set_xlabel('Width (pixels)')\nax[0, 0].set_ylabel('Count')\n\n# 2. Height Distribution\nsns.histplot(heights, ax=ax[0, 1], color='teal', kde=True, bins=20)\nax[0, 1].set_title('Raw Image Heights Distribution', fontsize=11)\nax[0, 1].set_xlabel('Height (pixels)')\nax[0, 1].set_ylabel('Count')\n\n# 3. Width vs Height Scatter\nsns.scatterplot(x=widths, y=heights, ax=ax[1, 0], alpha=0.5, color='purple')\nax[1, 0].set_title('Width vs. Height Scatter Plot', fontsize=11)\nax[1, 0].set_xlabel('Width (pixels)')\nax[1, 0].set_ylabel('Height (pixels)')\n\n# 4. Aspect Ratio Distribution\nsns.histplot(aspect_ratios, ax=ax[1, 1], color='coral', kde=True, bins=20)\nax[1, 1].set_title('Aspect Ratio (W / H) Distribution', fontsize=11)\nax[1, 1].set_xlabel('Aspect Ratio')\nax[1, 1].set_ylabel('Count')\n\nplt.tight_layout()\nplt.show()\n\n# Print Aspect Ratio summary\nunique_ratios = sorted(set([round(r, 2) for r in aspect_ratios]))\nprint(f\"Unique Aspect Ratios found across dataset: {unique_ratios}\")","metadata":{"execution":{"iopub.execute_input":"2026-09-19T18:51:12.246167Z","iopub.status.busy":"2026-09-19T18:51:12.245769Z","iopub.status.idle":"2026-09-19T18:51:52.282402Z","shell.execute_reply":"2026-09-19T18:51:52.281451Z","shell.execute_reply.started":"2026-09-19T18:51:12.246137Z"},"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"### 4. Overall Image Brightness & Illumination Bias Validation\n\nThis step checks if lighting differences are linked to disease severity, or if it's just random camera noise.\n\n* **Objective:** Calculate average image brightness, see how lighting varies across DR grades (0–4), and run an ANOVA test to check for correlation.\n* **Methodology:** \n  1. Measure mean pixel intensity for every image (0 = pitch black, 255 = pure white).\n  2. Plot overall brightness with a histogram + KDE curve.\n  3. Compare brightness across grades using boxplots and run a One-Way ANOVA test.\n* **Finding:** Average brightness sits around 58–63 across all grades. The ANOVA test proves lighting is just random camera variation, not a sign of disease.\n* **Takeaway:** Because dark and bright images happen in every class, we must apply contrast normalization (like Ben Graham's method or CLAHE) so the model doesn't get tricked by lighting.","metadata":{}},{"cell_type":"code","source":"# Function to calculate mean brightness across RGB channels for a single image\ndef get_mean_intensity(path):\n    img = cv2.imread(path)\n    if img is None:\n        return np.nan\n    img = cv2.cvtColor(img, cv2.COLOR_BGR2RGB)\n    return np.mean(img)\n\n# 1. Compute mean intensity with progress tracking\ntqdm.pandas(desc=\"Calculating Image Intensities\")\ndf['mean_intensity'] = df['file_path'].progress_apply(get_mean_intensity)\n\n# 2. Plot overall brightness distribution using Seaborn with KDE overlay\nplt.figure(figsize=(8, 4))\nsns.histplot(\n    data=df, \n    x='mean_intensity', \n    bins=30, \n    color='darkorange', \n    kde=True, \n    edgecolor='black'\n)\n\nplt.title('Raw Image Overall Brightness Distribution', fontsize=12)\nplt.xlabel('Mean Pixel Intensity (0 = Pitch Black, 255 = Pure White)', fontsize=10)\nplt.ylabel('Image Count', fontsize=10)\nplt.grid(axis='y', linestyle='--', alpha=0.5)\nplt.tight_layout()\nplt.show()","metadata":{"execution":{"iopub.execute_input":"2026-09-19T18:59:44.277095Z","iopub.status.busy":"2026-09-19T18:59:44.276321Z","iopub.status.idle":"2026-09-19T19:08:15.417946Z","shell.execute_reply":"2026-09-19T19:08:15.416846Z","shell.execute_reply.started":"2026-09-19T18:59:44.277051Z"},"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"code","source":"# 1. Visualize Brightness Distribution Across Diagnosis Grades\nplt.figure(figsize=(10, 5))\n\n# Boxplot shows the median, quartiles, and outliers per grade\nsns.boxplot(\n    x='diagnosis', \n    y='mean_intensity', \n    data=df, \n    showmeans=True,\n    meanprops={\"marker\": \"o\", \"markerfacecolor\": \"red\", \"markeredgecolor\": \"red\"}\n)\n\n# Overlay individual data points (jittered) to see density\nsns.stripplot(\n    x='diagnosis', \n    y='mean_intensity', \n    data=df, \n    color='black', \n    alpha=0.15, \n    jitter=0.2\n)\n\nplt.title('Image Brightness (Mean Intensity) Across Retinopathy Severity Grades', fontsize=12)\nplt.xlabel('Diagnosis Grade (0: No DR, 1: Mild, 2: Moderate, 3: Severe, 4: Proliferative)', fontsize=10)\nplt.ylabel('Mean Pixel Intensity (0–255)', fontsize=10)\nplt.grid(axis='y', linestyle='--', alpha=0.5)\n\nplt.tight_layout()\nplt.show()\n\n# 2. Statistical Correlation Check (One-Way ANOVA Test)\n# Group intensities by diagnosis grade\ngrade_groups = [group['mean_intensity'].dropna().values for _, group in df.groupby('diagnosis')]\n\n# Perform One-Way ANOVA test\nf_stat, p_val = stats.f_oneway(*grade_groups)\n\nprint(\"--- Statistical Analysis (One-Way ANOVA) ---\")\nprint(f\"F-Statistic: {f_stat:.4f}\")\nprint(f\"p-value:     {p_val:.4e}\")\n\nif p_val > 0.05:\n    print(\"\\nTakeaway: No statistically significant brightness difference across grades (p > 0.05).\")\n    print(\"Lighting variations are independent of disease severity—safe to apply uniform CLAHE/normalization.\")\nelse:\n    print(\"\\nTakeaway: Statistically significant brightness difference detected across grades (p <= 0.05).\")\n    print(\"Careful: Normalize lighting across all grades so the model doesn't learn brightness as a shortcut feature.\")","metadata":{"execution":{"iopub.execute_input":"2026-09-19T19:36:14.725837Z","iopub.status.busy":"2026-09-19T19:36:14.725506Z","iopub.status.idle":"2026-09-19T19:36:15.098877Z","shell.execute_reply":"2026-09-19T19:36:15.097582Z","shell.execute_reply.started":"2026-09-19T19:36:14.725769Z"},"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"### 5. RGB Channel Spectrum & Signal-to-Noise Profile\n\nThis section breaks down the color spectrum to see which channel carries the most useful information for spotting retinal lesions.\n\n* **Objective:** Calculate separate RGB intensity distributions (`red_mean`, `green_mean`, `blue_mean`) across all images.\n* **Methodology:** Extract channel-wise averages using OpenCV and plot their density curves using Seaborn's `kdeplot`.\n* **Finding:** \n  * **Red Channel:** Saturated and dominant (reflecting vascular background tissue).\n  * **Blue Channel:** Deficient with low signal-to-noise ratio.\n  * **Green Channel:** Sits right in the middle, offering optimal structural contrast for microaneurysms and hemorrhages.\n* **Takeaway:** The **Green channel** is the primary signal carrier for Diabetic Retinopathy detection, making green-plane extraction or contrast tuning essential during preprocessing.","metadata":{}},{"cell_type":"code","source":"# Register tqdm with pandas for progress tracking\ntqdm.pandas(desc=\"Extracting RGB Channel Means\")\n\ndef get_rgb_channel_means(path):\n    img = cv2.imread(path)\n    if img is None:\n        return np.nan, np.nan, np.nan\n    \n    img = cv2.cvtColor(img, cv2.COLOR_BGR2RGB)\n    \n    # Calculate mean across height and width for each channel (0: R, 1: G, 2: B)\n    return img[:, :, 0].mean(), img[:, :, 1].mean(), img[:, :, 2].mean()\n\n# 1. Apply function across dataset and expand into separate DataFrame columns\nrgb_means = df['file_path'].progress_apply(get_rgb_channel_means)\ndf[['red_mean', 'green_mean', 'blue_mean']] = pd.DataFrame(rgb_means.tolist(), index=df.index)\n\n# 2. Plot Channel Density Distributions\nplt.figure(figsize=(10, 5))\nsns.kdeplot(df['red_mean'], color='red', label='Red Channel', fill=True, alpha=0.3)\nsns.kdeplot(df['green_mean'], color='green', label='Green Channel', fill=True, alpha=0.3)\nsns.kdeplot(df['blue_mean'], color='blue', label='Blue Channel', fill=True, alpha=0.3)\n\nplt.title('Color Channel Mean Distribution Across Retinal Fundus Images', fontsize=12)\nplt.xlabel('Mean Pixel Intensity (0 - 255)', fontsize=10)\nplt.ylabel('Density', fontsize=10)\nplt.legend(title='Channels')\nplt.grid(axis='x', linestyle='--', alpha=0.5)\n\nplt.tight_layout()\nplt.show()\n\n# 3. Print channel metric summary\nprint(\"--- RGB Channel Averages ---\")\nprint(f\"Red Channel Mean:   {df['red_mean'].mean():.2f}\")\nprint(f\"Green Channel Mean: {df['green_mean'].mean():.2f}\")\nprint(f\"Blue Channel Mean:  {df['blue_mean'].mean():.2f}\")","metadata":{"execution":{"iopub.execute_input":"2026-09-19T19:27:32.781165Z","iopub.status.busy":"2026-09-19T19:27:32.776475Z","iopub.status.idle":"2026-09-19T19:36:14.715398Z","shell.execute_reply":"2026-09-19T19:36:14.714046Z","shell.execute_reply.started":"2026-09-19T19:27:32.781036Z"},"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"### 6. Tissue Coverage Ratio & Dead Space Audit\n\nThis section measures how much of each image contains actual retinal tissue versus empty black background margins.\n\n* **Objective:** Calculate the percentage of active ocular frame coverage (`eye_tissue_ratio`) across all images to quantify background waste.\n* **Methodology:** Threshold grayscale pixels (>10 intensity) to separate foreground eye tissue from black padding, then plot the distribution using Seaborn `histplot` with KDE.\n* **Finding:** Frame coverage varies significantly across clinics, leaving large amounts of unused black border space in many images.\n* **Takeaway:** Automated bounding-box / contour cropping is required to strip away dead space and focus model capacity directly on the eye globe.","metadata":{}},{"cell_type":"code","source":"# Register tqdm with pandas for progress bar tracking\ntqdm.pandas(desc=\"Calculating Eye Tissue Ratio\")\n\ndef get_eye_tissue_ratio(path):\n    img = cv2.imread(path, cv2.IMREAD_GRAYSCALE)\n    if img is None:\n        return np.nan\n        \n    # Threshold: Consider pixels > 10 as foreground \"eye content\"\n    eye_pixels = np.sum(img > 10)\n    total_pixels = img.size\n    \n    # Return percentage of active eye tissue\n    return (eye_pixels / total_pixels) * 100\n\n# 1. Apply calculation across dataset with progress tracking\ndf['eye_tissue_ratio'] = df['file_path'].progress_apply(get_eye_tissue_ratio)\n\n# 2. Plot Tissue Ratio Distribution\nplt.figure(figsize=(8, 4))\nsns.histplot(\n    data=df, \n    x='eye_tissue_ratio', \n    color='teal', \n    kde=True, \n    bins=25,\n    edgecolor='black'\n)\n\nplt.title('Ocular Globe Area vs. Frame Size (% Eye Content - Full Dataset)', fontsize=12)\nplt.xlabel('Percentage of Image Covered by Eye Tissue (%)', fontsize=10)\nplt.ylabel('Image Count', fontsize=10)\nplt.grid(axis='y', linestyle='--', alpha=0.5)\n\nplt.tight_layout()\nplt.show()\n\n# 3. Print summary metrics\nprint(\"--- Eye Tissue Coverage Metrics ---\")\nprint(f\"Average Eye Frame Coverage: {df['eye_tissue_ratio'].mean():.2f}%\")\nprint(f\"Min Eye Coverage Found:     {df['eye_tissue_ratio'].min():.2f}%\")\nprint(f\"Max Eye Coverage Found:     {df['eye_tissue_ratio'].max():.2f}%\")","metadata":{"execution":{"iopub.execute_input":"2026-09-19T19:36:15.101426Z","iopub.status.busy":"2026-09-19T19:36:15.101069Z","iopub.status.idle":"2026-09-19T19:44:20.737302Z","shell.execute_reply":"2026-09-19T19:44:20.736248Z","shell.execute_reply.started":"2026-09-19T19:36:15.101395Z"},"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"---\n\n## 🚀 Handoff Blueprint: Preprocessing & Data Augmentation Guidelines\n\nBased on the empirical findings from this EDA stage, the next phase (**Preprocessing & Augmentation**) should implement the following pipeline transformations prior to model training:\n\n### 1. ✂️ Preprocessing Pipeline Recommendations\n* **Automated Contour Cropping:** Use `eye_tissue_ratio` insights to crop out static black borders (dead space). Threshold the grayscale image to create a bounding box around the ocular globe before resizing.\n* **Aspect-Ratio Preserving Resize:** Standardize image resolutions (e.g., $512 \\times 512$ or $224 \\times 224$) using **letterbox padding** to prevent horizontal/vertical stretching of spherical retinal structures.\n* **Illumination Normalization:** Apply **Ben Graham's method** (local average subtraction) or **CLAHE** (Contrast Limited Adaptive Histogram Equalization) across all inputs to neutralize global camera lighting artifacts.\n* **Green-Channel Contrast Priority:** Focus contrast adjustments or CLAHE on the Green channel, as EDA proved it holds the primary signal-to-noise ratio for micro-lesions.\n\n---\n\n### 2. 🔄 Data Augmentation Strategy\n* **Geometrical Transformations:** Apply random horizontal and vertical flips, as well as rotations ($0^\\circ–360^\\circ$). Retinal lesions retain their clinical validity regardless of orientation.\n* **Avoid Heavy Color/Brightness Jitter:** Do not apply extreme brightness or hue augmentations, as contrast normalization will handle exposure variance, and hue shifts can alter vessel/lesion fidelity.\n* **Class Imbalance Mitigation:** Leverage the class counts from Section 2 to set up **Weighted Random Sampling** in your DataLoader or configure class weights inside the loss function (e.g., Focal Loss or Weighted Cross-Entropy).","metadata":{}},{"cell_type":"markdown","source":"# 🧪 Retinal Image Preprocessing & Augmentation\n\n**Team:** MedVision - Group 4\n\n**Project:** Diabetic Retinopathy Detection and Grading Using Retinal Images\n\n**Role:** Shada — Image Preprocessing & Data Augmentation\n\n**Stage:** Phase 2 — Image Preprocessing, Augmentation & Pipeline Validation\n\n\n","metadata":{}},{"cell_type":"markdown","source":"## 📌 Executive Summary\n\nFollowing the exploratory data analysis and data quality assessment, this phase focuses on developing and validating the image preprocessing and augmentation pipeline for the APTOS 2019 retinal image dataset.\n\nThe preprocessing stage evaluates:\n- Black border removal\n- Aspect-ratio-preserving resizing with padding\n- Local contrast enhancement using CLAHE\n\nThe augmentation stage evaluates:\n- Horizontal flipping\n- Small-angle rotation\n- Brightness and contrast variation\n\nEach transformation was visually evaluated across different diabetic retinopathy grades to assess image quality and preservation of clinically relevant retinal structures.\n\nThe final preprocessing pipeline was selected based on visual quality, preservation of retinal information, and consistency across the dataset.","metadata":{}},{"cell_type":"markdown","source":"## 1. Library Initialization & Environment Setup\n\nThis section imports the libraries required for image preprocessing, augmentation experiments, visualization, and progress monitoring.","metadata":{}},{"cell_type":"code","source":"import os\nimport cv2\nimport numpy as np\nimport pandas as pd\nimport matplotlib.pyplot as plt\n\nfrom PIL import Image\nfrom torchvision import transforms","metadata":{},"outputs":[],"execution_count":null},{"cell_type":"code","source":"# Dataset paths\ntrain_csv = \"/kaggle/input/competitions/aptos2019-blindness-detection/train.csv\"\nimage_dir = \"/kaggle/input/competitions/aptos2019-blindness-detection/train_images\"\n\n# Load dataset\ndf = pd.read_csv(train_csv)\n\n# Create full image paths\ndf[\"file_path\"] = df[\"id_code\"].apply(\n    lambda x: os.path.join(image_dir, f\"{x}.png\")\n)\n\nprint(f\"Total images: {len(df)}\")","metadata":{},"outputs":[],"execution_count":null},{"cell_type":"code","source":"import cv2\nimport matplotlib.pyplot as plt\n\nsample_paths = df['file_path'].sample(6, random_state=42)\n\nplt.figure(figsize=(15, 8))\n\nfor i, path in enumerate(sample_paths):\n    img = cv2.imread(path)\n    img = cv2.cvtColor(img, cv2.COLOR_BGR2RGB)\n\n    plt.subplot(2, 3, i + 1)\n    plt.imshow(img)\n    plt.axis('off')\n\nplt.tight_layout()\nplt.show()","metadata":{},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"## 2. Black Border Cropping\n\n### Objective\n\nRemove unnecessary black borders surrounding the retinal field while preserving the visible retinal structures.\n\n### Method\n\nA grayscale threshold is used to identify non-black pixels. The bounding rectangle of the detected foreground region is then used to crop the image.\n\nThe cropping method is evaluated visually across randomly selected retinal images to verify that important retinal structures are preserved.","metadata":{}},{"cell_type":"code","source":"import cv2\nimport matplotlib.pyplot as plt\n\ndef crop_black_borders(img, threshold=10):\n    gray = cv2.cvtColor(img, cv2.COLOR_RGB2GRAY)\n\n    mask = gray > threshold\n\n    coords = cv2.findNonZero(mask.astype('uint8'))\n\n    if coords is None:\n        return img\n\n    x, y, w, h = cv2.boundingRect(coords)\n\n    return img[y:y+h, x:x+w]\n\n\nsample_paths = df['file_path'].sample(6, random_state=42)\n\nplt.figure(figsize=(14, 12))\n\nfor i, path in enumerate(sample_paths):\n    img = cv2.imread(path)\n    img = cv2.cvtColor(img, cv2.COLOR_BGR2RGB)\n\n    cropped = crop_black_borders(img)\n\n    # Original image\n    plt.subplot(6, 2, 2*i + 1)\n    plt.imshow(img)\n    plt.axis('off')\n    plt.title(\"Original\")\n\n    # Cropped image\n    plt.subplot(6, 2, 2*i + 2)\n    plt.imshow(cropped)\n    plt.axis('off')\n    plt.title(\"After Cropping\")\n\nplt.tight_layout()\nplt.show()","metadata":{},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"## 3. Aspect-Ratio-Preserving Resize with Padding\n\n### Objective\n\nStandardize all retinal images to a fixed input size while preserving their original aspect ratio.\n\n### Method\n\nEach image is first cropped to remove unnecessary black borders. It is then resized proportionally so that neither the width nor the height is distorted.\n\nThe resized image is centered on a black 224×224 canvas, preserving the original geometric proportions.\n\n### Evaluation\n\nThe transformation was visually inspected on randomly selected retinal images to verify that the retinal field remains proportionally preserved after resizing.","metadata":{}},{"cell_type":"code","source":"import cv2\nimport numpy as np\nimport matplotlib.pyplot as plt\n\ndef resize_with_padding(img, target_size=(224, 224)):\n    target_w, target_h = target_size\n\n    h, w = img.shape[:2]\n\n    # Calculate scaling factor while preserving aspect ratio\n    scale = min(target_w / w, target_h / h)\n\n    new_w = int(w * scale)\n    new_h = int(h * scale)\n\n    # Resize the image\n    resized = cv2.resize(\n        img,\n        (new_w, new_h),\n        interpolation=cv2.INTER_AREA\n    )\n\n    # Create a black canvas with the target dimensions\n    padded = np.zeros(\n        (target_h, target_w, 3),\n        dtype=np.uint8\n    )\n\n    # Center the resized image on the canvas\n    x_offset = (target_w - new_w) // 2\n    y_offset = (target_h - new_h) // 2\n\n    padded[\n        y_offset:y_offset + new_h,\n        x_offset:x_offset + new_w\n    ] = resized\n\n    return padded","metadata":{},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"### Finding\n\nVisual comparison showed that both 224×224 and 512×512 preserved the main retinal structures in the tested samples, with no major visual difference observed.\n\n### Decision\n\nA target size of 224×224 was selected for the baseline preprocessing pipeline, while preserving the original aspect ratio through proportional resizing and padding.","metadata":{}},{"cell_type":"code","source":"fig, axes = plt.subplots(3, 6, figsize=(18, 10))\n\nfor i, path in enumerate(sample_paths):\n    img = cv2.imread(path)\n    img = cv2.cvtColor(img, cv2.COLOR_BGR2RGB)\n\n    cropped = crop_black_borders(img)\n\n    resized_224 = resize_with_padding(\n        cropped,\n        target_size=(224, 224)\n    )\n\n    resized_512 = resize_with_padding(\n        cropped,\n        target_size=(512, 512)\n    )\n\n    axes[0, i].imshow(cropped)\n    axes[0, i].set_title(\"Cropped\")\n    axes[0, i].axis(\"off\")\n\n    axes[1, i].imshow(resized_224)\n    axes[1, i].set_title(\"224 × 224\")\n    axes[1, i].axis(\"off\")\n\n    axes[2, i].imshow(resized_512)\n    axes[2, i].set_title(\"512 × 512\")\n    axes[2, i].axis(\"off\")\n\nplt.tight_layout()\nplt.show()","metadata":{},"outputs":[],"execution_count":null},{"cell_type":"code","source":"import os\n\nprocessed_dir = \"/kaggle/working/processed_images\"\nos.makedirs(processed_dir, exist_ok=True)\n\nprint(processed_dir)","metadata":{},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"### Model Input Normalization\n\nNormalization using ImageNet mean and standard deviation may be applied during model training if ImageNet-pretrained weights are used.\n\nThis transformation is considered a model input step and is not applied to the saved preprocessed PNG images.","metadata":{}},{"cell_type":"code","source":"def apply_clahe(img, clip_limit=2.0, tile_grid_size=(8, 8)):\n    lab = cv2.cvtColor(img, cv2.COLOR_RGB2LAB)\n\n    l_channel, a_channel, b_channel = cv2.split(lab)\n\n    clahe = cv2.createCLAHE(\n        clipLimit=clip_limit,\n        tileGridSize=tile_grid_size\n    )\n\n    l_enhanced = clahe.apply(l_channel)\n\n    enhanced_lab = cv2.merge([\n        l_enhanced,\n        a_channel,\n        b_channel\n    ])\n\n    enhanced = cv2.cvtColor(\n        enhanced_lab,\n        cv2.COLOR_LAB2RGB\n    )\n\n    return enhanced","metadata":{},"outputs":[],"execution_count":null},{"cell_type":"code","source":"grade4 = df[df['diagnosis'] == 4].sample(\n    n=3,\n    random_state=42\n)\n\nfig, axes = plt.subplots(3, 3, figsize=(12, 12))\n\nfor i, (_, row) in enumerate(grade4.iterrows()):\n\n    img = cv2.imread(row['file_path'])\n    img = cv2.cvtColor(img, cv2.COLOR_BGR2RGB)\n\n    cropped = crop_black_borders(img)\n\n    enhanced = apply_clahe(cropped)\n\n    # Original\n    axes[i, 0].imshow(cropped)\n    axes[i, 0].set_title(\"After Crop\")\n    axes[i, 0].axis(\"off\")\n\n    # CLAHE\n    axes[i, 1].imshow(enhanced)\n    axes[i, 1].set_title(\"After CLAHE\")\n    axes[i, 1].axis(\"off\")\n\n    # الفرق في السطوع\n    axes[i, 2].imshow(\n        np.abs(\n            enhanced.astype(float) -\n            cropped.astype(float)\n        ).mean(axis=2),\n        cmap=\"gray\"\n    )\n    axes[i, 2].set_title(\"Change Map\")\n    axes[i, 2].axis(\"off\")\n\nplt.tight_layout()\nplt.show()","metadata":{},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"### Finding\n\nVisual inspection showed that CLAHE enhanced local contrast and improved the visibility of retinal structures in the tested Grade 4 samples.\n\nThe change maps confirmed that the transformation primarily affected image intensity rather than altering the image structure.\n\n### Decision\n\nCLAHE was retained as a candidate preprocessing step with:\n\n- Clip limit: 2.0\n- Tile grid size: 8×8\n\nThe final effect on model performance should be evaluated during training.","metadata":{}},{"cell_type":"code","source":"# اختيار 3 صور من كل Grade\nsample_df = (\n    df.groupby('diagnosis', group_keys=False)\n      .sample(n=3, random_state=42)\n      .reset_index(drop=True)\n)\n\nfig, axes = plt.subplots(\n    len(sample_df), 2,\n    figsize=(10, 4 * len(sample_df))\n)\n\nfor i, (_, row) in enumerate(sample_df.iterrows()):\n\n    # قراءة الصورة\n    img = cv2.imread(row['file_path'])\n    img = cv2.cvtColor(img, cv2.COLOR_BGR2RGB)\n\n    # إزالة الحدود السوداء\n    cropped = crop_black_borders(img)\n\n    # تطبيق CLAHE\n    enhanced = apply_clahe(cropped)\n\n    # قبل CLAHE\n    axes[i, 0].imshow(cropped)\n    axes[i, 0].set_title(\n        f\"Grade {row['diagnosis']} - Before CLAHE\"\n    )\n    axes[i, 0].axis(\"off\")\n\n    # بعد CLAHE\n    axes[i, 1].imshow(enhanced)\n    axes[i, 1].set_title(\n        f\"Grade {row['diagnosis']} - After CLAHE\"\n    )\n    axes[i, 1].axis(\"off\")\n\nplt.tight_layout()\nplt.show()","metadata":{},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"### Multi-grade Evaluation\n\nCLAHE was visually evaluated on representative samples from all five diabetic retinopathy grades.\n\nThe transformation improved local contrast and enhanced the visibility of retinal structures across the tested samples, without obvious visual degradation.\n\n### Decision\n\nCLAHE was retained as part of the baseline preprocessing pipeline with:\n\n- Clip limit: 2.0\n- Tile grid size: 8×8\n\nModel-based evaluation is still required to determine its effect on classification performance.","metadata":{}},{"cell_type":"code","source":"grade4 = df[df['diagnosis'] == 4].sample(\n    n=3,\n    random_state=42\n)\n\nfig, axes = plt.subplots(\n    3, 3,\n    figsize=(12, 12)\n)\n\nfor i, (_, row) in enumerate(grade4.iterrows()):\n\n    img = cv2.imread(row['file_path'])\n    img = cv2.cvtColor(img, cv2.COLOR_BGR2RGB)\n\n    cropped = crop_black_borders(img)\n\n    clahe_2 = apply_clahe(\n        cropped,\n        clip_limit=2.0\n    )\n\n    clahe_15 = apply_clahe(\n        cropped,\n        clip_limit=1.5\n    )\n\n    axes[i, 0].imshow(cropped)\n    axes[i, 0].set_title(\"Before CLAHE\")\n    axes[i, 0].axis(\"off\")\n\n    axes[i, 1].imshow(clahe_2)\n    axes[i, 1].set_title(\"CLAHE 2.0\")\n    axes[i, 1].axis(\"off\")\n\n    axes[i, 2].imshow(clahe_15)\n    axes[i, 2].set_title(\"CLAHE 1.5\")\n    axes[i, 2].axis(\"off\")\n\nplt.tight_layout()\nplt.show()","metadata":{},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"### CLAHE Parameter Comparison\n\nTwo CLAHE clip limits, 1.5 and 2.0, were visually compared on representative Grade 4 samples.\n\nThe difference between the two settings was relatively small on normal images, while the 2.0 setting provided slightly clearer enhancement in darker samples without obvious visual degradation.\n\n### Decision\n\nA clip limit of 2.0 was selected for the baseline preprocessing pipeline.\n\n**Tile grid size:** 8×8","metadata":{}},{"cell_type":"code","source":"def preprocess_image(path):\n    img = cv2.imread(path)\n    if img is None:\n        return None\n\n    img = cv2.cvtColor(img, cv2.COLOR_BGR2RGB)\n\n    cropped = crop_black_borders(img)\n\n    resized = resize_with_padding(\n        cropped,\n        target_size=(224, 224)\n    )\n\n    processed = apply_clahe(\n        resized,\n        clip_limit=2.0,\n        tile_grid_size=(8, 8)\n    )\n\n    return processed","metadata":{},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"### Final Multi-grade Visual Check\n\nThe final preprocessing pipeline was visually checked on one representative image from each of the five diabetic retinopathy grades.\n\nThe comparison was used to verify that the preprocessing steps preserve the main retinal structures while producing standardized 224×224 outputs.","metadata":{}},{"cell_type":"markdown","source":"## Final Preprocessing Pipeline\n\nBased on the visual evaluation and parameter comparisons, the selected preprocessing pipeline is:\n\n**Crop black borders → Resize with padding to 224×224 → CLAHE**\n\n- **Cropping:** removes unnecessary black borders while preserving the retinal field.\n- **Resize + Padding:** preserves the original aspect ratio while standardizing all images to 224×224.\n- **CLAHE:** applied to the LAB L channel using `clipLimit=2.0` and `tileGridSize=(8, 8)` to improve local contrast.\n\nThese preprocessing choices are preliminary and should be validated through model-based experiments.","metadata":{}},{"cell_type":"code","source":"sample_df = (\n    df.groupby('diagnosis', group_keys=False)\n      .sample(n=1, random_state=42)\n      .reset_index(drop=True)\n)\n\nfig, axes = plt.subplots(5, 2, figsize=(8, 18))\n\nfor i, (_, row) in enumerate(sample_df.iterrows()):\n    img = cv2.imread(row['file_path'])\n    img = cv2.cvtColor(img, cv2.COLOR_BGR2RGB)\n\n    processed = preprocess_image(row['file_path'])\n\n    axes[i, 0].imshow(img)\n    axes[i, 0].set_title(\n        f\"Grade {row['diagnosis']} - Original\"\n    )\n    axes[i, 0].axis(\"off\")\n\n    axes[i, 1].imshow(processed)\n    axes[i, 1].set_title(\n        f\"Grade {row['diagnosis']} - Processed\"\n    )\n    axes[i, 1].axis(\"off\")\n\nplt.tight_layout()\nplt.show()","metadata":{},"outputs":[],"execution_count":null},{"cell_type":"code","source":"from torchvision import transforms\nfrom PIL import Image\nimport matplotlib.pyplot as plt\n\nhorizontal_flip = transforms.RandomHorizontalFlip(p=1.0)\n\nrotation = transforms.RandomRotation(degrees=15)\n\nbrightness_contrast = transforms.ColorJitter(\n    brightness=0.1,\n    contrast=0.1\n)\n\nsample_df = (\n    df.groupby('diagnosis', group_keys=False)\n      .sample(n=1, random_state=42)\n      .reset_index(drop=True)\n)\n\nfig, axes = plt.subplots(\n    5, 4,\n    figsize=(12, 18)\n)\n\nfor i, (_, row) in enumerate(sample_df.iterrows()):\n\n    img = Image.open(row['file_path']).convert(\"RGB\")\n\n    augmented_images = [\n        (\"Original\", img),\n        (\"Horizontal Flip\", horizontal_flip(img)),\n        (\"Rotation ±15°\", rotation(img)),\n        (\"Brightness + Contrast\", brightness_contrast(img))\n    ]\n\n    for j, (title, image) in enumerate(augmented_images):\n\n        axes[i, j].imshow(image)\n        axes[i, j].set_title(\n            f\"Grade {row['diagnosis']} - {title}\"\n        )\n        axes[i, j].axis(\"off\")\n\nplt.tight_layout()\nplt.show()","metadata":{},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"### Multi-grade Visual Evaluation\n\nThe selected augmentation candidates were visually evaluated on representative samples from all five diabetic retinopathy grades.\n\nThe tested augmentations were:\n\n- Horizontal Flip\n- Rotation (±15°)\n- Brightness and Contrast Adjustment (0.1)\n\nThe transformations preserved the main retinal structures in the tested samples without obvious visual degradation.\n\n### Decision\n\nThese three augmentations were retained as candidates for model training.\n\nAugmentation should be applied only to the training data after the train/validation split or within each cross-validation training fold.","metadata":{}},{"cell_type":"code","source":"import matplotlib.pyplot as plt\nfrom torchvision import transforms\nfrom PIL import Image\n\npath = df['file_path'].iloc[0]\nimg = Image.open(path).convert(\"RGB\")\n\naugmentations = {\n    \"Original\": transforms.Compose([]),\n\n    \"Horizontal Flip\": transforms.Compose([\n        transforms.RandomHorizontalFlip(p=1.0)\n    ]),\n\n    \"Small Rotation\": transforms.Compose([\n        transforms.RandomRotation(degrees=15)\n    ]),\n\n    \"Small Zoom\": transforms.Compose([\n        transforms.RandomResizedCrop(\n            size=img.size[::-1],\n            scale=(0.9, 1.0),\n            ratio=(0.95, 1.05)\n        )\n    ])\n}\n\nplt.figure(figsize=(14, 10))\n\nfor i, (name, transform) in enumerate(augmentations.items()):\n\n    augmented = transform(img)\n\n    plt.subplot(2, 2, i + 1)\n    plt.imshow(augmented)\n    plt.axis(\"off\")\n    plt.title(name)\n\nplt.tight_layout()\nplt.show()","metadata":{},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"### Rejected Augmentation\n\n**RandomResizedCrop (Small Zoom):**  \nVisual inspection showed noticeable loss of retinal details in the tested sample. Therefore, this transformation was not retained as a baseline augmentation candidate.\n\n**Decision:** Rejected due to potential loss of clinically relevant retinal details.","metadata":{}},{"cell_type":"code","source":"from tqdm import tqdm\n\nprint(f\"Processing {len(df)} images...\")\n\nprocessed_count = 0\nfailed_count = 0\n\nfor _, row in tqdm(df.iterrows(), total=len(df), desc=\"Preprocessing images\"):\n    processed = preprocess_image(row[\"file_path\"])\n\n    if processed is None:\n        failed_count += 1\n        continue\n\n    output_path = os.path.join(\n        processed_dir,\n        f\"{row['id_code']}.png\"\n    )\n\n    # Convert RGB to BGR before saving with OpenCV\n    cv2.imwrite(\n        output_path,\n        cv2.cvtColor(processed, cv2.COLOR_RGB2BGR)\n    )\n\n    processed_count += 1\n\nprint(f\"Processed successfully: {processed_count}\")\nprint(f\"Failed: {failed_count}\")\nprint(f\"Output directory: {processed_dir}\")","metadata":{},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"This section verifies the output of the complete preprocessing pipeline by checking the processed image dimensions, data type, number of generated images, and output consistency.\n\nThe expected output is 224×224×3 with `uint8` data type for all successfully processed images.","metadata":{}},{"cell_type":"code","source":"output_files = [\n    f for f in os.listdir(processed_dir)\n    if f.lower().endswith(\".png\")\n]\n\nprint(f\"Number of processed images: {len(output_files)}\")\n\nsample_output = os.path.join(processed_dir, output_files[0])\ncheck_img = cv2.imread(sample_output)\n\nprint(f\"Sample output shape: {check_img.shape}\")\nprint(f\"Sample output dtype: {check_img.dtype}\")","metadata":{},"outputs":[],"execution_count":null},{"cell_type":"code","source":"bad_outputs = []\n\nfor filename in tqdm(output_files, desc=\"Checking output sizes\"):\n    path = os.path.join(processed_dir, filename)\n    img = cv2.imread(path)\n\n    if img is None or img.shape[:2] != (224, 224):\n        bad_outputs.append(filename)\n\nprint(f\"Invalid outputs: {len(bad_outputs)}\")","metadata":{},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"## Non-Dark Pixel Ratio Analysis\n\n### Objective\n\nEstimate how much of each image contains non-dark pixels compared with the full image frame.\n\n### Method\n\nFor each image, grayscale pixels with intensity greater than 10 were considered non-dark pixels. The percentage of these pixels relative to the total image area was calculated.\n\nThis metric provides a simple descriptive measure of how much of the image frame is occupied by visible content. It is not a true retinal tissue segmentation measure.\n\n### Finding\n\nAcross the dataset, the non-dark pixel ratio ranged from approximately 46.81% to 99.97%, with an average of approximately 76.42%.\n\n### Interpretation\n\nThe variation indicates that the amount of non-dark content differs substantially across images. This supports the need to standardize the image field before model training.\n\n### Limitation\n\nThis metric is based only on grayscale intensity thresholding and should not be interpreted as an exact measurement of retinal tissue area.","metadata":{}},{"cell_type":"code","source":"# Register tqdm with pandas for progress tracking\ntqdm.pandas(desc=\"Extracting RGB Channel Means\")\n\ndef get_rgb_channel_means(path):\n    img = cv2.imread(path)\n    if img is None:\n        return np.nan, np.nan, np.nan\n    \n    img = cv2.cvtColor(img, cv2.COLOR_BGR2RGB)\n    \n    # Calculate mean across height and width for each channel (0: R, 1: G, 2: B)\n    return img[:, :, 0].mean(), img[:, :, 1].mean(), img[:, :, 2].mean()\n\n# 1. Apply function across dataset and expand into separate DataFrame columns\nrgb_means = df['file_path'].progress_apply(get_rgb_channel_means)\ndf[['red_mean', 'green_mean', 'blue_mean']] = pd.DataFrame(rgb_means.tolist(), index=df.index)\n\n# 2. Plot Channel Density Distributions\nplt.figure(figsize=(10, 5))\nsns.kdeplot(df['red_mean'], color='red', label='Red Channel', fill=True, alpha=0.3)\nsns.kdeplot(df['green_mean'], color='green', label='Green Channel', fill=True, alpha=0.3)\nsns.kdeplot(df['blue_mean'], color='blue', label='Blue Channel', fill=True, alpha=0.3)\n\nplt.title('Color Channel Mean Distribution Across Retinal Fundus Images', fontsize=12)\nplt.xlabel('Mean Pixel Intensity (0 - 255)', fontsize=10)\nplt.ylabel('Density', fontsize=10)\nplt.legend(title='Channels')\nplt.grid(axis='x', linestyle='--', alpha=0.5)\n\nplt.tight_layout()\nplt.show()\n\n# 3. Print channel metric summary\nprint(\"--- RGB Channel Averages ---\")\nprint(f\"Red Channel Mean:   {df['red_mean'].mean():.2f}\")\nprint(f\"Green Channel Mean: {df['green_mean'].mean():.2f}\")\nprint(f\"Blue Channel Mean:  {df['blue_mean'].mean():.2f}\")","metadata":{},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"## RGB Channel Analysis\n\n### Objective\n\nAnalyze the average intensity of the red, green, and blue channels across the retinal fundus images.\n\n### Method\n\nThe mean pixel intensity of each RGB channel was calculated for every image in the dataset and summarized across the full dataset.\n\n### Finding\n\nThe average channel intensities were:\n\n- Red: 105.53\n- Green: 56.36\n- Blue: 18.79\n\n### Interpretation\n\nThe RGB channel distributions provide descriptive information about the color characteristics of the retinal images and their overall intensity patterns.\n\nThese measurements are exploratory and are not interpreted as evidence that any individual color channel is diagnostically more important for diabetic retinopathy.\n\n### Decision\n\nThe RGB channel statistics were retained as part of the exploratory dataset analysis.","metadata":{}},{"cell_type":"code","source":"# 1. Visualize Brightness Distribution Across Diagnosis Grades\nplt.figure(figsize=(10, 5))\n\n# Boxplot shows the median, quartiles, and outliers per grade\nsns.boxplot(\n    x='diagnosis', \n    y='mean_intensity', \n    data=df, \n    showmeans=True,\n    meanprops={\"marker\": \"o\", \"markerfacecolor\": \"red\", \"markeredgecolor\": \"red\"}\n)\n\n# Overlay individual data points (jittered) to see density\nsns.stripplot(\n    x='diagnosis', \n    y='mean_intensity', \n    data=df, \n    color='black', \n    alpha=0.15, \n    jitter=0.2\n)\n\nplt.title('Image Brightness (Mean Intensity) Across Retinopathy Severity Grades', fontsize=12)\nplt.xlabel('Diagnosis Grade (0: No DR, 1: Mild, 2: Moderate, 3: Severe, 4: Proliferative)', fontsize=10)\nplt.ylabel('Mean Pixel Intensity (0–255)', fontsize=10)\nplt.grid(axis='y', linestyle='--', alpha=0.5)\n\nplt.tight_layout()\nplt.show()\n\n# 2. Statistical Correlation Check (One-Way ANOVA Test)\n# Group intensities by diagnosis grade\ngrade_groups = [group['mean_intensity'].dropna().values for _, group in df.groupby('diagnosis')]\n\n# Perform One-Way ANOVA test\nf_stat, p_val = stats.f_oneway(*grade_groups)\n\nprint(\"--- Statistical Analysis (One-Way ANOVA) ---\")\nprint(f\"F-Statistic: {f_stat:.4f}\")\nprint(f\"p-value:     {p_val:.4e}\")\n\nif p_val > 0.05:\n    print(\"\\nTakeaway: No statistically significant brightness difference across grades (p > 0.05).\")\n    print(\"Lighting variations are independent of disease severity—safe to apply uniform CLAHE/normalization.\")\nelse:\n    print(\"\\nTakeaway: Statistically significant brightness difference detected across grades (p <= 0.05).\")\n    print(\"Careful: Normalize lighting across all grades so the model doesn't learn brightness as a shortcut feature.\")","metadata":{},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"## Brightness Distribution Across Diagnosis Grades\n\n### Objective\n\nExamine whether the distribution of mean image intensity differs across the five diabetic retinopathy grades.\n\n### Method\n\nMean image intensity was compared across diagnosis grades using a boxplot with individual observations.\n\nA one-way ANOVA was then performed to test whether the group means differed statistically.\n\n### Finding\n\nThe ANOVA produced:\n\n- F-statistic: 15.4079\n- p-value: 1.6751 × 10⁻¹²\n\nThe result indicates that mean image intensity differs significantly across at least one pair of diagnosis grades.\n\n### Interpretation\n\nThis analysis does not establish that image brightness causes differences in diabetic retinopathy severity. Instead, it indicates that illumination characteristics vary across the dataset and should be considered during preprocessing.\n\n### Decision\n\nIllumination normalization was therefore evaluated as a preprocessing option, leading to the visual evaluation of CLAHE.","metadata":{}},{"cell_type":"code","source":"# Function to calculate mean brightness across RGB channels for a single image\ndef get_mean_intensity(path):\n    img = cv2.imread(path)\n    if img is None:\n        return np.nan\n    img = cv2.cvtColor(img, cv2.COLOR_BGR2RGB)\n    return np.mean(img)\n\n# 1. Compute mean intensity with progress tracking\ntqdm.pandas(desc=\"Calculating Image Intensities\")\ndf['mean_intensity'] = df['file_path'].progress_apply(get_mean_intensity)\n\n# 2. Plot overall brightness distribution using Seaborn with KDE overlay\nplt.figure(figsize=(8, 4))\nsns.histplot(\n    data=df, \n    x='mean_intensity', \n    bins=30, \n    color='darkorange', \n    kde=True, \n    edgecolor='black'\n)\n\nplt.title('Raw Image Overall Brightness Distribution', fontsize=12)\nplt.xlabel('Mean Pixel Intensity (0 = Pitch Black, 255 = Pure White)', fontsize=10)\nplt.ylabel('Image Count', fontsize=10)\nplt.grid(axis='y', linestyle='--', alpha=0.5)\nplt.tight_layout()\nplt.show()","metadata":{},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"### Finding\n\nThe dataset shows variation in overall image brightness across the retinal images.\n\nThis exploratory analysis provides a baseline description of the illumination characteristics before preprocessing.","metadata":{}},{"cell_type":"code","source":"widths, heights, aspect_ratios = [], [], []\n\n# Extract image dimensions from headers\nfor path in df['file_path']:\n    with Image.open(path) as img:\n        w, h = img.size\n        widths.append(w)\n        heights.append(h)\n        aspect_ratios.append(w / h)\n\n# Plot Spatial Geometry Distributions\nfig, ax = plt.subplots(2, 2, figsize=(14, 8))\n\n# 1. Width Distribution\nsns.histplot(widths, ax=ax[0, 0], color='teal', kde=True, bins=20)\nax[0, 0].set_title('Raw Image Widths Distribution', fontsize=11)\nax[0, 0].set_xlabel('Width (pixels)')\nax[0, 0].set_ylabel('Count')\n\n# 2. Height Distribution\nsns.histplot(heights, ax=ax[0, 1], color='teal', kde=True, bins=20)\nax[0, 1].set_title('Raw Image Heights Distribution', fontsize=11)\nax[0, 1].set_xlabel('Height (pixels)')\nax[0, 1].set_ylabel('Count')\n\n# 3. Width vs Height Scatter\nsns.scatterplot(x=widths, y=heights, ax=ax[1, 0], alpha=0.5, color='purple')\nax[1, 0].set_title('Width vs. Height Scatter Plot', fontsize=11)\nax[1, 0].set_xlabel('Width (pixels)')\nax[1, 0].set_ylabel('Height (pixels)')\n\n# 4. Aspect Ratio Distribution\nsns.histplot(aspect_ratios, ax=ax[1, 1], color='coral', kde=True, bins=20)\nax[1, 1].set_title('Aspect Ratio (W / H) Distribution', fontsize=11)\nax[1, 1].set_xlabel('Aspect Ratio')\nax[1, 1].set_ylabel('Count')\n\nplt.tight_layout()\nplt.show()\n\n# Print Aspect Ratio summary\nunique_ratios = sorted(set([round(r, 2) for r in aspect_ratios]))\nprint(f\"Unique Aspect Ratios found across dataset: {unique_ratios}\")","metadata":{},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"## Image Dimensions and Aspect Ratio Analysis\n\n### Objective\n\nExamine the spatial dimensions and aspect ratios of the retinal images before preprocessing.\n\n### Method\n\nThe width and height of each image were extracted, and the aspect ratio was calculated as:\n\nAspect Ratio = Width / Height\n\nThe distributions were visualized to assess the geometric variability of the dataset.\n\n### Finding\n\nThe dataset contains images with varying spatial dimensions and aspect ratios. The rounded unique aspect ratios ranged from approximately 1.00 to 1.51.\n\n### Decision\n\nBecause the original aspect ratios vary across images, direct resizing to a square input could distort retinal structures.\n\nTherefore, the preprocessing pipeline uses proportional resizing followed by padding to 224×224, preserving the original aspect ratio.","metadata":{}},{"cell_type":"code","source":"# Plot Class Distribution\n\nplt.figure(figsize=(8, 4))\n\nax = sns.countplot(x='diagnosis', data=df)\n\nplt.title('Class Distribution: DR Severity Grades (0–4)', fontsize=12)\nplt.xlabel(\n    'Diagnosis Grade (0: No DR, 1: Mild, 2: Moderate, 3: Severe, 4: Proliferative)',\n    fontsize=10\n)\nplt.ylabel('Image Count', fontsize=10)\n\ntotal = len(df)\n\nfor p in ax.patches:\n    height = p.get_height()\n    percentage = (height / total) * 100\n\n    ax.annotate(\n        f'{int(height)}\\n({percentage:.1f}%)',\n        (p.get_x() + p.get_width() / 2., height),\n        ha='center',\n        va='bottom',\n        xytext=(0, 3),\n        textcoords='offset points',\n        fontsize=9\n    )\n\nplt.ylim(0, max([p.get_height() for p in ax.patches]) * 1.15)\n\nplt.tight_layout()\nplt.show()\n\n\n# Class distribution summary\n\nprint(f\"Total Dataset Size: {len(df)} images\\n\")\nprint(\"Class Counts & Percentages:\")\n\nbreakdown = pd.DataFrame({\n    'Count': df['diagnosis'].value_counts(),\n    'Percentage (%)': np.round(\n        df['diagnosis'].value_counts(normalize=True) * 100,\n        2\n    )\n}).sort_index()\n\ndisplay(breakdown)","metadata":{},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"## Class Distribution Analysis\n\n### Objective\n\nExamine the distribution of diabetic retinopathy severity grades in the APTOS 2019 training dataset.\n\n### Finding\n\nThe dataset is imbalanced across the five diagnosis grades, with Grade 0 representing the largest class and Grades 1 and 3 having substantially fewer samples.\n\nThe class distribution was retained as an important dataset characteristic for consideration during model training.\n\n### Decision\n\nNo class-balancing strategy is fixed at the preprocessing stage.\n\nPotential approaches such as class weighting or training-data augmentation should be evaluated separately during model development.","metadata":{}},{"cell_type":"code","source":"# Function to scan all image paths for file or image-level corruption\ndef scan_for_corrupted_images(df, min_brightness_threshold=5):\n    corrupted_files = []\n    \n    print(f\"Scanning {len(df)} images for corruption and errors...\\n\")\n    \n    for idx, row in tqdm(df.iterrows(), total=len(df)):\n        path = row['file_path']\n        \n        # 1. Check if the file actually exists on disk\n        if not os.path.exists(path):\n            corrupted_files.append({\n                'index': idx, 'file_path': path, 'reason': 'File Not Found'\n            })\n            continue\n            \n        # 2. Check for zero-byte / empty files\n        if os.path.getsize(path) == 0:\n            corrupted_files.append({\n                'index': idx, 'file_path': path, 'reason': 'Zero-Byte File (Empty)'\n            })\n            continue\n            \n        # 3. Attempt to read the image via OpenCV\n        img = cv2.imread(path)\n        if img is None:\n            corrupted_files.append({\n                'index': idx, 'file_path': path, 'reason': 'Unreadable / Corrupted Image'\n            })\n            continue\n            \n        # 4. Check for extreme low-intensity \"blackout\" images\n        mean_val = img.mean()\n        if mean_val < min_brightness_threshold:\n            corrupted_files.append({\n                'index': idx, 'file_path': path, 'reason': f'Severe Blackout Image (Mean Pixel = {mean_val:.2f})'\n            })\n\n    # Summary Output\n    corrupted_df = pd.DataFrame(corrupted_files)\n    \n    if len(corrupted_df) == 0:\n        print(\"\\nSuccess! All images are valid, readable, and non-corrupted.\")\n    else:\n        print(f\"\\nWarning: Found {len(corrupted_df)} problematic files!\")\n        \n    return corrupted_df\n\n# Execute scan and output findings report\ncorrupted_report = scan_for_corrupted_images(df)\n\nif not corrupted_report.empty:\n    display(corrupted_report)","metadata":{},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"## Data Integrity Check\n\n### Objective\n\nVerify that the retinal images satisfy basic file and readability requirements before preprocessing.\n\n### Checks Performed\n\nEach image was checked for:\n\n- File existence\n- Zero-byte files\n- Successful image decoding with OpenCV\n- Extremely low mean pixel intensity (threshold < 5)\n\n### Finding\n\nAll 3,662 images passed the defined integrity checks, with no problematic files detected.\n\n### Interpretation\n\nThe dataset was suitable for proceeding with the preprocessing pipeline under the defined integrity criteria.\n\nThis check does not guarantee the absence of every possible type of image corruption.","metadata":{}},{"cell_type":"code","source":"# Load dataset\ndf = pd.read_csv('/kaggle/input/competitions/aptos2019-blindness-detection/train.csv')\n\nimage_dir = '/kaggle/input/competitions/aptos2019-blindness-detection/train_images'\n\n# Add a column with the direct file path for each image\ndf['file_path'] = df['id_code'].apply(lambda x: os.path.join(image_dir, f\"{x}.png\"))","metadata":{},"outputs":[],"execution_count":null},{"cell_type":"code","source":"# System and File Handling\nimport os\n\n# Data Manipulation, Numerical Processing & Statistical Testing\nimport numpy as np\nimport pandas as pd\nfrom scipy import stats\n\n# Computer Vision & Image Processing\nimport cv2\nfrom PIL import Image\n\n# Data Visualization\nimport matplotlib.pyplot as plt\nimport seaborn as sns\n\n# Progress Monitoring\nfrom tqdm import tqdm","metadata":{},"outputs":[],"execution_count":null}]}