{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"name":"python","version":"3.12.12","mimetype":"text/x-python","codemirror_mode":{"name":"ipython","version":3},"pygments_lexer":"ipython3","nbconvert_exporter":"python","file_extension":".py"},"kaggle":{"accelerator":"gpu","dataSources":[{"sourceType":"datasetVersion","sourceId":5789588,"datasetId":3325988,"databundleVersionId":5866258}],"dockerImageVersionId":31287,"isInternetEnabled":false,"language":"python","sourceType":"notebook","isGpuEnabled":true}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"# 🚀 Project #10: HubMap AI Engine (Vasculature Segmentation)\n**Autonomous Detection and Pixel-Level Mapping of Human Vasculature**\n\n**Project Objective:** To develop a deep learning-based semantic segmentation engine capable of autonomously identifying vascular structures (blood vessels) within microscopic tissue imagery. This project focuses on medical image processing and autonomous mapping through custom neural architectures.\n\n---\n\n### 🏗️ Project Architecture: 10-Step Development Pipeline\n\n1.  **Defining the Project Goal:** Building a high-precision semantic segmentation engine for the HubMap dataset.\n2.  **Exploratory Data Analysis (EDA):** Visualizing microscopic tissue samples and their corresponding ground-truth vascular masks.\n3.  **Path & Column Selection:** Mapping the exact directory structures for `images` and `masks` within the Kaggle input environment.\n4.  **Data Manipulation:** Utilizing `map` and `lambda` functions to optimize high-resolution imagery into 128x128 feature sets in memory.\n5.  **Data Cleaning:** Ensuring dataset integrity by filtering out mismatched pairs and verifying file accessibility.\n6.  **Feature Engineering:** Implementing Feature Scaling by normalizing pixel values from the 0-255 range to 0.0-1.0.\n7.  **Encoding Strategy:** Preparing 2D binary masks as the target output instead of traditional categorical encoding.\n8.  **Data Splitting (X and y):** Partitioning the data into Training and Testing sets using `train_test_split` to prevent data leakage.\n9.  **Model Training (Fit & Predict):** Executing a Sequential CNN architecture fortified with `BatchNormalization` and `Dropout` layers for autonomous mapping.\n10. **Performance Audit:** Evaluating model success through a \"Visual Audit,\" comparing predicted vascular maps against true biological labels.\n\n","metadata":{}},{"cell_type":"code","source":"import os\nimport cv2\nimport glob\nimport numpy as np\nimport matplotlib.pyplot as plt\nfrom sklearn.model_selection import train_test_split\nimport tensorflow as tf\nfrom tensorflow.keras.models import Sequential\nfrom tensorflow.keras.layers import Conv2D, InputLayer, Dropout, BatchNormalization\nfrom tensorflow.keras.callbacks import EarlyStopping\nimport warnings\nwarnings.filterwarnings('ignore')","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-04-21T09:45:16.986992Z","iopub.execute_input":"2026-04-21T09:45:16.987273Z","iopub.status.idle":"2026-04-21T09:45:16.991914Z","shell.execute_reply.started":"2026-04-21T09:45:16.987250Z","shell.execute_reply":"2026-04-21T09:45:16.991331Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"# 2. Read and Explore Data (EDA) & 3. Select Columns\nIMG_DIR = '/kaggle/input/datasets/aneesh10/hubmap-hacking-the-human-vasculature-processed/images/images/'\nMASK_DIR = '/kaggle/input/datasets/aneesh10/hubmap-hacking-the-human-vasculature-processed/masks/masks/'","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-04-21T09:45:18.064700Z","iopub.execute_input":"2026-04-21T09:45:18.065348Z","iopub.status.idle":"2026-04-21T09:45:18.068304Z","shell.execute_reply.started":"2026-04-21T09:45:18.065322Z","shell.execute_reply":"2026-04-21T09:45:18.067772Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"# Selecting the first 800 images to prevent Kaggle RAM crashes (Memory Management)\nimg_paths = sorted(glob.glob(IMG_DIR + '*.*'))[:800]\nmask_paths = sorted(glob.glob(MASK_DIR + '*.*'))[:800]","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-04-21T09:45:18.760071Z","iopub.execute_input":"2026-04-21T09:45:18.760293Z","iopub.status.idle":"2026-04-21T09:45:18.771992Z","shell.execute_reply.started":"2026-04-21T09:45:18.760274Z","shell.execute_reply":"2026-04-21T09:45:18.771435Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"# 4. Data Manipulation & 5. Cleaning \n# to resize all images and masks to 128x128 natively in memory.\nX_raw = list(map(lambda p: cv2.resize(cv2.cvtColor(cv2.imread(p), cv2.COLOR_BGR2RGB), (128, 128)), img_paths))\ny_raw = list(map(lambda p: cv2.resize(cv2.imread(p, cv2.IMREAD_GRAYSCALE), (128, 128)), mask_paths))","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-04-21T09:45:45.855959Z","iopub.execute_input":"2026-04-21T09:45:45.856588Z","iopub.status.idle":"2026-04-21T09:45:57.298075Z","shell.execute_reply.started":"2026-04-21T09:45:45.856559Z","shell.execute_reply":"2026-04-21T09:45:57.297307Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"# 6. Feature Engineering & 7. Encoding\n# Pixel Normalization (Scaling pixel values from 0-255 to 0.0-1.0)\n# We use binary masks (0 and 1) instead of One-Hot Encoding for segmentation.\nX_train_full = np.array(X_raw) / 255.0\ny_train_full = np.expand_dims(np.array(y_raw) / 255.0, axis=-1)","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-04-21T09:46:03.977573Z","iopub.execute_input":"2026-04-21T09:46:03.978054Z","iopub.status.idle":"2026-04-21T09:46:04.128054Z","shell.execute_reply.started":"2026-04-21T09:46:03.978030Z","shell.execute_reply":"2026-04-21T09:46:04.127326Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"# 8. Split Data into X and y\ntrain_images, test_images, train_labels, test_labels = train_test_split(\n    X_train_full, y_train_full, test_size=0.2, random_state=42\n)","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-04-21T09:46:05.801509Z","iopub.execute_input":"2026-04-21T09:46:05.802078Z","iopub.status.idle":"2026-04-21T09:46:05.929575Z","shell.execute_reply.started":"2026-04-21T09:46:05.802050Z","shell.execute_reply":"2026-04-21T09:46:05.929005Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"# 9. Train and Predict (Fit-Predict)\nmodel = Sequential([\n    InputLayer(input_shape=(128, 128, 3)),\n    \n    # --- CNN Block 1 ---\n    Conv2D(filters=32, kernel_size=(3,3), activation='relu', padding='same'),\n    BatchNormalization(),\n    Dropout(0.2),\n    \n    # --- CNN Block 2 ---\n    Conv2D(filters=64, kernel_size=(3,3), activation='relu', padding='same'),\n    BatchNormalization(),\n    Dropout(0.2),\n    \n    # --- Output Layer (Segmentation Mask) ---\n    # Creating a 2D mask matrix instead of using Flatten/Dense\n    Conv2D(filters=1, kernel_size=(3,3), activation='sigmoid', padding='same')\n])\n\nmodel.compile(optimizer='adam', loss='binary_crossentropy', metrics=['accuracy'])\n\nearly_stop = EarlyStopping(monitor='val_loss', patience=10)","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-04-21T09:46:23.445883Z","iopub.execute_input":"2026-04-21T09:46:23.446155Z","iopub.status.idle":"2026-04-21T09:46:23.492898Z","shell.execute_reply.started":"2026-04-21T09:46:23.446133Z","shell.execute_reply":"2026-04-21T09:46:23.492370Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"# Training the model (fit)\nmodel.fit(train_images, train_labels, \n          validation_data=(test_images, test_labels), \n          epochs=20, \n          batch_size=16, \n          callbacks=[early_stop])\n\n# Predictions (predict)\nindex = 1\nimage_to_predict = test_images[index].reshape(1, 128, 128, 3) \npredicted_mask = model.predict(image_to_predict)","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-04-21T09:46:26.369310Z","iopub.execute_input":"2026-04-21T09:46:26.370029Z","iopub.status.idle":"2026-04-21T09:46:46.520753Z","shell.execute_reply.started":"2026-04-21T09:46:26.370001Z","shell.execute_reply":"2026-04-21T09:46:46.520019Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"\n# 10. Measure Model Accuracy (Visual Audit)\n# In segmentation, visual plotting often replaces standard classification reports.\nprint(\"\\n--- Step 10: Visualizing Segmentation Results ---\")\n\nplt.figure(figsize=(12, 4))\n\nplt.subplot(1, 3, 1)\nplt.imshow(test_images[index])\nplt.title(\"1. Original Microscopic Image\")\nplt.axis('off')\n\nplt.subplot(1, 3, 2)\nplt.imshow(test_labels[index], cmap='gray')\nplt.title(\"2. True Blood Vessels\")\nplt.axis('off')\n\nplt.subplot(1, 3, 3)\nplt.imshow(predicted_mask[0], cmap='gray')\nplt.title(\"3. Predicted Blood Vessels\")\nplt.axis('off')\n\nplt.tight_layout()\nplt.show()","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-04-21T09:46:54.886880Z","iopub.execute_input":"2026-04-21T09:46:54.887361Z","iopub.status.idle":"2026-04-21T09:46:55.287431Z","shell.execute_reply.started":"2026-04-21T09:46:54.887335Z","shell.execute_reply":"2026-04-21T09:46:55.286686Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"# 11. Save the Model\nmodel.save('HubMap_Vessels_Model.keras')\n","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-04-21T09:51:44.211972Z","iopub.execute_input":"2026-04-21T09:51:44.212794Z","iopub.status.idle":"2026-04-21T09:51:44.278060Z","shell.execute_reply.started":"2026-04-21T09:51:44.212763Z","shell.execute_reply":"2026-04-21T09:51:44.277551Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"# 🏁 Project #10 Final Audit: HubMap Vasculature Segmentation\n\n## 🚀 LIVE DEPLOYMENT\nThe HubMap AI engine has been successfully serialized and deployed as a real-time medical image segmentation service. You can test the autonomous vascular mapping engine here:\n### 🔗 [LIVE ENGINE: HubMap AI Engine on Hugging Face](https://huggingface.co/spaces/Ironside35/HubMap-AI-Engine)\n\n---\n\n## 🧠 Architectural Conclusion & Visual Analysis\n\n### 1. Data Pipeline Integrity\nThe visual audit confirms that our custom functional data pipeline successfully extracted and mapped the complex HubMap dataset. The microscopic tissue structures (Image 1) align perfectly with the ground truth binary masks (Image 2), proving that our `map/lambda` optimization loaded the arrays without any spatial distortion or memory overflow.\n\n### 2. The Probabilistic Output (Understanding the Blur)\nIn **Image 3 (Predicted Blood Vessels)**, the output appears as a hazy, grayscale heatmap rather than a crisp black-and-white mask. **This is an expected architectural result.** Since we utilized a simplified `Sequential` CNN (as per baseline requirements) instead of a complex U-Net Decoder, the model outputs \"probabilities\" via the `sigmoid` activation. The gray areas represent the model's uncertainty, effectively acting as a \"probability map\" of where the vasculature is likely located.\n\n### 3. Future Engineering (The U-Net Requirement)\nThis project visually demonstrates a core concept in Computer Vision: While a standard Sequential CNN can *localize* features, it lacks the skip-connections and upsampling power needed for sharp *pixel-perfect segmentation*. To upgrade this probabilistic heatmap into a crisp medical mask, the next phase of development would involve transitioning from this Sequential baseline to a full **U-Net Architecture**.\n\n","metadata":{}},{"cell_type":"code","source":"","metadata":{"trusted":true},"outputs":[],"execution_count":null}]}