{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"name":"python","version":"3.10.12","mimetype":"text/x-python","codemirror_mode":{"name":"ipython","version":3},"pygments_lexer":"ipython3","nbconvert_exporter":"python","file_extension":".py"},"kaggle":{"accelerator":"none","dataSources":[{"sourceId":92860,"databundleVersionId":11074257,"sourceType":"competition"}],"dockerImageVersionId":30886,"isInternetEnabled":false,"language":"python","sourceType":"notebook","isGpuEnabled":false}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"## 🚀 Environment Setup for Image Segmentation  \n\nThis code sets up the environment for an image segmentation task using **PyTorch** and **Segmentation Models PyTorch (SMP)**.  \n\n### **🔹 Key Components**  \n- **Data Handling**: `pandas`, `numpy`, `os` for file and data management  \n- **Deep Learning**: `torch`, `torchvision.transforms` for model training and preprocessing  \n- **Image Processing**: `PIL.Image`, `cv2` for handling and augmenting images  \n- **Segmentation Models**: `segmentation_models_pytorch` (SMP) for segmentation architectures  \n- **CUDA Support**: Detects if a **GPU** is available and assigns the device  \n","metadata":{}},{"cell_type":"code","source":"import numpy as np \nimport pandas as pd \nimport os\nimport torch\nfrom torch.utils.data import Dataset, DataLoader\nfrom torchvision import transforms\nfrom PIL import Image\nimport torch.nn as nn\nimport torch.optim as optim\nimport torchvision.transforms as transforms\nimport segmentation_models_pytorch as smp\nimport torch\nimport cv2\nfrom torchvision import transforms\n\n# Setting device\ndevice = torch.device(\"cuda\" if torch.cuda.is_available() else \"cpu\")\nprint(f\"Using device: {device}\")","metadata":{"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"## 📂 Dataset Organization and File Management  \n\nDefines **file paths** for the training, validation, and test datasets, ensuring proper organization.  \n\n### **🔹 Key Features**  \n- **Dataset Paths**:  \n  - Stores images in `images/train`, `images/val`, and `images/test`  \n  - Stores corresponding masks in `annotations/train`, `annotations/val`, and `annotations/test`  \n- **Sorting Mechanism**:  \n  - Retrieves image and mask files using `os.listdir()`  \n  - Filters images (`.jpg`) and masks (`.png`)  \n  - Ensures files are **sorted** to maintain correct alignment between images and masks ","metadata":{}},{"cell_type":"code","source":"# Define dataset paths\nbase_path = r\"E:\\slicemyface\"\ntrain_image_path = os.path.join(base_path, \"images/train\")\nval_image_path = os.path.join(base_path, \"images/val\")\ntest_image_path = os.path.join(base_path, \"images/test\")\n\ntrain_mask_path = os.path.join(base_path, \"annotations/train\")\nval_mask_path = os.path.join(base_path, \"annotations/val\")\ntest_mask_path = os.path.join(base_path, \"annotations/test\")\n\n# Function to get sorted file paths\ndef get_file_paths(image_path, mask_path):\n    image_files = sorted([os.path.join(image_path, f) for f in os.listdir(image_path) if f.endswith(\".jpg\")])\n    mask_files = sorted([os.path.join(mask_path, f) for f in os.listdir(mask_path) if f.endswith(\".png\")])\n    return image_files, mask_files\n    \n# Get file paths\ntrain_image_files, train_mask_files = get_file_paths(train_image_path, train_mask_path)\nval_image_files, val_mask_files = get_file_paths(val_image_path, val_mask_path)\ntest_image_files, test_mask_files = get_file_paths(test_image_path, test_mask_path) ","metadata":{"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"# 📊 Verifying Dataset Integrity  \n\nThis snippet prints the **count of images and masks** for the training, validation, and test datasets to ensure they are properly loaded and aligned.  \n","metadata":{}},{"cell_type":"code","source":"# Print counts to verify\nprint(f\"Train Images: {len(train_image_files)}, Masks: {len(train_mask_files)}\")\nprint(f\"Val Images: {len(val_image_files)}, Masks: {len(val_mask_files)}\")\nprint(f\"Test Images: {len(test_image_files)}, Masks: {len(test_mask_files)}\")","metadata":{"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"## 🖼️ Image & Mask Preprocessing for Segmentation  \n\nThis code provides essential functions to **load, resize, and convert images and masks** into tensors, making them compatible for deep learning models in **semantic segmentation** tasks.  \n\n---\n\n### 📏 Global Configuration  \n- **IMG_SIZE = 256**: Standardizes all images and masks to **256×256 pixels** for uniformity.  \n\n```python\nIMG_SIZE = 256  # Resizing all images to (256, 256)\n","metadata":{}},{"cell_type":"code","source":"IMG_SIZE = 256  # Resizing all images to (256, 256)\n\n\ndef load_image(image_path):\n    img = Image.open(image_path).convert(\"RGB\")  # Open image as RGB\n    transform = transforms.Compose([\n        transforms.Resize((IMG_SIZE, IMG_SIZE)),\n        transforms.ToTensor(),  # Convert to tensor and normalize [0,1]\n    ])\n    return transform(img)\n\ndef load_mask(mask_path):\n    mask = Image.open(mask_path).convert(\"L\")  # Open mask as grayscale\n    transform = transforms.Compose([\n        transforms.Resize((IMG_SIZE, IMG_SIZE), interpolation=Image.NEAREST), \n        transforms.PILToTensor(),  \n        transforms.Lambda(lambda x: x.long())  \n    ])\n    return transform(mask)\n\ndef torch_load_data(image_path, mask_path):\n    img = load_image(image_path)\n    mask = load_mask(mask_path)\n    return img, mask\n","metadata":{"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"## 🏷️ Custom Dataset Class for Facial Segmentation  \n\nDefines a **PyTorch Dataset** for handling **facial image segmentation**. It efficiently loads, processes, and returns **images, masks, and image IDs**, making it suitable for training deep learning models.  \n\n---\n\n### ⚙️ Batch Size Configuration  \n- **BATCH_SIZE = 8** ensures a balanced memory load while training.  \n\n```python\nBATCH_SIZE = 8\n","metadata":{}},{"cell_type":"code","source":"BATCH_SIZE = 8\n\nclass FacialSegmentationDataset(Dataset):\n    def __init__(self, image_files, mask_files, image_transform=None, mask_transform=None):\n        self.image_files = image_files\n        self.mask_files = mask_files\n        self.image_transform = image_transform\n        self.mask_transform = mask_transform\n\n    def __len__(self):\n        return len(self.image_files)\n\n    def __getitem__(self, idx):\n        image = Image.open(self.image_files[idx]).convert(\"RGB\")\n        mask = Image.open(self.mask_files[idx])  # Keep as single-channel grayscale\n        \n        # Extract image ID from the filename (assuming filenames are unique)\n        image_id = self.image_files[idx].split('/')[-1].split('.')[0]  # Extract file name without extension\n        \n        if self.image_transform:\n            image = self.image_transform(image)\n        if self.mask_transform:\n            mask = self.mask_transform(mask)\n\n        return image, mask, image_id  \n","metadata":{"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"## 🖼️ Verifying Mask Processing  \n\nThis snippet **loads and verifies a segmentation mask** from the dataset to ensure correct preprocessing.  \n\n---","metadata":{}},{"cell_type":"code","source":"mask_path = train_mask_files[0]\n\nmask = load_mask(mask_path)\nprint(mask.shape)  # Should be (IMG_SIZE, IMG_SIZE)\nprint(mask.unique())  # Should contain only valid class labels\n","metadata":{"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"## 🎨 Data Augmentation and Preprocessing  \n\nThis section **defines transformations** for both images and segmentation masks.  \n\n---\n\n### 🖼️ **Image Transformations**  \nThe following augmentations are applied to training images:  \n\n- 🔹 **Resize** → `(256, 256)` consistent input size.  \n- 🔹 **Color Jitter** → Adjusts brightness, contrast, saturation, and hue.  \n- 🔹 **Random Horizontal & Vertical Flip** → Increases diversity of perspectives.  \n- 🔹 **Random Affine Transform** → Adds small rotations, translations, and scaling variations.  \n- 🔹 **Gaussian Blur** → Simulates natural noise in images.  \n- 🔹 **ToTensor & Normalize** → Converts to tensor and normalizes based on ImageNet statistics.  \n","metadata":{}},{"cell_type":"code","source":"# Define transformations\n\nimage_transform = transforms.Compose([\n    transforms.Resize((IMG_SIZE, IMG_SIZE)),\n    transforms.ColorJitter(brightness=0.3, contrast=0.3, saturation=0.3, hue=0.2), \n    transforms.RandomHorizontalFlip(p=0.5), \n    transforms.RandomVerticalFlip(p=0.5),\n    transforms.RandomAffine(degrees=15, translate=(0.1, 0.1), scale=(0.8, 1.2)),\n    transforms.GaussianBlur(kernel_size=3),\n    transforms.ToTensor(),\n    transforms.Normalize(mean=[0.485, 0.456, 0.406], std=[0.229, 0.224, 0.225]),\n])\n\nmask_transform = transforms.Compose([\n    transforms.Resize((IMG_SIZE, IMG_SIZE), interpolation=Image.NEAREST),\n    transforms.PILToTensor(),  \n    transforms.Lambda(lambda x: x.long())  \n])","metadata":{"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"## 📂 Creating Datasets & DataLoaders  \n\nThis section **initializes datasets and DataLoaders** for training, validation, and testing.\n\n---\n\n### 🗂️ **Dataset Initialization**  \nWe create instances of `FacialSegmentationDataset` for:  \n✔️ **Training** → Uses augmented images and masks.  \n✔️ **Validation** → Evaluates performance on unseen data.  \n✔️ **Testing** → Assesses final model performance.","metadata":{}},{"cell_type":"code","source":"# Create dataset and DataLoader\ntrain_dataset = FacialSegmentationDataset(train_image_files, train_mask_files, image_transform, mask_transform)\nval_dataset = FacialSegmentationDataset(val_image_files, val_mask_files, image_transform, mask_transform)\ntest_dataset = FacialSegmentationDataset(test_image_files, test_mask_files, image_transform, mask_transform)\n\ntrain_loader = DataLoader(train_dataset, batch_size=BATCH_SIZE, shuffle=True, num_workers=0, pin_memory=True)\nval_loader = DataLoader(val_dataset, batch_size=BATCH_SIZE, shuffle=False, num_workers=0, pin_memory=True)\ntest_loader = DataLoader(test_dataset, batch_size=1, shuffle=False, num_workers=0, pin_memory=True)","metadata":{"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"## 🏗️ Initializing the DeepLabV3+ Model  \n\nThis section **sets up the DeepLabV3+ model** for semantic segmentation.\n\n---\n\n### 🧠 **Model Architecture**  \nWe use **DeepLabV3+**, a powerful segmentation model with:  \n✔️ **ResNet-34 Encoder** → Pretrained on ImageNet for feature extraction.  \n✔️ **Atrous Spatial Pyramid Pooling (ASPP)** → Captures multi-scale context.  \n✔️ **Decoder with Upsampling** → Produces high-resolution masks.","metadata":{}},{"cell_type":"code","source":"num_classes=9\nmodel = smp.DeepLabV3Plus(\n    encoder_name=\"resnet34\",  \n    encoder_weights=\"imagenet\",\n    in_channels=3\n)\n\nprint(f\"Model initialized with {num_classes} classes.\"","metadata":{"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"# 🎯 Multiclass Dice Loss for Semantic Segmentation\n\n## 🔹 Function Overview  \nComputes the **Dice loss** for multi-class segmentation, measuring the overlap between predictions and ground truth masks.  ","metadata":{}},{"cell_type":"code","source":"\nimport torch.nn.functional as F\n\n\ndef dice_loss_multiclass(pred, target, smooth=1e-6):\n    pred = torch.softmax(pred, dim=1)\n    num_classes = pred.shape[1]  \n\n    if target.dim() == 4:  # If (N, 1, H, W), remove extra dimension\n        target = target.squeeze(1)\n    \n    # Convert target to one-hot encoding\n    target_one_hot = F.one_hot(target, num_classes=num_classes)  \n    target_one_hot = target_one_hot.permute(0, 3, 1, 2).float()  \n    \n    intersection = (pred * target_one_hot).sum(dim=(2, 3))\n    union = pred.sum(dim=(2, 3)) + target_one_hot.sum(dim=(2, 3))\n    \n    dice_score = (2. * intersection + smooth) / (union + smooth)\n    \n    return 1 - dice_score.mean()","metadata":{"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"# 📊 Model Evaluation with Dice Score  \n\n## 🔹 Overview  \nThis function evaluates a segmentation model using the **Dice Score**, a metric that measures the overlap between predicted and ground truth masks. The function processes an entire dataset, computing Dice scores batch-wise and returning the **average Dice score** as the final evaluation metric.\n","metadata":{}},{"cell_type":"code","source":"import torch.nn.functional as F\nfrom torch.utils.data import DataLoader\nimport segmentation_models_pytorch as smp\nfrom tqdm import tqdm\n\n\ndef evaluate_model(model, dataloader, num_classes, device):\n    model.eval()\n    dice_scores = []\n    \n    with torch.no_grad():\n        for images, masks, image_ids in tqdm(dataloader, desc=\"Evaluating\"): \n            images, masks = images.to(device), masks.to(device)\n            outputs = model(images)\n            dice = 1 - dice_loss_multiclass(outputs, masks, num_classes).item() \n            dice_scores.append(dice)\n    \n    return sum(dice_scores) / len(dice_scores)\n","metadata":{"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"# 🚀 Model Training with Dice Loss & Adaptive Learning Rate  \n\n## 🔹 Overview  \nThis function trains a deep learning segmentation model using **Dice Loss** as the optimization criterion. It employs an **AdamW optimizer** and an **adaptive learning rate scheduler** to improve convergence. The training process monitors validation Dice scores and saves the best-performing model.\n\n## 🔹 Training Workflow  \n1. **Model Setup** → Moves the model to the appropriate device (`CPU` or `CUDA`).  \n2. **Optimizer & Scheduler** → Uses `AdamW` with weight decay and `StepLR` to dynamically adjust learning rate.  \n3. **Batch-wise Training**  \n4. **Validation Checkpointing** → After each epoch:   \n5. **Completion Notification** → Prints final status when training is done.","metadata":{}},{"cell_type":"code","source":"import torch.optim as optim\nfrom tqdm import tqdm\n\nimport torch.optim as optim\nfrom tqdm import tqdm\n\ndef train_model(model, train_loader, val_loader, num_classes, num_epochs=10, lr=0.0001, device=device):\n    optimizer = optim.AdamW(model.parameters(), lr=lr, weight_decay=1e-4)\n\n\n    scheduler = torch.optim.lr_scheduler.StepLR(optimizer, step_size=10, gamma=0.5)\n\n    best_val_dice = 0.0  \n\n    for epoch in range(num_epochs):\n        model.train()\n        train_loss = 0.0\n        \n        for images, masks, image_ids in tqdm(train_loader, desc=f\"Epoch {epoch+1}/{num_epochs} - Training\"):  # Unpacking image_id\n            images, masks = images.to(device), masks.to(device)\n            optimizer.zero_grad()\n            \n            outputs = model(images)\n            loss = dice_loss_multiclass(outputs, masks, num_classes)  # Ensure num_classes is passed here\n            loss.backward()\n            optimizer.step()\n            \n            train_loss += loss.item()\n        \n        avg_train_loss = train_loss / len(train_loader)\n        avg_val_dice = evaluate_model(model, val_loader, num_classes, device)  # Updated evaluate_model also unpacks 3 values\n\n        # Step the scheduler based on validation Dice score\n        scheduler.step(avg_val_dice)\n\n        print(f\"Epoch {epoch+1}/{num_epochs}: Train Loss: {avg_train_loss:.4f}, Val Dice: {avg_val_dice:.4f}, LR: {optimizer.param_groups[0]['lr']:.6f}\")\n\n        # Save the best model\n        if avg_val_dice > best_val_dice:\n            best_val_dice = avg_val_dice\n            torch.save(model.state_dict(), \"Model.pth\")\n            print(\"Model saved!\")\n\n    print(\"Training complete.\")\n    \n# Move model to the appropriate device\nmodel.to(device)\n\n# Training loop execution\ntrain_model(model, train_loader, val_loader, num_classes, num_epochs=10, lr=0.0001, device=device)","metadata":{"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"# 🔄 Convert Binary Mask to Run-Length Encoding (RLE)  \n\n## 📌 Overview  \nThis function converts a **binary segmentation mask** into **Run-Length Encoding (RLE)**, a compact format commonly used in image segmentation tasks. RLE represents the mask as pairs of (start position, length), making it efficient for storage and submission in competitions like **Kaggle’s segmentation challenges**.\n\n## 🔹 Function Logic  \n1. **Flatten Mask 🖼️** → Converts the 2D mask into a 1D array (`pixels`).  \n2. **RLE Encoding 🔢** → Iterates through `use_mask`, recording start indices and lengths of consecutive nonzero regions.  ","metadata":{}},{"cell_type":"code","source":"# Function to convert a binary mask to RLE\ndef mask_to_rle(mask):\n    pixels = mask.flatten()\n    use_mask = pixels > 0  \n    rle = []\n    last_val = 0\n    for i, val in enumerate(use_mask, start=1):\n        if val and not last_val:\n            rle.append(i)\n        if not val and last_val:\n            rle.append(i - rle[-1])\n        last_val = val\n    if last_val:\n        rle.append(len(use_mask) - rle[-1])\n    return \" \".join(map(str, rle)) if rle else \"\"\n","metadata":{"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"# 🖼️ Visualizing Segmentation Results  \n\n## 📌 Overview  \nThis function **visualizes** a sample from the dataset by displaying:  \n1️⃣ The **input image** 🏞️  \n2️⃣ The **true segmentation mask** 🎭  \n3️⃣ The **predicted mask** by the model 🧠\n\n## 🔹 Function Logic  \n1. **Convert Image Tensor** → Changes (C, H, W) format to (H, W, C) for proper visualization.  \n2. **Flatten Mask Tensor** → Removes unnecessary dimensions.  \n3. **Use `imshow()` for Display** → Shows the **original image**, **true mask**, and **predicted mask**.  \n4. **Removes Unnecessary Conversions** → Ensures predicted mask remains in correct format.  \n","metadata":{}},{"cell_type":"code","source":"import matplotlib.pyplot as plt\nimport numpy as np\n\ndef visualize_sample(image, mask, predicted_mask):\n\n    image = image.permute(1, 2, 0).numpy()  \n    mask = mask.squeeze().numpy()          \n\n\n    # Plot original image, true mask, and predicted mask\n    fig, axes = plt.subplots(1, 3, figsize=(12, 4))\n    axes[0].imshow(image)  \n    axes[0].set_title(\"Input Image\")\n    axes[0].axis(\"off\")\n\n    axes[1].imshow(mask, cmap=\"gray\")  \n    axes[1].set_title(\"True Mask\")\n    axes[1].axis(\"off\")\n\n    axes[2].imshow(predicted_mask, cmap=\"gray\")  # Use directly\n    axes[2].set_title(\"Predicted Mask\")\n    axes[2].axis(\"off\")\n\n    plt.show()\n","metadata":{"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"# 📊 Evaluating the Model on Test Data  \n\n## 📌 Overview  \nThis function **evaluates** the trained segmentation model on a test dataset and **generates a CSV submission file**.  \n\n## 🔹 Key Features   \n- **Saves Results to CSV 📂** → Generates a submission file for competitions or analysis.  \n- **One-Time Visualization 👀** → Displays one sample prediction for validation.\n\n ## 📌 Example CSV Output  \n| id         | predicted              |  \n|------------|------------------------|  \n| `img_001` | `3 5 7 9 11 13 ...`     |  \n| `img_002` | `2 4 6 8 10 12 ...`     |  ","metadata":{}},{"cell_type":"code","source":"import os\nimport torch\nimport pandas as pd\nfrom tqdm import tqdm\n\ndef evaluate_test_model(model, test_loader, num_classes, device, output_file=\"submission_5.csv\"):\n    model.eval()\n    results = []\n    visualized = False  # Flag to track if visualization has been done\n\n    with torch.no_grad():\n        for images, masks, image_ids in tqdm(test_loader, desc=\"Evaluating Test Data\"):\n            images = images.to(device)\n            outputs = model(images)\n            predicted_masks = torch.argmax(outputs, dim=1).cpu().numpy()  # Convert to class mask\n            \n            for i, image_id in enumerate(image_ids):\n                rle_encoded_mask = mask_to_rle(predicted_masks[i])\n                results.append({\"id\": image_id, \"predicted\": rle_encoded_mask})\n\n                if not visualized:\n                    visualize_sample(images[i].cpu(), masks[i].cpu(), predicted_masks[i])\n                    visualized = True  # Set flag to True\n    \n    df = pd.DataFrame(results)\n    df.to_csv(output_file, index=False)\n    print(f\"Submission saved to {output_file}\")\n    \n# Move model to the appropriate device\nmodel.to(device)\n\n# Call the function to evaluate test data\nevaluate_test_model(model, test_loader, num_classes, device)","metadata":{"trusted":true},"outputs":[],"execution_count":null}]}