{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"pygments_lexer":"ipython3","nbconvert_exporter":"python","version":"3.6.4","file_extension":".py","codemirror_mode":{"name":"ipython","version":3},"name":"python","mimetype":"text/x-python"}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"","metadata":{}},{"cell_type":"markdown","source":"# Imports","metadata":{}},{"cell_type":"code","source":"from glob import glob\nimport pandas as pd\nimport json\nimport numpy as np\nimport matplotlib.pyplot as plt","metadata":{"execution":{"iopub.status.busy":"2023-05-27T17:21:46.670789Z","iopub.execute_input":"2023-05-27T17:21:46.671876Z","iopub.status.idle":"2023-05-27T17:21:46.676685Z","shell.execute_reply.started":"2023-05-27T17:21:46.671833Z","shell.execute_reply":"2023-05-27T17:21:46.675716Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Take a Glance","metadata":{}},{"cell_type":"markdown","source":"","metadata":{}},{"cell_type":"code","source":"\n\nprint(\"The number of train images = \",len(glob(\"/kaggle/input/hubmap-hacking-the-human-vasculature/train/*\")))\n","metadata":{"execution":{"iopub.status.busy":"2023-05-27T17:21:46.678710Z","iopub.execute_input":"2023-05-27T17:21:46.679300Z","iopub.status.idle":"2023-05-27T17:21:46.719117Z","shell.execute_reply.started":"2023-05-27T17:21:46.679265Z","shell.execute_reply":"2023-05-27T17:21:46.718038Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"df_tile = pd.read_csv(\"/kaggle/input/hubmap-hacking-the-human-vasculature/tile_meta.csv\")\ndf_tile.head()","metadata":{"execution":{"iopub.status.busy":"2023-05-27T17:21:46.721315Z","iopub.execute_input":"2023-05-27T17:21:46.721664Z","iopub.status.idle":"2023-05-27T17:21:46.740404Z","shell.execute_reply.started":"2023-05-27T17:21:46.721630Z","shell.execute_reply":"2023-05-27T17:21:46.739480Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Based on the data info: **tile_meta.csv** Metadata for each image.\n* **source_wsi** Identifies the WSI this tile was extracted from.\n* **{i|j}** The location of the upper-left corner within the WSI where the tile was extracted.\n* **dataset** The dataset this tile belongs to, as described above.","metadata":{}},{"cell_type":"code","source":"df_wsi = pd.read_csv(\"/kaggle/input/hubmap-hacking-the-human-vasculature/wsi_meta.csv\")\ndf_wsi","metadata":{"execution":{"iopub.status.busy":"2023-05-27T17:21:46.746593Z","iopub.execute_input":"2023-05-27T17:21:46.746876Z","iopub.status.idle":"2023-05-27T17:21:46.764599Z","shell.execute_reply.started":"2023-05-27T17:21:46.746852Z","shell.execute_reply":"2023-05-27T17:21:46.763770Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Base on data info: **wsi_meta.csv** Metadata for the **Whole Slide Images (WSI)** the tiles were extracted from.\n* **source_wsi** Identifies the WSI.\n* **age, sex, race, height, weight, and bmi** demographic information about the tissue donor.","metadata":{}},{"cell_type":"code","source":"polygons = \"/kaggle/input/hubmap-hacking-the-human-vasculature/polygons.jsonl\"\n\nimport json\n\n# Open the JSONL file\ndata = []\nwith open(polygons, 'r') as file:\n    for line in file:\n        item = json.loads(line)\n        data.append(item)\n\n# Convert JSONL data to DataFrame\ndf = pd.DataFrame(data)\n\n# Display the DataFrame\ndf.head()","metadata":{"execution":{"iopub.status.busy":"2023-05-27T17:21:46.768677Z","iopub.execute_input":"2023-05-27T17:21:46.768942Z","iopub.status.idle":"2023-05-27T17:21:50.156698Z","shell.execute_reply.started":"2023-05-27T17:21:46.768920Z","shell.execute_reply":"2023-05-27T17:21:50.155703Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"**polygons.jsonl** Polygonal segmentation masks in JSONL format, available for Dataset 1 and Dataset 2. Each line gives JSON annotations for a single image with:\n* **id** Identifies the corresponding image in train/\n* **annotations** A list of mask annotations with:\n* **type** Identifies the type of structure annotated:\n    1. **blood_vessel** The target structure. **Your goal in this competition** is to predict these kinds of masks on the test set.\n    2. **glomerulus** A capillary ball structure in the kidney. These parts of the images were excluded from blood vessel annotation. You should ensure none of your test set predictions occur within glomerulus structures as they will be counted as false positives. Annotations are provided for test set tiles.\n    3. **unsure** A structure the expert annotators cannot confidently distinguish as a blood vessel.\n    4. **coordinates** A list of polygon coordinates defining the segmentation mask.","metadata":{}},{"cell_type":"code","source":"# df[\"annotations\"][0]","metadata":{"execution":{"iopub.status.busy":"2023-05-27T17:21:50.159069Z","iopub.execute_input":"2023-05-27T17:21:50.160052Z","iopub.status.idle":"2023-05-27T17:21:50.165067Z","shell.execute_reply.started":"2023-05-27T17:21:50.160015Z","shell.execute_reply":"2023-05-27T17:21:50.164220Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"import matplotlib.pyplot as plt\nfrom matplotlib.patches import Polygon\n\n# Extract Mask Annotations\nmasks = []\nfor _, row in df.iterrows():\n    annotations = row['annotations']\n    for annotation in annotations:\n        if annotation['type'] == 'blood_vessel':\n            coordinates = annotation['coordinates']\n            masks.append(coordinates)\n\n# Visualize the Masks\nfor i, mask in enumerate(masks):\n    plt.figure()\n    ax = plt.gca()\n    for polygon in mask:\n        poly = Polygon(polygon, edgecolor='r', facecolor='none')\n        ax.add_patch(poly)\n    plt.title('Mask')\n    plt.axis('scaled')\n    plt.show()\n    if i == 1:\n        break\n    \n","metadata":{"execution":{"iopub.status.busy":"2023-05-27T17:21:50.166344Z","iopub.execute_input":"2023-05-27T17:21:50.167458Z","iopub.status.idle":"2023-05-27T17:21:50.714765Z","shell.execute_reply.started":"2023-05-27T17:21:50.167425Z","shell.execute_reply":"2023-05-27T17:21:50.713720Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# creat a list of images path\nimages_path = glob(\"/kaggle/input/hubmap-hacking-the-human-vasculature/train/*\")\nimages_path[:3]","metadata":{"execution":{"iopub.status.busy":"2023-05-27T17:21:50.717385Z","iopub.execute_input":"2023-05-27T17:21:50.717846Z","iopub.status.idle":"2023-05-27T17:21:50.748686Z","shell.execute_reply.started":"2023-05-27T17:21:50.717809Z","shell.execute_reply":"2023-05-27T17:21:50.747700Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Dataset and DataLoader","metadata":{}},{"cell_type":"code","source":"import torch\nfrom torch.utils.data import Dataset\nfrom PIL import Image\nimport os\nimport json\nimport numpy as np\nfrom skimage.draw import polygon2mask\n\nclass CustomDataset(Dataset):\n    def __init__(self, image_dir, labels_file, transform=None):\n        # Open the labels file and load the JSON data\n        with open(labels_file, 'r') as json_file:\n            # Read each line in the JSON file and parse it as JSON\n            self.json_labels = [json.loads(line) for line in json_file]\n\n        # Set the image directory, labels file, and transformation\n        self.image_dir = image_dir\n        self.transform = transform\n\n    def __len__(self):\n        # Return the number of samples in the dataset\n        return len(self.json_labels)\n\n    def __getitem__(self, idx):\n        # Load image\n        image_path = os.path.join(self.image_dir, f\"{self.json_labels[idx]['id']}.tif\")\n        image = Image.open(image_path)\n\n        # Initialize mask\n        mask = np.zeros((512, 512), dtype=np.float32)\n\n        # Process annotations\n        for annot in self.json_labels[idx]['annotations']:\n            cords = annot['coordinates']\n            if annot['type'] == \"blood_vessel\":\n                # Iterate over the coordinates of the annotation\n                for cord in cords:\n                    # Extract the x and y coordinates from the coordinates list\n                    rr, cc = np.array([i[1] for i in cord]), np.asarray([i[0] for i in cord])\n                    # Set the corresponding pixels in the mask to 1\n                    mask[rr, cc] = 1\n\n        # Convert PIL Image and mask to PyTorch tensor\n        image = torch.tensor(np.array(image), dtype=torch.float32).permute(2, 0, 1)  # Shape: [C, H, W]\n        mask = torch.tensor(mask, dtype=torch.float32)\n\n        if self.transform:\n            # Apply the transformation to the image if provided\n            image = self.transform(image)\n\n        return image, mask\n","metadata":{"execution":{"iopub.status.busy":"2023-05-27T17:21:50.758080Z","iopub.execute_input":"2023-05-27T17:21:50.758689Z","iopub.status.idle":"2023-05-27T17:21:50.773672Z","shell.execute_reply.started":"2023-05-27T17:21:50.758649Z","shell.execute_reply":"2023-05-27T17:21:50.772854Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"import matplotlib.pyplot as plt\n\n# Set the figure size\nplt.figure(figsize=(12, 8))\n\nimage_folder = \"/kaggle/input/hubmap-hacking-the-human-vasculature/train\"\n# Create an instance of the CustomDataset\ndataset = CustomDataset(image_dir=image_folder, labels_file=polygons)\n\n# Define the number of samples to visualize\nnum_samples = 8\n\n# Calculate the number of rows and columns based on the number of samples\nnum_rows = (num_samples + 3) // 4  # Each row contains 4 samples\nnum_cols = min(num_samples, 4)\n\n# Visualize the samples\nfor i in range(num_samples):\n    # Get the image and mask for the current sample\n    image, mask = dataset[i]\n\n    # Convert the image tensor to a numpy array and transpose the dimensions\n    image = image.permute(1, 2, 0).numpy() / 255\n\n    # Calculate the subplot index based on the current sample\n    subplot_index = i + 1\n\n    # Plot the image\n    plt.subplot(num_rows, 2 * num_cols, 2 * subplot_index - 1)\n    plt.imshow(image)\n    plt.axis('off')\n    plt.title('Image')\n\n    # Calculate the subplot index for the mask\n    mask_subplot_index = (subplot_index - 1) % num_samples + 1\n\n    # Plot the mask\n    plt.subplot(num_rows, 2 * num_cols, 2 * subplot_index)\n    plt.imshow(mask, cmap='gray')\n    plt.axis('off')\n    plt.title('Mask')\n\n# Adjust the spacing between subplots\nplt.tight_layout(pad=0.2)\n\n# Display the plot\nplt.show()\n","metadata":{"execution":{"iopub.status.busy":"2023-05-27T17:21:50.776814Z","iopub.execute_input":"2023-05-27T17:21:50.777100Z","iopub.status.idle":"2023-05-27T17:21:55.428691Z","shell.execute_reply.started":"2023-05-27T17:21:50.777078Z","shell.execute_reply":"2023-05-27T17:21:55.425860Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Training Phase","metadata":{}},{"cell_type":"code","source":"import torch\nimport torch.nn as nn\n\nclass SegmentationModel(nn.Module):\n    def __init__(self):\n        super(SegmentationModel, self).__init__()\n\n        # Encoder\n        self.encoder = nn.Sequential(\n            nn.Conv2d(3, 64, kernel_size=3, padding=1),\n            nn.ReLU(inplace=True),\n            nn.Conv2d(64, 64, kernel_size=3, padding=1),\n            nn.ReLU(inplace=True),\n            nn.MaxPool2d(kernel_size=2, stride=2)\n        )\n\n        # Decoder\n        self.decoder = nn.Sequential(\n            nn.Conv2d(64, 64, kernel_size=3, padding=1),\n            nn.ReLU(inplace=True),\n            nn.Conv2d(64, 64, kernel_size=3, padding=1),\n            nn.ReLU(inplace=True),\n            nn.ConvTranspose2d(64, 1, kernel_size=2, stride=2)\n        )\n\n    def forward(self, x):\n        # Encoder\n        x = self.encoder(x)\n\n        # Decoder\n        x = self.decoder(x)\n\n        # Apply a sigmoid activation to the output for binary segmentation\n        x = torch.sigmoid(x)\n\n        return x\n\n\n","metadata":{"execution":{"iopub.status.busy":"2023-05-27T17:21:55.430197Z","iopub.execute_input":"2023-05-27T17:21:55.431109Z","iopub.status.idle":"2023-05-27T17:21:55.442272Z","shell.execute_reply.started":"2023-05-27T17:21:55.431075Z","shell.execute_reply":"2023-05-27T17:21:55.441221Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"import torch\nimport torch.nn as nn\nimport torch.optim as optim\nfrom torch.utils.data import DataLoader\nimport time\n\n\n# Initialize the dataset and dataloader\nimage_folder = \"/kaggle/input/hubmap-hacking-the-human-vasculature/train\"\nlabels_file = \"/kaggle/input/hubmap-hacking-the-human-vasculature/polygons.jsonl\"\ndataset = CustomDataset(image_dir=image_folder, labels_file=labels_file)\ndataloader = DataLoader(dataset, batch_size=32, shuffle=True)\n\n# Initialize the segmentation model\nmodel = SegmentationModel()\n\n# Define the loss function and optimizer\ncriterion = nn.BCELoss()  # using binary segmentation (e.g., 0 for background, 1 for foreground)\noptimizer = optim.Adam(model.parameters(), lr=0.001)\n\n# Set the device (GPU if available, else CPU)\ndevice = torch.device(\"cuda\" if torch.cuda.is_available() else \"cpu\")\nprint(device)\nmodel.to(device)\n\n# Training loop\nnum_epochs = 10\n\nfor epoch in range(num_epochs):\n    model.train()  # Set the model in training mode\n    running_loss = 0.0\n    \n    start_time = time.time()\n    for images, masks in dataloader:\n        # Move images and masks to the device\n        images = images.to(device)\n        masks = masks.to(device)\n\n        # Zero the gradients\n        optimizer.zero_grad()\n\n        # Forward pass\n        outputs = model(images)\n\n        # Adjust the size of masks to match the size of outputs\n        masks = masks.unsqueeze(1)  # Add an extra dimension\n\n        # Calculate the loss\n        loss = criterion(outputs, masks)\n\n        # Backward pass and optimization\n        loss.backward()\n        optimizer.step()\n\n        running_loss += loss.item()\n\n    epoch_loss = running_loss / len(dataloader)\n    epoch_time = time.time() - start_time\n    print(f\"Epoch {epoch+1}/{num_epochs}, Loss: {epoch_loss:.4f}, Time: {epoch_time:.2f} seconds\")\n\n# After training, you can save the model\ntorch.save(model.state_dict(), \"model.pth\")\n","metadata":{"execution":{"iopub.status.busy":"2023-05-27T17:21:55.443541Z","iopub.execute_input":"2023-05-27T17:21:55.443852Z","iopub.status.idle":"2023-05-27T17:29:51.761546Z","shell.execute_reply.started":"2023-05-27T17:21:55.443824Z","shell.execute_reply":"2023-05-27T17:29:51.760393Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Submission","metadata":{}},{"cell_type":"code","source":"def mask_to_rle(mask):\n    \"\"\"\n    Convert a binary mask to the run-length encoding (RLE) format.\n    \n    Args:\n        mask (ndarray): Binary mask of shape (height, width).\n    \n    Returns:\n        str: Run-length encoding string.\n    \"\"\"\n    pixels = mask.flatten(order=\"F\")\n    pixels = np.concatenate([[0], pixels, [0]])\n    runs = np.where(pixels[1:] != pixels[:-1])[0] + 1\n    runs[1::2] -= runs[::2]\n    rle = \" \".join(str(x) for x in runs)\n    return rle\n","metadata":{"execution":{"iopub.status.busy":"2023-05-27T17:32:57.028736Z","iopub.execute_input":"2023-05-27T17:32:57.029104Z","iopub.status.idle":"2023-05-27T17:32:57.035654Z","shell.execute_reply.started":"2023-05-27T17:32:57.029076Z","shell.execute_reply":"2023-05-27T17:32:57.034304Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"import torch\nfrom torch.utils.data import DataLoader\nimport pandas as pd\nimport numpy as np\nfrom torchvision.transforms import transforms\nfrom PIL import Image\nimport os\n\n# Define the necessary functions and classes\n\nclass CustomTestDataset(Dataset):\n    def __init__(self, image_dir, transform=None):\n        self.image_dir = image_dir\n        self.transform = transform\n\n        # Get the list of test image file names\n        self.image_files = sorted(os.listdir(image_dir))\n\n    def __len__(self):\n        return len(self.image_files)\n\n    def __getitem__(self, idx):\n        # Load the test image\n        image_path = os.path.join(self.image_dir, self.image_files[idx])\n        image = Image.open(image_path)\n\n        if self.transform:\n            image = self.transform(image)\n\n        return image\n\n# Set the path to the test image directory\ntest_image_dir = \"/kaggle/input/hubmap-hacking-the-human-vasculature/test\"\n\n# Set the path to the output CSV file for submission\nsubmission_path = \"/kaggle/working/submission.csv\"\n\n# Set the device\ndevice = torch.device(\"cuda\" if torch.cuda.is_available() else \"cpu\")\n\n# Create the test dataset and data loader\ntest_transform = transforms.Compose([\n    transforms.ToTensor(),\n    transforms.Normalize(mean=[0.485, 0.456, 0.406], std=[0.229, 0.224, 0.225])\n])\ntest_dataset = CustomTestDataset(test_image_dir, transform=test_transform)\ntest_dataloader = DataLoader(test_dataset, batch_size=32, shuffle=False)\n\n# Create an instance of your model\nmodel = SegmentationModel()\nmodel.to(device)\n\n# Load the pre-trained weights\nmodel.load_state_dict(torch.load(\"model.pth\"))\n\n# Put the model in evaluation mode\nmodel.eval()\n\n# Generate predictions for the test data\npredictions = []\n\nwith torch.no_grad():\n    for images in test_dataloader:\n        images = images.to(device)\n        outputs = model(images)\n        predicted_masks = (outputs > 0.5).float()  # Convert probabilities to binary masks\n        predictions.extend(predicted_masks.cpu().numpy())\n\n# Format the predictions in the required Kaggle submission format\nsubmission_df = pd.DataFrame(columns=[\"id\", \"predicted\"])\nimage_ids = [os.path.splitext(file)[0] for file in test_dataset.image_files]\nfor image_id, prediction in zip(image_ids, predictions):\n    rle_encoded = mask_to_rle(prediction)\n    submission_df = submission_df.append({\"id\": image_id, \"predicted\": rle_encoded}, ignore_index=True)\n\n# Save the predictions to a CSV file\nsubmission_df.to_csv(submission_path, index=False)\n","metadata":{"execution":{"iopub.status.busy":"2023-05-27T17:32:59.404702Z","iopub.execute_input":"2023-05-27T17:32:59.405361Z","iopub.status.idle":"2023-05-27T17:32:59.450562Z","shell.execute_reply.started":"2023-05-27T17:32:59.405327Z","shell.execute_reply":"2023-05-27T17:32:59.449660Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<h2>Conclusion</h2>\n\n<p>In conclusion, this notebook has provided a comprehensive overview of the dataset and presented a baseline solution for the microvascular segmentation task in the Kaggle competition. Throughout this journey, I have delved into the intricacies of the data, explored various preprocessing techniques, and implemented a deep learning model using PyTorch.</p>\n\n<p>I am truly passionate about this project and deeply motivated to continue refining and enhancing the solution. The opportunity to contribute to the advancement of medical research, specifically in understanding the arrangement of blood vessels in human tissues, is both inspiring and humbling. By automating the segmentation of microvascular structures, we can make a significant impact on researchers' ability to analyze histology images efficiently.</p>\n\n<p>I would like to express my gratitude to the organizers of this competition for providing such a challenging and meaningful task. Additionally, I extend my appreciation to the Kaggle community for their valuable insights and support throughout this journey.</p>\n\n<p>Moving forward, I am committed to further optimizing the model, exploring advanced techniques, and collaborating with other passionate individuals to achieve even better results. Together, we can advance the field of medical image analysis and contribute to improving human health.</p>\n\n<p>Thank you for joining me on this exciting endeavor!</p>\n","metadata":{}}]}