{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"pygments_lexer":"ipython3","nbconvert_exporter":"python","version":"3.6.4","file_extension":".py","codemirror_mode":{"name":"ipython","version":3},"name":"python","mimetype":"text/x-python"}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"# Dataset Description\n\n### Data Description\n\nYour goal in this competition is to locate microvasculature structures (blood vessels) within human kidney histology slides.\n\nThe competition data comprises tiles extracted from five Whole Slide Images (WSI) split into two datasets. Tiles from Dataset 1 have annotations that have been expert reviewed. Dataset 2 comprises the remaining tiles from these same WSIs and contain sparse annotations that have not been expert reviewed.\n\n*     All of the test set tiles are from Dataset 1.\n*     Two of the WSIs make up the training set, two WSIs make up the public test set, and one WSI makes up the private test set.\n*     The training data includes Dataset 2 tiles from the public test WSI, but not from the private test WSI.\n\nWe also include, as Dataset 3, tiles extracted from an additional nine WSIs. These tiles have not been annotated. You may wish to apply semi- or self-supervised learning techniques on this data to support your predictions.\n\nNote that this is a Code Competition, in which the actual test set is hidden. In this version, we give some sample data drawn from the public test set to help you author your solutions. When your submission is scored, this example test data will be replaced with the full test set. There are about 650 tiles in the full test set.\n\nYou may find resources from the previous HuBMAP competitions useful as well:\n\n*     HuBMAP: Hacking the Kidney\n*     HuBMAP + HPA: Hacking the Human Body\n\nFiles and Field Descriptions\n\n* **{train|test}/** Folders containing TIFF images of the tiles. Each tile is 512x512 in size.\n* **polygons.jsonl** Polygonal segmentation masks in JSONL format, available for Dataset 1 and Dataset 2. Each line gives JSON annotations for a single image with:\n    * id Identifies the corresponding image in train/\n    * annotations A list of mask annotations with:\n    * type Identifies the type of structure annotated:\n        * blood_vessel The target structure. Your goal in this competition is to predict these kinds of masks on the test set.\n        * glomerulus A capillary ball structure in the kidney. These parts of the images were excluded from blood vessel annotation. You should ensure none of your test set predictions occur within glomerulus structures as they will be counted as false positives. Annotations are provided for test set tiles.\n        * unsure A structure the expert annotators cannot confidently distinguish as a blood vessel.\n        coordinates A list of polygon coordinates defining the segmentation mask.\n*     **tile_meta.csv** Metadata for each image.\n    * source_wsi Identifies the WSI this tile was extracted from.\n    * {i|j} The location of the upper-left corner within the WSI where the tile was extracted.\n    * dataset The dataset this tile belongs to, as described above.\n*     **wsi_meta.csv** Metadata for the Whole Slide Images the tiles were extracted from.\n    * source_wsi Identifies the WSI.\n    * age, sex, race, height, weight, and bmi demographic information about the tissue donor.\n*     **sample_submission.csv** A sample submission file in the correct format. See the Evaluation page for more details.\n\n","metadata":{}},{"cell_type":"markdown","source":"# Understanding the data set","metadata":{}},{"cell_type":"markdown","source":"### Having a look at wsi_meta.csv","metadata":{}},{"cell_type":"code","source":"import pandas as pd\nwsi_meta = pd.read_csv(\"/kaggle/input/hubmap-hacking-the-human-vasculature/wsi_meta.csv\")\nwsi_meta","metadata":{"execution":{"iopub.status.busy":"2023-05-24T16:52:47.246096Z","iopub.execute_input":"2023-05-24T16:52:47.246524Z","iopub.status.idle":"2023-05-24T16:52:47.268471Z","shell.execute_reply.started":"2023-05-24T16:52:47.246492Z","shell.execute_reply":"2023-05-24T16:52:47.267270Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"More on this (having a look at the data) later on.","metadata":{}},{"cell_type":"markdown","source":"### Having a look at an image\nThanks to Thomas Rochefort-Beaudoin https://www.kaggle.com/code/thomasrochefort/hubmap-simple-pytorch-dataloader its code could lead to this:","metadata":{}},{"cell_type":"code","source":"import torch\nfrom torch.utils.data import Dataset\nfrom PIL import Image\nimport os\nimport json\nimport numpy as np\nfrom skimage.draw import polygon2mask\n\nclass CustomDataset(Dataset):\n    def __init__(self, image_dir, labels_file, transform=None):\n        with open(labels_file, 'r') as json_file:\n            self.json_labels = [json.loads(line) for line in json_file]\n\n        self.image_dir = image_dir\n        self.transform = transform\n\n    def __len__(self):\n        return len(self.json_labels)\n\n    def __getitem__(self, idx):\n        # Load image\n        image_path = os.path.join(self.image_dir, f\"{self.json_labels[idx]['id']}.tif\")\n        image = Image.open(image_path)\n\n        # Initialize masks\n        mask_bv = np.zeros((512, 512), dtype=np.float32) # mask of blood_vessel\n        mask_gl = np.zeros((512, 512), dtype=np.float32) # mask of glomerulus  \n        mask_un = np.zeros((512, 512), dtype=np.float32) # mask of unsure\n\n        # Process annotations\n        for annot in self.json_labels[idx]['annotations']:\n            cords = annot['coordinates']\n            if annot['type'] == \"blood_vessel\":\n                for cord in cords:\n                    rr, cc = np.array([i[1] for i in cord]), np.asarray([i[0] for i in cord])\n                    mask_bv[rr, cc] = 1\n            elif annot['type'] == 'glomerulus':\n                for cord in cords:\n                    rr, cc = np.array([i[1] for i in cord]), np.asarray([i[0] for i in cord])\n                    mask_gl[rr, cc] = 1\n            elif annot['type'] == 'unsure':\n                for cord in cords:\n                    rr, cc = np.array([i[1] for i in cord]), np.asarray([i[0] for i in cord])\n                    mask_un[rr, cc] = 1\n            else:\n                print(\"error: wrong type\")\n                \n                \n\n        # Convert PIL Image and mask to PyTorch tensor\n        image = torch.tensor(np.array(image), dtype=torch.float32).permute(2, 0, 1)  # Shape: [C, H, W]\n        mask_bv = torch.tensor(mask_bv, dtype=torch.float32)\n        mask_gl = torch.tensor(mask_gl, dtype=torch.float32)\n        mask_un = torch.tensor(mask_un, dtype=torch.float32)\n\n        if self.transform:\n            image = self.transform(image)\n\n        return image, mask_bv, mask_gl, mask_un\n","metadata":{"execution":{"iopub.status.busy":"2023-05-24T16:52:47.278812Z","iopub.execute_input":"2023-05-24T16:52:47.279195Z","iopub.status.idle":"2023-05-24T16:52:47.300122Z","shell.execute_reply.started":"2023-05-24T16:52:47.279166Z","shell.execute_reply":"2023-05-24T16:52:47.298526Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"from torch.utils.data import DataLoader\n# Instantiate the dataset\ndataset = CustomDataset(image_dir='../input/hubmap-hacking-the-human-vasculature/train/', \n                        labels_file='../input/hubmap-hacking-the-human-vasculature/polygons.jsonl')\n\n# Instantiate the DataLoader and load a sample batch for viz:\ndataloader = DataLoader(dataset, batch_size=32, shuffle=True)\nimage, mask_bv, mask_gl, mask_un = next(iter(dataloader))\nprint(image.shape)\nprint(mask_bv.shape)\nprint(mask_gl.shape)\nprint(mask_un.shape)\n\n","metadata":{"execution":{"iopub.status.busy":"2023-05-24T16:52:47.306400Z","iopub.execute_input":"2023-05-24T16:52:47.306824Z","iopub.status.idle":"2023-05-24T16:52:52.831674Z","shell.execute_reply.started":"2023-05-24T16:52:47.306781Z","shell.execute_reply":"2023-05-24T16:52:52.830536Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"import matplotlib.pyplot as plt\n\ndef plot_image_and_mask(image, mask_bv, mask_gl=0, mask_un=0):\n    fig = plt.figure(figsize=(12, 6))\n\n    # Transpose the dimensions to height, width, channels for plotting\n    image = image.transpose((1, 2, 0))\n    \n    mask = mask_bv + mask_gl + mask_un\n    mask = mask-1\n    mask = -1*mask\n    \n    mask = np.array([mask]*3)\n    mask = mask.transpose((1, 2, 0))\n    \n    im = np.multiply(image, mask)\n    \n\n    # Plot the image\n    plt.imshow(im)\n    #plt.imshow(mask_nan, alpha=0.2)\n    plt.axis('off')\n    \n    plt.show()\n\nplot_image_and_mask(image[10].numpy()/255, mask_bv[10].numpy(), mask_gl[10].numpy(), mask_un[10].numpy() )","metadata":{"execution":{"iopub.status.busy":"2023-05-24T16:52:52.833788Z","iopub.execute_input":"2023-05-24T16:52:52.834205Z","iopub.status.idle":"2023-05-24T16:52:53.156356Z","shell.execute_reply.started":"2023-05-24T16:52:52.834168Z","shell.execute_reply":"2023-05-24T16:52:53.154963Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"The distinctions between the masks cannot be done using the plot_image_and_mask function. I'll have to see whether this is really an issue worth the trouble...","metadata":{}},{"cell_type":"markdown","source":"Feel free to use this work.","metadata":{}}]}