{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"name":"python","version":"3.10.13","mimetype":"text/x-python","codemirror_mode":{"name":"ipython","version":3},"pygments_lexer":"ipython3","nbconvert_exporter":"python","file_extension":".py"},"kaggle":{"accelerator":"gpu","dataSources":[{"sourceId":52279,"databundleVersionId":5822112,"sourceType":"competition"},{"sourceId":6021726,"sourceType":"datasetVersion","datasetId":3446188}],"dockerImageVersionId":30733,"isInternetEnabled":true,"language":"python","sourceType":"notebook","isGpuEnabled":true}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"<div class=\"alert alert-block alert-info\" style=\"font-size:20px; background-color: #cf6161; font-family:verdana; color: #701212; border: 3px #701212 solid\">\n    <b>Medical Instance Segmentation with YOLOv8</b>\n","metadata":{}},{"cell_type":"markdown","source":"<div class=\"alert alert-block alert-info\" style=\"font-size:17px; background-color: #d1a6ff; font-family:verdana; color: #533078; border: 3px #533078 solid\">\n    <b>Biological overview</b>\n    <br>Task is to segment medical images of kidney tissue samples. Human tissues are covered with many vessels. So task is to segment such vessels.<br>\n</div>","metadata":{}},{"cell_type":"markdown","source":"# Object Detection\n<div class=\"alert alert-block alert-info\" style=\"font-size:15px; background-color: #FFDAB9; font-family:verdana; color: #b86e1f; border: 2px #b86e1f solid\">\n    <b>Computer vision is a field of artificial intelligence that focuses on teaching computers to interpret and understand visual information. One popular and powerful technique used in computer vision for object detection is called YOLO, which stands for \"You Only Look Once\".\n\nYOLO aims to identify and locate objects in an image or video stream in real-time. Unlike traditional methods that rely on complex pipelines and multiple passes, YOLO takes a different approach by treating object detection as a single regression problem.\n\nThis algorithm divides the input image into a grid and predicts bounding boxes and class probabilities for objects within each grid cell. It simultaneously predicts the class labels and their corresponding bounding boxes, making it incredibly efficient and fast. YOLO is known for its real-time performance, enabling it to process images and videos at impressive speeds.\n \nBy leveraging deep convolutional neural networks, YOLO can learn to recognize a wide range of objects and accurately localize them within an image. It can detect multiple objects of different classes simultaneously, making it particularly useful for applications where real-time processing and high detection accuracy are crucial, such as autonomous driving, video surveillance, and robotics.</b>\n</div>","metadata":{}},{"cell_type":"markdown","source":"![](https://i.ibb.co/1MZNhNF/003.png)","metadata":{}},{"cell_type":"markdown","source":"<div class=\"alert alert-block alert-info\" style=\"font-size:15px; background-color: #d1a6ff; font-family:verdana; color: #533078; border: 2px #533078 solid\">\n    <b>Tissue samples</b>\n    <br>Renal cortex: the renal cortex is the outer portion of the kidney that contains round renal corpuscles enclosing glomerular tufts, balls of capillary loops. The renal corpuscle is the start of the nephron, through which the filtration of blood occurs. The renal cortex also contains proximal and distal convoluted tubules, PCTs and DCTs, which are also regions of the nephron. Between these tubular structures is a complex network of capillaries called the peritubular capillaries.<br>\n    <br>Renal medulla: the renal medulla is the inner portion of the kidney and is arranged in 8-15 renal pyramids containing linearly arranged tubules comprising the loops of Henle and ducts that gather products for excretion. The capillary network in the renal medulla consists of capillaries called vasa recta. The renal pyramids (medullary tissue) are divided by extensions of the renal cortex called renal columns. There are also projections of the renal medulla into the outer cortex called medullary rays.<br>\n    <br>Renal papilla (subsection of medulla): the broad bases of the pyramids connect to the renal cortex at the corticomedullary junctions while the tips form structures called the renal papilla, which project in the minor renal calyces where urine is collected.<br>\n</div>","metadata":{}},{"cell_type":"markdown","source":"<div class=\"alert alert-block alert-info\" style=\"font-size:15px; background-color: #b5e1ff; font-family:verdana; color: #011d82; border: 2px #011d82 solid\">\n    <b>What is instance segmentation?</b>\n    <br>Instance Segmentation is a unique form of image segmentation that deals with detecting and delineating each distinct instance of an object appearing in an image. Instance segmentation detects all instances of a class with the extra functionality of demarcating separate instances of any segment class. Hence, it is also referred to as incorporating object detection and semantic segmentation functionality. Watch video to learn more: <a href=instance segmentation>Instance segmentation | Tutorial</a> <br>\n</div>","metadata":{}},{"cell_type":"markdown","source":"<div class=\"alert alert-block alert-info\" style=\"font-size:15px; background-color: #b5e1ff; font-family:verdana; color: #011d82; border: 2px #011d82 solid\">\n    <b>Applications of Semantic Segmentation</b>\n<ol><li><strong>Medical Diagnostics: </strong>For detecting medical abnormalities in <a href=\"https://universe.roboflow.com/search?q=xray&amp;ref=blog.roboflow.com\">X-Rays, CT Scans, MRI Scans</a></li><li><strong>GeoSensing:</strong> For land usage mapping from <a href=\"https://roboflow.com/solutions/aerial?ref=blog.roboflow.com\">satellite imagery</a> and monitoring areas of deforestation and urbanization</li><li><strong>Autonomous Driving:</strong> For accurately <a href=\"https://universe.roboflow.com/browse/self-driving?ref=blog.roboflow.com\">detecting lanes, pedestrians, traffic signs, road</a>, sky and other vehicles on the road</li></ol><br>\n</div>","metadata":{}},{"cell_type":"markdown","source":"# Prepare data\n<div class=\"alert alert-block alert-info\" style=\"font-size:15px; background-color: #abffd1; font-family:verdana; color: #003819; border: 2px #003819 solid\">\n    <br>For training, we need a dataset, and not just any, but in the COCO format. COCO is an open database for object detection, it is huge and therefore it is used all over the world to evaluate new models according to various metrics. It plays a rather important role, and most detection models are pre-trained on it. That is why this dataset format is used in ultralytics. <br>\n</div>","metadata":{}},{"cell_type":"markdown","source":"<div class=\"alert alert-block alert-info\" style=\"font-size:20px; background-color: #abffd1; font-family:verdana; color: #003819; border: 2px #003819 solid\">\n    <b>Libraries for creating dataset</b>\n</div>","metadata":{}},{"cell_type":"code","source":"from itertools import chain\nimport json\nimport os\nimport shutil\nfrom tqdm.notebook import tqdm\nfrom colorama import Fore\nimport yaml\nimport numpy as np","metadata":{"execution":{"iopub.status.busy":"2024-07-22T15:34:56.816115Z","iopub.execute_input":"2024-07-22T15:34:56.816719Z","iopub.status.idle":"2024-07-22T15:34:56.943104Z","shell.execute_reply.started":"2024-07-22T15:34:56.816687Z","shell.execute_reply":"2024-07-22T15:34:56.942324Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<div class=\"alert alert-block alert-info\" style=\"font-size:20px; background-color: #abffd1; font-family:verdana; color: #003819; border: 2px #003819 solid\">\n    <b>Class for creating dataset</b>\n</div>","metadata":{}},{"cell_type":"code","source":"class COCODataset:\n    def __init__(self, images_dirpath: str, annotations_filepath: str, length: int = 1633):\n        self.train_size = None\n        self.val_size = None\n        self.length = length\n        self.classes = None\n        self.labels_counter = None\n        self.normalize = None\n        \n        self.images_dirpath = images_dirpath\n        self.annotations_filepath = annotations_filepath\n        self.dataset_dirpath = os.path.join(os.getcwd(), \"dataset\")\n        self.train_dirpath =  os.path.join(self.dataset_dirpath, \"train\")\n        self.val_dirpath =  os.path.join(self.dataset_dirpath, \"val\")\n        self.config_path = os.path.join(self.dataset_dirpath, \"coco.yaml\")\n\n        self.samples = self.parse_jsonl(annotations_filepath)\n        self.classes_dict = {\n            \"blood_vessel\": 0,\n            \"glomerulus\": 1,\n            \"unsure\": 2,\n        }\n\n    def __prepare_dirs(self) -> None:\n        if not os.path.exists(self.dataset_dirpath):\n            os.makedirs(os.path.join(self.train_dirpath, \"images\"), exist_ok=True)\n            os.makedirs(os.path.join(self.train_dirpath, \"labels\"), exist_ok=True)\n            os.makedirs(os.path.join(self.val_dirpath, \"images\"), exist_ok=True)\n            os.makedirs(os.path.join(self.val_dirpath, \"labels\"), exist_ok=True)\n        else:\n            raise RuntimeError(\"Dataset already exists!\")\n\n    def __define_splitratio(self) -> None:\n        self.train_size = round(self.length * self.train_size)\n        self.val_size = self.length - self.train_size\n        assert self.train_size + self.val_size == self.length\n\n    def parse_jsonl(self, path: str) -> list[dict, ...]:\n        with open(path, 'r') as json_file:\n            jsonl_samples = [\n                json.loads(line)\n                for line in tqdm(\n                    json_file, desc=\"Processing polygons\", total=self.length\n                )\n            ]\n        return jsonl_samples\n\n    def __define_paths(self, i: int) -> dict:\n        data_path = self.val_dirpath\n        if i < self.train_size:\n            data_path = self.train_dirpath\n        return {\n            \"images\": os.path.join(data_path, \"images\"),\n            \"labels\": os.path.join(data_path, \"labels\")\n        }\n\n    @staticmethod\n    def __get_label_path(paths_dict: dict, identifier: str) -> str:\n        return os.path.join(\n            paths_dict[\"labels\"],\n            f\"{identifier}.txt\"\n        )\n\n    @staticmethod\n    def __get_image_path(paths_dict: dict, identifier: str) -> str:\n        return os.path.join(\n            paths_dict[\"images\"],\n            f\"{identifier}.tif\"\n        )\n\n    def __copy_image(self, dst_path: str, identifier: str) -> str:\n        shutil.copyfile(\n            os.path.join(self.images_dirpath, f\"{identifier}.tif\"),\n            dst_path\n        )\n\n    def __copy_label(self, annotations: list, dst_path: str) -> None:\n        with open(dst_path, \"w\") as file:\n            for annotation in annotations:\n                coordinates = annotation[\"coordinates\"][0]\n                label = self.classes_dict[annotation[\"type\"]]\n                if label in self.classes:\n                    if coordinates:\n                        if self.normalize:\n                            coordinates = np.array(coordinates) / 512.0\n                        coordinates = \" \".join(map(str, chain(*coordinates)))\n                        file.write(f\"{label} {coordinates}\\n\")\n                        self.labels_counter += 1\n\n    def __splitfolders(self):\n        for i, line in tqdm(\n                enumerate(self.samples),\n                desc=\"Dataset creation\", total=self.length\n        ):\n            self.labels_counter = 0\n            identifier = line[\"id\"]\n            annotations = line[\"annotations\"]\n            paths_dict = self.__define_paths(i)\n\n            dst_image_path = self.__get_image_path(paths_dict, identifier)\n            dst_label_path = self.__get_label_path(paths_dict, identifier)\n\n            self.__copy_image(dst_image_path, identifier)\n            self.__copy_label(annotations, dst_label_path)\n\n            if self.labels_counter == 0:\n                os.remove(dst_image_path)\n                os.remove(dst_label_path)\n\n    def __count_dataset(self) -> dict:\n        train_images = len(os.listdir(os.path.join(self.train_dirpath, \"images\")))\n        train_labels = len(os.listdir(os.path.join(self.train_dirpath, \"labels\")))\n        val_images = len(os.listdir(os.path.join(self.val_dirpath, \"images\")))\n        val_labels = len(os.listdir(os.path.join(self.val_dirpath, \"labels\")))\n        return {\n            \"train_images\": train_images,\n            \"train_labels\": train_labels,\n            \"val_images\": val_images,\n            \"val_labels\": val_labels\n        }\n\n    @staticmethod\n    def __check_sanity(count_dict: dict) -> None:\n        assert count_dict[\"train_images\"] == count_dict[\"train_labels\"]\n        assert count_dict[\"val_images\"] == count_dict[\"val_labels\"]\n\n    def __finalizing(self, count_dict: dict) -> None:\n        assert os.path.exists(self.dataset_dirpath)\n\n        example_structure = [\n            \"dataset\",\n            \"train\", \"labels\", \"images\",\n            \"val\", \"labels\", \"images\"\n        ]\n\n        dir_bone = (\n            dirname.split(\"/\")[-1]\n            for dirname, _, filenames in os.walk(self.dataset_dirpath)\n            if dirname.split(\"/\")[-1] in example_structure\n        )\n\n        try:\n            print(\"\\n~ HuBMAP Dataset Structure ~\\n\")\n            print(\n            f\"\"\"\n          ├── {next(dir_bone)}\n          │   │\n          │   ├── {next(dir_bone)}\n          │   │   └── {next(dir_bone)}\n          │   │   └── {next(dir_bone)}\n          │   │\n          │   ├── {next(dir_bone)}\n          │   │   └── {next(dir_bone)}\n          │   │   └── {next(dir_bone)}\n            \"\"\"\n            )\n        except StopIteration as e:\n            print(e)\n        else:\n            print(Fore.GREEN + \"-> Success\")\n            print(Fore.GREEN + f\"Train dataset: {count_dict['train_images']}\\nVal dataset: {count_dict['val_images']}\")\n\n    def get_config(self) ->dict:\n        names = [\"blood_vessel\", \"glomerulus\", \"unsure\"]\n        return {\n            \"train\": str(self.train_dirpath),\n            \"val\": str(self.val_dirpath),\n            \"names\": [names[i] for i in self.classes]\n        }\n\n    @staticmethod\n    def display_config(config: dict) -> None:\n        print(Fore.BLACK + \"\\n~ HuBMAP Config Structure ~\\n\")\n        print(\n        f\"\"\"\n      │   │\n      │   ├── train\n      │   │   └── {config['train']}/images\n      │   │\n      │   │\n      │   ├── val\n      │   │   └── {config['val']}/images\n      │   │\n      │   │\n      │   ├── names\n      │   │   └── {' '.join(config['names'])}\n        \"\"\"\n        )\n        print(Fore.GREEN + \"-> Success\")\n        print(Fore.GREEN + f\"Number of classes: {len(config['names'])}\"\n                           f\"\\nClasses: {' '.join(config['names'])}\" \n              )\n\n    def write_config(self, config: dict) -> None:\n        with open(self.config_path, mode=\"w\") as f:\n            yaml.safe_dump(stream=f, data=config)\n\n    def __call__(self, train_size: float,\n                 classes: list[int, ...],\n                 make_config: bool = True,\n                 normalize: bool = True\n                ) -> None:\n        \n        self.train_size = train_size\n        self.classes = classes\n        self.normalize = normalize\n        \n        self.__define_splitratio()\n        self.__prepare_dirs()\n        self.__splitfolders()\n        count_dict = self.__count_dataset()\n        self.__check_sanity(count_dict)\n        self.__finalizing(count_dict)\n        \n        if make_config:\n            config = self.get_config()\n            self.write_config(config)\n            self.display_config(config)","metadata":{"execution":{"iopub.status.busy":"2024-07-22T15:34:56.945095Z","iopub.execute_input":"2024-07-22T15:34:56.945427Z","iopub.status.idle":"2024-07-22T15:34:56.989157Z","shell.execute_reply.started":"2024-07-22T15:34:56.945396Z","shell.execute_reply":"2024-07-22T15:34:56.988077Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<div class=\"alert alert-block alert-info\" style=\"font-size:20px; background-color: #abffd1; font-family:verdana; color: #003819; border: 2px #003819 solid\">\n    <b>Dataset creation</b>\n</div>","metadata":{}},{"cell_type":"code","source":"coco = COCODataset(\n    annotations_filepath=\"/kaggle/input/hubmap-hacking-the-human-vasculature/polygons.jsonl\",\n    images_dirpath=\"/kaggle/input/hubmap-hacking-the-human-vasculature/train\",\n) ","metadata":{"execution":{"iopub.status.busy":"2024-07-22T15:34:56.990891Z","iopub.execute_input":"2024-07-22T15:34:56.991242Z","iopub.status.idle":"2024-07-22T15:35:00.849343Z","shell.execute_reply.started":"2024-07-22T15:34:56.991212Z","shell.execute_reply":"2024-07-22T15:35:00.848262Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"coco(train_size=0.80, classes=[0, 1, 2])","metadata":{"execution":{"iopub.status.busy":"2024-07-22T15:35:00.852274Z","iopub.execute_input":"2024-07-22T15:35:00.852654Z","iopub.status.idle":"2024-07-22T15:35:33.387361Z","shell.execute_reply.started":"2024-07-22T15:35:00.852617Z","shell.execute_reply":"2024-07-22T15:35:33.386473Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"import shutil\nimport os\nimport sys\nfrom colorama import Fore","metadata":{"execution":{"iopub.status.busy":"2024-07-22T15:35:33.388778Z","iopub.execute_input":"2024-07-22T15:35:33.389271Z","iopub.status.idle":"2024-07-22T15:35:33.393286Z","shell.execute_reply.started":"2024-07-22T15:35:33.389238Z","shell.execute_reply":"2024-07-22T15:35:33.392474Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<div class=\"alert alert-block alert-info\" style=\"font-size:20px; background-color: #abffd1; font-family:verdana; color: #003819; border: 2px #003819 solid\">\n    <b>Setting up dependencies related to Ultralytics and Pycocotools</b>\n</div>","metadata":{}},{"cell_type":"code","source":"class SetupPipline:\n    def __init__(self, display: bool = True):\n        self.pycocotools = self.__pycocotools()\n        self.ultralytics = self.__ultralytics()\n        \n    @staticmethod\n    def __ultralytics() -> str:\n        sys.path.append(\"/kaggle/input/hubmap-tools-ultralytics-and-pycocotools/ultralytics/ultralytics\") \n        return \"successfully\"\n        \n    @staticmethod\n    def __pycocotools() -> str:\n        if not os.path.exists(\"/kaggle/working/packages\"):\n            shutil.copytree(\"/kaggle/input/hubmap-tools-ultralytics-and-pycocotools/pycocotools/pycocotools\", \"/kaggle/working/packages\")\n            os.chdir(\"/kaggle/working/packages/pycocotools-2.0.6/\")\n            os.system(\"python setup.py install\")\n            os.system(\"pip install . --no-index --find-links /kaggle/working/packages/\")\n            os.chdir(\"/kaggle/working\")\n            return \"successfully\"\n    \n    def display(self) -> None:\n        print(Fore.GREEN+f\"\\nPycocotools was installed {self.pycocotools}\")\n        print(f\"Ultralytics was installed {self.ultralytics}\"+Fore.WHITE)","metadata":{"execution":{"iopub.status.busy":"2024-07-22T15:35:33.394519Z","iopub.execute_input":"2024-07-22T15:35:33.394836Z","iopub.status.idle":"2024-07-22T15:35:33.405225Z","shell.execute_reply.started":"2024-07-22T15:35:33.394804Z","shell.execute_reply":"2024-07-22T15:35:33.404505Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"pipline = SetupPipline()","metadata":{"execution":{"iopub.status.busy":"2024-07-22T15:35:33.406224Z","iopub.execute_input":"2024-07-22T15:35:33.407356Z","iopub.status.idle":"2024-07-22T15:36:23.165168Z","shell.execute_reply.started":"2024-07-22T15:35:33.407331Z","shell.execute_reply":"2024-07-22T15:36:23.164383Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"pipline.display()","metadata":{"execution":{"iopub.status.busy":"2024-07-22T15:36:23.166259Z","iopub.execute_input":"2024-07-22T15:36:23.166552Z","iopub.status.idle":"2024-07-22T15:36:23.171374Z","shell.execute_reply.started":"2024-07-22T15:36:23.166527Z","shell.execute_reply":"2024-07-22T15:36:23.170412Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"from pycocotools import _mask as coco_mask \nfrom ultralytics import YOLO","metadata":{"execution":{"iopub.status.busy":"2024-07-22T15:36:23.172527Z","iopub.execute_input":"2024-07-22T15:36:23.172855Z","iopub.status.idle":"2024-07-22T15:36:30.400791Z","shell.execute_reply.started":"2024-07-22T15:36:23.172826Z","shell.execute_reply":"2024-07-22T15:36:30.400005Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Logging\n<div class=\"alert alert-block alert-info\" style=\"font-size:20px; background-color: #eec4ff; font-family:verdana; color: #511b66; border: 2px #511b66 solid\">\n    <br>Experiment control tools are a very important part of the pipeline. Its used to track learning and compare results with previous ones to solve the problem.<br>\n</div>","metadata":{}},{"cell_type":"markdown","source":"# ClearML Logging and Automation \n\n[ClearML](https://cutt.ly/yolov5-notebook-clearml) is completely integrated into YOLOv5 to track your experimentation, manage dataset versions and even remotely execute training runs. To enable ClearML (check cells above):\n\n- `pip install clearml`\n- run `clearml-init` to connect to a ClearML server (**deploy your own [open-source server](https://github.com/allegroai/clearml-server)**, or use our [free hosted server](https://cutt.ly/yolov5-notebook-clearml))\n\nYou'll get all the great expected features from an experiment manager: live updates, model upload, experiment comparison etc. but ClearML also tracks uncommitted changes and installed packages for example. Thanks to that ClearML Tasks (which is what we call experiments) are also reproducible on different machines! With only 1 extra line, we can schedule a YOLOv5 training task on a queue to be executed by any number of ClearML Agents (workers).\n\nYou can use ClearML Data to version your dataset and then pass it to YOLOv5 simply using its unique ID. This will help you keep track of your data without adding extra hassle. Explore the [ClearML Tutorial](https://github.com/ultralytics/yolov5/tree/master/utils/loggers/clearml) for details!\n\n<a href=\"https://cutt.ly/yolov5-notebook-clearml\">\n<img alt=\"ClearML Experiment Management UI\" src=\"https://github.com/thepycoder/clearml_screenshots/raw/main/scalars.jpg\" width=\"1280\"/></a>","metadata":{}},{"cell_type":"markdown","source":"# Weights and biases\n  \n <p align='center'> \n <a href=\"https://pypi.python.org/pypi/wandb\"><img src=\"https://img.shields.io/pypi/v/wandb\" /></a> \n <a href=\"https://anaconda.org/conda-forge/wandb\"><img src=\"https://img.shields.io/conda/vn/conda-forge/wandb\" /></a> \n <a href=\"https://circleci.com/gh/wandb/wandb\"><img src=\"https://img.shields.io/circleci/build/github/wandb/wandb/main\" /></a> \n <a href=\"https://codecov.io/gh/wandb/wandb\"><img src=\"https://img.shields.io/codecov/c/gh/wandb/wandb\" /></a> \n </p> \n <p align='center'> \n <a href=\"https://colab.research.google.com/github/wandb/examples/blob/master/colabs/intro/Intro_to_Weights_%26_Biases.ipynb\"><img src=\"https://colab.research.google.com/assets/colab-badge.svg\" /></a> \n </p> \n  \n Use W&B to build better models faster. Track and visualize all the pieces of your machine learning pipeline, from datasets to production machine learning models. Get started with W&B today, [sign up for a free account!](https://wandb.com?utm_source=github&utm_medium=code&utm_campaign=wandb&utm_content=readme) \n  \n 🎓 W&B is free for students, educators, and academic researchers. For more information, visit [https://wandb.ai/site/research](https://wandb.ai/site/research?utm_source=github&utm_medium=code&utm_campaign=wandb&utm_content=readme). \n  \n Want to use Weights & Biases for seamless collaboration between your ML or Data Science team? Looking for Production-grade MLOps at scale? Sign up to one of [our plans](https://wandb.ai/site/pricing) or [contact the Sales Team](https://wandb.ai/site/contact).\n","metadata":{}},{"cell_type":"markdown","source":"# YOLOv8 Architecture: A Deep Dive\n<div class=\"alert alert-block alert-info\" style=\"font-size:15px; background-color: #eec4ff; font-family:verdana; color: #511b66; border: 2px #511b66 solid\">\nYOLOv8 is the latest version of the YOLO AI model developed by Ultralytics, which has shown effectiveness in tackling tasks such as classification, object detection, and image segmentation. YOLOv8 models are fast, accurate, and easy to use, making them ideal for various object detection and image segmentation tasks. They can be trained on large datasets and run on diverse hardware platforms, from CPUs to GPUs. YOLOv8 detection models have no suffix and are the default YOLOv8 models, i.e. yolov8n.pt ,and are pre-trained on COCO. See <a href=\"https://docs.ultralytics.com/tasks/detect/\">Detection Docs</a> for full details.\n\nHere we provide a quick summary of impactful modeling updates and then we will look at the model's evaluation, which speaks for itself.\n\nThe following image made by GitHub user RangeKing shows a detailed visualisation of the network's architecture.\n</div>","metadata":{}},{"cell_type":"markdown","source":"![](https://blog.roboflow.com/content/images/size/w1000/2023/01/image-16.png)","metadata":{}},{"cell_type":"markdown","source":"# Training\n<div class=\"alert alert-block alert-info\" style=\"font-size:20px; background-color: #eec4ff; font-family:verdana; color: #511b66; border: 2px #511b66 solid\">    <br>Hypothesize hypotheses, select hypermaparmeters, train, and then check whether the hypothesis was correct using the resulting metrics.<br>\n</div>","metadata":{}},{"cell_type":"code","source":"def main():\n    model = YOLO(\"yolov8x-seg.pt\")\n    model.train(\n        # Project\n        project=\"HuBMAP\",\n        name=\"yolov8x-seg\",\n\n        # Random Seed parameters\n        deterministic=True,\n        seed=43,\n\n        # Data & model parameters\n        data=\"/kaggle/working/dataset/coco.yaml\", \n        save=True,\n        save_period=5,\n        pretrained=True,\n        imgsz=512,\n\n        # Training parameters\n        epochs=20,\n        batch=4,\n        workers=8,\n        val=True,\n        device=0,\n\n        # Optimization parameters\n        lr0=0.018,\n        patience=3,\n        optimizer=\"SGD\",\n        momentum=0.947,\n        weight_decay=0.0005,\n        close_mosaic=3,\n    )","metadata":{"execution":{"iopub.status.busy":"2024-07-22T15:36:30.403496Z","iopub.execute_input":"2024-07-22T15:36:30.403874Z","iopub.status.idle":"2024-07-22T15:36:30.410215Z","shell.execute_reply.started":"2024-07-22T15:36:30.403849Z","shell.execute_reply":"2024-07-22T15:36:30.409303Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<div class=\"alert alert-block alert-info\" style=\"font-size:15px; background-color: #eec4ff; font-family:verdana; color: #511b66; border: 2px #511b66 solid\">\n    <b>Please follow the <a href=https://wandb.ai/site>link</a>. Sign up and then paste your API key into the box that pops up below. Do not disclose your key to anyone. Here it is required so that you can track the training</b>\n</div>","metadata":{}},{"cell_type":"code","source":"if __name__ == \"__main__\":\n    main()","metadata":{"execution":{"iopub.status.busy":"2024-07-22T15:36:30.411283Z","iopub.execute_input":"2024-07-22T15:36:30.411553Z","iopub.status.idle":"2024-07-22T16:27:27.444108Z","shell.execute_reply.started":"2024-07-22T15:36:30.411529Z","shell.execute_reply":"2024-07-22T16:27:27.442815Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Predict\n\n<div class=\"alert alert-block alert-info\" style=\"font-size:20px; font-family:verdana;\">\n<b>The training has come to an end and now it is time to look at the results of the work of the neural network</b>\n<ul style=\"font-size:20px; font-family:verdana; line-height: 1.7em\">\n    <li>Choose any picture</li>\n    <li>Displaying the image</li>\n</ul>\n</div>","metadata":{}},{"cell_type":"code","source":"import matplotlib.pyplot as plt\nfrom PIL import Image","metadata":{"execution":{"iopub.status.busy":"2024-07-22T16:27:27.446106Z","iopub.execute_input":"2024-07-22T16:27:27.446629Z","iopub.status.idle":"2024-07-22T16:27:27.452438Z","shell.execute_reply.started":"2024-07-22T16:27:27.446589Z","shell.execute_reply":"2024-07-22T16:27:27.451540Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<div class=\"alert alert-block alert-info\" style=\"font-size:20px; font-family:verdana;\">\n    <b>Choose an image</b>\n</div>","metadata":{}},{"cell_type":"code","source":"dirlist = os.listdir(\"/kaggle/working/dataset/val/images\")\nprint(dirlist[:5])","metadata":{"execution":{"iopub.status.busy":"2024-07-22T16:27:27.453408Z","iopub.execute_input":"2024-07-22T16:27:27.453722Z","iopub.status.idle":"2024-07-22T16:27:27.968424Z","shell.execute_reply.started":"2024-07-22T16:27:27.453690Z","shell.execute_reply":"2024-07-22T16:27:27.967480Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<div class=\"alert alert-block alert-info\" style=\"font-size:20px; font-family:verdana;\">\n    <b>Displaying the image</b>\n</div>","metadata":{}},{"cell_type":"code","source":"model = YOLO(\"/kaggle/working/HuBMAP/yolov8x-seg/weights/best.pt\")\nhistory = model.predict(\"../working/dataset/val/images/ed6a92a9410c.tif\")[0]\nimage = history.plot()\nplt.imshow(image)\nplt.show()","metadata":{"execution":{"iopub.status.busy":"2024-07-22T16:27:27.970754Z","iopub.execute_input":"2024-07-22T16:27:27.971147Z","iopub.status.idle":"2024-07-22T16:27:29.637730Z","shell.execute_reply.started":"2024-07-22T16:27:27.971110Z","shell.execute_reply":"2024-07-22T16:27:29.636604Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Let's build the graphs \n<div class=\"alert alert-block alert-info\" style=\"font-size:20px; font-family:verdana;\">\n    <b>Losses & Recall & Precision & mAP</b>\n</div>","metadata":{}},{"cell_type":"code","source":"F1_curve = Image.open(\"/kaggle/working/HuBMAP/yolov8x-seg/results.png\")\nplt.figure(figsize=(15,20))\nplt.imshow(F1_curve)\nplt.show()","metadata":{"execution":{"iopub.status.busy":"2024-07-22T16:27:29.639178Z","iopub.execute_input":"2024-07-22T16:27:29.639502Z","iopub.status.idle":"2024-07-22T16:27:30.426391Z","shell.execute_reply.started":"2024-07-22T16:27:29.639474Z","shell.execute_reply":"2024-07-22T16:27:30.425250Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<div class=\"alert alert-block alert-info\" style=\"font-size:20px; font-family:verdana;\">\n    <b>Train Batch</b>\n</div>","metadata":{}},{"cell_type":"code","source":"P_curve = Image.open(\"/kaggle/working/HuBMAP/yolov8x-seg/train_batch5561.jpg\")\nplt.imshow(P_curve)\nplt.show()","metadata":{"execution":{"iopub.status.busy":"2024-07-22T16:27:30.427636Z","iopub.execute_input":"2024-07-22T16:27:30.427924Z","iopub.status.idle":"2024-07-22T16:27:30.817326Z","shell.execute_reply.started":"2024-07-22T16:27:30.427899Z","shell.execute_reply":"2024-07-22T16:27:30.816207Z"},"trusted":true},"execution_count":null,"outputs":[]}]}