{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"name":"python","version":"3.10.10","mimetype":"text/x-python","codemirror_mode":{"name":"ipython","version":3},"pygments_lexer":"ipython3","nbconvert_exporter":"python","file_extension":".py"},"kaggle":{"accelerator":"gpu","dataSources":[{"sourceId":52279,"databundleVersionId":5822112,"sourceType":"competition"},{"sourceId":6021726,"sourceType":"datasetVersion","datasetId":3446188}],"dockerImageVersionId":30512,"isInternetEnabled":true,"language":"python","sourceType":"notebook","isGpuEnabled":true}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"![](https://i.pinimg.com/originals/51/ae/4b/51ae4b928ea815934b562d907ec5dd91.gif)","metadata":{}},{"cell_type":"markdown","source":"<div class=\"alert alert-block alert-info\" style=\"font-size:20px; background-color: #f3fc9a; font-family:verdana; color: #252908; border: 3px #252908 solid\">\n    <b>⚕️ MaskRCNN from scratch | Medical Segmentation</b>\n    <br>Hi friends! 😉<br>\n    Today we will dive into the problem of segmentation of instances in a medical context, and we will also use the MaskRCNN model. I already have work on the topic of instance segmentation, but I believe that a real expert should be able to use not only ready-made pipelines like ultralytics, but also be able to build a pipeline, design a model and train it. It is for this purpose that we are gathered here. Many parts with analysis or explanation of data and tasks are taken from previous work, since I didn’t want to reinvent the wheel (but I changed the gif with bears 🐼) I hope you can learn something new 🙃\n</div>","metadata":{}},{"cell_type":"markdown","source":"![](https://66.media.tumblr.com/dcae971a5a6e5de34e084e3cf744178e/tumblr_omhrqxDQY91qg39ewo1_500.gif)","metadata":{}},{"cell_type":"markdown","source":"<div class=\"alert alert-block alert-info\" style=\"font-size:20px; background-color: #d1a6ff; font-family:verdana; color: #533078; border: 3px #533078 solid\">\n    <b>Biological overview ⚕️</b>\n    <br>Our task is to segment medical images of kidney tissue samples. Human tissues are covered with many vessels. Our task is to segment such vessels.<br>\n</div>","metadata":{}},{"cell_type":"markdown","source":"![](https://i.ibb.co/1MZNhNF/003.png)","metadata":{}},{"cell_type":"markdown","source":"<div class=\"alert alert-block alert-info\" style=\"font-size:20px; background-color: #d1a6ff; font-family:verdana; color: #533078; border: 2px #533078 solid\">\n    <b>Tissue samples 🧪</b>\n    <br>Renal cortex: the renal cortex is the outer portion of the kidney that contains round renal corpuscles enclosing glomerular tufts, balls of capillary loops. The renal corpuscle is the start of the nephron, through which the filtration of blood occurs. The renal cortex also contains proximal and distal convoluted tubules, PCTs and DCTs, which are also regions of the nephron. Between these tubular structures is a complex network of capillaries called the peritubular capillaries.<br>\n    <br>Renal medulla: the renal medulla is the inner portion of the kidney and is arranged in 8-15 renal pyramids containing linearly arranged tubules comprising the loops of Henle and ducts that gather products for excretion. The capillary network in the renal medulla consists of capillaries called vasa recta. The renal pyramids (medullary tissue) are divided by extensions of the renal cortex called renal columns. There are also projections of the renal medulla into the outer cortex called medullary rays.<br>\n    <br>Renal papilla (subsection of medulla): the broad bases of the pyramids connect to the renal cortex at the corticomedullary junctions while the tips form structures called the renal papilla, which project in the minor renal calyces where urine is collected.<br>\n</div>","metadata":{}},{"cell_type":"markdown","source":"<div class=\"alert alert-block alert-info\" style=\"font-size:20px; background-color: #b5e1ff; font-family:verdana; color: #011d82; border: 2px #011d82 solid\">\n    <b>What is instance segmentation?</b>\n    <br>Instance Segmentation is a unique form of image segmentation that deals with detecting and delineating each distinct instance of an object appearing in an image. Instance segmentation detects all instances of a class with the extra functionality of demarcating separate instances of any segment class. Hence, it is also referred to as incorporating object detection and semantic segmentation functionality. Watch video to learn more: <a href=instance segmentation>Instance segmentation | Tutorial</a> <br>\n</div>","metadata":{}},{"cell_type":"markdown","source":"<div class=\"alert alert-block alert-info\" style=\"font-size:20px; background-color: #b5e1ff; font-family:verdana; color: #011d82; border: 2px #011d82 solid\">\n    <b>Applications of Semantic Segmentation</b>\n<ol><li><strong>Medical Diagnostics: </strong>For detecting medical abnormalities in <a href=\"https://universe.roboflow.com/search?q=xray&amp;ref=blog.roboflow.com\">X-Rays, CT Scans, MRI Scans</a></li><li><strong>GeoSensing:</strong> For land usage mapping from <a href=\"https://roboflow.com/solutions/aerial?ref=blog.roboflow.com\">satellite imagery</a> and monitoring areas of deforestation and urbanization</li><li><strong>Autonomous Driving:</strong> For accurately <a href=\"https://universe.roboflow.com/browse/self-driving?ref=blog.roboflow.com\">detecting lanes, pedestrians, traffic signs, road</a>, sky and other vehicles on the road</li></ol><br>\n</div>\n\n","metadata":{}},{"cell_type":"markdown","source":"<div class=\"alert alert-block alert-info\" style=\"font-size:20px; background-color: #bdc9bf; font-family:verdana; color: #2b2b2b; border: 2px #2b2b2b solid\">\n    <b>For determinism 🤗 It's a joke, but we still need reproducibility of experiments</b>\n</div>","metadata":{}},{"cell_type":"code","source":"import numpy as np\nimport random\nimport torch\nimport os","metadata":{"execution":{"iopub.status.busy":"2024-01-11T17:21:07.516602Z","iopub.execute_input":"2024-01-11T17:21:07.517387Z","iopub.status.idle":"2024-01-11T17:21:10.473659Z","shell.execute_reply.started":"2024-01-11T17:21:07.517356Z","shell.execute_reply":"2024-01-11T17:21:10.472872Z"},"trusted":true},"execution_count":1,"outputs":[]},{"cell_type":"code","source":"def seed_everythig(seed):\n    np.random.seed(seed)\n    random.seed(seed)\n    os.environ[\"GLOBALSEED\"] = str(seed)\n    torch.manual_seed(seed)\n    torch.cuda.manual_seed(seed)\n    if torch.cuda.is_available():\n        torch.cuda.manual_seed_all(seed)\n        torch.backends.cudnn.deterministic = True","metadata":{"execution":{"iopub.status.busy":"2024-01-11T17:21:10.475312Z","iopub.execute_input":"2024-01-11T17:21:10.475927Z","iopub.status.idle":"2024-01-11T17:21:10.481115Z","shell.execute_reply.started":"2024-01-11T17:21:10.475894Z","shell.execute_reply":"2024-01-11T17:21:10.480257Z"},"trusted":true},"execution_count":2,"outputs":[]},{"cell_type":"code","source":"seed_everythig(42)","metadata":{"execution":{"iopub.status.busy":"2024-01-11T17:21:10.482115Z","iopub.execute_input":"2024-01-11T17:21:10.482369Z","iopub.status.idle":"2024-01-11T17:21:10.519244Z","shell.execute_reply.started":"2024-01-11T17:21:10.482346Z","shell.execute_reply":"2024-01-11T17:21:10.518433Z"},"trusted":true},"execution_count":3,"outputs":[]},{"cell_type":"markdown","source":"# Prepare data\n<div class=\"alert alert-block alert-info\" style=\"font-size:20px; background-color: #abffd1; font-family:verdana; color: #003819; border: 2px #003819 solid\">\n    <b>It's time to think about data 📚</b>\n    <br>For training, we need a dataset, and not just any, but in the COCO format. COCO is an open database for object detection, it is huge and therefore it is used all over the world to evaluate new models according to various metrics. It plays a rather important role, and most detection models are pre-trained on it. That is why this dataset format is used in ultralytics. You don't need to worry about this, as I have already prepared a class for data processing. I advise you to skip the cell below and immediately proceed to the next chapter.<br>\n</div>","metadata":{}},{"cell_type":"markdown","source":"<div class=\"alert alert-block alert-info\" style=\"font-size:20px; background-color: #abffd1; font-family:verdana; color: #003819; border: 2px #003819 solid\">\n    <b>Libraries for creating dataset</b>\n</div>","metadata":{}},{"cell_type":"code","source":"from itertools import chain\nimport json\nimport os\nimport shutil\nfrom tqdm.notebook import tqdm\nfrom colorama import Fore\nimport yaml","metadata":{"execution":{"iopub.status.busy":"2024-01-11T17:21:10.520287Z","iopub.execute_input":"2024-01-11T17:21:10.520544Z","iopub.status.idle":"2024-01-11T17:21:10.609139Z","shell.execute_reply.started":"2024-01-11T17:21:10.520521Z","shell.execute_reply":"2024-01-11T17:21:10.608311Z"},"trusted":true},"execution_count":4,"outputs":[]},{"cell_type":"markdown","source":"<div class=\"alert alert-block alert-info\" style=\"font-size:20px; background-color: #abffd1; font-family:verdana; color: #003819; border: 2px #003819 solid\">\n    <b>Class for creating dataset</b>\n</div>","metadata":{}},{"cell_type":"code","source":"class COCODataset:\n    def __init__(self, images_dirpath: str, annotations_filepath: str, length: int = 1633):\n        self.train_size = None\n        self.val_size = None\n        self.length = length\n        self.classes = None\n        self.labels_counter = None\n        self.normalize = None\n        \n        self.images_dirpath = images_dirpath\n        self.annotations_filepath = annotations_filepath\n        self.dataset_dirpath = os.path.join(os.getcwd(), \"dataset\")\n        self.train_dirpath =  os.path.join(self.dataset_dirpath, \"train\")\n        self.val_dirpath =  os.path.join(self.dataset_dirpath, \"val\")\n        self.config_path = os.path.join(self.dataset_dirpath, \"coco.yaml\")\n\n        self.samples = self.parse_jsonl(annotations_filepath)\n        self.classes_dict = {\n            \"background\": 0,\n            \"blood_vessel\": 1,\n            \"glomerulus\": 2,\n            \"unsure\": 3,\n        }\n\n    def __prepare_dirs(self) -> None:\n        if not os.path.exists(self.dataset_dirpath):\n            os.makedirs(os.path.join(self.train_dirpath, \"images\"), exist_ok=True)\n            os.makedirs(os.path.join(self.train_dirpath, \"labels\"), exist_ok=True)\n            os.makedirs(os.path.join(self.val_dirpath, \"images\"), exist_ok=True)\n            os.makedirs(os.path.join(self.val_dirpath, \"labels\"), exist_ok=True)\n        else:\n            raise RuntimeError(\"Dataset already exists!\")\n\n    def __define_splitratio(self) -> None:\n        self.train_size = round(self.length * self.train_size)\n        self.val_size = self.length - self.train_size\n        assert self.train_size + self.val_size == self.length\n\n    def parse_jsonl(self, path: str) -> list[dict, ...]:\n        with open(path, 'r') as json_file:\n            jsonl_samples = [\n                json.loads(line)\n                for line in tqdm(\n                    json_file, desc=\"Processing polygons\", total=self.length\n                )\n            ]\n        return jsonl_samples\n\n    def __define_paths(self, i: int) -> dict:\n        data_path = self.val_dirpath\n        if i < self.train_size:\n            data_path = self.train_dirpath\n        return {\n            \"images\": os.path.join(data_path, \"images\"),\n            \"labels\": os.path.join(data_path, \"labels\")\n        }\n\n    @staticmethod\n    def __get_label_path(paths_dict: dict, identifier: str) -> str:\n        return os.path.join(\n            paths_dict[\"labels\"],\n            f\"{identifier}.txt\"\n        )\n\n    @staticmethod\n    def __get_image_path(paths_dict: dict, identifier: str) -> str:\n        return os.path.join(\n            paths_dict[\"images\"],\n            f\"{identifier}.tif\"\n        )\n\n    def __copy_image(self, dst_path: str, identifier: str) -> str:\n        shutil.copyfile(\n            os.path.join(self.images_dirpath, f\"{identifier}.tif\"),\n            dst_path\n        )\n\n    def __copy_label(self, annotations: list, dst_path: str) -> None:\n        with open(dst_path, \"w\") as file:\n            for annotation in annotations:\n                coordinates = annotation[\"coordinates\"][0]\n                label = self.classes_dict[annotation[\"type\"]]\n                if label in self.classes:\n                    if coordinates:\n                        if self.normalize:\n                            coordinates = np.array(coordinates) / 512.0\n                        coordinates = \" \".join(map(str, chain(*coordinates)))\n                        file.write(f\"{label} {coordinates}\\n\")\n                        self.labels_counter += 1\n\n    def __splitfolders(self):\n        for i, line in tqdm(\n                enumerate(self.samples),\n                desc=\"Dataset creation\", total=self.length\n        ):\n            self.labels_counter = 0\n            identifier = line[\"id\"]\n            annotations = line[\"annotations\"]\n            paths_dict = self.__define_paths(i)\n\n            dst_image_path = self.__get_image_path(paths_dict, identifier)\n            dst_label_path = self.__get_label_path(paths_dict, identifier)\n\n            self.__copy_image(dst_image_path, identifier)\n            self.__copy_label(annotations, dst_label_path)\n\n            if self.labels_counter == 0:\n                os.remove(dst_image_path)\n                os.remove(dst_label_path)\n\n    def __count_dataset(self) -> dict:\n        train_images = len(os.listdir(os.path.join(self.train_dirpath, \"images\")))\n        train_labels = len(os.listdir(os.path.join(self.train_dirpath, \"labels\")))\n        val_images = len(os.listdir(os.path.join(self.val_dirpath, \"images\")))\n        val_labels = len(os.listdir(os.path.join(self.val_dirpath, \"labels\")))\n        return {\n            \"train_images\": train_images,\n            \"train_labels\": train_labels,\n            \"val_images\": val_images,\n            \"val_labels\": val_labels\n        }\n\n    @staticmethod\n    def __check_sanity(count_dict: dict) -> None:\n        assert count_dict[\"train_images\"] == count_dict[\"train_labels\"]\n        assert count_dict[\"val_images\"] == count_dict[\"val_labels\"]\n\n    def __finalizing(self, count_dict: dict) -> None:\n        assert os.path.exists(self.dataset_dirpath)\n\n        example_structure = [\n            \"dataset\",\n            \"train\", \"labels\", \"images\",\n            \"val\", \"labels\", \"images\"\n        ]\n\n        dir_bone = (\n            dirname.split(\"/\")[-1]\n            for dirname, _, filenames in os.walk(self.dataset_dirpath)\n            if dirname.split(\"/\")[-1] in example_structure\n        )\n\n        try:\n            print(\"\\n~ HuBMAP Dataset Structure ~\\n\")\n            print(\n            f\"\"\"\n          ├── {next(dir_bone)}\n          │   │\n          │   ├── {next(dir_bone)}\n          │   │   └── {next(dir_bone)}\n          │   │   └── {next(dir_bone)}\n          │   │\n          │   ├── {next(dir_bone)}\n          │   │   └── {next(dir_bone)}\n          │   │   └── {next(dir_bone)}\n            \"\"\"\n            )\n        except StopIteration as e:\n            print(e)\n        else:\n            print(Fore.GREEN + \"-> Success\")\n            print(Fore.GREEN + f\"Train dataset: {count_dict['train_images']}\\nVal dataset: {count_dict['val_images']}\")\n\n    def get_config(self) ->dict:\n        names = [\"background\", \"blood_vessel\", \"glomerulus\", \"unsure\"]\n        return {\n            \"train\": str(self.train_dirpath),\n            \"val\": str(self.val_dirpath),\n            \"names\": [names[i] for i in self.classes]\n        }\n\n    @staticmethod\n    def display_config(config: dict) -> None:\n        print(Fore.BLACK + \"\\n~ HuBMAP Config Structure ~\\n\")\n        print(\n        f\"\"\"\n      │   │\n      │   ├── train\n      │   │   └── {config['train']}/images\n      │   │\n      │   │\n      │   ├── val\n      │   │   └── {config['val']}/images\n      │   │\n      │   │\n      │   ├── names\n      │   │   └── {' '.join(config['names'])}\n        \"\"\"\n        )\n        print(Fore.GREEN + \"-> Success\")\n        print(Fore.GREEN + f\"Number of classes: {len(config['names'])}\"\n                           f\"\\nClasses: {' '.join(config['names'])}\" \n              )\n\n    def write_config(self, config: dict) -> None:\n        with open(self.config_path, mode=\"w\") as f:\n            yaml.safe_dump(stream=f, data=config)\n\n    def __call__(self, train_size: float,\n                 classes: list[int, ...],\n                 make_config: bool = True,\n                 normalize: bool = True\n                ) -> None:\n        \n        self.train_size = train_size\n        self.classes = classes\n        self.normalize = normalize\n        \n        self.__define_splitratio()\n        self.__prepare_dirs()\n        self.__splitfolders()\n        count_dict = self.__count_dataset()\n        self.__check_sanity(count_dict)\n        self.__finalizing(count_dict)\n        \n        if make_config:\n            config = self.get_config()\n            self.write_config(config)\n            self.display_config(config)","metadata":{"execution":{"iopub.status.busy":"2024-01-11T17:21:10.612864Z","iopub.execute_input":"2024-01-11T17:21:10.613547Z","iopub.status.idle":"2024-01-11T17:21:10.647514Z","shell.execute_reply.started":"2024-01-11T17:21:10.613515Z","shell.execute_reply":"2024-01-11T17:21:10.646606Z"},"trusted":true},"execution_count":5,"outputs":[]},{"cell_type":"markdown","source":"<div class=\"alert alert-block alert-info\" style=\"font-size:20px; background-color: #abffd1; font-family:verdana; color: #003819; border: 2px #003819 solid\">\n    <b>Let's fire up the data 🔥</b>\n</div>","metadata":{}},{"cell_type":"code","source":"coco = COCODataset(\n    annotations_filepath=\"/kaggle/input/hubmap-hacking-the-human-vasculature/polygons.jsonl\",\n    images_dirpath=\"/kaggle/input/hubmap-hacking-the-human-vasculature/train\",\n) ","metadata":{"execution":{"iopub.status.busy":"2024-01-11T17:21:10.648353Z","iopub.execute_input":"2024-01-11T17:21:10.648625Z","iopub.status.idle":"2024-01-11T17:21:14.859319Z","shell.execute_reply.started":"2024-01-11T17:21:10.648595Z","shell.execute_reply":"2024-01-11T17:21:14.858412Z"},"trusted":true},"execution_count":6,"outputs":[{"output_type":"display_data","data":{"text/plain":"Processing polygons:   0%|          | 0/1633 [00:00<?, ?it/s]","application/vnd.jupyter.widget-view+json":{"version_major":2,"version_minor":0,"model_id":"0729812239aa441e93a5fcf4a4813fac"}},"metadata":{}}]},{"cell_type":"code","source":"coco(train_size=0.85, classes=[1, 2, 3], normalize=False)","metadata":{"execution":{"iopub.status.busy":"2024-01-11T17:21:14.86093Z","iopub.execute_input":"2024-01-11T17:21:14.861273Z","iopub.status.idle":"2024-01-11T17:21:42.36505Z","shell.execute_reply.started":"2024-01-11T17:21:14.861241Z","shell.execute_reply":"2024-01-11T17:21:42.364132Z"},"trusted":true},"execution_count":7,"outputs":[{"output_type":"display_data","data":{"text/plain":"Dataset creation:   0%|          | 0/1633 [00:00<?, ?it/s]","application/vnd.jupyter.widget-view+json":{"version_major":2,"version_minor":0,"model_id":"bd5933415bc44c4f858371b31411fdce"}},"metadata":{}},{"name":"stdout","text":"\n~ HuBMAP Dataset Structure ~\n\n\n          ├── dataset\n          │   │\n          │   ├── train\n          │   │   └── images\n          │   │   └── labels\n          │   │\n          │   ├── val\n          │   │   └── images\n          │   │   └── labels\n            \n\u001b[32m-> Success\n\u001b[32mTrain dataset: 1388\nVal dataset: 245\n\u001b[30m\n~ HuBMAP Config Structure ~\n\n\n      │   │\n      │   ├── train\n      │   │   └── /kaggle/working/dataset/train/images\n      │   │\n      │   │\n      │   ├── val\n      │   │   └── /kaggle/working/dataset/val/images\n      │   │\n      │   │\n      │   ├── names\n      │   │   └── blood_vessel glomerulus unsure\n        \n\u001b[32m-> Success\n\u001b[32mNumber of classes: 3\nClasses: blood_vessel glomerulus unsure\n","output_type":"stream"}]},{"cell_type":"markdown","source":"# About Mask-RCNN\n<div class=\"alert alert-block alert-info\" style=\"font-size:20px; background-color: #a1e6c4; font-family:verdana; color: #004725; border: 2px #004725 solid\">\n    <b>Understanding Architecture</b>\n    <br>Here I would like to share information about the Mask-RCNN architecture and its development<br>\n</div>","metadata":{}},{"cell_type":"markdown","source":"![](https://media.tenor.com/J5zw2jZ6TsYAAAAC/ice-bear-math-lady.gif)","metadata":{}},{"cell_type":"markdown","source":"<div class=\"alert alert-block alert-info\" style=\"font-size:20px; background-color: #a1e6c4; font-family:verdana; color: #004725; border: 2px #004725 solid\">\n    <b>R-CNN (Regions With CNNs)</b>\n    <br>The R-CNN (Regions With CNNs) network architecture was developed by a team at UC Berkley to apply Convolution Neural Networks to the object detection problem. The approaches to solving such problems that existed at that time approached the maximum of their capabilities and it was not possible to significantly improve their performance.<br>\n    <br>CNNs have been good at image classification, and in this network they have essentially been applied to do the same thing. To do this, not the entire image was fed to the input of the CNN, but regions previously selected in a different way, in which presumably there are some objects. At that time there were several such approaches, the authors chose Selective Search, although they indicate that there are no special reasons for preferring it.<br>\n    <br>A ready-made architecture, CaffeNet (AlexNet), was also used as a CNN network. Such neural networks, like others for the ImageNet image set, classify into 1000 classes. R-CNN was designed to detect objects of fewer classes (N=20 or 200), so the last CaffeNet classification layer was replaced with a layer with N+1 outputs (with an additional class for the background).<br>\n    <br>Selective Search returned about 2000 regions of different sizes and aspect ratios, but CaffeNet accepts images of a fixed size of 227x227 pixels as input, so they had to be modified before submitting the regions to the network input. To do this, the image from the region was enclosed in the smallest enclosing square. Along that (smaller) side on which the fields were formed, several “context” (surrounding the region) pixels of the image were added, the rest of the field was not filled with anything. The resulting square was scaled to 227x227 and fed to the CaffeNet input.<br>\n    <br>Even though the CNN was trained to recognize N+1 classes, it ended up being used only to extract a fixed 4096-dimensional feature vector. N linear SVMs were engaged in the direct determination of the object in the image, each of which carried out a binary classification according to its type of objects, determining whether there is one in the transferred region or not. In the original document, the whole procedure is illustrated by the following diagram:<br>\n</div>","metadata":{}},{"cell_type":"markdown","source":"![](https://habrastorage.org/r/w1560/webt/6i/zz/rh/6izzrhggbprdgbpbqz-qa7lfreu.jpeg)","metadata":{}},{"cell_type":"markdown","source":"<div class=\"alert alert-block alert-info\" style=\"font-size:20px; background-color: #a1e6c4; font-family:verdana; color: #004725; border: 2px #004725 solid\">\n    <b>Fast R-CNN</b>\n    <br>Despite the good results, the performance of R-CNN was still poor, especially for networks deeper than CaffeNet (such as VGG16). In addition, training the bounding box regressor and SVM required a large number of features to be saved to disk, so it was expensive in terms of storage size. The authors of Fast R-CNN suggested speeding up the process with a couple of modifications:<br>\n    <ul>\n        <li>Pass through CNN not each of the 2000 candidate regions separately, but the entire image. The proposed regions are then superimposed on the resulting overall feature map;</li>\n        <li>Instead of training three models (CNN, SVM, bbox regressor) independently, combine all training procedures into one.</li>\n    </ul>\n    <br>The transformation of features that fell into different regions to a fixed size was performed using the RoIPooling procedure. The region window of width w and height h was divided into a grid with H × W cells of size h/H × w/W. (Document authors used W=H=7). For each such cell, Max Pooling was performed to select only one value, thus giving the resulting H×W feature matrix.<br>\n    <br>Binary SVMs were not used, instead the selected features were passed to a fully connected layer and then to two parallel layers: softmax with K+1 outputs (one for each class + 1 for the background) and a bounding box regressor. The general network architecture looks like this:<br>\n</div>","metadata":{}},{"cell_type":"markdown","source":"![](https://habrastorage.org/r/w1560/webt/yt/e-/9u/yte-9u2kp27hykwg9m1n98w6upg.png)","metadata":{}},{"cell_type":"markdown","source":"<div class=\"alert alert-block alert-info\" style=\"font-size:20px; background-color: #a1e6c4; font-family:verdana; color: #004725; border: 2px #004725 solid\">\n    <b>Faster R-CNN</b>\n    <br>After the improvements made in Fast R-CNN, the mechanism for generating candidate regions turned out to be the bottleneck of the neural network. In 2015, the team at Microsoft Research was able to make this phase much faster. They proposed to calculate the regions not from the original image, but again from the feature map obtained from CNN. To do this, a module called the Region Proposal Network (RPN) was added. The new architecture looks like this:<br>\n</div>","metadata":{}},{"cell_type":"markdown","source":"![](https://habrastorage.org/r/w1560/webt/jf/-3/22/jf-3224hkrudwxi1_7oogb1avg0.png)","metadata":{}},{"cell_type":"markdown","source":"<div class=\"alert alert-block alert-info\" style=\"font-size:20px; background-color: #a1e6c4; font-family:verdana; color: #004725; border: 2px #004725 solid\">\nWithin the framework of RPN, a “mini-neural network” with a small (3x3) window slides along the extracted CNN features. The values obtained with its help are transferred to two parallel fully connected layers: box-regression layer (reg) and box-classification layer (cls). The outputs of these layers are based on the so-called anchors: k frames for each position of the sliding window, with different sizes and aspect ratios. The reg layer for each such anchor produces 4 coordinates that adjust the position of the enclosing frame; cls-layer produces two numbers each - the probabilities that the frame contains at least some object or that it does not. This is illustrated in the document as follows:\n</div>","metadata":{}},{"cell_type":"markdown","source":"![](https://habrastorage.org/r/w1560/webt/kj/kj/oz/kjkjozq2gq70fuvf77dbml8-_nw.png)","metadata":{}},{"cell_type":"markdown","source":"<div class=\"alert alert-block alert-info\" style=\"font-size:20px; background-color: #a1e6c4; font-family:verdana; color: #004725; border: 2px #004725 solid\">\n    <b>Mask R-CNN</b>\n<br>Mask R-CNN develops the Faster R-CNN architecture by adding another branch that predicts the position of the mask covering the found object, and thus solves the instance segmentation problem. The mask is simply a rectangular matrix, in which 1 at some position means that the corresponding pixel belongs to an object of a given class, 0 means that the pixel does not belong to the object.<br>\n</div>","metadata":{}},{"cell_type":"markdown","source":"![](https://habrastorage.org/r/w1560/webt/n3/cs/tp/n3cstpty6ktfwhw6vklswah1rxk.png)","metadata":{}},{"cell_type":"markdown","source":"<div class=\"alert alert-block alert-info\" style=\"font-size:20px; background-color: #a1e6c4; font-family:verdana; color: #004725; border: 2px #004725 solid\">\n    <b>Do you feel? That feeling of realizing what you're using 🙃</b>\n<br>This is a squeeze that, in principle, should give you an idea of ​​​​the history of the development of Mask-RCNN and its architecture. I know I didn’t mention the intricacies of training and composite loss functions, but I decided to leave this to your interest so as not to drag out the notebook<br>\n</div>","metadata":{}},{"cell_type":"code","source":"import torchvision\nfrom torchvision.models.detection.faster_rcnn import FastRCNNPredictor\nfrom torchvision.models.detection.mask_rcnn import MaskRCNNPredictor","metadata":{"execution":{"iopub.status.busy":"2024-01-11T17:21:42.366192Z","iopub.execute_input":"2024-01-11T17:21:42.36648Z","iopub.status.idle":"2024-01-11T17:21:42.599296Z","shell.execute_reply.started":"2024-01-11T17:21:42.366455Z","shell.execute_reply":"2024-01-11T17:21:42.598612Z"},"trusted":true},"execution_count":8,"outputs":[]},{"cell_type":"code","source":"def get_model(num_classes: int):\n    # load an instance segmentation model pre-trained on COCO\n    model = torchvision.models.detection.maskrcnn_resnet50_fpn(weights=\"DEFAULT\")\n\n    # get number of input features for the classifier\n    in_features = model.roi_heads.box_predictor.cls_score.in_features\n    # replace the pre-trained head with a new one\n    model.roi_heads.box_predictor = FastRCNNPredictor(in_features, num_classes)\n\n    # now get the number of input features for the mask classifier\n    in_features_mask = model.roi_heads.mask_predictor.conv5_mask.in_channels\n    hidden_layer = 256\n    # and replace the mask predictor with a new one\n    model.roi_heads.mask_predictor = MaskRCNNPredictor(in_features_mask,\n                                                       hidden_layer,\n                                                       num_classes)\n\n    return model","metadata":{"execution":{"iopub.status.busy":"2024-01-11T17:21:42.600344Z","iopub.execute_input":"2024-01-11T17:21:42.600644Z","iopub.status.idle":"2024-01-11T17:21:42.606471Z","shell.execute_reply.started":"2024-01-11T17:21:42.600617Z","shell.execute_reply":"2024-01-11T17:21:42.605635Z"},"trusted":true},"execution_count":9,"outputs":[]},{"cell_type":"code","source":"model = get_model(4)","metadata":{"execution":{"iopub.status.busy":"2024-01-11T17:21:42.607693Z","iopub.execute_input":"2024-01-11T17:21:42.608001Z","iopub.status.idle":"2024-01-11T17:21:44.383775Z","shell.execute_reply.started":"2024-01-11T17:21:42.607972Z","shell.execute_reply":"2024-01-11T17:21:44.38299Z"},"trusted":true},"execution_count":10,"outputs":[{"name":"stderr","text":"Downloading: \"https://download.pytorch.org/models/maskrcnn_resnet50_fpn_coco-bf2d0c1e.pth\" to /root/.cache/torch/hub/checkpoints/maskrcnn_resnet50_fpn_coco-bf2d0c1e.pth\n100%|██████████| 170M/170M [00:00<00:00, 284MB/s] \n","output_type":"stream"}]},{"cell_type":"markdown","source":"<div class=\"alert alert-block alert-info\" style=\"font-size:20px; background-color: #ffd37a; font-family:verdana; color: #543800; border: 2px #543800 solid\">\n    <b>What about transforming our data?</b>\n<br>Firstly, we absolutely need to bring the pictures to a single size and also normalize them. It would also not be bad to augment our data. But here's the problem. After all, we are now working not with an image and its single mask as in semantic segmentation, but with individual instances, their masks and boxes. In this case, we can ignore color transformations, but all spatial transformations change both box and mask coordinates. In this case, we must take this into account. A wonderful library from the community comes to the rescue - albumentations. It will allow you to quickly and easily take into account all the transformations. Otherwise, we had to read it ourselves.<br>\n</div>","metadata":{}},{"cell_type":"markdown","source":"![](https://www.researchgate.net/publication/319413978/figure/fig2/AS:533727585333249@1504261980375/Data-augmentation-using-semantic-preserving-transformation-for-SBIR.png)","metadata":{}},{"cell_type":"code","source":"import albumentations as A\nfrom albumentations.pytorch.transforms import ToTensorV2","metadata":{"execution":{"iopub.status.busy":"2024-01-11T17:21:44.38497Z","iopub.execute_input":"2024-01-11T17:21:44.385248Z","iopub.status.idle":"2024-01-11T17:21:45.867551Z","shell.execute_reply.started":"2024-01-11T17:21:44.385225Z","shell.execute_reply":"2024-01-11T17:21:45.866416Z"},"trusted":true},"execution_count":11,"outputs":[]},{"cell_type":"markdown","source":"<div class=\"alert alert-block alert-info\" style=\"font-size:20px; background-color: #ffd37a; font-family:verdana; color: #543800; border: 2px #543800 solid\">\n    <b>Transforms for training</b>\n</div>","metadata":{}},{"cell_type":"code","source":"train_transforms = [\n    A.Resize(512, 512, p=1), \n    A.HorizontalFlip(p=0.5),\n    A.VerticalFlip(p=0.5),\n    A.RandomBrightnessContrast(p=0.45),\n    A.HueSaturationValue(p=0.35),\n    \n    A.OneOf([\n        A.MotionBlur(),\n        A.Blur(blur_limit=3),\n        A.MedianBlur(blur_limit=3),\n        A.GaussNoise()\n        ], p=0.10\n    ),\n                         \n    A.Normalize(),\n    ToTensorV2()\n]","metadata":{"execution":{"iopub.status.busy":"2024-01-11T17:21:45.868979Z","iopub.execute_input":"2024-01-11T17:21:45.869336Z","iopub.status.idle":"2024-01-11T17:21:45.877064Z","shell.execute_reply.started":"2024-01-11T17:21:45.869303Z","shell.execute_reply":"2024-01-11T17:21:45.876139Z"},"trusted":true},"execution_count":12,"outputs":[]},{"cell_type":"markdown","source":"<div class=\"alert alert-block alert-info\" style=\"font-size:20px; background-color: #ffd37a; font-family:verdana; color: #543800; border: 2px #543800 solid\">\n    <b>And for validation</b>\n</div>","metadata":{}},{"cell_type":"code","source":"val_transforms = [\n    A.Resize(512, 512, p=1), \n    A.Normalize(),\n    ToTensorV2()\n]","metadata":{"execution":{"iopub.status.busy":"2024-01-11T17:21:45.878316Z","iopub.execute_input":"2024-01-11T17:21:45.878729Z","iopub.status.idle":"2024-01-11T17:21:45.890337Z","shell.execute_reply.started":"2024-01-11T17:21:45.878692Z","shell.execute_reply":"2024-01-11T17:21:45.889538Z"},"trusted":true},"execution_count":13,"outputs":[]},{"cell_type":"markdown","source":"# Dataset\n<div class=\"alert alert-block alert-info\" style=\"font-size:20px; background-color: #859cd6; font-family:verdana; color: #001342; border: 2px #001342 solid\">\n    <b>It's time to take care of our dataset</b>\n    <br>It should act as an interface for getting images and a dictionary with target data. What I'm talking about?<br>\n</div>","metadata":{}},{"cell_type":"markdown","source":"<div class=\"alert alert-block alert-info\" style=\"font-size:20px; background-color: #859cd6; font-family:verdana; color: #001342; border: 2px #001342 solid\">\n<ul class=\"simple\">\n<li><p>image: a PIL Image of size <code class=\"docutils literal notranslate\"><span class=\"pre\">(H,</span> <span class=\"pre\">W)</span></code></p></li>\n<li><p>target: a dict containing the following fields</p>\n<ul>\n<li><p><code class=\"docutils literal notranslate\"><span class=\"pre\">boxes</span> <span class=\"pre\">(FloatTensor[N,</span> <span class=\"pre\">4])</span></code>: the coordinates of the <code class=\"docutils literal notranslate\"><span class=\"pre\">N</span></code>\nbounding boxes in <code class=\"docutils literal notranslate\"><span class=\"pre\">[x0,</span> <span class=\"pre\">y0,</span> <span class=\"pre\">x1,</span> <span class=\"pre\">y1]</span></code> format, ranging from <code class=\"docutils literal notranslate\"><span class=\"pre\">0</span></code>\nto <code class=\"docutils literal notranslate\"><span class=\"pre\">W</span></code> and <code class=\"docutils literal notranslate\"><span class=\"pre\">0</span></code> to <code class=\"docutils literal notranslate\"><span class=\"pre\">H</span></code></p></li>\n<li><p><code class=\"docutils literal notranslate\"><span class=\"pre\">labels</span> <span class=\"pre\">(Int64Tensor[N])</span></code>: the label for each bounding box. <code class=\"docutils literal notranslate\"><span class=\"pre\">0</span></code> represents always the background class.</p></li>\n<li><p><code class=\"docutils literal notranslate\"><span class=\"pre\">image_id</span> <span class=\"pre\">(Int64Tensor[1])</span></code>: an image identifier. It should be\nunique between all the images in the dataset, and is used during\nevaluation</p></li>\n<li><p><code class=\"docutils literal notranslate\"><span class=\"pre\">area</span> <span class=\"pre\">(Tensor[N])</span></code>: The area of the bounding box. This is used\nduring evaluation with the COCO metric, to separate the metric\nscores between small, medium and large boxes.</p></li>\n<li><p><code class=\"docutils literal notranslate\"><span class=\"pre\">iscrowd</span> <span class=\"pre\">(UInt8Tensor[N])</span></code>: instances with iscrowd=True will be\nignored during evaluation.</p></li>\n<li><p>(optionally) <code class=\"docutils literal notranslate\"><span class=\"pre\">masks</span> <span class=\"pre\">(UInt8Tensor[N,</span> <span class=\"pre\">H,</span> <span class=\"pre\">W])</span></code>: The segmentation\nmasks for each one of the objects</p></li>\n<li><p>(optionally) <code class=\"docutils literal notranslate\"><span class=\"pre\">keypoints</span> <span class=\"pre\">(FloatTensor[N,</span> <span class=\"pre\">K,</span> <span class=\"pre\">3])</span></code>: For each one of\nthe N objects, it contains the K keypoints in\n<code class=\"docutils literal notranslate\"><span class=\"pre\">[x,</span> <span class=\"pre\">y,</span> <span class=\"pre\">visibility]</span></code> format, defining the object. visibility=0\nmeans that the keypoint is not visible. Note that for data\naugmentation, the notion of flipping a keypoint is dependent on\nthe data representation, and you should probably adapt\n<code class=\"docutils literal notranslate\"><span class=\"pre\">references/detection/transforms.py</span></code> for your new keypoint\nrepresentation</p></li>\n</ul>\n</li>\n</ul>\n</div>","metadata":{}},{"cell_type":"code","source":"from torch.utils.data import Dataset\nimport albumentations as A\nimport torchvision.transforms as T\nimport cv2\nimport os\nimport yaml\nfrom typing import Literal, Any, Union\nfrom tqdm import tqdm\nimport json","metadata":{"execution":{"iopub.status.busy":"2024-01-11T17:21:45.894757Z","iopub.execute_input":"2024-01-11T17:21:45.895101Z","iopub.status.idle":"2024-01-11T17:21:45.900434Z","shell.execute_reply.started":"2024-01-11T17:21:45.895077Z","shell.execute_reply":"2024-01-11T17:21:45.899553Z"},"trusted":true},"execution_count":14,"outputs":[]},{"cell_type":"code","source":"class HuBMAPDataset(Dataset):\n    def __init__(self,\n                 stage: Literal[\"train\", \"val\"],\n                 config_path: str,\n                 transforms: Union[A.Compose, T.Compose] = None,\n                 *args, **kwargs\n                 ):\n        \n        self.config = self.load_config(config_path)\n        self.images_dirpath = None\n        self.labels_dirpath = None\n        self.__define_paths(stage)\n        \n        self.names = self.config[\"names\"]\n        self.num_classes = len(self.names) \n        self.samples = os.listdir(self.labels_dirpath)\n        \n        if transforms:\n            self.bbox_params = {\n                \"format\":\"pascal_voc\",\n                \"min_area\": 0,\n                \"min_visibility\": 0,\n                \"label_fields\": [\"category_id\"]\n            }\n            self.transforms = A.Compose(transforms, bbox_params=self.bbox_params)\n        \n    @staticmethod\n    def load_config(path: str) -> dict:\n        with open(path, mode=\"r\") as f:\n            data = yaml.load(stream=f, Loader=yaml.SafeLoader)\n        return data\n\n    def __len__(self) -> int:\n        return len(self.samples)\n\n    def __getitem__(self, idx: int) -> tuple[torch.Tensor, dict]:\n        filename = self.samples[idx].split(\".\")[0]\n        paths = self._get_paths(filename)\n        image = cv2.imread(paths[\"image\"], cv2.COLOR_BGR2RGB)  \n        target = self._get_target(paths[\"label\"])\n        target[\"image_id\"] = torch.tensor([idx])\n        if self.transforms:\n            image, target = self.transform(image, target)\n        return image, target\n    \n    def transform(self, image: np.ndarray, target: dict) -> tuple[torch.Tensor, dict]:\n        transformed = self.transforms(\n            image=image, masks=target[\"masks\"],\n            bboxes=target[\"boxes\"], \n            category_id=target[\"labels\"]\n        )\n    \n        image = transformed[\"image\"]\n        target[\"masks\"] = torch.as_tensor(\n            np.array(list(map(np.array, transformed[\"masks\"])), dtype=np.uint8)\n        ) \n        \n        target[\"labels\"] = torch.tensor(transformed[\"category_id\"])\n        target[\"boxes\"] = torch.as_tensor(transformed[\"bboxes\"], dtype=torch.float32)\n        target[\"area\"] = self.__get_area(target[\"boxes\"])\n        return image, target\n        \n    def __define_paths(self, stage: Literal[\"train\", \"val\"]) -> None:\n        data_dirpath = self.config[stage]\n        self.images_dirpath = os.path.join(data_dirpath, \"images\")\n        self.labels_dirpath = os.path.join(data_dirpath, \"labels\")\n        \n    def _get_paths(self, filename: str) -> dict:\n        image_path = os.path.join(self.images_dirpath, f\"{filename}.tif\")\n        label_path = os.path.join(self.labels_dirpath, f\"{filename}.txt\")\n        return {\n            \"image\": image_path,\n            \"label\": label_path\n        }\n    \n    @staticmethod\n    def _get_target_sample() -> dict:\n        return {\n            \"boxes\": [],\n            \"masks\": [],\n            \"area\": [],\n            \"labels\": [],\n            \"iscrowd\": None,\n            \"image_id\": None\n        }\n\n    def _get_target(self, annotations_path: str) -> dict:\n        target = self._get_target_sample()\n\n        with open(annotations_path, \"r\") as file:\n            for line in file:\n                label = int(line[0])\n                coordinates = np.array(list(map(int, line[1:].split()))).reshape(1, -1, 2)\n                mask = self.__get_mask(label, coordinates)\n                box = self.__get_box(mask)\n                target[\"masks\"].append(mask)    \n                target[\"boxes\"].append(box)\n                target[\"labels\"].append(label)\n                \n        num_objs = len(target[\"labels\"])\n        target[\"iscrowd\"] = torch.zeros((num_objs,), dtype=torch.int64)\n        return target \n\n    @staticmethod\n    def __get_mask(label: int, coordinates: np.ndarray) -> np.ndarray:\n        mask = np.zeros((512, 512), dtype=np.uint8)\n        return cv2.fillPoly(\n            mask, pts=coordinates,\n            color=(label, label, label)\n        )\n    \n    @staticmethod\n    def __get_box(mask: np.ndarray) -> list[np.ndarray, ...]:\n        pos = np.nonzero(mask)\n        xmin = np.min(pos[1])\n        xmax = np.max(pos[1])\n        ymin = np.min(pos[0])\n        ymax = np.max(pos[0])\n        return [xmin, ymin, xmax, ymax]\n    \n    @staticmethod\n    def __get_area(boxes: list[list, ...]) -> torch.Tensor:\n         return (boxes[:, 3] - boxes[:, 1]) * (boxes[:, 2] - boxes[:, 0])","metadata":{"execution":{"iopub.status.busy":"2024-01-11T17:21:45.90135Z","iopub.execute_input":"2024-01-11T17:21:45.901617Z","iopub.status.idle":"2024-01-11T17:21:45.926553Z","shell.execute_reply.started":"2024-01-11T17:21:45.901586Z","shell.execute_reply":"2024-01-11T17:21:45.925736Z"},"trusted":true},"execution_count":15,"outputs":[]},{"cell_type":"markdown","source":"<div class=\"alert alert-block alert-info\" style=\"font-size:20px; background-color: #859cd6; font-family:verdana; color: #001342; border: 2px #001342 solid\">\n    <b>Dataset for training</b>\n</div>","metadata":{}},{"cell_type":"code","source":"train_dataset = HuBMAPDataset(\n    stage=\"train\",\n    config_path=coco.config_path,\n    transforms=train_transforms\n)","metadata":{"execution":{"iopub.status.busy":"2024-01-11T17:21:45.92767Z","iopub.execute_input":"2024-01-11T17:21:45.927933Z","iopub.status.idle":"2024-01-11T17:21:45.94207Z","shell.execute_reply.started":"2024-01-11T17:21:45.927905Z","shell.execute_reply":"2024-01-11T17:21:45.941253Z"},"trusted":true},"execution_count":16,"outputs":[]},{"cell_type":"markdown","source":"<div class=\"alert alert-block alert-info\" style=\"font-size:20px; background-color: #859cd6; font-family:verdana; color: #001342; border: 2px #001342 solid\">\n    <b>Dataset for validation</b>\n</div>","metadata":{}},{"cell_type":"code","source":"val_dataset = HuBMAPDataset(\n    stage=\"val\",\n    config_path=coco.config_path,\n    transforms=val_transforms\n)","metadata":{"execution":{"iopub.status.busy":"2024-01-11T17:21:45.94302Z","iopub.execute_input":"2024-01-11T17:21:45.943281Z","iopub.status.idle":"2024-01-11T17:21:45.951671Z","shell.execute_reply.started":"2024-01-11T17:21:45.943259Z","shell.execute_reply":"2024-01-11T17:21:45.950896Z"},"trusted":true},"execution_count":17,"outputs":[]},{"cell_type":"markdown","source":"# Dataloaders\n<div class=\"alert alert-block alert-info\" style=\"font-size:20px; background-color: #b7ebea; font-family:verdana; color: #233d3d; border: 2px #233d3d solid\">\n    <b>And how to manipulate the unloading of data?</b>\n    <br>Imagine that for training the neural network there are mastaba data, which should remain less in the RAM for a short while and stay in the video memory for a longer time. This is quite costly in terms of storage. But what to do? Make a lazy iterator. At least before, the output was like this, as soon as the loader issued the batch, it immediately appeared in memory, and disappeared from the iterator, and at the same time, the iterator itself contained only information about the object, and not the object itself. However, now we do not need to worry about such low-level things and we are engaged in the Dataloader sports class. You, in turn, now know why it is needed.<br>\n</div>","metadata":{}},{"cell_type":"code","source":"from torch.utils.data import DataLoader","metadata":{"execution":{"iopub.status.busy":"2024-01-11T17:21:45.952672Z","iopub.execute_input":"2024-01-11T17:21:45.95292Z","iopub.status.idle":"2024-01-11T17:21:45.961172Z","shell.execute_reply.started":"2024-01-11T17:21:45.952898Z","shell.execute_reply":"2024-01-11T17:21:45.960489Z"},"trusted":true},"execution_count":18,"outputs":[]},{"cell_type":"markdown","source":"<div class=\"alert alert-block alert-info\" style=\"font-size:20px; background-color: #b7ebea; font-family:verdana; color: #233d3d; border: 2px #233d3d solid\">\n    <b>Dataloader for training</b>\n</div>","metadata":{}},{"cell_type":"code","source":"train_dataloader = DataLoader(\n    dataset=train_dataset,\n    batch_size=16, \n    shuffle=True,\n    pin_memory=True,\n    num_workers=4,\n    collate_fn=lambda x: tuple(zip(*x))\n)","metadata":{"execution":{"iopub.status.busy":"2024-01-11T17:21:45.962253Z","iopub.execute_input":"2024-01-11T17:21:45.962843Z","iopub.status.idle":"2024-01-11T17:21:45.971251Z","shell.execute_reply.started":"2024-01-11T17:21:45.962812Z","shell.execute_reply":"2024-01-11T17:21:45.970507Z"},"trusted":true},"execution_count":19,"outputs":[]},{"cell_type":"markdown","source":"<div class=\"alert alert-block alert-info\" style=\"font-size:20px; background-color: #b7ebea; font-family:verdana; color: #233d3d; border: 2px #233d3d solid\">\n    <b>And for validation</b>\n</div>","metadata":{}},{"cell_type":"code","source":"val_dataloader = DataLoader(\n    dataset=val_dataset,\n    batch_size=16, \n    shuffle=False,\n    pin_memory=True,\n    num_workers=4,\n    collate_fn=lambda x: tuple(zip(*x))\n)","metadata":{"execution":{"iopub.status.busy":"2024-01-11T17:21:45.972292Z","iopub.execute_input":"2024-01-11T17:21:45.972682Z","iopub.status.idle":"2024-01-11T17:21:45.981583Z","shell.execute_reply.started":"2024-01-11T17:21:45.972652Z","shell.execute_reply":"2024-01-11T17:21:45.980772Z"},"trusted":true},"execution_count":20,"outputs":[]},{"cell_type":"markdown","source":"# Logging\n<div class=\"alert alert-block alert-info\" style=\"font-size:20px; background-color: #eec4ff; font-family:verdana; color: #511b66; border: 2px #511b66 solid\">\n    <b>And now a little about the control of experiments</b>\n    <br>Experiment control tools are a very important part of your pipeline. With the help of them, you can track learning and compare results with previous ones to solve your problem. It is very convenient and allows you not to get confused.<br>\n</div>","metadata":{}},{"cell_type":"markdown","source":"# Weights and biases\n  \n <p align='center'> \n <a href=\"https://pypi.python.org/pypi/wandb\"><img src=\"https://img.shields.io/pypi/v/wandb\" /></a> \n <a href=\"https://anaconda.org/conda-forge/wandb\"><img src=\"https://img.shields.io/conda/vn/conda-forge/wandb\" /></a> \n <a href=\"https://circleci.com/gh/wandb/wandb\"><img src=\"https://img.shields.io/circleci/build/github/wandb/wandb/main\" /></a> \n <a href=\"https://codecov.io/gh/wandb/wandb\"><img src=\"https://img.shields.io/codecov/c/gh/wandb/wandb\" /></a> \n </p> \n <p align='center'> \n <a href=\"https://colab.research.google.com/github/wandb/examples/blob/master/colabs/intro/Intro_to_Weights_%26_Biases.ipynb\"><img src=\"https://colab.research.google.com/assets/colab-badge.svg\" /></a> \n </p> \n  \n Use W&B to build better models faster. Track and visualize all the pieces of your machine learning pipeline, from datasets to production machine learning models. Get started with W&B today, [sign up for a free account!](https://wandb.com?utm_source=github&utm_medium=code&utm_campaign=wandb&utm_content=readme) \n  \n 🎓 W&B is free for students, educators, and academic researchers. For more information, visit [https://wandb.ai/site/research](https://wandb.ai/site/research?utm_source=github&utm_medium=code&utm_campaign=wandb&utm_content=readme). \n  \n Want to use Weights & Biases for seamless collaboration between your ML or Data Science team? Looking for Production-grade MLOps at scale? Sign up to one of [our plans](https://wandb.ai/site/pricing) or [contact the Sales Team](https://wandb.ai/site/contact).\n","metadata":{}},{"cell_type":"markdown","source":"<div class=\"alert alert-block alert-info\" style=\"font-size:20px; background-color: #eec4ff; font-family:verdana; color: #511b66; border: 2px #511b66 solid\">\n    <b>Please follow the <a href=https://wandb.ai/site>link</a>. Sign up and then paste your API key into the box that pops up below. Do not disclose your key to anyone. Here it is required so that you can track the training</b>\n</div>","metadata":{}},{"cell_type":"markdown","source":"![](https://i.ibb.co/q1xnxnD/2023-07-01-06-42-42.png)","metadata":{}},{"cell_type":"code","source":"import wandb\n\nwandb.login() \n\nrun = wandb.init(\n    # Set the project where this run will be logged\n    project=\"HuBMAP-Mask-RCNN\",\n    # Track hyperparameters and run metadata\n    config={\n        \"learning_rate\": 0.0035,\n        \"epochs\": 15,\n    })","metadata":{"_kg_hide-output":false,"execution":{"iopub.status.busy":"2024-01-11T17:21:45.982616Z","iopub.execute_input":"2024-01-11T17:21:45.982895Z","iopub.status.idle":"2024-01-11T17:25:23.730999Z","shell.execute_reply.started":"2024-01-11T17:21:45.982865Z","shell.execute_reply":"2024-01-11T17:25:23.729837Z"},"trusted":true},"execution_count":21,"outputs":[{"name":"stderr","text":"\u001b[34m\u001b[1mwandb\u001b[0m: Logging into wandb.ai. (Learn how to deploy a W&B server locally: https://wandb.me/wandb-server)\n\u001b[34m\u001b[1mwandb\u001b[0m: You can find your API key in your browser here: https://wandb.ai/authorize\n\u001b[34m\u001b[1mwandb\u001b[0m: Paste an API key from your profile and hit enter, or press ctrl+c to quit:","output_type":"stream"},{"output_type":"stream","name":"stdin","text":"  \n"},{"name":"stderr","text":"\u001b[34m\u001b[1mwandb\u001b[0m: Paste an API key from your profile and hit enter, or press ctrl+c to quit:","output_type":"stream"},{"output_type":"stream","name":"stdin","text":"  ···\n"},{"name":"stderr","text":"\u001b[34m\u001b[1mwandb\u001b[0m: \u001b[32m\u001b[41mERROR\u001b[0m API key must be 40 characters long, yours was 3\n\u001b[34m\u001b[1mwandb\u001b[0m: Logging into wandb.ai. (Learn how to deploy a W&B server locally: https://wandb.me/wandb-server)\n\u001b[34m\u001b[1mwandb\u001b[0m: You can find your API key in your browser here: https://wandb.ai/authorize\n\u001b[34m\u001b[1mwandb\u001b[0m: Paste an API key from your profile and hit enter, or press ctrl+c to quit:","output_type":"stream"},{"output_type":"stream","name":"stdin","text":"  ································\n"},{"name":"stderr","text":"\u001b[34m\u001b[1mwandb\u001b[0m: \u001b[32m\u001b[41mERROR\u001b[0m API key must be 40 characters long, yours was 32\n\u001b[34m\u001b[1mwandb\u001b[0m: Logging into wandb.ai. (Learn how to deploy a W&B server locally: https://wandb.me/wandb-server)\n\u001b[34m\u001b[1mwandb\u001b[0m: You can find your API key in your browser here: https://wandb.ai/authorize\n\u001b[34m\u001b[1mwandb\u001b[0m: Paste an API key from your profile and hit enter, or press ctrl+c to quit:","output_type":"stream"},{"output_type":"stream","name":"stdin","text":"  ····································································\n"},{"name":"stderr","text":"\u001b[34m\u001b[1mwandb\u001b[0m: \u001b[32m\u001b[41mERROR\u001b[0m API key must be 40 characters long, yours was 68\n\u001b[34m\u001b[1mwandb\u001b[0m: Logging into wandb.ai. (Learn how to deploy a W&B server locally: https://wandb.me/wandb-server)\n\u001b[34m\u001b[1mwandb\u001b[0m: You can find your API key in your browser here: https://wandb.ai/authorize\n\u001b[34m\u001b[1mwandb\u001b[0m: Paste an API key from your profile and hit enter, or press ctrl+c to quit:","output_type":"stream"},{"output_type":"stream","name":"stdin","text":"  ············\n"},{"name":"stderr","text":"\u001b[34m\u001b[1mwandb\u001b[0m: \u001b[32m\u001b[41mERROR\u001b[0m API key must be 40 characters long, yours was 12\n\u001b[34m\u001b[1mwandb\u001b[0m: Logging into wandb.ai. (Learn how to deploy a W&B server locally: https://wandb.me/wandb-server)\n\u001b[34m\u001b[1mwandb\u001b[0m: You can find your API key in your browser here: https://wandb.ai/authorize\n\u001b[34m\u001b[1mwandb\u001b[0m: Paste an API key from your profile and hit enter, or press ctrl+c to quit:","output_type":"stream"},{"output_type":"stream","name":"stdin","text":"  ································\n"},{"name":"stderr","text":"\u001b[34m\u001b[1mwandb\u001b[0m: \u001b[32m\u001b[41mERROR\u001b[0m API key must be 40 characters long, yours was 32\n\u001b[34m\u001b[1mwandb\u001b[0m: Logging into wandb.ai. (Learn how to deploy a W&B server locally: https://wandb.me/wandb-server)\n\u001b[34m\u001b[1mwandb\u001b[0m: You can find your API key in your browser here: https://wandb.ai/authorize\n\u001b[34m\u001b[1mwandb\u001b[0m: Paste an API key from your profile and hit enter, or press ctrl+c to quit:","output_type":"stream"},{"output_type":"stream","name":"stdin","text":"  ········································\n"},{"name":"stderr","text":"\u001b[34m\u001b[1mwandb\u001b[0m: Appending key for api.wandb.ai to your netrc file: /root/.netrc\n\u001b[34m\u001b[1mwandb\u001b[0m: W&B API key is configured. Use \u001b[1m`wandb login --relogin`\u001b[0m to force relogin\nwandb: ERROR Error while calling W&B API: user is not logged in (<Response [401]>)\nTraceback (most recent call last):\n  File \"/opt/conda/lib/python3.10/site-packages/wandb/sdk/wandb_init.py\", line 1152, in init\n    run = wi.init()\n  File \"/opt/conda/lib/python3.10/site-packages/wandb/sdk/wandb_init.py\", line 768, in init\n    raise error\nwandb.errors.AuthenticationError: The API key you provided is either invalid or missing.  If the `WANDB_API_KEY` environment variable is set, make sure it is correct. Otherwise, to resolve this issue, you may try running the 'wandb login --relogin' command. If you are using a local server, make sure that you're using the correct hostname. If you're not sure, you can try logging in again using the 'wandb login --relogin --host [hostname]' command.(Error 401: Unauthorized)\n\nDuring handling of the above exception, another exception occurred:\n\nTraceback (most recent call last):\n  File \"/opt/conda/lib/python3.10/site-packages/wandb/sdk/wandb_init.py\", line 1160, in init\n    getcaller()\n  File \"/opt/conda/lib/python3.10/site-packages/wandb/sdk/wandb_init.py\", line 835, in getcaller\n    src, line, func, stack = logger.findCaller(stack_info=True)\n  File \"/root/.local/lib/python3.10/site-packages/log.py\", line 42, in findCaller\n    sio = io.StringIO()\nNameError: name 'io' is not defined\n","output_type":"stream"},{"traceback":["\u001b[0;31m---------------------------------------------------------------------------\u001b[0m","\u001b[0;31mAuthenticationError\u001b[0m                       Traceback (most recent call last)","File \u001b[0;32m/opt/conda/lib/python3.10/site-packages/wandb/sdk/wandb_init.py:1152\u001b[0m, in \u001b[0;36minit\u001b[0;34m(job_type, dir, config, project, entity, reinit, tags, group, name, notes, magic, config_exclude_keys, config_include_keys, anonymous, mode, allow_val_change, resume, force, tensorboard, sync_tensorboard, monitor_gym, save_code, id, settings)\u001b[0m\n\u001b[1;32m   1151\u001b[0m \u001b[38;5;28;01mtry\u001b[39;00m:\n\u001b[0;32m-> 1152\u001b[0m     run \u001b[38;5;241m=\u001b[39m \u001b[43mwi\u001b[49m\u001b[38;5;241;43m.\u001b[39;49m\u001b[43minit\u001b[49m\u001b[43m(\u001b[49m\u001b[43m)\u001b[49m\n\u001b[1;32m   1153\u001b[0m     except_exit \u001b[38;5;241m=\u001b[39m wi\u001b[38;5;241m.\u001b[39msettings\u001b[38;5;241m.\u001b[39m_except_exit\n","File \u001b[0;32m/opt/conda/lib/python3.10/site-packages/wandb/sdk/wandb_init.py:768\u001b[0m, in \u001b[0;36m_WandbInit.init\u001b[0;34m(self)\u001b[0m\n\u001b[1;32m    767\u001b[0m         \u001b[38;5;28mself\u001b[39m\u001b[38;5;241m.\u001b[39mteardown()\n\u001b[0;32m--> 768\u001b[0m     \u001b[38;5;28;01mraise\u001b[39;00m error\n\u001b[1;32m    770\u001b[0m \u001b[38;5;28;01massert\u001b[39;00m run_result \u001b[38;5;129;01mis\u001b[39;00m \u001b[38;5;129;01mnot\u001b[39;00m \u001b[38;5;28;01mNone\u001b[39;00m  \u001b[38;5;66;03m# for mypy\u001b[39;00m\n","\u001b[0;31mAuthenticationError\u001b[0m: The API key you provided is either invalid or missing.  If the `WANDB_API_KEY` environment variable is set, make sure it is correct. Otherwise, to resolve this issue, you may try running the 'wandb login --relogin' command. If you are using a local server, make sure that you're using the correct hostname. If you're not sure, you can try logging in again using the 'wandb login --relogin --host [hostname]' command.(Error 401: Unauthorized)","\nDuring handling of the above exception, another exception occurred:\n","\u001b[0;31mNameError\u001b[0m                                 Traceback (most recent call last)","File \u001b[0;32m/opt/conda/lib/python3.10/site-packages/wandb/sdk/wandb_init.py:1160\u001b[0m, in \u001b[0;36minit\u001b[0;34m(job_type, dir, config, project, entity, reinit, tags, group, name, notes, magic, config_exclude_keys, config_include_keys, anonymous, mode, allow_val_change, resume, force, tensorboard, sync_tensorboard, monitor_gym, save_code, id, settings)\u001b[0m\n\u001b[1;32m   1157\u001b[0m \u001b[38;5;28;01mif\u001b[39;00m \u001b[38;5;129;01mnot\u001b[39;00m (\n\u001b[1;32m   1158\u001b[0m     wandb\u001b[38;5;241m.\u001b[39mwandb_agent\u001b[38;5;241m.\u001b[39m_is_running() \u001b[38;5;129;01mand\u001b[39;00m \u001b[38;5;28misinstance\u001b[39m(e, \u001b[38;5;167;01mKeyboardInterrupt\u001b[39;00m)\n\u001b[1;32m   1159\u001b[0m ):\n\u001b[0;32m-> 1160\u001b[0m     \u001b[43mgetcaller\u001b[49m\u001b[43m(\u001b[49m\u001b[43m)\u001b[49m\n\u001b[1;32m   1161\u001b[0m \u001b[38;5;28;01massert\u001b[39;00m logger\n","File \u001b[0;32m/opt/conda/lib/python3.10/site-packages/wandb/sdk/wandb_init.py:835\u001b[0m, in \u001b[0;36mgetcaller\u001b[0;34m()\u001b[0m\n\u001b[1;32m    834\u001b[0m     \u001b[38;5;28;01mreturn\u001b[39;00m \u001b[38;5;28;01mNone\u001b[39;00m\n\u001b[0;32m--> 835\u001b[0m src, line, func, stack \u001b[38;5;241m=\u001b[39m \u001b[43mlogger\u001b[49m\u001b[38;5;241;43m.\u001b[39;49m\u001b[43mfindCaller\u001b[49m\u001b[43m(\u001b[49m\u001b[43mstack_info\u001b[49m\u001b[38;5;241;43m=\u001b[39;49m\u001b[38;5;28;43;01mTrue\u001b[39;49;00m\u001b[43m)\u001b[49m\n\u001b[1;32m    836\u001b[0m \u001b[38;5;28mprint\u001b[39m(\u001b[38;5;124m\"\u001b[39m\u001b[38;5;124mProblem at:\u001b[39m\u001b[38;5;124m\"\u001b[39m, src, line, func)\n","File \u001b[0;32m~/.local/lib/python3.10/site-packages/log.py:42\u001b[0m, in \u001b[0;36m_Logger.findCaller\u001b[0;34m(self, stack_info, stacklevel)\u001b[0m\n\u001b[1;32m     41\u001b[0m \u001b[38;5;28;01mif\u001b[39;00m stack_info:\n\u001b[0;32m---> 42\u001b[0m     sio \u001b[38;5;241m=\u001b[39m \u001b[43mio\u001b[49m\u001b[38;5;241m.\u001b[39mStringIO()\n\u001b[1;32m     43\u001b[0m     sio\u001b[38;5;241m.\u001b[39mwrite(\u001b[38;5;124m'\u001b[39m\u001b[38;5;124mStack (most recent call last):\u001b[39m\u001b[38;5;130;01m\\n\u001b[39;00m\u001b[38;5;124m'\u001b[39m)\n","\u001b[0;31mNameError\u001b[0m: name 'io' is not defined","\nThe above exception was the direct cause of the following exception:\n","\u001b[0;31mError\u001b[0m                                     Traceback (most recent call last)","Cell \u001b[0;32mIn[21], line 5\u001b[0m\n\u001b[1;32m      1\u001b[0m \u001b[38;5;28;01mimport\u001b[39;00m \u001b[38;5;21;01mwandb\u001b[39;00m\n\u001b[1;32m      3\u001b[0m wandb\u001b[38;5;241m.\u001b[39mlogin() \n\u001b[0;32m----> 5\u001b[0m run \u001b[38;5;241m=\u001b[39m \u001b[43mwandb\u001b[49m\u001b[38;5;241;43m.\u001b[39;49m\u001b[43minit\u001b[49m\u001b[43m(\u001b[49m\n\u001b[1;32m      6\u001b[0m \u001b[43m    \u001b[49m\u001b[38;5;66;43;03m# Set the project where this run will be logged\u001b[39;49;00m\n\u001b[1;32m      7\u001b[0m \u001b[43m    \u001b[49m\u001b[43mproject\u001b[49m\u001b[38;5;241;43m=\u001b[39;49m\u001b[38;5;124;43m\"\u001b[39;49m\u001b[38;5;124;43mHuBMAP-Mask-RCNN\u001b[39;49m\u001b[38;5;124;43m\"\u001b[39;49m\u001b[43m,\u001b[49m\n\u001b[1;32m      8\u001b[0m \u001b[43m    \u001b[49m\u001b[38;5;66;43;03m# Track hyperparameters and run metadata\u001b[39;49;00m\n\u001b[1;32m      9\u001b[0m \u001b[43m    \u001b[49m\u001b[43mconfig\u001b[49m\u001b[38;5;241;43m=\u001b[39;49m\u001b[43m{\u001b[49m\n\u001b[1;32m     10\u001b[0m \u001b[43m        \u001b[49m\u001b[38;5;124;43m\"\u001b[39;49m\u001b[38;5;124;43mlearning_rate\u001b[39;49m\u001b[38;5;124;43m\"\u001b[39;49m\u001b[43m:\u001b[49m\u001b[43m \u001b[49m\u001b[38;5;241;43m0.0035\u001b[39;49m\u001b[43m,\u001b[49m\n\u001b[1;32m     11\u001b[0m \u001b[43m        \u001b[49m\u001b[38;5;124;43m\"\u001b[39;49m\u001b[38;5;124;43mepochs\u001b[39;49m\u001b[38;5;124;43m\"\u001b[39;49m\u001b[43m:\u001b[49m\u001b[43m \u001b[49m\u001b[38;5;241;43m15\u001b[39;49m\u001b[43m,\u001b[49m\n\u001b[1;32m     12\u001b[0m \u001b[43m    \u001b[49m\u001b[43m}\u001b[49m\u001b[43m)\u001b[49m\n","File \u001b[0;32m/opt/conda/lib/python3.10/site-packages/wandb/sdk/wandb_init.py:1190\u001b[0m, in \u001b[0;36minit\u001b[0;34m(job_type, dir, config, project, entity, reinit, tags, group, name, notes, magic, config_exclude_keys, config_include_keys, anonymous, mode, allow_val_change, resume, force, tensorboard, sync_tensorboard, monitor_gym, save_code, id, settings)\u001b[0m\n\u001b[1;32m   1188\u001b[0m             wandb\u001b[38;5;241m.\u001b[39mtermerror(\u001b[38;5;124m\"\u001b[39m\u001b[38;5;124mAbnormal program exit\u001b[39m\u001b[38;5;124m\"\u001b[39m)\n\u001b[1;32m   1189\u001b[0m             os\u001b[38;5;241m.\u001b[39m_exit(\u001b[38;5;241m1\u001b[39m)\n\u001b[0;32m-> 1190\u001b[0m         \u001b[38;5;28;01mraise\u001b[39;00m Error(\u001b[38;5;124m\"\u001b[39m\u001b[38;5;124mAn unexpected error occurred\u001b[39m\u001b[38;5;124m\"\u001b[39m) \u001b[38;5;28;01mfrom\u001b[39;00m \u001b[38;5;21;01merror_seen\u001b[39;00m\n\u001b[1;32m   1191\u001b[0m \u001b[38;5;28;01mreturn\u001b[39;00m run\n","\u001b[0;31mError\u001b[0m: An unexpected error occurred"],"ename":"Error","evalue":"An unexpected error occurred","output_type":"error"}]},{"cell_type":"markdown","source":"# Training\n<div class=\"alert alert-block alert-info\" style=\"font-size:20px; background-color: #bcf5a9; font-family:verdana; color: #1c3d11; border: 2px #1c3d11 solid\">\n    <b>Finally, we can start learning</b>\n    <br>Yes, yes I know. I was just talking about high-level and how abstraction can be annoying, but look around. We use pure pytorch without wrappers. However, all work with losses is registered under the hood of the model class. The model itself is just as seriously provided to us and everyone can assemble it as a lego constructor. In general, you do not need to be afraid of abstraction, there is no escape from it, but I am against high-level if it allows you to remain in the dark and continue using the tool.<br>\n</div>","metadata":{}},{"cell_type":"code","source":"from tqdm.notebook import tqdm\nimport torch.nn as nn\nfrom torch import optim","metadata":{"execution":{"iopub.status.busy":"2024-01-11T17:25:23.731899Z","iopub.status.idle":"2024-01-11T17:25:23.732232Z","shell.execute_reply.started":"2024-01-11T17:25:23.732068Z","shell.execute_reply":"2024-01-11T17:25:23.732084Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"class Trainer:\n    def __init__(self,\n                 model: nn.Module,\n                train_dataloader: DataLoader,\n                 val_dataloader: DataLoader,\n                 early_stop: dict = {\"monitor\": \"loss_mask\", \"patience\": 5},\n                 save_every_epoch: int = 1,\n                 save_dirpath: str = \"/kaggle/working/runs\"\n                ):\n        \n        # Callbacks | Early stoping & Model checkpoint\n        self.patience = early_stop[\"patience\"]\n        self.monitor = early_stop[\"monitor\"]\n        self.track_list = []\n        self.save_every_epoch = save_every_epoch\n        self.save_dirpath = save_dirpath\n        \n        self.device = \"cuda\" if torch.cuda.is_available() else \"cpu\"\n        self.train_dataloader = train_dataloader\n        self.val_dataloader = val_dataloader\n        self.train_batches = len(train_dataloader)\n        self.val_batches =  len(val_dataloader)\n        \n        self.model = model\n        self.setup_model()\n        \n        self.optim_dict = self.configure_optimizers()\n        self.optimizer = self.optim_dict[\"optimizer\"]\n        self.lr_scheduler = self.optim_dict[\"lr_scheduler\"]\n        \n        self.step_outputs = {\n            \"loss_objectness\": 0,\n            \"loss_mask\": 0,\n            \"loss_classifier\": 0,\n            \"loss_rpn_box_reg\": 0,\n            \"loss_box_reg\": 0,\n            \"loss\": 0\n        }\n        \n    def configure_optimizers(self) -> dict:\n        # construct an optimizer\n        params = [\n            p\n            for p in self.model.parameters()\n            if p.requires_grad\n        ]\n\n        optimizer = optim.SGD(\n            params,\n            lr=0.0018,\n            momentum=0.938,\n            weight_decay=0.00053\n        )\n\n        # and a learning rate scheduler\n        lr_scheduler = torch.optim.lr_scheduler.StepLR(\n            optimizer,\n            step_size=3,\n            gamma=0.1\n        )\n        return {\n            \"lr_scheduler\": lr_scheduler,\n            \"optimizer\": optimizer\n        }\n    \n    def setup_model(self) -> None:\n        for param in self.model.parameters():\n            param.requires_grad = True\n        self.model.to(self.device)\n        self.model.train()\n    \n    def to_device(self, batch: tuple) -> tuple:\n        images, targets = batch\n        images = list(image.to(self.device) for image in images)\n\n        targets = [\n            {key: value.to(self.device) \n             for key, value in target.items()}\n            for target in targets\n        ]\n        \n        return images, targets\n    \n    def training_step(self, batch) -> dict:\n        images, targets = self.to_device(batch)\n        self.optimizer.zero_grad() \n        outputs = self.model(images, targets)\n        loss = sum([loss for loss in outputs.values()])\n        outputs[\"loss\"] = loss\n        loss.backward()\n        self.optimizer.step()\n        self.lr_scheduler.step()\n        return outputs\n    \n    def validation_step(self, batch) -> dict:\n        images, targets = self.to_device(batch)\n        with torch.no_grad():\n            outputs = self.model(images, targets)\n            loss = sum([loss for loss in outputs.values()])\n            outputs[\"loss\"] = loss\n        return outputs\n    \n    def shared_epoch_end(self, stage: str, epoch: int) -> float:\n        tracked_loss = self.step_outputs[self.monitor]    \n        loss_objectness = self.step_outputs[\"loss_objectness\"]\n        loss_mask = self.step_outputs[\"loss_mask\"]\n        loss_classifier = self.step_outputs[\"loss_classifier\"]\n        loss_rpn_box_reg = self.step_outputs[\"loss_rpn_box_reg\"]\n        loss_box_reg = self.step_outputs[\"loss_box_reg\"]\n        loss = self.step_outputs[\"loss\"]\n        \n        wandb.log({\n            f\"{stage}_loss_objectness\": loss_objectness,\n            f\"{stage}_loss_mask\": loss_mask,\n            f\"{stage}_loss_classifier\": loss_classifier,\n            f\"{stage}_loss_rpn_box_reg\": loss_rpn_box_reg,\n            f\"{stage}_loss_box_reg\": loss_box_reg,\n            f\"{stage}_loss\": loss    \n        })\n        \n        print(\n            f\"\"\"\n            || End {epoch} {stage} epoch ||\n            loss_objectness: {loss_objectness:.2f}\n            loss_mask: {loss_mask:.2f}\n            loss_classifier: {loss_classifier:.2f}\n            loss_rpn_box_reg: {loss_rpn_box_reg:.2f}\n            loss_box_reg: {loss_box_reg:.2f} \n            loss: {loss:.2f}\\n\n            \"\"\"\n        )\n              \n        self.step_outputs = self.step_outputs.fromkeys(self.step_outputs, 0)\n        if stage == \"val\":\n            return tracked_loss\n              \n    def on_train_epoch_end(self, epoch: int) -> None:\n        return self.shared_epoch_end(stage=\"train\", epoch=epoch)\n\n    def on_validation_epoch_end(self, epoch: int) -> None:\n        tracked_loss = self.shared_epoch_end(stage=\"val\", epoch=epoch)\n        patience = 0\n        \n        if epoch > self.patience:\n            last_tracked = list(reversed(self.track_list))[:self.patience]\n            for i in last_tracked:\n                if i <= tracked_loss:\n                    patience += 1\n                    \n        self.track_list.append(tracked_loss)\n        return tracked_loss, (self.patience - patience)\n\n    def train(self, max_epochs: int) -> None:\n        \n        for epoch in range(1, max_epochs + 1):\n            for batch_idx, batch in tqdm(enumerate(self.train_dataloader, 1), desc=\"Training\", total=self.train_batches, colour=\"#068e58\"):\n                outputs = self.training_step(batch)\n                for key, value in outputs.items():\n                    self.step_outputs[key] += float(value.detach().cpu().numpy()) / self.train_batches\n            self.on_train_epoch_end(epoch)\n\n            for batch_idx, batch in tqdm(enumerate(val_dataloader, 1), desc=\"Validation\", total=self.val_batches, colour=\"#013385\"):\n                outputs = self.validation_step(batch)\n                for key, value in outputs.items():\n                    self.step_outputs[key] += float(value.detach().cpu().numpy()) / self.val_batches\n            tracked_loss, patience = self.on_validation_epoch_end(epoch)\n            \n            if epoch % self.save_every_epoch == 0:\n                if not os.path.exists(self.save_dirpath):\n                    os.mkdir(self.save_dirpath)\n                path = os.path.join(self.save_dirpath, f\"epoch_{epoch}_{self.monitor}_{tracked_loss:.2f}.pt\")\n                torch.save(model.state_dict(), path) \n                print(\"\\nThe model passed the save checkpoint successfully!\\n\")\n                \n            if patience == 0:\n                print(\"Our patience has run out! Model training stopped beforehand.\")\n                break\n","metadata":{"execution":{"iopub.status.busy":"2024-01-11T17:25:23.733636Z","iopub.status.idle":"2024-01-11T17:25:23.733946Z","shell.execute_reply.started":"2024-01-11T17:25:23.73379Z","shell.execute_reply":"2024-01-11T17:25:23.733805Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"trainer = Trainer(\n    model=model,\n    train_dataloader=train_dataloader,\n    val_dataloader=val_dataloader,\n    early_stop = {\"monitor\": \"loss_mask\", \"patience\": 5},\n    save_every_epoch=1\n)","metadata":{"execution":{"iopub.status.busy":"2024-01-11T17:25:23.734885Z","iopub.status.idle":"2024-01-11T17:25:23.735195Z","shell.execute_reply.started":"2024-01-11T17:25:23.735041Z","shell.execute_reply":"2024-01-11T17:25:23.735056Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"trainer.train(max_epochs=15)","metadata":{"_kg_hide-output":true,"execution":{"iopub.status.busy":"2024-01-11T17:25:23.73681Z","iopub.status.idle":"2024-01-11T17:25:23.737176Z","shell.execute_reply.started":"2024-01-11T17:25:23.736994Z","shell.execute_reply":"2024-01-11T17:25:23.73701Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<div class=\"alert alert-block alert-info\" style=\"font-size:20px; background-color: #bcf5a9; font-family:verdana; color: #1c3d11; border: 2px #1c3d11 solid\">\n    <b>Wrapper class for Mask-RCNN model</b>\n    <br>So it will be more convenient to look at the fruits of our labors, as well as make a submit<br>\n</div>","metadata":{}},{"cell_type":"code","source":"class BestMaskRCNN:\n    def __init__(self, \n                 model: nn.Module,\n                 checkpoint_path: str,\n                 conf: float = 0.05\n                ):\n        self.device = \"cuda\" if torch.cuda.is_available() else \"cpu\"\n        self.checkpoint_path = checkpoint_path\n        self.model = self.get_model(model)\n        self.conf = conf\n        self.transforms = A.Compose([\n            A.Resize(512, 512, p=1), \n            A.Normalize(),\n            ToTensorV2()\n        ])\n        \n    def to_cpu(self, outputs: dict) -> tuple:\n        outputs = {\n            key: value.cpu().numpy()\n            for key, value in outputs.items()\n        }       \n        return outputs\n    \n    def get_model(self, model: nn.Module) -> nn.Module:\n        checkpoint = torch.load(self.checkpoint_path) \n        model.load_state_dict(checkpoint) \n        model.eval()\n        model.to(self.device)\n        for param in model.parameters():\n            param.requires_grad = False\n        return model\n    \n    def get_image(self, path: str):\n        image = cv2.imread(path, cv2.COLOR_BGR2RGB)  \n        return self.transforms(image=image)[\"image\"]\n    \n    def forward(self, path: str):\n        image = self.get_image(path)\n        \n        image = torch.as_tensor(\n            np.expand_dims(image.numpy(), 0),\n            dtype=image.dtype\n        ).to(self.device)\n        \n        outputs = self.model(image)[0]\n        outputs = self.to_cpu(outputs)\n        return outputs\n    \n    def __call__(self, source: str) -> list[dict, ...]:\n        sublist = []\n        result = self.forward(source)\n\n        for i in range(len(result[\"masks\"])):\n            conf = round(float(result[\"scores\"][i]), 2)\n            mask = result[\"masks\"][i].transpose(1,2,0)\n\n            if int(result[\"labels\"][i]) == 1 and conf >= self.conf:\n                sublist.append({\"mask\": mask, \"confidence\": conf})\n            else:\n                continue\n        return sublist","metadata":{"execution":{"iopub.status.busy":"2024-01-11T17:25:23.738736Z","iopub.status.idle":"2024-01-11T17:25:23.739052Z","shell.execute_reply.started":"2024-01-11T17:25:23.738891Z","shell.execute_reply":"2024-01-11T17:25:23.738905Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"model = BestMaskRCNN(\n    model=get_model(4),\n    checkpoint_path=\"/kaggle/working/runs/epoch_2_loss_mask_0.53.pt\"\n)","metadata":{"execution":{"iopub.status.busy":"2024-01-11T17:25:23.740046Z","iopub.status.idle":"2024-01-11T17:25:23.740375Z","shell.execute_reply.started":"2024-01-11T17:25:23.740212Z","shell.execute_reply":"2024-01-11T17:25:23.740228Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<div class=\"alert alert-block alert-info\" style=\"font-size:20px; background-color: #f2ffb0; font-family:verdana; color: #3c450e; border: 2px #3c450e solid\">\n    <b>And how can I use the right tools if the internet is banned in the competition? 🙁</b>\n    <br>In the competition, many have difficulty installing the necessary tools and using them, because the Internet is prohibited, but it is prohibited only so that the test data is not stolen (they are uploaded to the test folder during the submission of the result). This means that you can use the tools you want, but you must add them as data. You can find my ultralytics and pycocotools dataset here. Importantly, the Internet is still present in this notebook to load the model, however, after training it, you will save the model, add it to the data and easily use it for forecasting without the Internet.<br>\n</div>","metadata":{}},{"cell_type":"code","source":"import shutil\nimport os\nimport sys\nfrom colorama import Fore","metadata":{"execution":{"iopub.status.busy":"2024-01-11T17:25:23.741897Z","iopub.status.idle":"2024-01-11T17:25:23.742209Z","shell.execute_reply.started":"2024-01-11T17:25:23.742057Z","shell.execute_reply":"2024-01-11T17:25:23.742071Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"class SetupPipline:\n    def __init__(self, display: bool = True):\n        self.pycocotools = self.__pycocotools()\n        \n    @staticmethod\n    def __pycocotools() -> str:\n        if not os.path.exists(\"/kaggle/working/packages\"):\n            shutil.copytree(\"/kaggle/input/hubmap-tools-ultralytics-and-pycocotools/pycocotools/pycocotools\", \"/kaggle/working/packages\")\n            os.chdir(\"/kaggle/working/packages/pycocotools-2.0.6/\")\n            os.system(\"python setup.py install\")\n            os.system(\"pip install . --no-index --find-links /kaggle/working/packages/\")\n            os.chdir(\"/kaggle/working\")\n            return \"successfully\"\n    \n    def display(self) -> None:\n        print(Fore.GREEN+f\"\\nPycocotools was installed {self.pycocotools}\")","metadata":{"execution":{"iopub.status.busy":"2024-01-11T17:25:23.743123Z","iopub.status.idle":"2024-01-11T17:25:23.743443Z","shell.execute_reply.started":"2024-01-11T17:25:23.743284Z","shell.execute_reply":"2024-01-11T17:25:23.743299Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"pipline = SetupPipline()","metadata":{"_kg_hide-output":true,"execution":{"iopub.status.busy":"2024-01-11T17:25:23.744445Z","iopub.status.idle":"2024-01-11T17:25:23.744801Z","shell.execute_reply.started":"2024-01-11T17:25:23.744617Z","shell.execute_reply":"2024-01-11T17:25:23.744633Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"pipline.display()","metadata":{"execution":{"iopub.status.busy":"2024-01-11T17:25:23.746043Z","iopub.status.idle":"2024-01-11T17:25:23.746344Z","shell.execute_reply.started":"2024-01-11T17:25:23.746194Z","shell.execute_reply":"2024-01-11T17:25:23.746208Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<div class=\"alert alert-block alert-info\" style=\"font-size:20px; background-color: #ffc369; font-family:verdana; color: #663500; border: 2px #663500 solid\">\n    <b>What about submission? 😸</b>\n    <br>We have come a long way, done data analysis, trained the model, but what about the submission? We want to send the results, right? Well, I had difficulties with this, and therefore, in order to make life easier for myself, it is possible to help someone, I wrote several classes that allow you to quickly submit without delving into the difficulties that I encountered.<br>\n</div>","metadata":{}},{"cell_type":"markdown","source":"<div class=\"alert alert-block alert-info\" style=\"font-size:20px; background-color: #ffc98f; font-family:verdana; color: #b86e1f; border: 2px #b86e1f solid\">\n    <b>Encode Binary Mask</b>\n</div>","metadata":{}},{"cell_type":"code","source":"import base64\nfrom pycocotools import _mask as coco_mask\nimport typing as t\nimport zlib\nimport pandas as pd\nimport torchvision.transforms as T\nfrom PIL import Image","metadata":{"execution":{"iopub.status.busy":"2024-01-11T17:25:23.747416Z","iopub.status.idle":"2024-01-11T17:25:23.747767Z","shell.execute_reply.started":"2024-01-11T17:25:23.747599Z","shell.execute_reply":"2024-01-11T17:25:23.747621Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"class EncodeBinaryMask:\n    @staticmethod\n    def __checking_mask(mask: np.ndarray) -> np.ndarray:\n        if mask.dtype != np.bool:\n            raise ValueError(\n                \"expects a binary mask, received dtype == %s\" %\n                mask.dtype\n            )\n        return mask\n\n    @staticmethod\n    def __convert_mask(mask: np.ndarray):\n        mask_to_encode = mask.astype(np.uint8)\n        mask_to_encode = np.asfortranarray(mask_to_encode)\n        return mask_to_encode\n\n    @staticmethod\n    def __compress_encode(encoded_mask) -> t.Text:\n        binary_str = zlib.compress(encoded_mask, zlib.Z_BEST_COMPRESSION)\n        base64_str = base64.b64encode(binary_str)\n        return base64_str\n\n    def __call__(self, mask: np.ndarray) -> t.Text:\n        mask = self.__checking_mask(mask)\n        mask_to_encode = self.__convert_mask(mask)\n        encoded_mask = coco_mask.encode(mask_to_encode)[0][\"counts\"]\n        base64_str = self.__compress_encode(encoded_mask)\n        return base64_str","metadata":{"execution":{"iopub.status.busy":"2024-01-11T17:25:23.749657Z","iopub.status.idle":"2024-01-11T17:25:23.749964Z","shell.execute_reply.started":"2024-01-11T17:25:23.749811Z","shell.execute_reply":"2024-01-11T17:25:23.749826Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<div class=\"alert alert-block alert-info\" style=\"font-size:20px; background-color: #ffc98f; font-family:verdana; color: #b86e1f; border: 2px #b86e1f solid\">\n    <b>Submission</b>\n</div>","metadata":{}},{"cell_type":"code","source":"class Submission:\n    def __init__(self, dirpath: str, model: torch.nn.Module):\n        self.__eval_transforms = self.get_transforms()\n        self.__model = model\n        self.__encoder = EncodeBinaryMask()\n        self.__dirpath = dirpath\n        self.__filenames = os.listdir(dirpath)\n        self.height = 512\n        self.width = 512\n        \n        self.__submission_dict = {\n            \"id\": [],\n            \"height\": [],\n            \"width\": [],\n            \"prediction_string\": []\n        }\n        \n        self.submission = None\n    \n    @staticmethod\n    def get_transforms():\n        return T.Compose([\n            T.ToTensor(),\n            T.Resize(size=(512, 512)),\n            T.Normalize(mean=[0.485, 0.456, 0.406],\n                        std=[0.229, 0.224, 0.225])\n        ])\n\n    def __len__(self):\n        return len(self.__filenames)\n\n    def __get_columns(self) -> None:\n        for filename in self.__filenames:\n            path = self.__get_image_path(filename)\n            masks = self.__forward(path)\n            identifier, height, width, prediction_string = self.__get_cells(filename, masks)\n            self.__update_columns(identifier, height, width, prediction_string)\n\n    def __update_columns(self, identifier: str, height: int, width: int, prediction_string: str) -> None:\n        self.__submission_dict[\"id\"].append(identifier)\n        self.__submission_dict[\"height\"].append(height)\n        self.__submission_dict[\"width\"].append(width)\n        self.__submission_dict[\"prediction_string\"].append(prediction_string)\n\n    def __get_cells(self, filename: str, masks: list):\n        prediction_string = \"\"\n        prediction_string = self.__get_prediction_string(masks, prediction_string)\n        identifier = filename.split(\".\")[0]\n        return identifier, self.height, self.width, prediction_string\n\n    def __get_prediction_string(self, masks: list, prediction_string: str) -> str:\n        if masks:\n            for outputs in masks:\n                mask = outputs[\"mask\"]\n                mask = np.where(mask > 0.5, 1, 0).astype(np.bool)\n                base64_str = self.__encoder(mask)\n                confidence = outputs[\"confidence\"]\n                prediction_string += f\"0 {confidence} {base64_str.decode('utf-8')} \"\n        else:\n            return \"\"\n        return prediction_string\n\n    def __get_image_path(self, filename: str) -> str:\n        return os.path.join(\n            self.__dirpath, filename\n        )\n\n    def __get_image(self, path: str) -> torch.Tensor:\n        image = Image.open(path)\n        image = np.asarray(image)\n        image = self.__eval_transforms(image)\n        return image\n\n    def __forward(self, image: torch.tensor) -> list:\n        masks = self.__model(image) \n        return masks \n\n    def submit(self) -> None:\n        if not self.submission:\n            self.__get_columns()\n            self.submission = pd.DataFrame(self.__submission_dict)\n            self.submission = self.submission.set_index('id')\n            self.submission.to_csv(\"submission.csv\")","metadata":{"execution":{"iopub.status.busy":"2024-01-11T17:25:23.751055Z","iopub.status.idle":"2024-01-11T17:25:23.751357Z","shell.execute_reply.started":"2024-01-11T17:25:23.751206Z","shell.execute_reply":"2024-01-11T17:25:23.75122Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"__TEST_PATH = \"/kaggle/input/hubmap-hacking-the-human-vasculature/test\"\nsub = Submission(dirpath=__TEST_PATH, model=model)\nsub.submit()","metadata":{"_kg_hide-output":true,"execution":{"iopub.status.busy":"2024-01-11T17:25:23.752217Z","iopub.status.idle":"2024-01-11T17:25:23.752535Z","shell.execute_reply.started":"2024-01-11T17:25:23.752377Z","shell.execute_reply":"2024-01-11T17:25:23.752393Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"sub.submission.head()","metadata":{"execution":{"iopub.status.busy":"2024-01-11T17:25:23.753904Z","iopub.status.idle":"2024-01-11T17:25:23.754334Z","shell.execute_reply.started":"2024-01-11T17:25:23.754114Z","shell.execute_reply":"2024-01-11T17:25:23.754135Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# TO DO\n<div class=\"alert alert-block alert-info\" style=\"font-size:20px; background-color: #ccbda7; font-family:verdana; color: #b86e1f; border: 2px #b86e1f solid\">\n    <b>Will this competition be over?</b>\n    <br>I'm really already quite tired of this competition, but I hope my work is already enough so that they can at least help someone. Why am I saying this? Just so you keep in mind that there are moments that I just didn’t have the strength for and I hope that you can handle them yourself. Namely: use non-max suppression, with lower scores, filter the masks to keep the most likely pixels. I will not return to this notebook, but I would like to know if you have any difficulties.<br>\n</div>","metadata":{}},{"cell_type":"markdown","source":"<div class=\"alert alert-block alert-info\" style=\"font-size:20px; background-color: #ccbda7; font-family:verdana; color: #b86e1f; border: 2px #b86e1f solid\">\n    <b>Thanks a lot for making it to the end 🙃</b>\n    <br>I hope you have found this work useful. I will be very grateful if you vote for this work, if it really helped you. Good luck <br>\n</div>","metadata":{}},{"cell_type":"markdown","source":"![](https://i.pinimg.com/originals/05/01/1c/05011ce8b4b326ed62c70f3eab93f913.gif)","metadata":{}}]}