{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"pygments_lexer":"ipython3","nbconvert_exporter":"python","version":"3.6.4","file_extension":".py","codemirror_mode":{"name":"ipython","version":3},"name":"python","mimetype":"text/x-python"}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"![](https://i.ibb.co/7t6wKN7/000.png)","metadata":{}},{"cell_type":"markdown","source":"<div class=\"alert alert-block alert-info\" style=\"font-size:20px; background-color: #ffa3a3; font-family:verdana; color: #8a0f0f; border: 2px #ff0303 solid\">\n    <b>HuBMAP - Hacking the Human Vasculature 🩸</b>\n    <br>Hello, Dear friends 🙃<br>\n    Research in the field of medicine has always been directed towards the field of machine learning, and today we are going to one of such studies. Grab a cup of tea and let's get started ☕️\n</div>","metadata":{}},{"cell_type":"markdown","source":"<div class=\"alert alert-block alert-info\" style=\"font-size:20px; background-color: #ffa578; font-family:verdana; color: #8a0f0f; border: 2px #ff0303 solid\">\n    <b>Skip this text if you're short on time. You won't miss anything</b>\n    <br>\n    I have a goal 🎯 and I am stubbornly moving towards it ✊, namely to get a master before the end of June. Only if this notebook really helps you, I will be glad if you can also see my other works 🧡<br>\n</div>","metadata":{}},{"cell_type":"markdown","source":"<div class=\"alert alert-block alert-info\" style=\"font-size:20px; background-color: #f0c7c7; font-family:verdana; color: #a63c3c; border: 2px #a63c3c solid\">\n    <b>Introduction 👁</b>\n    <br>The proper functioning of your body's organs and tissues depends on the interaction, spatial organization, and specialization of your cells—all 37 trillion of them. With so many cells, determining their functions and relationships is a monumental undertaking.<br>\n    <br>Current efforts to map cells involve the Vasculature Common Coordinate Framework (VCCF), which uses the blood vasculature in the human body as the primary navigation system. The VCCF crosses all scale levels--from the whole body to the single cell level--and provides a unique way to identify cellular locations using capillary structures as an address. However, the gaps in what researchers know about microvasculature lead to gaps in the VCCF. If we could automatically segment microvasculature arrangements, researchers could use the real-world tissue data to begin to fill in those gaps and map out the vasculature.<br>\n    <br>Competition host Human BioMolecular Atlas Program (HuBMAP) hopes to develop an open and global platform to map healthy cells in the human body. Using the latest molecular and cellular biology technologies, HuBMAP researchers are studying the connections that cells have with each other throughout the body.<br>\n    <br>There are still many unknowns regarding microvasculature, but your Machine Learning insights could enable researchers to use the available tissue data to augment their understanding of how these small vessels are arranged throughout the body. Ultimately, you'll be helping to pave the way towards building a Vascular Common Coordinate Framework (VCCF) and a Human Reference Atlas (HRA), which will identify how the relationships between cells can affect our health.<br>\n</div>","metadata":{}},{"cell_type":"markdown","source":"<div class=\"alert alert-block alert-info\" style=\"font-size:20px; background-color: #ffa3a3; font-family:verdana; color: #8a0f0f; border: 2px #ff0303 solid\">\n    <b>Exploratory Data Analysis  📚</b>\n</div>","metadata":{}},{"cell_type":"markdown","source":"<div class=\"alert alert-block alert-info\" style=\"font-size:20px; background-color: #f0c7c7; font-family:verdana; color: #a63c3c; border: 2px #a63c3c solid\"><h1>Data Description</h1>\n<p>Your goal in this competition is to locate microvasculature structures (blood vessels) within human kidney histology slides.</p>\n<p>The competition data comprises tiles extracted from five Whole Slide Images (WSI) split into two datasets. Tiles from Dataset 1 have annotations that have been expert reviewed. Dataset 2 comprises the remaining tiles from these same WSIs and contain sparse annotations that have not been expert reviewed.</p>\n<ul>\n<li>All of the test set tiles are from Dataset 1.</li>\n<li>Two of the WSIs make up the training set, two WSIs make up the public test set, and one WSI makes up the private test set.</li>\n<li>The training data includes Dataset 2 tiles from the <em>public</em> test WSI, but <em>not</em> from the <em>private</em> test WSI.</li>\n</ul>\n<p>We also include, as Dataset 3, tiles extracted from an additional nine WSIs. These tiles have not been annotated. You may wish to apply semi- or self-supervised learning techniques on this data to support your predictions.</p>\n<p>Note that this is a <a target=\"_blank\" href=\"https://www.kaggle.com/competitions/hubmap-hacking-the-human-vasculature/overview/code-requirements\"><strong>Code Competition</strong></a>, in which the actual test set is hidden. In this version, we give some sample data drawn from the public test set to help you author your solutions. When your submission is scored, this example test data will be replaced with the full test set. There are about 650 tiles in the full test set.</p>\n<p>You may find resources from the previous HuBMAP competitions useful as well:</p>\n<ul>\n<li><a target=\"_blank\" href=\"https://www.kaggle.com/competitions/hubmap-kidney-segmentation/\">HuBMAP: Hacking the Kidney</a></li>\n<li><a target=\"_blank\" href=\"https://www.kaggle.com/competitions/hubmap-organ-segmentation/\">HuBMAP + HPA: Hacking the Human Body</a></li>\n</ul>\n<h2>Files and Field Descriptions</h2>\n<ul>\n<li><strong>{train|test}/</strong> Folders containing TIFF images of the tiles. Each tile is 512x512 in size.</li>\n<li><strong>polygons.jsonl</strong> Polygonal segmentation masks in JSONL format, available for Dataset 1 and Dataset 2. Each line gives JSON annotations for a single image with:<ul>\n<li><code>id</code> Identifies the corresponding image in <strong>train/</strong></li>\n<li><code>annotations</code>  A list of mask annotations with:</li>\n<li><code>type</code> Identifies the type of structure annotated:<ul>\n<li><code>blood_vessel</code> The target structure. Your goal in this competition is to predict these kinds of masks on the test set.</li>\n<li><code>glomerulus</code> A capillary ball structure in the kidney. These parts of the images were excluded from blood vessel annotation. You should ensure none of your test set predictions occur within glomerulus structures as they will be counted as false positives. Annotations are provided for test set tiles.</li>\n<li><code>unsure</code> A structure the expert annotators cannot confidently distinguish as a blood vessel.</li></ul></li>\n<li><code>coordinates</code> A list of polygon coordinates defining the segmentation mask.</li></ul></li>\n<li><strong>tile_meta.csv</strong> Metadata for each image.<ul>\n<li><code>source_wsi</code> Identifies the WSI this tile was extracted from.</li>\n<li><code>{i|j}</code> The location of the upper-left corner within the WSI where the tile was extracted.</li>\n<li><code>dataset</code> The dataset this tile belongs to, as described above.</li></ul></li>\n<li><strong>wsi_meta.csv</strong> Metadata for the Whole Slide Images the tiles were extracted from.<ul>\n<li><code>source_wsi</code> Identifies the WSI.</li>\n<li><code>age</code>, <code>sex</code>, <code>race</code>, <code>height</code>, <code>weight</code>, and <code>bmi</code> demographic information about the tissue donor.</li></ul></li>\n<li><strong>sample_submission.csv</strong> A sample submission file in the correct format. See the <a target=\"_blank\" href=\"https://www.kaggle.com/competitions/hubmap-hacking-the-human-vasculature/overview/evaluation\"><strong>Evaluation</strong></a> page for more details.</li>\n</ul></div>","metadata":{}},{"cell_type":"markdown","source":"<div class=\"alert alert-block alert-info\" style=\"font-size:20px; background-color: #ffa3a3; font-family:verdana; color: #8a0f0f; border: 2px #ff0303 solid\">\n    <b>Biological Overview 🔬</b>\n</div>","metadata":{}},{"cell_type":"markdown","source":"![](https://i.gifer.com/2zFp.gif)","metadata":{}},{"cell_type":"markdown","source":"![](https://i.ibb.co/1MZNhNF/003.png)","metadata":{}},{"cell_type":"markdown","source":"<div class=\"alert alert-block alert-info\" style=\"font-size:20px; background-color: #f0c7c7; font-family:verdana; color: #a63c3c; border: 2px #a63c3c solid\">\n    <b>Tissue samples 🧪</b>\n    <br>Renal cortex: the renal cortex is the outer portion of the kidney that contains round renal corpuscles enclosing glomerular tufts, balls of capillary loops. The renal corpuscle is the start of the nephron, through which the filtration of blood occurs. The renal cortex also contains proximal and distal convoluted tubules, PCTs and DCTs, which are also regions of the nephron. Between these tubular structures is a complex network of capillaries called the peritubular capillaries.<br>\n    <br>Renal medulla: the renal medulla is the inner portion of the kidney and is arranged in 8-15 renal pyramids containing linearly arranged tubules comprising the loops of Henle and ducts that gather products for excretion. The capillary network in the renal medulla consists of capillaries called vasa recta. The renal pyramids (medullary tissue) are divided by extensions of the renal cortex called renal columns. There are also projections of the renal medulla into the outer cortex called medullary rays.<br>\n    <br>Renal papilla (subsection of medulla): the broad bases of the pyramids connect to the renal cortex at the corticomedullary junctions while the tips form structures called the renal papilla, which project in the minor renal calyces where urine is collected.<br>\n</div>","metadata":{}},{"cell_type":"markdown","source":"<div class=\"alert alert-block alert-info\" style=\"font-size:20px; background-color: #ffa3a3; font-family:verdana; color: #8a0f0f; border: 2px #ff0303 solid\">\n    <b>EDA Libraries ⚒</b>\n</div>","metadata":{}},{"cell_type":"code","source":"import matplotlib.pyplot as plt \nimport numpy as np\nimport os \nimport pandas as pd \nfrom tqdm import tqdm, notebook \nfrom collections import Counter\nimport warnings ","metadata":{"execution":{"iopub.status.busy":"2023-08-10T19:58:01.89956Z","iopub.execute_input":"2023-08-10T19:58:01.900349Z","iopub.status.idle":"2023-08-10T19:58:02.062633Z","shell.execute_reply.started":"2023-08-10T19:58:01.90031Z","shell.execute_reply":"2023-08-10T19:58:02.061158Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"warnings.filterwarnings(\"ignore\")","metadata":{"execution":{"iopub.status.busy":"2023-08-10T19:58:02.065651Z","iopub.execute_input":"2023-08-10T19:58:02.066201Z","iopub.status.idle":"2023-08-10T19:58:02.072193Z","shell.execute_reply.started":"2023-08-10T19:58:02.066158Z","shell.execute_reply":"2023-08-10T19:58:02.070738Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<div class=\"alert alert-block alert-info\" style=\"font-size:20px; background-color: #ffa3a3; font-family:verdana; color: #8a0f0f; border: 2px #ff0303 solid\">\n    <b>Pathes 🧶</b>\n</div>","metadata":{}},{"cell_type":"code","source":"# Files \n__SAMPLE_SBMISSION_PATH = \"/kaggle/input/hubmap-hacking-the-human-vasculature/sample_submission.csv\"\n__TILE_META_PATH = \"/kaggle/input/hubmap-hacking-the-human-vasculature/tile_meta.csv\"\n__WSI_META_PATH = \"/kaggle/input/hubmap-hacking-the-human-vasculature/wsi_meta.csv\"\n__ANNOTATION_PATH = \"/kaggle/input/hubmap-hacking-the-human-vasculature/polygons.jsonl\"\n__DATASET_CONFIG_PATH = \"/kaggle/working/classes_config.csv\" # Custom file\n\n# Folders\n__TRAIN_PATH = \"/kaggle/input/hubmap-hacking-the-human-vasculature/train\"\n__TEST_PATH = \"/kaggle/input/hubmap-hacking-the-human-vasculature/test\"","metadata":{"execution":{"iopub.status.busy":"2023-08-10T19:58:02.074546Z","iopub.execute_input":"2023-08-10T19:58:02.075084Z","iopub.status.idle":"2023-08-10T19:58:02.092446Z","shell.execute_reply.started":"2023-08-10T19:58:02.075039Z","shell.execute_reply":"2023-08-10T19:58:02.09141Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"sample_submission = pd.read_csv(__SAMPLE_SBMISSION_PATH)\ntile_meta = pd.read_csv(__TILE_META_PATH)\nwsi_meta = pd.read_csv(__WSI_META_PATH)","metadata":{"execution":{"iopub.status.busy":"2023-08-10T19:58:02.095429Z","iopub.execute_input":"2023-08-10T19:58:02.095828Z","iopub.status.idle":"2023-08-10T19:58:02.160929Z","shell.execute_reply.started":"2023-08-10T19:58:02.095795Z","shell.execute_reply":"2023-08-10T19:58:02.159538Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<div class=\"alert alert-block alert-info\" style=\"font-size:20px; background-color: #ffa3a3; font-family:verdana; color: #8a0f0f; border: 2px #ff0303 solid\">\n    <b>Sample submission 📩</b>\n</div>","metadata":{}},{"cell_type":"markdown","source":"<div class=\"alert alert-block alert-info\" style=\"font-size:20px; background-color: #f0c7c7; font-family:verdana; color: #a63c3c; border: 2px #a63c3c solid\">\n    Submissions are evaluated by computing the Average Precision over confidence scores. It is identical to the <a target=\"_blank\" href=\"https://www.kaggle.com/c/open-images-2019-instance-segmentation/overview/evaluation\">OpenImages Instance Segmentation Challenge evaluation</a>, though with only a single class. The OpenImages version of the metric is <a rel=\"noreferrer nofollow\" target=\"_blank\" href=\"https://storage.googleapis.com/openimages/web/evaluation.html#instance_segmentation_eval\">described in detail here</a>. See also <a rel=\"noreferrer nofollow\" target=\"_blank\" href=\"https://github.com/tensorflow/models/blob/master/research/object_detection/g3doc/challenge_evaluation.md#instance-segmentation-track\">this tutorial</a> on running the evaluation in Python.</p>\n<p>Segmentation is calculated using IoU with a threshold of <code>0.6</code>.</p>\n</div>","metadata":{}},{"cell_type":"code","source":"sample_submission","metadata":{"execution":{"iopub.status.busy":"2023-08-10T19:58:02.162929Z","iopub.execute_input":"2023-08-10T19:58:02.163403Z","iopub.status.idle":"2023-08-10T19:58:02.19672Z","shell.execute_reply.started":"2023-08-10T19:58:02.163362Z","shell.execute_reply":"2023-08-10T19:58:02.195423Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<div class=\"sc-dmLtQE jWwnsR\"><p>\n<h2>Submission File</h2>\n<p>For each image in the test set, you must predict a list of instance segmentation masks and their associated detection score (<code>Confidence</code>). The <code>submission.csv</code> file uses the following format:</p>\n<pre><code><span class=\"hljs-attribute\">id</span>,height,width,prediction_string\n<span class=\"hljs-attribute\">72e40acccadf</span>,<span class=\"hljs-number\">512</span>,<span class=\"hljs-number\">512</span>,<span class=\"hljs-number\">0</span> <span class=\"hljs-number\">1</span>.<span class=\"hljs-number\">0</span> eNoLTDAwyrM3yI/PMwcAE94DZA==\n</code></pre>\n<p>where <code>prediction_string</code> has the format <code>0 {confidence} {EncodedMask}</code>. Note that the metric has several \"boilerplate\" values needed to adapt it to this competition; namely, the <code>height</code>, <code>width</code>, and the leading <code>0</code> in <code>prediction_string</code>, which ordinarily is a class label.</p>\n<p>Separate prediction strings multiple instance masks for the same image with a space, like so:</p>\n<pre><code><span class=\"hljs-attribute\">id</span>,height,width,prediction_string\n<span class=\"hljs-attribute\">72e40acccadf</span>,<span class=\"hljs-number\">512</span>,<span class=\"hljs-number\">512</span>,<span class=\"hljs-number\">0</span> <span class=\"hljs-number\">1</span>.<span class=\"hljs-number\">0</span> eNoLTDAwyrM3yI/PMwcAE94DZA== <span class=\"hljs-number\">0</span> <span class=\"hljs-number\">0</span>.<span class=\"hljs-number\">5</span> eAndnnDS1A/mdmkE35Ek9d\n</code></pre>\n<p>The binary segmentation masks are <a rel=\"noreferrer nofollow\" target=\"_blank\" href=\"https://en.wikipedia.org/wiki/Run-length_encoding\">run-length encoded</a> (RLE), <a rel=\"noreferrer nofollow\" target=\"_blank\" href=\"https://en.wikipedia.org/wiki/Zlib\">zlib</a> compressed, and <a rel=\"noreferrer nofollow\" target=\"_blank\" href=\"https://en.wikipedia.org/wiki/Base64\">base64</a> encoded to be used in text format as <code>EncodedMask</code>. Specifically, we use the Coco masks RLE encoding/decoding (see the <code>encode</code> method of <a rel=\"noreferrer nofollow\" target=\"_blank\" href=\"http://cocodataset.org/#download\">COCO’s mask API</a>), the zlib compression/decompression (<a rel=\"noreferrer nofollow\" target=\"_blank\" href=\"https://www.ietf.org/rfc/rfc1950.txt\">RFC1950</a>), and vanilla base64 encoding.</p>","metadata":{}},{"cell_type":"markdown","source":"<div class=\"alert alert-block alert-info\" style=\"font-size:20px; background-color: #ffa3a3; font-family:verdana; color: #8a0f0f; border: 2px #ff0303 solid\">\n    <b>Tile meta</b>\n</div>","metadata":{}},{"cell_type":"markdown","source":"<div class=\"alert alert-block alert-info\" style=\"font-size:20px; background-color: #f0c7c7; font-family:verdana; color: #a63c3c; border: 2px #a63c3c solid\">\n    <br>The competition data includes small pieces called \"tiles\" that come from five big pictures called \"Whole Slide Images\" (WSI). These WSIs are split into two groups called datasets. In Dataset 1, the tiles have been looked at by experts who reviewed and marked them. In Dataset 2, the tiles are from the same big pictures but they don't have as many marks, and the marks they do have haven't been reviewed by experts.<br>\n</div>","metadata":{}},{"cell_type":"code","source":"tile_meta","metadata":{"execution":{"iopub.status.busy":"2023-08-10T19:58:31.937043Z","iopub.execute_input":"2023-08-10T19:58:31.937463Z","iopub.status.idle":"2023-08-10T19:58:31.955346Z","shell.execute_reply.started":"2023-08-10T19:58:31.937436Z","shell.execute_reply":"2023-08-10T19:58:31.953793Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"dataset_count = tile_meta['dataset'].value_counts()\n\nplt.bar(list(map(str, dataset_count.index)), dataset_count.values)\n\nplt.xlabel('Dataset Number')\nplt.ylabel('Number of Examples')\nplt.title('Number of examples in each dataset')\n\nplt.show()","metadata":{"execution":{"iopub.status.busy":"2023-08-10T19:58:35.082799Z","iopub.execute_input":"2023-08-10T19:58:35.083218Z","iopub.status.idle":"2023-08-10T19:58:35.36913Z","shell.execute_reply.started":"2023-08-10T19:58:35.083186Z","shell.execute_reply":"2023-08-10T19:58:35.367926Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<div class=\"alert alert-block alert-info\" style=\"font-size:20px; background-color: #ffa3a3; font-family:verdana; color: #8a0f0f; border: 2px #ff0303 solid\">\n    <b>Wsi meta</b>\n</div>","metadata":{}},{"cell_type":"markdown","source":"<div class=\"alert alert-block alert-info\" style=\"font-size:20px; background-color: #f0c7c7; font-family:verdana; color: #a63c3c; border: 2px #a63c3c solid\">\n    <br>Whole Slide Images (WSIs), also known as virtual slides or digital slides, refer to high-resolution digital representations of entire histopathology glass slides. These slides are typically generated by scanning glass slides using specialized slide scanners.<br>\n    <br>We should expect 14 unique sources of WSIs where 5 have been used for annotations by either a expert or non expert which then get put into dataset 1 and 2 respectively, and the other 9 correspond to WSIs for dataset 3 which have no annotations. Although we should expect 14 we see that there is only thirteen. A keen eye would see that number 5 is missing. This is probably because the 5th WSI belongs to the first dataset and was removed from the dataset to be used as the test set.<br>\n</div>","metadata":{}},{"cell_type":"markdown","source":"## Distribution of WSI by datasets\n![](https://i.ibb.co/6Y1zJvv/001.png)","metadata":{}},{"cell_type":"markdown","source":"## Legend \n![](https://i.ibb.co/gghC5LZ/002.png)","metadata":{}},{"cell_type":"code","source":"wsi_meta","metadata":{"execution":{"iopub.status.busy":"2023-08-10T19:58:41.572986Z","iopub.execute_input":"2023-08-10T19:58:41.573397Z","iopub.status.idle":"2023-08-10T19:58:41.588793Z","shell.execute_reply.started":"2023-08-10T19:58:41.573366Z","shell.execute_reply":"2023-08-10T19:58:41.587932Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"sex_counts = wsi_meta.sex.value_counts()\nrace_counts = wsi_meta.race.value_counts()\n\nfig, (ax1, ax2) = plt.subplots(1, 2)\n\nax1.pie(sex_counts.values, labels=sex_counts.index, autopct='%1.1f%%')\nax1.set_title(\"Distribution of Sexes\")\n\nax2.pie(race_counts.values, labels=race_counts.index, autopct='%1.1f%%')\nax2.set_title(\"Distribution of races\")\n\nplt.show()","metadata":{"execution":{"iopub.status.busy":"2023-08-10T19:58:42.325223Z","iopub.execute_input":"2023-08-10T19:58:42.326415Z","iopub.status.idle":"2023-08-10T19:58:42.673685Z","shell.execute_reply.started":"2023-08-10T19:58:42.326376Z","shell.execute_reply":"2023-08-10T19:58:42.671736Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<div class=\"alert alert-block alert-info\" style=\"font-size:20px; background-color: #ffa3a3; font-family:verdana; color: #8a0f0f; border: 2px #ff0303 solid\">\n    <b>Polygons-annotation 🎯</b>\n</div>","metadata":{}},{"cell_type":"markdown","source":"****\n![](https://i.ibb.co/z4BQ2jH/Screenshot-from-2023-06-11-15-19-16.png)","metadata":{}},{"cell_type":"markdown","source":"<div class=\"alert alert-block alert-info\" style=\"font-size:20px; background-color: #f0c7c7; font-family:verdana; color: #a63c3c; border: 2px #a63c3c solid\">\nPolygonal segmentation masks in JSONL format, available for Dataset 1 and Dataset 2. Each line gives JSON annotations for a single image with:<ul>\n<li><code>id</code> Identifies the corresponding image in <strong>train/</strong></li>\n<li><code>annotations</code>  A list of mask annotations with:</li>\n<li><code>type</code> Identifies the type of structure annotated:<ul>\n<li><code>blood_vessel</code> The target structure. Your goal in this competition is to predict these kinds of masks on the test set.</li>\n<li><code>glomerulus</code> A capillary ball structure in the kidney. These parts of the images were excluded from blood vessel annotation. You should ensure none of your test set predictions occur within glomerulus structures as they will be counted as false positives. Annotations are provided for test set tiles.</li>\n<li><code>unsure</code> A structure the expert annotators cannot confidently distinguish as a blood vessel.</li></ul></li>\n<li><code>coordinates</code> A list of polygon coordinates defining the segmentation mask.</li></ul></li>\n</div>\n\n****","metadata":{}},{"cell_type":"markdown","source":"<div class=\"alert alert-block alert-info\" style=\"font-size:20px; background-color: #ffa3a3; font-family:verdana; color: #8a0f0f; border: 2px #ff0303 solid\">\n    <b>Setup Pipeline ⚙️</b>\n</div>","metadata":{}},{"cell_type":"markdown","source":"<div class=\"alert alert-block alert-info\" style=\"font-size:20px; background-color: #f0c7c7; font-family:verdana; color: #a63c3c; border: 2px #a63c3c solid\">\n    And now we will install pycocotools. \n    <br>This tool is needed for encoding masks and their subsequent evaluation.<br>\n    We will also set a random number generator for all frameworks.\n</div>","metadata":{}},{"cell_type":"code","source":"import torch\nimport random\n\nimport shutil\nimport os\nimport sys\nfrom colorama import Fore","metadata":{"execution":{"iopub.status.busy":"2023-08-10T19:58:50.29348Z","iopub.execute_input":"2023-08-10T19:58:50.293899Z","iopub.status.idle":"2023-08-10T19:58:53.6014Z","shell.execute_reply.started":"2023-08-10T19:58:50.293869Z","shell.execute_reply":"2023-08-10T19:58:53.600515Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"class SetupPipline:   \n    @staticmethod\n    def __pycocotools() -> None:\n        if not os.path.exists(\"/kaggle/working/packages\"):\n            shutil.copytree(\"/kaggle/input/hubmap-tools-ultralytics-and-pycocotools/pycocotools/pycocotools\", \"/kaggle/working/packages\")\n            os.chdir(\"/kaggle/working/packages/pycocotools-2.0.6/\")\n            os.system(\"python setup.py install\")\n            os.system(\"pip install . --no-index --find-links /kaggle/working/packages/\")\n            os.chdir(\"/kaggle/working\")\n\n    @staticmethod\n    def seed_everything(seed: int) -> None:\n        random.seed(seed)\n        np.random.seed(seed)\n        torch.manual_seed(seed)\n        torch.cuda.manual_seed(seed)\n        torch.backends.cudnn.deterministic = True\n    \n    def __call__(self, seed: int = 42, pycoco: bool = True) -> None:\n        if pycoco:\n            self.__pycocotools()\n        if seed:\n            self.seed_everything(seed)","metadata":{"execution":{"iopub.status.busy":"2023-06-25T17:18:28.657946Z","iopub.execute_input":"2023-06-25T17:18:28.65912Z","iopub.status.idle":"2023-06-25T17:18:28.669396Z","shell.execute_reply.started":"2023-06-25T17:18:28.659074Z","shell.execute_reply":"2023-06-25T17:18:28.668191Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<div class=\"alert alert-block alert-info\" style=\"font-size:20px; background-color: #f0c7c7; font-family:verdana; color: #a63c3c; border: 2px #a63c3c solid\">\n    For determinism and stuff like that 🙃\n    <br><br>\n    If the cell swears at something, then do not pay attention. This does not affect operation\n</div>","metadata":{}},{"cell_type":"code","source":"%%capture\nsetup = SetupPipline()\nSEED: int = 42\nsetup(seed=SEED, pycoco=True)","metadata":{"_kg_hide-input":false,"_kg_hide-output":true,"execution":{"iopub.status.busy":"2023-08-10T19:58:53.603246Z","iopub.execute_input":"2023-08-10T19:58:53.604365Z","iopub.status.idle":"2023-08-10T19:58:54.079799Z","shell.execute_reply.started":"2023-08-10T19:58:53.604319Z","shell.execute_reply":"2023-08-10T19:58:54.078601Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<div class=\"alert alert-block alert-info\" style=\"font-size:20px; background-color: #ffa3a3; font-family:verdana; color: #8a0f0f; border: 2px #ff0303 solid\">\n    <b>Dataset libraries</b>\n</div>","metadata":{}},{"cell_type":"code","source":"from torch.utils.data import Dataset\nimport cv2\nimport yaml\nimport json\nfrom PIL import Image","metadata":{"execution":{"iopub.status.busy":"2023-08-10T19:58:54.433284Z","iopub.execute_input":"2023-08-10T19:58:54.433777Z","iopub.status.idle":"2023-08-10T19:58:54.647152Z","shell.execute_reply.started":"2023-08-10T19:58:54.433738Z","shell.execute_reply":"2023-08-10T19:58:54.646029Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<div class=\"alert alert-block alert-info\" style=\"font-size:20px; background-color: #ffa3a3; font-family:verdana; color: #8a0f0f; border: 2px #ff0303 solid\">\n    <b>Dataset config</b>\n</div>","metadata":{}},{"cell_type":"code","source":"dataset_config = {\n    \"background\": {\n        \"apply_mask\": None,\n        \"label\": 0,\n        \"rgb\": (0, 0, 0),\n        \"loss_weight\": None\n    },\n    \"blood_vessel\": {\n        \"apply_mask\": True,\n        \"label\": 1,\n        \"rgb\": (255, 8, 8),\n        \"loss_weight\": None\n    },\n    \"glomerulus\": {\n        \"apply_mask\": True,\n        \"label\": 2,\n        \"rgb\": (8, 12, 255),\n        \"loss_weight\": None\n    },\n    \"unsure\": {\n        \"apply_mask\": True,\n        \"label\": 3,\n        \"rgb\": (8, 255, 20),\n        \"loss_weight\": None\n    }\n}","metadata":{"execution":{"iopub.status.busy":"2023-08-10T19:58:56.325266Z","iopub.execute_input":"2023-08-10T19:58:56.325653Z","iopub.status.idle":"2023-08-10T19:58:56.333234Z","shell.execute_reply.started":"2023-08-10T19:58:56.325624Z","shell.execute_reply":"2023-08-10T19:58:56.332228Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"def write_config(data: dict[dict, ...], path: str) -> None:\n    with open(path, mode=\"w\") as f:\n        yaml.safe_dump(stream=f, data=data)\n        \nwrite_config(dataset_config, __DATASET_CONFIG_PATH)","metadata":{"execution":{"iopub.status.busy":"2023-08-10T19:58:58.587588Z","iopub.execute_input":"2023-08-10T19:58:58.588499Z","iopub.status.idle":"2023-08-10T19:58:58.598269Z","shell.execute_reply.started":"2023-08-10T19:58:58.588462Z","shell.execute_reply":"2023-08-10T19:58:58.597136Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<div class=\"alert alert-block alert-info\" style=\"font-size:20px; background-color: #ffa3a3; font-family:verdana; color: #8a0f0f; border: 2px #ff0303 solid\">\n    <b>Dataset 🪢</b>\n</div>","metadata":{}},{"cell_type":"code","source":"class HuBMAPDataset(Dataset):\n    def __init__(self, \n                 annotation_path: str,\n                 image_path: str, \n                 config_path: str):\n        self.__image_path = image_path\n        self.__samples = self.parse_jsonl(annotation_path)\n        self.__config = self.load_config(config_path)\n    \n    def __len__(self) -> int:\n        return len(self.__samples)\n\n    def __getitem__(self, idx: int) -> tuple[np.ndarray, np.ndarray]:\n        id, image = self.__get_image(idx)\n        mask = self.__get_mask(idx)\n        # image = torch.tensor(image, dtype=torch.float32).permute(2, 0, 1)\n        # mask = torch.tensor(mask, dtype=torch.float32) \n        return id, image, mask\n        \n    @staticmethod\n    def parse_jsonl(path: str) -> list[dict, ...]:\n        with open(path, 'r') as json_file:\n            jsonl_labels = [\n                json.loads(line)\n                for line in notebook.tqdm(\n                    json_file, desc=\"Processing polygons\", total=1633\n                )\n            ]\n        return jsonl_labels\n\n    @staticmethod\n    def load_config(path: str) -> dict:\n        with open(path, mode=\"r\") as f:\n            data = yaml.load(stream=f, Loader=yaml.SafeLoader)\n        return data\n    \n    def __get_image_path(self, id: str) -> str:\n        path = os.path.join(\n            self.__image_path, f\"{id}.tif\"\n        )\n        return path\n    \n    def __get_image(self, idx: int) -> np.ndarray:\n        id = self.__samples[idx][\"id\"]\n        image_path = self.__get_image_path(id)\n        image = Image.open(image_path)\n        image = np.asarray(image)\n        return id, image\n            \n    def __get_mask(self, idx: int) -> np.ndarray:\n        mask = np.zeros((512, 512), dtype=np.uint8)\n        annotations = self.__samples[idx][\"annotations\"]\n        \n        for vessel in annotations:\n            vessel_type = vessel[\"type\"] \n            config = self.__config[vessel_type]\n            \n            if config[\"apply_mask\"]:\n                coordinates = np.array(vessel[\"coordinates\"])\n                mask = cv2.fillPoly(\n                    mask, pts=coordinates,\n                    color=config[\"rgb\"]\n                )\n        return mask","metadata":{"execution":{"iopub.status.busy":"2023-08-10T20:06:03.069572Z","iopub.execute_input":"2023-08-10T20:06:03.070022Z","iopub.status.idle":"2023-08-10T20:06:03.086932Z","shell.execute_reply.started":"2023-08-10T20:06:03.069988Z","shell.execute_reply":"2023-08-10T20:06:03.085466Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<div class=\"alert alert-block alert-info\" style=\"font-size:20px; background-color: #ffa3a3; font-family:verdana; color: #8a0f0f; border: 2px #ff0303 solid\">\n    <b>Let's check dataset</b>\n</div>","metadata":{}},{"cell_type":"code","source":"dataset = HuBMAPDataset(__ANNOTATION_PATH, __TRAIN_PATH, __DATASET_CONFIG_PATH)","metadata":{"execution":{"iopub.status.busy":"2023-08-10T20:06:05.654222Z","iopub.execute_input":"2023-08-10T20:06:05.654622Z","iopub.status.idle":"2023-08-10T20:06:10.022082Z","shell.execute_reply.started":"2023-08-10T20:06:05.654592Z","shell.execute_reply":"2023-08-10T20:06:10.020696Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"id, image, mask = dataset[0]\n\nfig, (ax1, ax2) = plt.subplots(1, 2)\n\nax1.imshow(image)\nax2.imshow(mask)\nplt.title(f\"Segmentation for {id}\")\nplt.show()","metadata":{"execution":{"iopub.status.busy":"2023-08-10T20:06:12.95637Z","iopub.execute_input":"2023-08-10T20:06:12.956767Z","iopub.status.idle":"2023-08-10T20:06:13.43877Z","shell.execute_reply.started":"2023-08-10T20:06:12.956738Z","shell.execute_reply":"2023-08-10T20:06:13.437756Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"id, image, mask = dataset[0]\n\nfig, ax = plt.subplots(1, 1)\nplt.title(f\"Segmentation for {id}\")\n# Display the image\nax.imshow(image, cmap='gray') # You may want to use a different colormap for the image\n\n# Display the mask with a different colormap and some transparency\nax.imshow(mask, cmap='jet', alpha=0.5) # You can adjust the alpha for desired transparency\n\nplt.show()","metadata":{"execution":{"iopub.status.busy":"2023-08-10T20:06:23.686277Z","iopub.execute_input":"2023-08-10T20:06:23.686669Z","iopub.status.idle":"2023-08-10T20:06:24.168492Z","shell.execute_reply.started":"2023-08-10T20:06:23.686642Z","shell.execute_reply":"2023-08-10T20:06:24.167139Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<div class=\"alert alert-block alert-info\" style=\"font-size:20px; background-color: #ffa3a3; font-family:verdana; color: #8a0f0f; border: 2px #ff0303 solid\">\n    <b>Submission 💌</b>\n</div>","metadata":{}},{"cell_type":"markdown","source":"<div class=\"alert alert-block alert-info\" style=\"font-size:20px; background-color: #f0c7c7; font-family:verdana; color: #a63c3c; border: 2px #a63c3c solid\">\n    <b>id, height, width, prediction_string</b>\n    <ul style=\"font-size:20px; font-family:verdana; line-height: 1.7em\">\n        <li>id - identifier of the image that came to the input of the neural network.</li>\n        <li>height - image height (constant value equal to 512).</li> \n        <li>width - image width (constant value equal to 512).</li> \n        <li>prediction_string:\n            <ul style=\"font-size:20px; font-family:verdana; line-height: 1.7em\">\n                <li>label - label class constant value equal to 0 (at least as I understand it, correct me if I'm wrong)</li>\n                <li>confidence - the confidence of the model that this object or instance belongs to the target class. That is, your model will give a mask at the output, where each mask will have the probability that this particular mask belongs to the desired class. It's like the probability of a model belonging to an object to a class. In fact, this is a probabilistic membership score that always appears when it comes to classification.</li>  \n            </ul>\n        </li> \n    </ul>\n</div>\n\n","metadata":{}},{"cell_type":"markdown","source":"<div class=\"alert alert-block alert-info\" style=\"font-size:20px; background-color: #f0c7c7; font-family:verdana; color: #a63c3c; border: 2px #a63c3c solid\">\n    <b>Important 🙃</b>\n    <br>I removed the full implementation so that you understand that you must override the forward method, since I do not know what exactly your neural network returns and this could lead to confusion. In general, the class turned out to be quite comfortable. You only need to pass the model and the path to the test directory, and then call the submit() method. Please let me know if something doesn't work!<br>\n</div>","metadata":{}},{"cell_type":"code","source":"import base64\nimport numpy as np\nimport torch\nfrom pycocotools import _mask as coco_mask\nimport typing as t\nimport zlib\nimport pandas as pd\nimport torchvision.transforms as T","metadata":{"execution":{"iopub.status.busy":"2023-06-25T17:19:33.533436Z","iopub.execute_input":"2023-06-25T17:19:33.53403Z","iopub.status.idle":"2023-06-25T17:19:33.791334Z","shell.execute_reply.started":"2023-06-25T17:19:33.533992Z","shell.execute_reply":"2023-06-25T17:19:33.789819Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"class EncodeBinaryMask:\n    @staticmethod\n    def __checking_mask(mask: np.ndarray) -> np.ndarray:\n        if mask.dtype != np.bool:\n            raise ValueError(\n                \"expects a binary mask, received dtype == %s\" %\n                mask.dtype\n            )\n        return mask\n\n    @staticmethod\n    def __convert_mask(mask: np.ndarray):\n        mask_to_encode = mask.astype(np.uint8)\n        mask_to_encode = np.asfortranarray(mask_to_encode)\n        return mask_to_encode\n\n    @staticmethod\n    def __compress_encode(encoded_mask) -> t.Text:\n        binary_str = zlib.compress(encoded_mask, zlib.Z_BEST_COMPRESSION)\n        base64_str = base64.b64encode(binary_str)\n        return base64_str\n\n    def __call__(self, mask: np.ndarray) -> t.Text:\n        mask = self.__checking_mask(mask)\n        mask_to_encode = self.__convert_mask(mask)\n        encoded_mask = coco_mask.encode(mask_to_encode)[0][\"counts\"]\n        base64_str = self.__compress_encode(encoded_mask)\n        return base64_str\n","metadata":{"execution":{"iopub.status.busy":"2023-06-25T17:19:33.792891Z","iopub.execute_input":"2023-06-25T17:19:33.793236Z","iopub.status.idle":"2023-06-25T17:19:33.802956Z","shell.execute_reply.started":"2023-06-25T17:19:33.793208Z","shell.execute_reply":"2023-06-25T17:19:33.801686Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"class Submission:\n    def __init__(self, dirpath: str, model: torch.nn.Module):\n        self.__eval_transforms = self.get_transforms()\n        self.__model = model\n        self.__encoder = EncodeBinaryMask()\n        self.__dirpath = dirpath\n        self.__filenames = os.listdir(dirpath)\n        \n        self.__submission_dict = {\n            \"id\": [],\n            \"height\": [],\n            \"width\": [],\n            \"prediction_string\": []\n        }\n        \n        self.submission = None\n    \n    @staticmethod\n    def get_transforms():\n        return T.Compose([\n            T.ToTensor(),\n            T.Resize(size=(512, 512)),\n            T.Normalize(mean=[0.485, 0.456, 0.406],\n                        std=[0.229, 0.224, 0.225])\n        ])\n\n    def __len__(self):\n        return len(self.__filenames)\n\n    def __get_columns(self) -> None:\n        for filename in self.__filenames:\n            path = self.__get_image_path(filename)\n            _, image = self.__get_image(path)\n            masks = self.__forward(image)\n            identifier, height, width, prediction_string = self.__get_cells(filename, masks)\n            self.__update_columns(identifier, height, width, prediction_string)\n\n    def __update_columns(self, identifier: str, height: int, width: int, prediction_string: str) -> None:\n        self.__submission_dict[\"id\"].append(identifier)\n        self.__submission_dict[\"height\"].append(height)\n        self.__submission_dict[\"width\"].append(width)\n        self.__submission_dict[\"prediction_string\"].append(prediction_string)\n\n    def __get_cells(self, filename: str, masks: list):\n        prediction_string = \"\"\n        prediction_string = self.__get_prediction_string(masks, prediction_string)\n        identifier = filename.split(\".\")[0]\n        height, width = mask.shape\n        return identifier, height, width, prediction_string\n\n    def __get_prediction_string(self, masks: list, prediction_string: str) -> str:\n        \"\"\"\n        If the neural network did not find the target structure,\n        then we will return an empty prediction_string.\n        \n        However, the columns:\n            id, height, width - must be present for all test images!!!\n        \"\"\"\n        \n        if masks:\n            for outputs in masks:\n                mask = outputs[\"mask\"].detach().permute(1,2,0).cpu().numpy()\n                mask = np.where(mask > 0.5, 1, 0).astype(np.bool)\n                base64_str = self.__encoder(mask)\n                confidence = outputs[\"confidence\"]\n                prediction_string += f\"0 {confidence} {base64_str.decode('utf-8')} \"\n        return prediction_string\n\n    def __get_image_path(self, filename: str) -> str:\n        return os.path.join(\n            self.__dirpath, filename\n        )\n\n    def __get_image(self, path: str) -> torch.Tensor:\n        image = Image.open(path)\n        image = np.asarray(image)\n        image = self.__eval_transforms(image)\n        return image\n\n    # You must implement this function to work with your custom neural network!\n    def __forward(self, image: torch.tensor) ->list:\n        \"\"\"\n        This fuction should return list with masks & confidence.\n        And each mask we shoulde encode (see EncodeBinaryMask & sample_submission file) \n        \n        outputs <- model(image) \n        outputs -> [\n                    {\"mask\": mask1, \"confidence\": confidence1},\n                    {\"mask\": mask2, \"confidence\": confidence2}, \n                    ...,\n                    {\"mask\": maskN, \"confidence\": confidenceN}\n                   ]\n                                  \n        Example:\n        \"\"\"\n        masks = self.__model(image) # Mask must have shape 1x512x512\n        return masks \n\n    def submit(self) -> None:\n        if not self.submission:\n            self.__get_columns()\n            self.submission = pd.DataFrame(self.__submission_dict)\n            self.submission = self.submission.set_index('id')\n            self.submission.to_csv(\"submission.csv\")\n","metadata":{"execution":{"iopub.status.busy":"2023-06-25T17:19:33.804726Z","iopub.execute_input":"2023-06-25T17:19:33.805117Z","iopub.status.idle":"2023-06-25T17:19:33.82827Z","shell.execute_reply.started":"2023-06-25T17:19:33.805083Z","shell.execute_reply":"2023-06-25T17:19:33.826934Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<div class=\"alert alert-block alert-info\" style=\"font-size:20px; background-color: #ffa3a3; font-family:verdana; color: #8a0f0f; border: 2px #ff0303 solid\">\n    <b>Sample model</b>\n</div>","metadata":{}},{"cell_type":"code","source":"class MyBestModel:\n    @staticmethod\n    def generate_masks(num_masks: int) -> list[dict, ...]:\n        masks = []\n        for _ in range(num_masks):\n            mask = torch.randint(0, 2, (1, 512, 512))\n            confidence = round(float(torch.rand(1)[0]), 2)\n            masks.append({\"mask\": mask, \"confidence\": confidence})\n        return masks\n        \n    def __call__(self, image) -> list[dict, ...]:\n        num_masks = torch.randint(1, 5, (1, 1))\n        masks = self.generate_masks(num_masks)\n        return masks","metadata":{"execution":{"iopub.status.busy":"2023-06-25T17:19:33.830304Z","iopub.execute_input":"2023-06-25T17:19:33.830771Z","iopub.status.idle":"2023-06-25T17:19:33.846939Z","shell.execute_reply.started":"2023-06-25T17:19:33.830729Z","shell.execute_reply":"2023-06-25T17:19:33.845798Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"model = MyBestModel()","metadata":{"execution":{"iopub.status.busy":"2023-06-25T17:19:33.848793Z","iopub.execute_input":"2023-06-25T17:19:33.849361Z","iopub.status.idle":"2023-06-25T17:19:33.863233Z","shell.execute_reply.started":"2023-06-25T17:19:33.849322Z","shell.execute_reply":"2023-06-25T17:19:33.862264Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"sub = Submission(dirpath=__TEST_PATH, model=model)\nsub.submit()","metadata":{"execution":{"iopub.status.busy":"2023-06-25T17:19:33.864713Z","iopub.execute_input":"2023-06-25T17:19:33.865722Z","iopub.status.idle":"2023-06-25T17:19:34.127071Z","shell.execute_reply.started":"2023-06-25T17:19:33.865658Z","shell.execute_reply":"2023-06-25T17:19:34.125945Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"sub.submission.head()","metadata":{"execution":{"iopub.status.busy":"2023-06-25T17:19:34.128745Z","iopub.execute_input":"2023-06-25T17:19:34.12913Z","iopub.status.idle":"2023-06-25T17:19:34.140821Z","shell.execute_reply.started":"2023-06-25T17:19:34.129097Z","shell.execute_reply":"2023-06-25T17:19:34.139655Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<div class=\"alert alert-block alert-info\" style=\"font-size:20px; background-color: #ffa3a3; font-family:verdana; color: #8a0f0f; border: 2px #ff0303 solid\">\n    <b>Good luck winning the competition! 🥇</b>\n    <br>I hope this work will help you not to waste too much time and you will quickly be able to understand the data and start the competition. Good luck, friend 🙃<br>\n</div>","metadata":{}},{"cell_type":"markdown","source":"![](https://media.tenor.com/BP2ZMnJf14EAAAAC/winner-winner-chicken-dinner.gif)","metadata":{}}]}