{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"pygments_lexer":"ipython3","nbconvert_exporter":"python","version":"3.6.4","file_extension":".py","codemirror_mode":{"name":"ipython","version":3},"name":"python","mimetype":"text/x-python"}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"# Problem definition:\nThe goal in this competition is to locate microvasculature structures (blood vessels) within human kidney histology slides. \n\nThe competition data includes small pieces called \"tiles\" that come from five big pictures called \"Whole Slide Images\" (WSI). We have to map the 'Tiles' into capillaries, arterioles, and venules. ","metadata":{}},{"cell_type":"markdown","source":"# DATA\n1. {train|test}/ Folders containing TIFF images of the tiles. Each tile is 512x512 in size. \n   1. Number of Train Images: 7033\n   2. Number of Test Images: 1\n   \n\n2. polygons.jsonl- Polygonal segmentation masks in JSONL format, available for Dataset 1 and Dataset 2. Each line gives JSON annotations for a single image with:\n   1. id- Identifies the corresponding image in train/\n   2. annotations- A list of mask annotations with:\n   3. type- Identifies the type of structure annotated:\n   \n      i) blood_vessel- The target structure. Your goal in this competition is to predict these kinds of masks on the test set.\n      \n      ii) glomerulus - A capillary ball structure in the kidney. These parts of the images were excluded from blood vessel annotation. \n      \n      You should ensure none of your test set predictions occur within glomerulus structures as they will be counted as false positives. Annotations are provided for test set tiles in the hidden version of the dataset.\n      \n      iii) unsure-  A structure the expert annotators cannot confidently distinguish as a blood vessel.\n      \n   4. coordinates A list of polygon coordinates defining the segmentation mask.\n   \n   \n3. tile_meta.csv\n   \n   \n4. wsi_meta.csv- Metadata for the Whole Slide Images the tiles were extracted from.\n   1. source_wsi Identifies the WSI.\n   2. age, sex, race, height, weight, and bmi demographic information about the tissue donor.\n   \n   \n5. sample_submission.csv-  For each image in the test set, you must predict a list of instance segmentation masks and their associated detection score (Confidence). The submission.csv file uses the following format:\n\nid,height,width,prediction_string\n72e40acccadf,512,512,0 1.0 eNoLTDAwyrM3yI/PMwcAE94DZA==\nwhere prediction_string has the format 0 {confidence} {EncodedMask}. Note that the metric has several \"boilerplate\" values needed to adapt it to this competition; namely, the height, width, and the leading 0 in prediction_string, which ordinarily is a class label.\n\nSeparate prediction strings multiple instance masks for the same image with a space, like so:\n\nid,height,width,prediction_string\n72e40acccadf,512,512,0 1.0 eNoLTDAwyrM3yI/PMwcAE94DZA== 0 0.5 eAndnnDS1A/mdmkE35Ek9d\nThe binary segmentation masks are run-length encoded (RLE), zlib compressed, and base64 encoded to be used in text format as EncodedMask. Specifically, we use the Coco masks RLE encoding/decoding (see the encode method of COCO’s mask API), the zlib compression/decompression (RFC1950), and vanilla base64 encoding.","metadata":{}},{"cell_type":"code","source":"#import libraies\nfrom skimage.io import imread, imshow\nimport seaborn as sns\nimport matplotlib.pylab as plt\nimport os\nimport glob\nfrom glob import glob\nimport numpy as np\nimport pandas as pd\n%matplotlib inline\n\nfrom PIL import Image\n\nfrom IPython.display import display","metadata":{"execution":{"iopub.status.busy":"2023-07-15T11:28:05.566608Z","iopub.execute_input":"2023-07-15T11:28:05.567049Z","iopub.status.idle":"2023-07-15T11:28:06.761375Z","shell.execute_reply.started":"2023-07-15T11:28:05.567015Z","shell.execute_reply":"2023-07-15T11:28:06.760027Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# take a look at the data set directories:\nlist(os.listdir('/kaggle/input/hubmap-hacking-the-human-vasculature'))\n","metadata":{"execution":{"iopub.status.busy":"2023-07-15T11:28:06.930824Z","iopub.execute_input":"2023-07-15T11:28:06.931677Z","iopub.status.idle":"2023-07-15T11:28:06.943204Z","shell.execute_reply.started":"2023-07-15T11:28:06.931619Z","shell.execute_reply":"2023-07-15T11:28:06.942211Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# define path of the 2D PAS-stained histology Tif images:\npath = '/kaggle/input/hubmap-hacking-the-human-vasculature/'\ntrain_path = '/kaggle/input/hubmap-hacking-the-human-vasculature/train/'\ntest_path = '/kaggle/input/hubmap-hacking-the-human-vasculature/test/'","metadata":{"execution":{"iopub.status.busy":"2023-07-15T11:28:08.134057Z","iopub.execute_input":"2023-07-15T11:28:08.134919Z","iopub.status.idle":"2023-07-15T11:28:08.140157Z","shell.execute_reply.started":"2023-07-15T11:28:08.134875Z","shell.execute_reply":"2023-07-15T11:28:08.139406Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# get the number of tif images in train and test directories:\ntrain_images_count = len(glob(os.path.join(train_path,'*.tif')))\ntest_images_count = len(glob(os.path.join(test_path,'*.tif')))\nprint('Number of Train Images: {}'.format(train_images_count))\nprint('Number of Test Images: {}'.format(test_images_count))\n","metadata":{"execution":{"iopub.status.busy":"2023-07-15T11:28:09.142787Z","iopub.execute_input":"2023-07-15T11:28:09.143535Z","iopub.status.idle":"2023-07-15T11:28:09.342866Z","shell.execute_reply.started":"2023-07-15T11:28:09.143489Z","shell.execute_reply":"2023-07-15T11:28:09.341592Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Get polygons data:\npolygons_df = pd.read_json(os.path.join(path,\"polygons.jsonl\"), lines=True)\ndisplay(polygons_df.head(10))\npolygons_df.describe()","metadata":{"execution":{"iopub.status.busy":"2023-07-15T11:28:10.166826Z","iopub.execute_input":"2023-07-15T11:28:10.167310Z","iopub.status.idle":"2023-07-15T11:28:16.142504Z","shell.execute_reply.started":"2023-07-15T11:28:10.167265Z","shell.execute_reply":"2023-07-15T11:28:16.141074Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# take a look at sample submission:\nsubmission = pd.read_csv(path+'sample_submission.csv')\nsubmission.head()","metadata":{"execution":{"iopub.status.busy":"2023-07-15T11:28:16.144917Z","iopub.execute_input":"2023-07-15T11:28:16.146131Z","iopub.status.idle":"2023-07-15T11:28:16.167704Z","shell.execute_reply.started":"2023-07-15T11:28:16.146082Z","shell.execute_reply":"2023-07-15T11:28:16.165972Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"#wsi metadata\nwsi_meta_df = pd.read_csv(path + \"wsi_meta.csv\")\ndisplay(wsi_meta_df.head(10))\nwsi_meta_df.describe()","metadata":{"execution":{"iopub.status.busy":"2023-07-15T11:28:16.168889Z","iopub.execute_input":"2023-07-15T11:28:16.169916Z","iopub.status.idle":"2023-07-15T11:28:16.215599Z","shell.execute_reply.started":"2023-07-15T11:28:16.169885Z","shell.execute_reply":"2023-07-15T11:28:16.214745Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# tile metadata:\ntile_meta_df = pd.read_csv(path + \"tile_meta.csv\")\ndisplay(tile_meta_df.head(10))\ntile_meta_df.describe()","metadata":{"execution":{"iopub.status.busy":"2023-07-15T11:28:16.217634Z","iopub.execute_input":"2023-07-15T11:28:16.218119Z","iopub.status.idle":"2023-07-15T11:28:16.270562Z","shell.execute_reply.started":"2023-07-15T11:28:16.218092Z","shell.execute_reply":"2023-07-15T11:28:16.269381Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Tile Meta Files Exploration\n\ntile_meta.csv- Metadata for each image. The hidden version of this file also contains metadata for the test set tiles.\n\n1. source_wsi- Identifies the WSI this tile was extracted from. There are 13 sources.\n\n2. {i|j}- The location of the upper-left corner within the WSI where the tile was extracted.\n\n3. dataset- The dataset this tile belongs to, as described above. The data comes from five big pictures called \"Whole Slide Images\" (WSI). These WSIs are split into three groups called datasets.\n   1. In Dataset 1, 6% of the tiles have been looked at by experts who reviewed and marked them.\n   2. In Dataset 2, the tiles are from the same big pictures but they don't have as many marks, and the marks they do have haven't been reviewed by experts.\n   3. We also include, as Dataset 3, tiles extracted from an additional nine WSIs. These tiles have not been annotated. You may wish to apply semi- or self-supervised learning techniques on this data to support your predictions.\n","metadata":{}},{"cell_type":"code","source":"tile_meta_df","metadata":{"execution":{"iopub.status.busy":"2023-07-15T11:28:16.272933Z","iopub.execute_input":"2023-07-15T11:28:16.273874Z","iopub.status.idle":"2023-07-15T11:28:16.289844Z","shell.execute_reply.started":"2023-07-15T11:28:16.273823Z","shell.execute_reply":"2023-07-15T11:28:16.288427Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"tile_meta_df.nunique()","metadata":{"execution":{"iopub.status.busy":"2023-07-15T11:28:16.291333Z","iopub.execute_input":"2023-07-15T11:28:16.291800Z","iopub.status.idle":"2023-07-15T11:28:16.314003Z","shell.execute_reply.started":"2023-07-15T11:28:16.291760Z","shell.execute_reply":"2023-07-15T11:28:16.312890Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"dataset_counts = tile_meta_df.dataset.value_counts()\nplt.pie(dataset_counts, labels=dataset_counts.index, autopct=lambda pct: f'{pct:.1f}% ({int(pct * sum(dataset_counts)/100)})')\n\nplt.title(\"Distribution of Dataset\")\n\nplt.show()","metadata":{"execution":{"iopub.status.busy":"2023-07-15T11:29:00.025418Z","iopub.execute_input":"2023-07-15T11:29:00.025877Z","iopub.status.idle":"2023-07-15T11:29:00.203511Z","shell.execute_reply.started":"2023-07-15T11:29:00.025844Z","shell.execute_reply":"2023-07-15T11:29:00.201878Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"tile_meta_df.source_wsi.value_counts()\n#plot source_wsi distribution:\nsns.countplot(data=tile_meta_df, x='source_wsi', hue = 'dataset')\n\n# Set plot labels and title\nplt.xlabel('source_wsi')\nplt.ylabel('Count')\nplt.title('distribution of source_wsi Variable')\n\n# Show the plot\nplt.show()","metadata":{"execution":{"iopub.status.busy":"2023-07-15T11:29:03.445542Z","iopub.execute_input":"2023-07-15T11:29:03.446286Z","iopub.status.idle":"2023-07-15T11:29:03.924070Z","shell.execute_reply.started":"2023-07-15T11:29:03.446218Z","shell.execute_reply":"2023-07-15T11:29:03.923010Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"clearly:\n1. Dataset 1 -comes from WSI 1 and 2.\n2. Dataset 2 - comes from 1,2,3,4 WSI\n3. Dataest 3 - comes from 6-14 WSI sources.\n","metadata":{}},{"cell_type":"markdown","source":"# WSI metadata:\n\nWhole Slide Images (WSIs), also known as virtual slides or digital slides, refer to high-resolution digital representations of entire histopathology glass slides. These slides are typically generated by scanning glass slides using specialized slide scanners.\n\nWe should expect 14 unique sources of WSIs where 5 have been used for annotations by either a expert or non expert which then get put into dataset 1 and 2 respectively, and the other 9 correspond to WSIs for dataset 3 which have no annotations. Although we should expect 14 we see that there is only thirteen. \n\nA keen eye would see that number 5 is missing. This is probably because the 5th WSI belongs to the first dataset and was removed from the dataset to be used as the test set.","metadata":{}},{"cell_type":"code","source":"wsi_meta_df","metadata":{"execution":{"iopub.status.busy":"2023-07-15T11:29:20.583960Z","iopub.execute_input":"2023-07-15T11:29:20.585192Z","iopub.status.idle":"2023-07-15T11:29:20.600962Z","shell.execute_reply.started":"2023-07-15T11:29:20.585142Z","shell.execute_reply":"2023-07-15T11:29:20.598607Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"As can be seen above, only 4 wsi are present in the available data out of the 5 mentioned.","metadata":{}},{"cell_type":"code","source":"sex_counts = wsi_meta_df.sex.value_counts()\nrace_counts = wsi_meta_df.race.value_counts()\n\nfig, (ax1, ax2) = plt.subplots(1, 2)\n\nax1.pie(sex_counts.values, labels=sex_counts.index, autopct='%1.1f%%')\nax1.set_title(\"Distribution of Sexes\")\n\nax2.pie(race_counts.values, labels=race_counts.index, autopct='%1.1f%%')\nax2.set_title(\"Distribution of races\")\n\nplt.show()","metadata":{"execution":{"iopub.status.busy":"2023-07-15T11:29:22.816722Z","iopub.execute_input":"2023-07-15T11:29:22.817162Z","iopub.status.idle":"2023-07-15T11:29:23.042160Z","shell.execute_reply.started":"2023-07-15T11:29:22.817123Z","shell.execute_reply":"2023-07-15T11:29:23.040533Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"#Age Distribution of WSI sample:\nplt.hist(wsi_meta_df['age'], bins = 5 )\nplt.xlabel('Age')\nplt.ylabel('distribution')\nplt.title('Histogram of age')\nplt.show()","metadata":{"execution":{"iopub.status.busy":"2023-07-15T11:33:11.773687Z","iopub.execute_input":"2023-07-15T11:33:11.774138Z","iopub.status.idle":"2023-07-15T11:33:12.113776Z","shell.execute_reply.started":"2023-07-15T11:33:11.774105Z","shell.execute_reply":"2023-07-15T11:33:12.112498Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Polygonal segmentation\nPolygonal segmentation masks in JSONL format, available for Dataset 1 and Dataset 2. Each line gives JSON annotations for a single image with:\n1. id-  Identifies the corresponding image in train/\n2. annotations-  A list of mask annotations with:\n3. type- Identifies the type of structure annotated:\n   1. blood_vessel: The target structure. Your goal in this competition is to predict these kinds of masks on the test set.\n   2. glomerulus: A capillary ball structure in the kidney. These parts of the images were excluded from blood vessel annotation. You should ensure none of your test set predictions occur within glomerulus structures as they will be counted as false positives. Annotations are provided for test set tiles.\n   3. unsure: A structure the expert annotators cannot confidently distinguish as a blood vessel.\n4. coordinates A list of polygon coordinates defining the segmentation mask.","metadata":{}},{"cell_type":"code","source":"polygons_df","metadata":{"execution":{"iopub.status.busy":"2023-07-14T12:32:41.255192Z","iopub.execute_input":"2023-07-14T12:32:41.255755Z","iopub.status.idle":"2023-07-14T12:32:41.365432Z","shell.execute_reply.started":"2023-07-14T12:32:41.255716Z","shell.execute_reply":"2023-07-14T12:32:41.364028Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Annotated vs Non Annotated data\nassert polygons_df.id.isin(tile_meta_df.id).all()\ntotal_images_count = len(tile_meta_df.id)\nannotated_images_count = len(polygons_df.id)\nnot_annotated_images_count = total_images_count - annotated_images_count\n\nannotation_counts_df = pd.DataFrame({\n    \"annotated\": [annotated_images_count],\n    \"not_annotated\": [not_annotated_images_count]\n})\n\nprint(f\"Not all the images provided are annotated. Here are the counts for total images {total_images_count} we have only {annotated_images_count} annotated and {not_annotated_images_count} are not annotated. it would be interesting to see the metadata and see how many of the datasets are annotated.\")\ndisplay(annotation_counts_df)\n","metadata":{"execution":{"iopub.status.busy":"2023-07-14T12:34:15.002864Z","iopub.execute_input":"2023-07-14T12:34:15.003313Z","iopub.status.idle":"2023-07-14T12:34:15.021704Z","shell.execute_reply.started":"2023-07-14T12:34:15.003282Z","shell.execute_reply":"2023-07-14T12:34:15.020570Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"#distribution of kind of annatation:\nannotations_list = polygons_df['annotations'].explode()\n# Extract the 'type' field from each dictionary\ntypes = annotations_list.apply(lambda x: x['type'])\n# Count the occurrences of each 'type'\ntype_counts = types.value_counts()\nplt.pie(type_counts, labels=type_counts.index, autopct=lambda pct: f'{pct:.1f}% ({int(pct * sum(type_counts)/100)})')\n\n# Add a title\nplt.title(\"annotation Counts\")\n\n# Display the chart\nplt.show()\nplt.close()","metadata":{"execution":{"iopub.status.busy":"2023-07-14T12:33:35.618357Z","iopub.execute_input":"2023-07-14T12:33:35.618875Z","iopub.status.idle":"2023-07-14T12:33:35.806123Z","shell.execute_reply.started":"2023-07-14T12:33:35.618836Z","shell.execute_reply":"2023-07-14T12:33:35.804741Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"fig, (ax0, ax1) = plt.subplots(nrows=1, ncols=2, figsize=(10,5))\n\nsource_wsi_counts_1 = tile_meta_df[tile_meta_df.id.isin(polygons_df.id) & (tile_meta_df.dataset==1)].source_wsi.value_counts()\nvalues = source_wsi_counts_1.values\nlabels = [f\"wsi src {x}\" for x in source_wsi_counts_1.index]\nax0.pie(values,labels=labels, autopct=lambda pct: f'{pct:.1f}% ({int(pct * sum(source_wsi_counts_1)/100)})')\nax0.set_title('Counts images from the wsi source in DATASET 1')\n\nsource_wsi_counts_2 = tile_meta_df[tile_meta_df.id.isin(polygons_df.id) & (tile_meta_df.dataset==2)].source_wsi.value_counts()\nvalues = source_wsi_counts_2.values\nlabels = [f\"wsi src {x}\" for x in source_wsi_counts_2.index]\nax1.pie(values,labels=labels, autopct=lambda pct: f'{pct:.1f}% ({int(pct * sum(source_wsi_counts_1)/100)})')\nax1.set_title('Counts images from the wsi source in DATASET 2')\n\nplt.show()","metadata":{"execution":{"iopub.status.busy":"2023-07-14T12:35:38.548145Z","iopub.execute_input":"2023-07-14T12:35:38.548515Z","iopub.status.idle":"2023-07-14T12:35:38.884061Z","shell.execute_reply.started":"2023-07-14T12:35:38.548487Z","shell.execute_reply":"2023-07-14T12:35:38.882287Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Visualize slides:","metadata":{}},{"cell_type":"code","source":"# From https://www.kaggle.com/code/leonidkulyk/eda-hubmap-hhv-interactive-annotations\ndef get_cartesian_coords(coords, img_height=512):\n    coords_array = np.array(coords).squeeze()\n    xs = coords_array[:, 0]\n    ys = coords_array[:, 1]\n    \n    return xs, ys","metadata":{"execution":{"iopub.status.busy":"2023-07-14T11:34:46.516290Z","iopub.execute_input":"2023-07-14T11:34:46.516655Z","iopub.status.idle":"2023-07-14T11:34:46.522900Z","shell.execute_reply.started":"2023-07-14T11:34:46.516611Z","shell.execute_reply":"2023-07-14T11:34:46.521665Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"import json\nwith open(\"../input/hubmap-hacking-the-human-vasculature/polygons.jsonl\") as f:\n    data = f.read()\n    \n    \nres = []\nfor file in data.splitlines():\n    d = json.loads(file)\n    res.append(d)","metadata":{"execution":{"iopub.status.busy":"2023-07-14T11:34:47.324191Z","iopub.execute_input":"2023-07-14T11:34:47.324564Z","iopub.status.idle":"2023-07-14T11:34:50.296399Z","shell.execute_reply.started":"2023-07-14T11:34:47.324536Z","shell.execute_reply":"2023-07-14T11:34:50.294945Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"d = res[500]\nimg_id = d[\"id\"]\npath = f\"../input/hubmap-hacking-the-human-vasculature/train/{img_id}.tif\"\nimg = imread(path)","metadata":{"execution":{"iopub.status.busy":"2023-07-14T11:34:50.298683Z","iopub.execute_input":"2023-07-14T11:34:50.299139Z","iopub.status.idle":"2023-07-14T11:34:50.356734Z","shell.execute_reply.started":"2023-07-14T11:34:50.299099Z","shell.execute_reply":"2023-07-14T11:34:50.355806Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"fig, ax = plt.subplots(1, 1, figsize=(10, 10))\nax.imshow(img)\nfor e in d[\"annotations\"]:\n    if e[\"type\"] == \"blood_vessel\":\n        coordinates = e[\"coordinates\"]\n        xs, ys = get_cartesian_coords(coordinates)\n        ax.plot(xs, ys, c=\"red\")\nfig.show()","metadata":{"execution":{"iopub.status.busy":"2023-07-14T11:34:50.358205Z","iopub.execute_input":"2023-07-14T11:34:50.358839Z","iopub.status.idle":"2023-07-14T11:34:51.307361Z","shell.execute_reply.started":"2023-07-14T11:34:50.358800Z","shell.execute_reply":"2023-07-14T11:34:51.306209Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"","metadata":{},"execution_count":null,"outputs":[]}]}