{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"pygments_lexer":"ipython3","nbconvert_exporter":"python","version":"3.6.4","file_extension":".py","codemirror_mode":{"name":"ipython","version":3},"name":"python","mimetype":"text/x-python"}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"### Objective of this competition:\n\nIn this competition, we’ll identify and segment functional tissue units (FTUs) across five human organs. We'll build the model using a dataset of tissue section images. We can see the further analysis of the dataset in this 1st notebook.\n\n* **The five human organs we are segmenting are**:\n                *     Prostate\n                *     Spleen\n                *     Lung\n                *     Kidney\n                *     Large Intestine\n                \n### Dataset Information::\nThe goal of this competition is to identify the locations of each functional tissue unit (FTU) in biopsy slides from several different organs..\n\nThe training annotations are provided as RLE-encoded masks, and the images are in 16-bit, grayscale, PNG format.\n\nThe underlying data includes imagery from different sources prepared with different protocols at a variety of resolutions, reflecting typical challenges for working with medical data.\n\nThis competition uses data from two different consortia, the Human Reference Atlas (HPA) and Human BioMolecular Atlas Program (HuBMAP).\n\nThe data is sourced as following:\n\n * The training dataset consists of data from public HPA data\n * The public test set is a combination of private HPA data and HuBMAP data\n * The private test set contains only HuBMAP data.\n","metadata":{}},{"cell_type":"markdown","source":"***FILE INFORMATION***\n\n[train|test.csv]\n\n* id - The image ID.\n* organ - The organ that the biopsy sample was taken from.\n* data_source - Whether the image was provided by Hubamp or HPA.\n* img_height - The height of the image in pixels.\n* img_width - The width of the image in pixels.\n* pixel_size - The height/width of a single pixel from this image in micrometers. All HPA images have a pixel size of 0.4 µm. For Hubmap imagery the pixel size is 0.5 µm for kidney, 0.2290 µm for large intestine, 0.7562 µm for lung, 0.4945 µm for spleen, and 6.263 µm for prostate.\n* tissue_thickness - The thickness of the biopsy sample in micrometers. All HPA images have a thickness of 4 µm. The Hubmap samples have tissue slice thicknesses 10 µm for kidney, 8 µm for large intestine, 4 µm for spleen, 5 µm for lung, and 5 µm for prostate.\n* rle - The target column. A run length encoded copy of the annotations. Provided for the training set only.\n* age - The patient's age in years. Provided for the training set only.\n* sex - The sex of the patient. Provided for the training set only.\n\n\n[sample_submission.csv]\n\n* id - The image ID.\n* rle - A run length encoded mask of the FTUs in the image.\n\n\n[train|test_images]\n\n**The images:**\n\n* Expect roughly 550 images in the hidden test set.\n* All images used have at least one FTU.\n* All tissue data used in this competition is from healthy donors that pathologists identified as pathologically unremarkable tissue.\n\n**HPA details:**\n\n* All HPA images are 3000 x 3000 pixels with a tissue area within the image around 2500 x 2500 pixels.\n* HPA samples were stained with antibodies visualized with 3,3'-diaminobenzidine (DAB) and counterstained with hematoxylin.\n\n**HuBMAP details:**\n\n* The Hubmap images range in size from 4500x4500 down to 160x160 pixels.\n* HuBMAP images were prepared using Periodic acid-Schiff (PAS)/hematoxylin and eosin (H&E) stains.\n\n\n[train_annotations]\n\n**The annotations:**\n\n* Provided in the format of points that define the boundaries of the polygon masks of the FTUs","metadata":{}},{"cell_type":"code","source":"print(\"\\n\\tVERSION INFORMATION\")\n# Machine Learning and Data Science Imports\nimport tensorflow as tf; print(f\"\\t\\t– TENSORFLOW VERSION: {tf.__version__}\");\nimport tensorflow_hub as tfhub; print(f\"\\t\\t– TENSORFLOW HUB VERSION: {tfhub.__version__}\");\nimport tensorflow_addons as tfa; print(f\"\\t\\t– TENSORFLOW ADDONS VERSION: {tfa.__version__}\");\nimport tensorflow_io as tfio; print(f\"\\t\\t– TENSORFLOW I/O VERSION: {tfio.__version__}\");\nimport pandas as pd; pd.options.mode.chained_assignment = None;\nimport numpy as np; print(f\"\\t\\t– NUMPY VERSION: {np.__version__}\");\n\n\n# Built In Imports\nfrom kaggle_datasets import KaggleDatasets\nfrom collections import Counter\nfrom datetime import datetime\nfrom glob import glob\nimport warnings\nimport requests\nimport hashlib\nimport imageio\nimport IPython\nimport sklearn\nimport urllib\nimport zipfile\nimport pickle\nimport random\nimport shutil\nimport string\nimport json\nimport math\nimport time\nimport gzip\nimport ast\nimport sys\nimport io\nimport os\nimport gc\nimport re\n\n\n# Visualization Imports\nfrom matplotlib.colors import ListedColormap\nfrom matplotlib.patches import Rectangle\nimport matplotlib.patches as patches\nimport plotly.graph_objects as go\nimport matplotlib.pyplot as plt\nfrom tqdm.notebook import tqdm; tqdm.pandas();\nimport plotly.express as px\nimport seaborn as sns\nfrom PIL import Image, ImageEnhance\nimport matplotlib; print(f\"\\t\\t– MATPLOTLIB VERSION: {matplotlib.__version__}\");\nfrom matplotlib import animation, rc; rc('animation', html='jshtml')\nimport plotly\nimport PIL\nimport cv2\n\nimport plotly.io as pio\nprint(pio.renderers)\n\ndef seed_it_all(seed=7):\n    \"\"\" Attempt to be Reproducible \"\"\"\n    os.environ['PYTHONHASHSEED'] = str(seed)\n    random.seed(seed)\n    np.random.seed(seed)\n    tf.random.set_seed(seed)\n\n    \nprint(\"\\n\\n... IMPORTS COMPLETE ...\\n\")\n","metadata":{"execution":{"iopub.status.busy":"2022-08-21T09:33:08.208083Z","iopub.execute_input":"2022-08-21T09:33:08.208521Z","iopub.status.idle":"2022-08-21T09:33:08.225946Z","shell.execute_reply.started":"2022-08-21T09:33:08.208485Z","shell.execute_reply":"2022-08-21T09:33:08.224790Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Training & test files\nRandom_seed = 42\ndata_dir = '../input/hubmap-organ-segmentation'\ntrain_imgs_dir = os.path.join(data_dir,'train_images')\ntrain_annot_dir = os.path.join(data_dir,'train_annotations')\ntrain_file = os.path.join(data_dir,'train.csv')","metadata":{"execution":{"iopub.status.busy":"2022-08-21T09:33:13.306620Z","iopub.execute_input":"2022-08-21T09:33:13.307003Z","iopub.status.idle":"2022-08-21T09:33:13.313373Z","shell.execute_reply.started":"2022-08-21T09:33:13.306973Z","shell.execute_reply":"2022-08-21T09:33:13.312076Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"train_df = pd.read_csv(train_file)\ntrain_df.head(10)","metadata":{"execution":{"iopub.status.busy":"2022-08-21T09:33:16.664425Z","iopub.execute_input":"2022-08-21T09:33:16.664858Z","iopub.status.idle":"2022-08-21T09:33:16.827850Z","shell.execute_reply.started":"2022-08-21T09:33:16.664823Z","shell.execute_reply":"2022-08-21T09:33:16.826535Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# information of training directory\ntrain_df.info()","metadata":{"execution":{"iopub.status.busy":"2022-08-21T09:33:18.023716Z","iopub.execute_input":"2022-08-21T09:33:18.024148Z","iopub.status.idle":"2022-08-21T09:33:18.041630Z","shell.execute_reply.started":"2022-08-21T09:33:18.024111Z","shell.execute_reply":"2022-08-21T09:33:18.040797Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"##### Note : Total number of images are 350 in training directory.","metadata":{}},{"cell_type":"code","source":"# update dataframes \nsex_int = {'Male':0,'Female':1}\nint_2_sex = {v:k for k,v in sex_int.items()}\nage_divisor = 100\n\ntrain_df['sex'] = train_df['sex'].map(sex_int)\ntrain_df['age'] = train_df['age']/100\ntrain_df.head()","metadata":{"execution":{"iopub.status.busy":"2022-08-21T09:33:19.428420Z","iopub.execute_input":"2022-08-21T09:33:19.429379Z","iopub.status.idle":"2022-08-21T09:33:19.451414Z","shell.execute_reply.started":"2022-08-21T09:33:19.429326Z","shell.execute_reply":"2022-08-21T09:33:19.450046Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Adding the imagery path and annotation path in directory\ntrain_imgs = glob(os.path.join(train_imgs_dir, '*.tiff'), recursive=True)\ntrain_annotation = glob(os.path.join(train_annot_dir,'*.json'), recursive=True)\n\ntrain_img_col = {int(x[:-5].rsplit('/',1)[-1]):x for x in train_imgs}\ntrain_annot_col = {int(x[:-5].rsplit('/',1)[-1]):x for x in train_annotation}\n\ntrain_df.insert(3, 'img_path', train_df['id'].map(train_img_col))\ntrain_df.insert(4, 'label_path', train_df['id'].map(train_annot_col))\ntrain_df.head()\n","metadata":{"execution":{"iopub.status.busy":"2022-08-21T09:33:20.277338Z","iopub.execute_input":"2022-08-21T09:33:20.278354Z","iopub.status.idle":"2022-08-21T09:33:20.307893Z","shell.execute_reply.started":"2022-08-21T09:33:20.278306Z","shell.execute_reply":"2022-08-21T09:33:20.307105Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"train_df.describe()","metadata":{"execution":{"iopub.status.busy":"2022-08-21T09:33:20.755037Z","iopub.execute_input":"2022-08-21T09:33:20.755894Z","iopub.status.idle":"2022-08-21T09:33:20.792943Z","shell.execute_reply.started":"2022-08-21T09:33:20.755850Z","shell.execute_reply":"2022-08-21T09:33:20.791557Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# train_df[train_df['organ'] == 'largeintestine'].head()","metadata":{"execution":{"iopub.status.busy":"2022-08-21T09:33:21.413836Z","iopub.execute_input":"2022-08-21T09:33:21.414526Z","iopub.status.idle":"2022-08-21T09:33:21.419472Z","shell.execute_reply.started":"2022-08-21T09:33:21.414476Z","shell.execute_reply":"2022-08-21T09:33:21.418328Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# summarizing the train_df according to some sample ids\nsample_id = [10044, 10392, 10488, 10611, 10651]\ntrain_df[train_df['id'].isin(sample_id)]","metadata":{"execution":{"iopub.status.busy":"2022-08-21T09:33:21.892298Z","iopub.execute_input":"2022-08-21T09:33:21.892920Z","iopub.status.idle":"2022-08-21T09:33:21.914214Z","shell.execute_reply.started":"2022-08-21T09:33:21.892886Z","shell.execute_reply":"2022-08-21T09:33:21.913153Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"def rgb2hex(rgb_tuple):\n    r,g,b = rgb_tuple\n    def clamp(x): \n        return max(0, min(x, 255))\n    return \"#{0:02x}{1:02x}{2:02x}\".format(clamp(r), clamp(g), clamp(b))\nORGANS = ['kidney', 'largeintestine', 'lung', 'prostate', 'spleen']\n_COLOURS = [(230, 0, 73), (11, 180, 255), (80, 233, 145), (230, 216, 0), (155, 25, 245)]\nO2C_MAP = {_o:_c for _o,_c in zip(ORGANS, _COLOURS)}\nO2C_HEX_MAP = {_o:rgb2hex(_c) for _o,_c in O2C_MAP.items()}","metadata":{},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"## exploring the dataframe\nfig = px.histogram(train_df, \"organ\", color_discrete_map=O2C_HEX_MAP, color=\"organ\", title=\"<b>Number of Segmentation Masks Per Organ Type</b>\",)\nfig.show()","metadata":{"execution":{"iopub.status.busy":"2022-08-21T09:33:22.391387Z","iopub.execute_input":"2022-08-21T09:33:22.392407Z","iopub.status.idle":"2022-08-21T09:33:22.474400Z","shell.execute_reply.started":"2022-08-21T09:33:22.392363Z","shell.execute_reply":"2022-08-21T09:33:22.473028Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"## count of the value by sex of patient\nfig = px.histogram(train_df.sex.value_counts(), \"sex\", color_discrete_map=O2C_HEX_MAP, color=\"sex\", title=\"<b>Number of Segmentation Masks according to sex type of patient</b>\",)\nfig.show()","metadata":{"execution":{"iopub.status.busy":"2022-08-21T09:33:22.885269Z","iopub.execute_input":"2022-08-21T09:33:22.885709Z","iopub.status.idle":"2022-08-21T09:33:22.952996Z","shell.execute_reply.started":"2022-08-21T09:33:22.885658Z","shell.execute_reply":"2022-08-21T09:33:22.951816Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"## Visualizing the images\nsample_imgs = []\nsample_mask = []\n\n## decoding the rle \n# ref: https://www.kaggle.com/paulorzp/run-length-encode-and-decode\n# modified from: https://www.kaggle.com/inversion/run-length-decoding-quick-start\n\ndef rle_decode(mask_rle, shape, color=1):\n    \"\"\" TBD\n    \n    Args:\n        mask_rle (str): run-length as string frmated (start length)\n        shape (tuple of ints): (height, width) of array to return \n        \n    Returns:\n        Mask (np.array)\n            - 1 : indicating mask\n            - 0 : indicating background\n    \"\"\"\n    # Split the string by space , then convert it into a integer array\n    s = np.array(mask_rle.split(), dtype=int)\n    # Every even value is the start, every odd value is the \"run\" length\n    starts = s[0::2] - 1\n    lengths = s[1::2]\n    ends = starts + lengths\n    \n    # The image is actually flattened since RLE is a 1D \"run\"\n    if len(shape) == 3:\n        h,w,d = shape\n        img = np.zeros((h * w, d), dtype = np.float32)\n    else:\n        h ,w = shape\n        img = np.zeros((h*w,), dtype = np.float32)\n    \n    # The color here is actually just any integer you want!\n    for lo, hi in zip(starts, ends):\n        img[lo:hi] = color\n        \n    # Don't forget to change the image back to the origina; shape\n    return img.reshape(shape).T\n\n# https://www.kaggle.com/namgalielei/which-reshape-is-used-in-rle\ndef rle_decode_top_to_bot_first(mask_rle, shape):\n    \"\"\"\n    TBD\n    Args:\n        mask_rle (str): run-length as string formated (start length)\n        shape (tuple of ints):(height, width) of array to return\n        \n    Returns:\n        Mask(np.array)\n        ---1 indicating mask\n        ---0 indicating background\n        \"\"\"\n    s = mask_rle.split()\n    starts, lengths = [np.asarray(x, dtype=int) for x in (s[0:][::2], s[1:][::2])]\n    starts -= 1\n    ends = starts + lengths\n    img = np.zeros(shape[0]*shape[1], dtype=np.uint8)\n    for lo, hi in zip(starts, ends):\n        img[lo:hi] = 1\n    return img.reshape((shape[1], shape[0]), order='F').T # Reshape from top -> bottom first\n\n# ref.: https://www.kaggle.com/stainsby/fast-tested-rle\ndef rle_encode(img):\n    \"\"\"TBD\n    Args:\n        img(np.array):\n            - 1 indicating mask\n            - 0 indicating background\n            \n    Returns:\n        run length as string formated\n    \"\"\"\n    pxels = img.T.flatten()\n    pixels = np.concatenate([[0], pixels, [0]])\n    runs = np.where(pixels[1:] != pixels[:-1])[0] + 1\n    runs[1::2] -= runs[::2]\n    return ' '.join(str(x) for x in runs)\n\ndef flatten_l_o_l(nested_list):\n    \"\"\"Flatten a list of lists\"\"\"\n    return [item for sublist in nested_list for item in sublist]\n\ndef load_json_to_dict(json_path):\n    with open(json_path) as json_file:\n        data = json.load(json_file)\n    return data\n\ndef tf_decode_tiff(img_path, to_numpy=False, to_rgb=False):\n    img = tf.io.read_file(img_path)\n    img = tfio.experimental.image.decode_tiff(img)\n    \n    #optionals\n    if to_rgb: img = tfio.experimental.color.rgba_to_rgb(img)\n    if to_numpy: img = img.numpy()\n        \n    return img","metadata":{"execution":{"iopub.status.busy":"2022-08-21T09:33:23.490980Z","iopub.execute_input":"2022-08-21T09:33:23.491367Z","iopub.status.idle":"2022-08-21T09:33:23.508852Z","shell.execute_reply.started":"2022-08-21T09:33:23.491328Z","shell.execute_reply":"2022-08-21T09:33:23.507965Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"def get_overlay(img_path, organ, rle_str, img_shape, _alpha=0.999, _beta=0.35, _gamma=0):\n    _img = tf_decode_tiff(img_path, to_numpy=True, to_rgb=True).astype(np.float32)\n    _seg_rgb = (np.stack([rle_decode(rle_str, shape=img_shape, color=1),]*3,axis=-1)*O2C_MAP[organ]).astype(np.float32)\n    seg_overlay = cv2.addWeighted(src1=_img, alpha=_alpha,\n                                 src2=_seg_rgb, beta=_beta, gamma=_gamma)\n    return seg_overlay/255.  \n\ndef examine_id(df, ex_id=None, plot_overlay=True, print_meta=True, plot_original=False, plot_segmentation=False, _figsize=(20,20)):\n    \"\"\"Wrapper function to allow for easy visual exploration of an example\"\"\"\n    if ex_id is None:\n        ex_id = df['id'].sample(1).values[0]\n        print(f\"\\n... No id is given , random ID chosen: ID={ex_id}...\")\n    print(f\"\\n..ID ({ex_id}) Exploration started..\\n\\n\")\n    demo_ex = train_df[train_df[\"id\"] == ex_id].squeeze()\n    \n    if print_meta:\n        print(f\"\\n With id = `{ex_id}` (organ=`{demo_ex['organ']}`) - We have the following example to show \\n\")\n        display(demo_ex.to_frame())\n        \n    if plot_original:\n        print(f\"\\n\\n Original image plot...\\n\")\n        plt.figure(figsize=_figsize)\n        plt.imshow(tf_decode_tiff(demo_ex[\"img_path\"], to_numpy=True, to_rgb=True))\n        plt.title(f\"Original RGB Image for ID: {ex_id}\", fontweight=\"bold\")\n        plt.axis(False)\n        plt.show()\n        \n    if plot_segmentation:\n        print(f\"\\n\\n Segmentation Mask pot ({demo_ex['organ']})...\\n\")\n        plt.figure(figsize=_figsize)\n        plt.imshow((np.stack([rle_decode(demo_ex.rle, shape=(demo_ex.img_width, demo_ex.img_height), color=1),]*3, axis=-1)*O2C_MAP[demo_ex.organ]).astype(np.float32))\n        plt.title(f\"Segmentation Mask ({demo_ex.organ})\", fontweight='bold')\n        plt.show()\n        \n    if plot_overlay:\n        print(f\"\\n\\n... IMAGE WITH RGB SEGMENTATION MASK OVERLAY ({demo_ex['organ']}) ...\\n\")\n        seg_overlay = get_overlay(demo_ex.img_path, demo_ex.organ, demo_ex.rle, img_shape=(demo_ex.img_width, demo_ex.img_height))\n\n        plt.figure(figsize=_figsize)\n        plt.imshow(seg_overlay)\n        plt.title(f\"Segmentation Overlay id=`{ex_id}` (organ=`{demo_ex['organ']}`)\", fontweight=\"bold\")\n        handles = [Rectangle((0,0),1,1, color=(*[__c/255. for __c in _c], 0.5)) for _c in _COLOURS]\n        labels = ORGANS\n        plt.legend(handles,labels)\n        plt.axis(False)\n        plt.show()\n\n    print(\"\\n\\n... SINGLE ID EXPLORATION FINISHED ...\\n\\n\")","metadata":{"execution":{"iopub.status.busy":"2022-08-21T09:33:23.983009Z","iopub.execute_input":"2022-08-21T09:33:23.983726Z","iopub.status.idle":"2022-08-21T09:33:23.999164Z","shell.execute_reply.started":"2022-08-21T09:33:23.983674Z","shell.execute_reply":"2022-08-21T09:33:23.998162Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"\nfor _o in ORGANS:\n    print(f\"\\n\\n... RANDOM {_o.upper()} TRAIN EXAMPLE ...\\n\\n\")\n    examine_id(train_df[train_df.organ==_o])\n\n# print(\"\\n\\n\\n\\n... TEST IMAGE ...\\n\\n\")\n# examine_id(test_df, plot_overlay=False, plot_original=True)","metadata":{"execution":{"iopub.status.busy":"2022-08-21T09:33:24.729788Z","iopub.execute_input":"2022-08-21T09:33:24.730506Z","iopub.status.idle":"2022-08-21T09:33:47.968946Z","shell.execute_reply.started":"2022-08-21T09:33:24.730460Z","shell.execute_reply":"2022-08-21T09:33:47.967414Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Reference Notebook :\n\n* https://www.kaggle.com/code/abhinand05/hubmap-extensive-eda-what-are-we-hacking\n\n","metadata":{}},{"cell_type":"code","source":"","metadata":{},"execution_count":null,"outputs":[]}]}