{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"pygments_lexer":"ipython3","nbconvert_exporter":"python","version":"3.6.4","file_extension":".py","codemirror_mode":{"name":"ipython","version":3},"name":"python","mimetype":"text/x-python"}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"## HuBMAP Let's Visualize and Understand Dataset\n\n### Release note\n\n- Version1: First release\n\n- Version2: Reorganized and added information to easier understand.\n\n- Version3: Add how to visualize anatomical-structure.json.\n\n- Version4: Add how to visualize glomerulus_segmentation_file.\n\n- Version5: Add explanation for background.\n\n- Version6: Maintained with additional data.\n\n## Contents\n\n1. [Introduction](#1)\n1. [Loading and overviewing dataset](#2)\n1. [Read and show Kidney image data](#3)\n1. [Visualization of HuBMAP-20-dataset_information](#4)","metadata":{}},{"cell_type":"markdown","source":"<a id=\"1\"></a> <br>\n# <div class=\"alert alert-block alert-info\">Introduction</div>","metadata":{}},{"cell_type":"markdown","source":"## Goal\n\nGoal of this competition is development of a segmentation algorithm to identify the \"Glomerulus\" in the kidney.\n\nWe are given histological images of the kidney and annotation information representing the glomerular segmentation. Also we can use anatomical structure segmentation information and additional information (including anonymized patient data) about each image. ","metadata":{}},{"cell_type":"markdown","source":"## About Glomerulus\n\nThe glomerulus is one of the components of the \"nephron\". Nephron is said to be one million in one kidney. How many nephrons there are can be seen [later](#3) when you visualize the dataset. I'll put image of nephron from [reference[1]](#101). Nephron has roughly three components, glomerulus, bowman's capsule and tubule. \n\nGlomerulus is a mass of capillaries surrounded by a Bowman's capsule. The name comes from the fact that they look just like a hairball when viewed under a microscope. They act like filter paper.\nPlasma (the non-cellular components of blood) sent to the kidney is filtered out during its passage through the capillaries of the glomerulus. Some of it comes out of the Bowman's capsule as the original urine. \n\n<img src=\"https://upload.wikimedia.org/wikipedia/commons/thumb/1/18/Bowman%27s_capsule_and_glomerulus.svg/1280px-Bowman%27s_capsule_and_glomerulus.svg.png\" width=\"250\">","metadata":{}},{"cell_type":"markdown","source":"## Notes for new participants\n\nWe should be aware that notebooks that have been around for some time may not be in line with the current situation. We may want to pay attention to when the notebook was last run and the data used. \n\nThis competition [updated dataset](https://www.kaggle.com/c/hubmap-kidney-segmentation/discussion/224826), and [this discussion](https://www.kaggle.com/c/hubmap-kidney-segmentation/discussion/207884) will help us understand how it happened and the precautions we need to take with updating. ","metadata":{}},{"cell_type":"markdown","source":"<a id=\"2\"></a> <br>\n# <div class=\"alert alert-block alert-success\">Loading and overviewing dataset</div>\n\n## Load Library","metadata":{}},{"cell_type":"code","source":"import collections\nimport json\nimport os\nimport uuid\n\nimport cv2\nimport matplotlib.pyplot as plt\nimport numpy as np\nimport pandas as pd\nfrom PIL import Image, ImageDraw, ImageFilter\nimport tifffile as tiff \nimport seaborn as sns","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## Load data\n\nThere are three csv data and two directories.","metadata":{}},{"cell_type":"code","source":"!ls ../input/hubmap-kidney-segmentation/","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### train.csv\n\nThere are 8 training set. This csv includes ids corresponding to data in train directory. Also it has mask data in \"encoding\" column. This data is encoded with RLE encoding. ","metadata":{}},{"cell_type":"code","source":"train = pd.read_csv(\"../input/hubmap-kidney-segmentation/train.csv\")\ntrain.info()","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"train.head()","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### HuBMAP-20-dataset_information.csv\n\nThis file includes additional information (including anonymized patient data) about each image. I'll visualize this information in following part.","metadata":{}},{"cell_type":"code","source":"ds_info = pd.read_csv(\"../input/hubmap-kidney-segmentation/HuBMAP-20-dataset_information.csv\")\nds_info.info()","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"ds_info.head()","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### train directory\n\ntiff files are kidney image data. json files include unencoded annotations. ","metadata":{}},{"cell_type":"code","source":"!ls ../input/hubmap-kidney-segmentation/train","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"There are two kinds of json files. About glomerulus segmentation file files, I'll explain it in [here](#99), and about anatomical structure file in [here](#100).","metadata":{}},{"cell_type":"markdown","source":"### test directory","metadata":{}},{"cell_type":"code","source":"!ls ../input/hubmap-kidney-segmentation/test","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<a id=\"3\"></a> <br>\n# <div class=\"alert alert-block alert-warning\">Read and show Kidney image data</div>\n\nWe are given histological images of the kidney. These images are tiff format. We can load this data with tifffile module. Let's load and show them.","metadata":{}},{"cell_type":"code","source":"img_id_1 = \"aaa6a05cc\"\nimage_1 = tiff.imread('../input/hubmap-kidney-segmentation/train/' + img_id_1 + \".tiff\")\nprint(\"This image's id:\", img_id_1)\nimage_1.shape","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"plt.figure(figsize=(15, 15))\nplt.imshow(image_1)","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Glomerulus is...","metadata":{}},{"cell_type":"code","source":"plt.figure(figsize=(8,8))\nplt.imshow(image_1[5200:5600, 5600:6000, :])","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Some tiff files are saved by different shape. We can view them in the same way by reading and then taking the transposition.","metadata":{}},{"cell_type":"code","source":"img_id_4 = \"e79de561c\"\nimage_4 = tiff.imread('../input/hubmap-kidney-segmentation/train/' + img_id_4 + \".tiff\")\nprint(\"This image's id:\", img_id_4)\nimage_4.shape\nimage_4 = image_4[0][0].transpose(1, 2, 0)\nplt.figure(figsize=(10, 10))\nplt.imshow(image_4)","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## mask\n\nWe can decode mask from encoding column of train.csv.","metadata":{}},{"cell_type":"code","source":"# https://www.kaggle.com/paulorzp/rle-functions-run-lenght-encode-decode\ndef mask2rle(img):\n    '''\n    img: numpy array, 1 - mask, 0 - background\n    Returns run length as string formated\n    '''\n    pixels= img.T.flatten()\n    pixels = np.concatenate([[0], pixels, [0]])\n    runs = np.where(pixels[1:] != pixels[:-1])[0] + 1\n    runs[1::2] -= runs[::2]\n    return ' '.join(str(x) for x in runs)\n \ndef rle2mask(mask_rle, shape=(1600,256)):\n    '''\n    mask_rle: run-length as string formated (start length)\n    shape: (width,height) of array to return \n    Returns numpy array, 1 - mask, 0 - background\n\n    '''\n    s = mask_rle.split()\n    starts, lengths = [np.asarray(x, dtype=int) for x in (s[0:][::2], s[1:][::2])]\n    starts -= 1\n    ends = starts + lengths\n    img = np.zeros(shape[0]*shape[1], dtype=np.uint8)\n    for lo, hi in zip(starts, ends):\n        img[lo:hi] = 1\n    return img.reshape(shape).T","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"mask_1 = rle2mask(train[train[\"id\"]==img_id_1][\"encoding\"].iloc[-1], (image_1.shape[1], image_1.shape[0]))\nmask_1.shape","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Show the mask of the kidney image.","metadata":{}},{"cell_type":"code","source":"plt.figure(figsize=(10,10))\nplt.imshow(mask_1, cmap='coolwarm', alpha=0.5)","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"If we want to see with the image, ","metadata":{}},{"cell_type":"code","source":"plt.figure(figsize=(10,10))\nplt.imshow(image_1)\nplt.imshow(mask_1, cmap='coolwarm', alpha=0.5)","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Let's see one more example.","metadata":{}},{"cell_type":"code","source":"mask_4 = rle2mask(train[train[\"id\"]==img_id_4][\"encoding\"].iloc[-1], (image_4.shape[1], image_4.shape[0]))\nmask_4.shape\nplt.figure(figsize=(10,10))\nplt.imshow(image_4)\nplt.imshow(mask_4, cmap='coolwarm', alpha=0.5)","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"According to the [report](https://www.kaggle.com/c/hubmap-kidney-segmentation/discussion/198116) in the discussion, some annotations still be missing. For more information on the impact of this missing annotation data on prediction performance and data handling, [this discussion](https://www.kaggle.com/c/hubmap-kidney-segmentation/discussion/227616#1250442) is helpful.","metadata":{}},{"cell_type":"markdown","source":"## Processing\n\nTo create dataset for model training, we have to generate dataset from these pictures. I'll explain how to process and save image.","metadata":{}},{"cell_type":"markdown","source":"### Crop","metadata":{}},{"cell_type":"code","source":"plt.figure(figsize=(8,8))\nplt.imshow(image_1[5200:6200, 5000:6000, :])\nplt.imshow(mask_1[5200:6200, 5000:6000], cmap='coolwarm', alpha=0.5)","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### Rotate","metadata":{}},{"cell_type":"code","source":"plt.figure(figsize=(8,8))\nplt.imshow(np.rot90(image_1[5200:6200, 5000:6000, :]))\nplt.imshow(np.rot90(mask_1[5200:6200, 5000:6000]), cmap='coolwarm', alpha=0.5)","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### flip","metadata":{}},{"cell_type":"code","source":"plt.figure(figsize=(8,8))\nplt.imshow(np.fliplr(image_1[5200:6200, 5000:6000, :]))\nplt.imshow(np.fliplr(mask_1[5200:6200, 5000:6000]), cmap='coolwarm', alpha=0.5)","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### with filters\n\nWe can use filters with [ImageFilter Module](https://pillow.readthedocs.io/en/stable/reference/ImageFilter.html). The current version of Pillow provides the following set of predefined image enhancement filters:\n\n- BLUR\n- CONTOUR\n- DETAIL\n- EDGE_ENHANCE\n- EDGE_ENHANCE_MORE\n- EMBOSS\n- FIND_EDGES\n- SHARPEN\n- SMOOTH\n- SMOOTH_MORE\n\nAfter converting image to PIL image with Image.fromarray function, we can use image filter. And we can return it to np.array.","metadata":{}},{"cell_type":"code","source":"im_filterd = Image.fromarray(image_1)\nim_filterd = np.array(im_filterd.filter(ImageFilter.EDGE_ENHANCE_MORE))\nimage_1 = np.array(im_filterd)\n\nplt.figure(figsize=(8,8))\nplt.imshow(image_1[5200:6200, 5000:6000, :])\nplt.imshow(mask_1[5200:6200, 5000:6000], cmap='coolwarm', alpha=0.5)","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### Save as image\n\nWith pillow, we can save our processed kideny images and masks as image file. For create dataset, I'll try.","metadata":{}},{"cell_type":"code","source":"os.makedirs(f\"./image/{img_id_1}/\")\nos.makedirs(f\"./mask/{img_id_1}/\")","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"pil_img = Image.fromarray(image_1[5200:6200, 5000:6000, :])\nprint(pil_img.mode)\n\nimg_uuid = str(uuid.uuid4())\n\npil_img.save(f'./image/{img_id_1}/{img_id_1}_{img_uuid}.jpg')\nnp.save(f'./mask/{img_id_1}/{img_id_1}_{img_uuid}', mask_1[5200:6200, 5000:6000])","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## With Annotation json file\n\nWe have also two kinds of annotation files. I'll explain what information they have and how to visualize them.","metadata":{}},{"cell_type":"markdown","source":"<a id=\"99\"></a>\n### Glomerulus segmentation file\n\nAccording to the description of dataset, the same information as the rle-encoded mask is stored.","metadata":{}},{"cell_type":"code","source":"with open(\"../input/hubmap-kidney-segmentation/train/e79de561c.json\") as f:\n    e79de561c_json = json.load(f)\n    \nprint(\"lenght of json:\", len(e79de561c_json))\nprint(e79de561c_json[0])","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"I'll define utility function. By this function, we can get PIL.Image.Image instance with line.","metadata":{}},{"cell_type":"code","source":"def flatten(l):\n    for el in l:\n        if isinstance(el, collections.abc.Iterable) and not isinstance(el, (str, bytes)):\n            yield from flatten(el)\n        else:\n            yield el\n\ndef draw_structure(structures, im):\n    \"\"\"\n    anatomical_structure: list of points of anatomical_structure poligon.\n    im: numpy array of image read from tiff file.\n    \"\"\"\n    \n    im = Image.fromarray(im)\n    draw = ImageDraw.Draw(im)\n    for structure in structures:\n        structure_flatten = list(flatten(structure[\"geometry\"][\"coordinates\"][0]))\n        structure = []\n        for i in range(0, len(structure_flatten), 2):\n            structure.append(tuple(structure_flatten[i:i+2]))\n        \n        draw.line(structure, width=100, fill='Red')\n    return im","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"plt.figure(figsize=(8,8))\nimage_4_with_line = draw_structure(e79de561c_json, image_4)\nplt.imshow(image_4_with_line)","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<a id=\"100\"></a>\n### Anatomical structure file\n\nIn the same way with glomerulus segmentation file, we can show anatomical structure segmentations. This file contains anatomical structure segmentations. They are intended to help us identify the various parts of the tissue.","metadata":{}},{"cell_type":"code","source":"with open(f\"../input/hubmap-kidney-segmentation/train/{img_id_1}-anatomical-structure.json\") as f:\n    anatomical_structure_json = json.load(f)\n    \nanatomical_structure_json","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"plt.figure(figsize=(8,8))\nimage_1_with_line = draw_structure(anatomical_structure_json, image_1)\nplt.imshow(image_1_with_line)","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<a id=\"4\"></a> <br>\n# <div class=\"alert alert-block alert-success\">Visualization of HuBMAP-20-dataset_information</div>\n\nI'll try to visualize HuBMAP-20-dataset_information to easy understand.","metadata":{}},{"cell_type":"code","source":"ds_info.head()","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"ds_info.shape","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"There are 20 data. Each data has 16 colmuns.\n\n15 data are for training, and rest are test. It includes anonymized patient data.","metadata":{}},{"cell_type":"code","source":"def train_or_test(image_file):\n    id, _ = image_file.split(\".\")\n    if id in list(train[\"id\"]):\n        return \"train\"\n    else:\n        return \"test\"\n    \nds_info[\"category\"] = ds_info[\"image_file\"].map(train_or_test)","metadata":{"_kg_hide-input":true,"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"plt.style.use(\"Solarize_Light2\")","metadata":{"_kg_hide-input":true,"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"plt.figure(figsize=(15, 5))\ng = sns.countplot(data=ds_info, x=\"patient_number\", hue=\"category\", palette=sns.color_palette(\"Set2\", 8))\ng.set_title(\"Number of images per patient\")","metadata":{"_kg_hide-input":true,"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"fig, axes = plt.subplots(1, 2, figsize=(10,5), gridspec_kw=dict(wspace=0.1, hspace=0.6))\nfig.suptitle(\"race and ethnicity\", fontsize=15)\ng = sns.countplot(data=ds_info, x=\"race\", hue=\"category\", palette=sns.color_palette(\"Set2\", 8),ax=axes[0])\ng.set_title(\"distribution of race\", fontsize=12)\ng = sns.countplot(data=ds_info, x=\"ethnicity\", palette=sns.color_palette(\"Set2\", 8), hue=\"category\",ax=axes[1])\ng.set_title(\"distribution of ethnicity\", fontsize=12)","metadata":{"_kg_hide-input":true,"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"#Create figure and Axes. And set title.\nfig, axes = plt.subplots(2, 2, figsize=(10,6), gridspec_kw=dict(wspace=0.1, hspace=0.6))\nfig.suptitle(\"Sex and age\", fontsize=15)\n\n#Too check layout, I'll show text on each Axes.\ngs = axes[0, 1].get_gridspec()\naxes[0, 0].remove()\naxes[1, 0].remove()\n#Add gridspec we got\naxbig = fig.add_subplot(gs[:, 0])\n\ng = sns.countplot(data=ds_info, x=\"sex\", hue=\"category\", palette=sns.color_palette(\"Set2\", 8),ax=axbig)\ng.set_title(\"distribution of sex\", fontsize=12)\n\n#Add three plots.\ng = sns.distplot(ds_info[ds_info[\"category\"]==\"train\"][\"age\"], color=\"tomato\", kde=False, rug=False,ax=axes[0,1])\ng.set(xlim=(30,80))\ng.set(ylim=(0,3))\ng.set_title(\"distribution of age for train\", fontsize=12)\n\ng = sns.distplot(ds_info[ds_info[\"category\"]==\"test\"][\"age\"], color=\"teal\", kde=False, rug=False, ax=axes[1,1])\ng.set(xlim=(30,80))\ng.set(ylim=(0,3))\ng.set_title(\"distribution of age for test\", fontsize=12)","metadata":{"_kg_hide-input":true,"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"fig, axes = plt.subplots(2, 2, figsize=(10,10), gridspec_kw=dict(wspace=0.1, hspace=0.4))\nfig.suptitle(\"Physical information for train\", fontsize=15)\n\n\ng = sns.distplot(ds_info[ds_info[\"category\"]==\"train\"][\"weight_kilograms\"], color=\"tomato\", kde=False, rug=False, ax=axes[0,0])\ng.set(xlim=(55,135))\ng.set(ylim=(0,5))\ng.set_title(\"weight_kilograms\", fontsize=12)\n\ng = sns.distplot(ds_info[ds_info[\"category\"]==\"train\"][\"height_centimeters\"], color=\"tomato\", kde=False, rug=False, ax=axes[0,1])\ng.set(xlim=(155,195))\ng.set(ylim=(0,5))\ng.set_title(\"height_centimeters\", fontsize=12)\n\ng = sns.distplot(ds_info[ds_info[\"category\"]==\"train\"][\"bmi_kg/m^2\"], color=\"tomato\", kde=False, rug=False, ax=axes[1,0])\ng.set(xlim=(22,37.5))\ng.set(ylim=(0,5))\ng.set_title(\"bmi_kg/m^2\", fontsize=12)\n\ng = sns.countplot(ds_info[ds_info[\"category\"]==\"train\"][\"laterality\"], ax=axes[1,1])\ng.set_title(\"laterality\", fontsize=12)\n\n\nfig, axes = plt.subplots(2, 2, figsize=(10,10), gridspec_kw=dict(wspace=0.1, hspace=0.4))\nfig.suptitle(\"Physical information for test\", fontsize=15)\n\n\ng = sns.distplot(ds_info[ds_info[\"category\"]==\"test\"][\"weight_kilograms\"], color=\"teal\", kde=False, rug=False, ax=axes[0,0])\ng.set(xlim=(55,135))\ng.set(ylim=(0,5))\ng.set_title(\"weight_kilograms\", fontsize=12)\n\ng = sns.distplot(ds_info[ds_info[\"category\"]==\"test\"][\"height_centimeters\"], color=\"teal\", kde=False, rug=False, ax=axes[0,1])\ng.set(xlim=(155,195))\ng.set(ylim=(0,5))\ng.set_title(\"height_centimeters\", fontsize=12)\n\ng = sns.distplot(ds_info[ds_info[\"category\"]==\"test\"][\"bmi_kg/m^2\"], color=\"teal\", kde=False, rug=False, ax=axes[1,0])\ng.set(xlim=(22,37.5))\ng.set(ylim=(0,5))\ng.set_title(\"bmi_kg/m^2\", fontsize=12)\n\ng = sns.countplot(ds_info[ds_info[\"category\"]==\"test\"][\"laterality\"], ax=axes[1,1])\ng.set_title(\"laterality\", fontsize=12)","metadata":{"_kg_hide-input":true,"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"ds_info[\"Ratio_of_medulla_to_cortex\"] = ds_info[\"percent_medulla\"] / ds_info[\"percent_cortex\"] ","metadata":{"_kg_hide-input":true,"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"fig, axes = plt.subplots(1, 2, figsize=(10,5), gridspec_kw=dict(wspace=0.1, hspace=0.6))\nfig.suptitle(\"distribution of ratio of medulla to cortex\", fontsize=15)\ng = sns.distplot(ds_info[ds_info[\"category\"]==\"train\"][\"Ratio_of_medulla_to_cortex\"], color=\"tomato\",kde=False, rug=False, ax=axes[0])\ng.set(ylim=(0,5))\ng.set_title(\"train\", fontsize=12)\ng = sns.distplot(ds_info[ds_info[\"category\"]==\"test\"][\"Ratio_of_medulla_to_cortex\"], color=\"teal\", kde=False, rug=False, ax=axes[1])\ng.set(ylim=(0,5))\ng.set_title(\"test\", fontsize=12)","metadata":{"_kg_hide-input":true,"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"------------\n<a id=\"101\"></a> <br>\n## Reference\n\n[1] https://en.wikipedia.org/wiki/Glomerulus_(kidney) \n\n[2] https://www.kaggle.com/c/hubmap-kidney-segmentation/discussion/197552","metadata":{}}]}