{"cells":[{"metadata":{},"cell_type":"markdown","source":"<h2 style='background:purple; border:0; color:white;font-size:2em'><center> HuBMAP Hacking the Kidney </center></h2>\n\n<img src=\"https://storage.googleapis.com/kaggle-competitions/kaggle/22990/logos/header.png\" alt=\"HuBMAP\">"},{"metadata":{"_uuid":"8f2839f25d086af736a60e9eeb907d3b93b6e0e5","_cell_guid":"b1076dfc-b9ad-4769-8c92-a6c4dae69d19","trusted":true},"cell_type":"markdown","source":"1. [Competition purpose](#1)\n2. [Data Overview](#2)\n3. [Tiles](#3)\n\nThis notebook is using this notebook for the tiles making part : https://www.kaggle.com/iafoss/256x256-images \n\n\n\n\n<a id='1'></a>\n<h2 style='background:green; border:0; color:white;font-size:1.5em'><center> Competition purpose </center></h2>\n\n> Your challenge is to detect functional tissue units (FTUs) across different tissue preparation pipelines. An FTU is defined as a “three-dimensional block of cells centered around a capillary, such that each cell in this block is within diffusion distance from any other cell in the same block” (de Bono, 2013). The goal of this competition is the implementation of a successful and robust glomeruli FTU detector.\n\nThe dataset is comprised of 8 very large (>500MB - 5GB) TIFF files is huge reason why we are gonna use tiles. 8 big images are not fit for deep neural networks ! The training set includes annotations in both RLE-encoded and unencoded (JSON) forms. The annotations denote segmentations of glomeruli.\n\n\n## What we are prediciting?\n\nDevelop segmentation algorithms that identify glomeruli in the PAS stained microscopy data. Detect functional tissue units (FTUs) across different tissue preparation pipelines\n\n\n## Evaluation Metric: Dice Coefficient\n\nDice Coefficient is common in case our task involve **segmentation**. The Dice coefficient can be used to compare the pixel-wise agreement between a predicted segmentation and its corresponding ground truth. the Dice similarity coefficient for two sets X and Y is defined as:\n\n$$\\text{DC}(X, Y) = \\frac{2 \\times |X \\cap Y|}{|X| + |Y|}.$$\n\nwhere X is the predicted set of pixels and Y is the ground truth.\n\n<a id='2'></a>\n<h2 style='background:green; border:0; color:white;font-size:1.5em'><center> Overview </center></h2>\n"},{"metadata":{"_uuid":"d629ff2d2480ee46fbb7e2d37f6b5fab8052498a","collapsed":true,"_cell_guid":"79c7e3d0-c299-4dcb-8224-4455121ee9b0","trusted":false,"_kg_hide-input":true,"_kg_hide-output":true},"cell_type":"code","source":"# basic\nimport os, gc\nimport warnings\nimport numpy as np\nimport pandas as pd\n\n# visualize\nimport seaborn as sns\nimport matplotlib.pyplot as plt\n\n# reading tiff images\nimport tifffile as tiff \n\n%matplotlib inline\nwarnings.filterwarnings('ignore')\n\n# directory\nROOT = '../input/hubmap-kidney-segmentation/'","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"We have a **train and a test folders with .tiff images and annotations in JSON**. We have **train.csv** and **HuBMAP-20-dataset_information.csv** containing image, masks information and metadata respectively. \n\nWe have 8 images in the train dataset, with 2 caracteristics in the `train.csv`: id and encoding (RLE-encoded representation of the mask)."},{"metadata":{"trusted":true},"cell_type":"code","source":"train = pd.read_csv(f'{ROOT}train.csv')\ntrain","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"We have some additional metadata in the `HuBMAP-20-dataset_information.csv` file such as sex, age, weight, height, laterality, percentage of cortex and medulla within the kidney. Note that this file provides also metadata for the test dataset."},{"metadata":{"trusted":true},"cell_type":"code","source":"metadata = pd.read_csv(f'{ROOT}HuBMAP-20-dataset_information.csv')\nmetadata.head()","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"Finally, the images :"},{"metadata":{"trusted":true},"cell_type":"code","source":"example_image = tiff.imread(os.path.join(ROOT, 'train/2f6ecfcdf.tiff'))\nplt.figure(figsize=(16, 16))\nplt.imshow(example_image)\nplt.axis(\"off\")\nprint(f'Image Shape: {example_image.shape}')","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"<a id='3'></a>\n<h2 style='background:green; border:0; color:white;font-size:1.5em'><center> Tiles </center></h2>\n\nWe use the dataset used in this brilliant notebook : https://www.kaggle.com/iafoss/256x256-images\n\nWe first load the image tiles :"},{"metadata":{"trusted":true},"cell_type":"code","source":"TRAIN = '../input/256256-hubmap/train/'\nMASKS = '../input/256256-hubmap/masks/'\ntrain_images = os.listdir(TRAIN)\n\nfrom PIL import Image\nimport numpy as np\n\nplt.figure(figsize=(15,15))\nfor i in range(25):\n    plt.subplot(5,5,i+1)\n    img = Image.open(TRAIN + train_images[i])\n    img = np.array(img.getdata())\n    img = img.reshape((256,256,3))\n    plt.imshow(img)\n    plt.title(train_images[i])\n    plt.grid(False)\n    plt.axis(False)\nplt.show()","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"Then the corresponding masks :"},{"metadata":{"trusted":true},"cell_type":"code","source":"train_masks = os.listdir(MASKS)\nplt.figure(figsize=(15,15))\nfor i in range(25):\n    plt.subplot(5,5,i+1)\n    img = Image.open(MASKS + train_masks[i])\n    img = np.array(img.getdata())\n    img = img.reshape((256,256))\n    plt.imshow(img)\n    plt.title(train_images[i])\n    plt.grid(False)\n    plt.axis(False)\nplt.show()","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"And then we superpose the two :"},{"metadata":{"trusted":true},"cell_type":"code","source":"plt.figure(figsize=(15,15))\nfor i in range(25):\n    plt.subplot(5,5,i+1)\n    Image2_mask = Image.open(MASKS + train_masks[i])\n    Image2_mask = np.array(Image2_mask.getdata())\n    Image2_mask = Image2_mask.reshape((256,256))\n    img = Image.open(TRAIN + train_images[i])\n    img = np.array(img.getdata())\n    img = img.reshape((256,256,3))\n    plt.imshow(img)\n    plt.imshow(Image2_mask, alpha=0.5)\n    plt.title(train_images[i])\n    plt.grid(False)\n    plt.axis(False)\nplt.show()\n","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"# TO BE CONTINUED..."}],"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"pygments_lexer":"ipython3","nbconvert_exporter":"python","version":"3.6.4","file_extension":".py","codemirror_mode":{"name":"ipython","version":3},"name":"python","mimetype":"text/x-python"}},"nbformat":4,"nbformat_minor":4}