{"cells":[{"metadata":{},"cell_type":"markdown","source":"# HuBMAP - an analysis and methodology\n\n---\n\nHello there, and welcome to my notebook. This is a simple analysis and a potential guidelines for the HuBMAP: Hacking the Kidney competition currently on Kaggle. \n\nThis is the first time doing image segmentation on Kaggle so I need heavy feedback as to what I'm doing - so let's get started with a quick run-through of the images and data."},{"metadata":{},"cell_type":"markdown","source":"## Primary setup - load 256x256 images"},{"metadata":{"_kg_hide-input":true,"_kg_hide-output":true,"trusted":true},"cell_type":"code","source":"!unzip ../input/256x256-images/train.zip ","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"As you can see GM iafoss has processed the images into 256x256 tiles which makes it far easier to perform analysis on them. This is merely a preliminary step before carrying on."},{"metadata":{},"cell_type":"markdown","source":"# Introduction (from hosts)\n\nThis competition, “Hacking the Kidney,\" starts by mapping the human kidney at single cell resolution.\n\nYour challenge is to detect functional tissue units (FTUs) across different tissue preparation pipelines. An FTU is defined as a “three-dimensional block of cells centered around a capillary, such that each cell in this block is within diffusion distance from any other cell in the same block” (de Bono, 2013). The goal of this competition is the implementation of a successful and robust glomeruli FTU detector.\n\nYou will also have the opportunity to present your findings to a panel of judges for additional consideration. Successful submissions will construct the tools, resources, and cell atlases needed to determine how the relationships between cells can affect the health of an individual.\n\nAdvancements in HuBMAP will accelerate the world’s understanding of the relationships between cell and tissue organization and function and human health. These datasets and insights can be used by researchers in cell and tissue anatomy, pharmaceutical companies to develop therapies, or even parents to show their children the magnitude of the human body."},{"metadata":{},"cell_type":"markdown","source":"# Getting started"},{"metadata":{},"cell_type":"markdown","source":"First of all, we have to import the required libraries. "},{"metadata":{"_uuid":"8f2839f25d086af736a60e9eeb907d3b93b6e0e5","_cell_guid":"b1076dfc-b9ad-4769-8c92-a6c4dae69d19","trusted":true,"_kg_hide-input":true},"cell_type":"code","source":"import numpy as np, pandas as pd\nimport matplotlib.pyplot as plt, seaborn as sns\nimport cv2, tifffile\nimport os;list_ims = os.listdir('../working')\nfrom plotly.offline import iplot\nimport cufflinks\ncufflinks.go_offline()\ncufflinks.set_config_file(world_readable=True, theme='pearl')\n\n\"\"\"From https://www.kaggle.com/kool777/hubmap-extensive-eda\"\"\"\n\ndef mask2rle(img):\n    '''\n    img: numpy array, 1 - mask, 0 - background\n    Returns run length as string formated\n    '''\n    pixels= img.T.flatten()\n    pixels = np.concatenate([[0], pixels, [0]])\n    runs = np.where(pixels[1:] != pixels[:-1])[0] + 1\n    runs[1::2] -= runs[::2]\n    return ' '.join(str(x) for x in runs)\n \ndef rle2mask(mask_rle, shape=(1600,256)):\n    '''\n    mask_rle: run-length as string formated (start length)\n    shape: (width,height) of array to return \n    Returns numpy array, 1 - mask, 0 - background\n\n    '''\n    s = mask_rle.split()\n    starts, lengths = [np.asarray(x, dtype=int) for x in (s[0:][::2], s[1:][::2])]\n    starts -= 1\n    ends = starts + lengths\n    img = np.zeros(shape[0]*shape[1], dtype=np.uint8)\n    for lo, hi in zip(starts, ends):\n        img[lo:hi] = 1\n    return img.reshape(shape).T","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"Now we basically have to view an assortment of tiles prepared for us. "},{"metadata":{"_kg_hide-input":true,"trusted":true},"cell_type":"code","source":"def view_images(images, title = '', aug = None):\n    width = 6\n    height = 4\n    fig, axs = plt.subplots(height, width, figsize=(15,5))\n    for im in range(0, height * width):  \n        data = cv2.imread(images[im])\n        i = im // width\n        j = im % width\n        axs[i,j].imshow(data, cmap=plt.cm.bone) \n        axs[i,j].axis('off')\n\n    plt.suptitle(title)\n    plt.show()","execution_count":null,"outputs":[]},{"metadata":{"_kg_hide-input":true,"trusted":true},"cell_type":"code","source":"view_images(list_ims, title=\"First 20 256x256 tiles\")","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"Now we have to view a giant image at a glance, the image being the full scan from which we have obtained these small chunks. Also prepare to deal with memory issues if you want to use the huge image, it takes quite some time to read :-)"},{"metadata":{"_uuid":"d629ff2d2480ee46fbb7e2d37f6b5fab8052498a","_cell_guid":"79c7e3d0-c299-4dcb-8224-4455121ee9b0","trusted":true,"_kg_hide-input":true},"cell_type":"code","source":"im = tifffile.imread('../input/hubmap-kidney-segmentation/train/e79de561c.tiff')\nplt.figure(figsize=(16, 16))\nplt.imshow(im[0, 0, :, :, :].transpose(1, 2, 0))\nplt.axis(\"off\");","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"Zoom in, so that we can see some portions of the scan with better clarity and if we can potentially observe any shifts in the image as have been pointed out in the public discussion forum."},{"metadata":{"_kg_hide-input":true,"trusted":true},"cell_type":"code","source":"plt.figure(figsize=(10, 10))\nplt.imshow(im[0, 0, :, :, :].transpose(1, 2, 0)[7000:8500, 7000:8500])\nplt.axis(\"off\");","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"So you might be as confused as I am when it comes to these little things here, and apparently this is a zoomed-in visual of several glomeruli. Glomeruli are a tuft of blood capillaries in essence, and they act as a sieve, a filter for your blood as it passes through the nephrons in your kidney."},{"metadata":{"trusted":true},"cell_type":"code","source":"train = pd.read_csv('../input/hubmap-kidney-segmentation/train.csv')\nexample_mask = rle2mask(train['encoding'][4], (im.shape[4], im.shape[3]))","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"We can now try to view the image with the annotations created by Run Length Encoding to generate a mask for the image."},{"metadata":{"trusted":true},"cell_type":"code","source":"plt.figure(figsize=(15, 15))\nplt.imshow(im[0, 0, :, :, :].transpose(1, 2, 0))\nplt.imshow(example_mask, alpha=0.5)\nplt.title(\"Image + Mask\", fontsize=16);\nplt.axis(\"off\");","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"Zoom in and see if we can find any potential irregularities in the image?"},{"metadata":{"_kg_hide-input":true,"trusted":true},"cell_type":"code","source":"plt.figure(figsize=(15, 15))\nplt.imshow(im[0, 0, :, :, :].transpose(1, 2, 0)[7000:8500, 7000:8500])\nplt.imshow(example_mask[7000:8500, 7000:8500], alpha=0.5)\nplt.title(\"Image + Mask\", fontsize=16);\nplt.axis(\"off\");","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"Again we can see a shift between the image and the mask, so this will be another inaccuracy to deal with while modelling. On the forums it is estimated to be about 50 px which seems like a reasonable estimate."},{"metadata":{},"cell_type":"markdown","source":"# Metadata"},{"metadata":{},"cell_type":"markdown","source":"This is a brief analysis of the metadata files contained, there is very little but we can utilize it to get a better sense of what we are dealing with."},{"metadata":{},"cell_type":"markdown","source":"## Analysis by Race"},{"metadata":{"_kg_hide-input":true,"trusted":true},"cell_type":"code","source":"m = pd.read_csv('../input/hubmap-kidney-segmentation/HuBMAP-20-dataset_information.csv')\nm['race'].value_counts(normalize=True).iplot(kind='bar',\n                                                      yTitle='Race', \n                                                      linecolor='black', \n                                                      opacity=0.7,\n                                                      color='red',\n                                                      theme='pearl',\n                                                      bargap=0.8,\n                                                      gridcolor='white',\n                                                      title='Distribution of the Race column')","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"## Analysis by Gender"},{"metadata":{"trusted":true},"cell_type":"code","source":"m['sex'].value_counts(normalize=True).iplot(kind='bar',\n                                                      yTitle='Gender', \n                                                      linecolor='black', \n                                                      opacity=0.7,\n                                                      color='steelblue',\n                                                      theme='pearl',\n                                                      bargap=0.8,\n                                                      gridcolor='white',\n                                                      title='Distribution of the Sex column')","execution_count":null,"outputs":[]}],"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"pygments_lexer":"ipython3","nbconvert_exporter":"python","version":"3.6.4","file_extension":".py","codemirror_mode":{"name":"ipython","version":3},"name":"python","mimetype":"text/x-python"}},"nbformat":4,"nbformat_minor":4}