{"cells":[{"metadata":{},"cell_type":"markdown","source":"In this competition, we deal with TIFF images. I've never worked with this format before so my first question was: what library should I use to read this format? The most obvious choices were `Pillow` and `tifffile` packages. However, they didn't work: the former failed with `DecompressionBombError` (!), while the latter complained about NumPy array being non-contiguous for some image in the dataset.\n\n[After some search](https://stackoverflow.com/questions/7569553/working-with-tiffs-import-export-in-python-using-numpy), I stumbled upon the `osgeo` library. As the name suggests, the library was probably written to work with maps, satellite images, or something like that. But it has TIFF reading functionality built-in so why not to apply it for medical imagery, right? \n\nHere I show my approach to read the dataset images. This approach worked out for me right out of the box, without any errors that I encountered with other libraries. But if you know a better/faster way, please let me know!"},{"metadata":{"trusted":true},"cell_type":"code","source":"import gc\nimport glob\nimport json\nfrom os.path import basename, dirname, splitext\nimport numpy as np\nfrom osgeo import gdal","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"def human_readable_size(arr: np.ndarray) -> str:\n    \"\"\"Gets array's size as a verbose, human-readable string.\"\"\"\n    \n    n = arr.nbytes\n    for unit in ('bytes', 'Kb', 'Mb', 'Gb'):\n        if n >= 1024:\n            n /= 1024\n        else:\n            break\n    return f'{n:.3f} {unit}'","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"def read_tiff(path: str) -> np.ndarray:\n    \"\"\"Reads TIFF file.\"\"\"\n    \n    dataset = gdal.Open(path, gdal.GA_ReadOnly)\n    n_channels = dataset.RasterCount\n    width = dataset.RasterXSize\n    height = dataset.RasterYSize\n    image = np.zeros((n_channels, height, width), dtype=np.uint8)\n    for i in range(n_channels):\n        band = dataset.GetRasterBand(i+1)\n        channel = band.ReadAsArray()\n        image[i] = channel\n    return image","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"meta = []\n\nfor filename in glob.glob('/kaggle/input/hubmap-kidney-segmentation/**/*.tiff'):\n    print(f'Processing file: {filename}')\n    identifier, _ = splitext(basename(filename))\n    subset = basename(dirname(filename))\n    img = read_tiff(filename)\n    meta.append(dict(\n        identifier=identifier,\n        filename=filename,\n        subset=subset,\n        memory_bytes=img.nbytes,\n        memory_human_readable=human_readable_size(img),\n        image_shape=img.shape\n    ))\n    del img\n    gc.collect()","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"for info in meta:\n    print(\n        f'id: {info[\"identifier\"]}, '\n        f'memory: {info[\"memory_human_readable\"]:>10s}, '\n        f'shape: {info[\"image_shape\"]}'\n    )","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"with open('/kaggle/working/meta.json', 'w') as fp:\n    json.dump(meta, fp)","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"","execution_count":null,"outputs":[]}],"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"pygments_lexer":"ipython3","nbconvert_exporter":"python","version":"3.6.4","file_extension":".py","codemirror_mode":{"name":"ipython","version":3},"name":"python","mimetype":"text/x-python"}},"nbformat":4,"nbformat_minor":4}