{"cells":[{"metadata":{},"cell_type":"markdown","source":"* The data used in this competition includes 11 fresh frozen and 9 Formalin Fixed Paraffin Embedded (FFPE) PAS kidney images. \n* Glomeruli FTU annotations exist for all 20 tissue samples; some of these will be shared for training, and others will be used to judge submissions."},{"metadata":{},"cell_type":"markdown","source":"The Dataset is comprised of very large TIFF files.\n\n* The **training set** has **8** files.\n* The **public test set** has **5** files.\n* The **private test set** is larger than the public test set. I suppose there will be **7** files. \n\nThe train set includes annotations in both RLE-encoded and unencoded(JSON) forms. The annotations denote segmentations of glomeruli.\n\nBoth training and public test sets include anatomical structure segmentations. I suppose this can be used for pretraining."},{"metadata":{},"cell_type":"markdown","source":"JSON files are structured as follows\n* A `type` (`Feature`) and object type id (`PathAnnotationObject`). Note that these fields are the same between all files and do not offer signal.\n* A `geometry` containing a `Polygon` with `coordinates` for the feature's enclosing volume\n* Additional `properties`, including the name and color of the feature in the image.\n* The `IsLocked` field is the same across file types (locked for glomerulus, unlocked for anatomical structure) and is not signal-bearing."},{"metadata":{"_uuid":"d629ff2d2480ee46fbb7e2d37f6b5fab8052498a","_cell_guid":"79c7e3d0-c299-4dcb-8224-4455121ee9b0","trusted":true,"_kg_hide-input":true},"cell_type":"code","source":"import numpy as np\nimport pandas as pd\nimport matplotlib.pyplot as plt\nfrom PIL import Image\nimport tifffile as tiff\nimport cv2\nimport os\nfrom tqdm.notebook import tqdm\nimport zipfile","execution_count":null,"outputs":[]},{"metadata":{"_kg_hide-input":true,"trusted":true},"cell_type":"code","source":"from pathlib import Path\nPath.ls = lambda x: list(x.iterdir())","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"path = Path('/kaggle/input/hubmap-kidney-segmentation/')\npath.ls()","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"train_df = pd.read_csv(path/'train.csv')","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"# Understanding RLE"},{"metadata":{},"cell_type":"markdown","source":"The masks provided in the `train.csv` is in Running Length Encoding format. This encoding comes in pairs of pixel values as follows:\n1. The starting pixel.\n2. Number of pixels from the starting pixel. \n\nSo, to specify 10 pixels starting from pixel number 200 would be written as:\n>200 10\n\nAlso, the pixels are numbered from top to bottom and the left to right. This looks as follows:"},{"metadata":{"_kg_hide-input":true,"trusted":true},"cell_type":"code","source":"np.arange(0,25).reshape(5,5).T","execution_count":null,"outputs":[]}],"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"pygments_lexer":"ipython3","nbconvert_exporter":"python","version":"3.6.4","file_extension":".py","codemirror_mode":{"name":"ipython","version":3},"name":"python","mimetype":"text/x-python"}},"nbformat":4,"nbformat_minor":4}