{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"pygments_lexer":"ipython3","nbconvert_exporter":"python","version":"3.6.4","file_extension":".py","codemirror_mode":{"name":"ipython","version":3},"name":"python","mimetype":"text/x-python"}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"# Introduction\n> 📌**Note**: <br> This is a work in progress. If you liked or forked this notebook, please consider upvoting ⬆️⬆️ It encourages to keep posting relevant content\n\n\n<div style=\"border-radius:10px; border: #babab5 solid; padding: 15px; background-color: #e6f9ff; font-size:100%; \">\n    \nThis notebook covers the following:\n* General introduction    \n* Animation of the different kidney slices\n* Exploratory data analysis (WIP)","metadata":{"id":"QVc0j1rEi7x7"}},{"cell_type":"markdown","source":"# Competition Description\n<!-- <img src =\"https://www.energy.gov/sites/default/files/styles/full_article_width/public/Prosumer-Blog%20sans%20money-%201200%20x%20630-01_0.png?itok=2a3YSkUb\" width=600>\n\n> 📌**Note**:  Energy prosumers are individuals, businesses, or organizations that both consume and produce energy. This concept represents a shift from the traditional model where consumers simply purchase energy from utilities and rely on centralized power generation sources. Energy prosumers are actively involved in the energy ecosystem by generating their own electricity, typically through renewable energy sources like solar panels (or wind turbines, small-scale hydropower etc.). They also consume energy from the grid when their own generation is insufficient to meet their needs -->\n\n<div style=\"border-radius:10px; border: #babab5 solid; padding: 15px; background-color: #e6f9ff; font-size:100%; \">\n    \n* The goal of this competition is to segment blood vessels. You will create a model trained on 3D Hierarchical Phase-Contrast Tomography (HiP-CT) data from human kidneys to help complete a picture of vasculature throughout a body.\n    \n* Your work will better researchers' understanding of the size, shape, branching angles, and patterning of blood vessels in human tissue","metadata":{"id":"GfT3ur3ai7yB"}},{"cell_type":"markdown","source":"# Data Description\n> 📌**Note**:  This competition dataset comprises high-resolution 3D images of several kidneys together with 3D segmentation masks of their vasculature. Your task is to create segmentation masks for the kidney datasets in the test set. <br> <br>\nThe kidney images were obtained through Hierarchical Phase-Contrast Tomography (HiP-CT) imaging. HiP-CT is an imaging technique that obtains high-resolution (from 1.4 micrometers - 50 micrometers resolution) 3D data from ex vivo organs. See this Nature Methods article for more information.\n\n<div style=\"border-radius:10px; border: #babab5 solid; padding: 15px; background-color: #e6f9ff; font-size:100%; \">\n    \n# File and Field Information\n\n- **train/{dataset}/images** - Contains TIFF scans from several kidney datasets. Each image represents a 2D slice of a 3D volume. The slices run along the z-axis, with files enumerated from top to bottom. (The slices should be stacked vertically or depth-wise, in other words.)\n    \n- **train/{dataset}/labels** - Contains blood vessel segmentation masks in TIFF format for the images.  \n    The **{dataset}** folders comprise the following:\n    \n    - `kidney_1_dense` - The whole of a right kidney at 50um resolution. The entire 3D arterial vascular tree has been densely segmented, down to two generations from the glomeruli (i.e. the capillary bed). Uses beamline BM05.\n    - `kidney_1_voi` - A high-resolution subset of `kidney_1`, at 5.2um resolution.\n    - `kidney_2` - The whole of a kidney from another donor, at 50um resolution. Sparsely segmented (about 65%).\n    - `kidney_3_dense` - A portion (500 slices) of a kidney at 50.16um resolution using BM05. Densely segmented. Note that we provide all of the images for `kidney_3` in the `kidney_3_sparse/images` folder. This dataset accordingly has only a `labels` folder.\n    - `kidney_3_sparse` - The remainder of the segmentation masks for `kidney_3`. Sparsely segmented (about 85%).\n- **test/{dataset}/images** - Contains the TIFF scans for the test set. These scans may or may not use a different beamline or resolution from the scans used in the training set. The names of the datasets are `kidney_5` and `kidney_6`.\n    \n- **train\\_rles.csv** - The run length encoded segmentation masks for images in the training set.\n    \n    - `id` - A unique identifier for each slice, in the form `{dataset}_{slice}`.\n    - `rle` - The run length encoded mask for this slice.\n- **sample\\_submission.csv** - A sample submission file in the correct format. See the [Evaluation](https://www.kaggle.com/competitions/blood-vessel-segmentation/overview/evaluation) page for more details.\n    \n\nYou may find functions to encode and decode the segmentation masks in [this notebook](https://www.kaggle.com/code/paulorzp/run-length-encode-and-decode/script).\n\nPlease note that this is a [**Code Competition**](https://www.kaggle.com/competitions/blood-vessel-segmentation/overview/code-requirements). We give a few example images in the `test/` folder to help you author your submissions. These example images are not meant to be representative of the test set with regard to scan resolution, beamline, or other such qualities. When your submission is scored, this example test data will be replaced with the full test set. The full test set contains about 1500 TIFF images.","metadata":{"id":"jgy6HQ-Wi7yC"}},{"cell_type":"markdown","source":"# Install & imports","metadata":{"id":"aBUs90zXi7yD"}},{"cell_type":"code","source":"#General\nimport pandas as pd\nimport numpy as np\nimport json\nimport os\nfrom colorama import Fore, Style, init;\n\n# Visualization\nimport seaborn as sns\nimport matplotlib.pyplot as plt\nfrom PIL import Image\nimport imageio\nimport cv2\nimport matplotlib.animation as animation\nfrom IPython.display import HTML\nfrom colorama import Fore, Style, init;\n\n\n# Modeling\nimport torch\n\n# Options\npd.set_option('display.max_columns', 100)","metadata":{"_uuid":"8f2839f25d086af736a60e9eeb907d3b93b6e0e5","_cell_guid":"b1076dfc-b9ad-4769-8c92-a6c4dae69d19","id":"xsbKmUBpi7yE","execution":{"iopub.status.busy":"2023-11-09T19:59:47.619084Z","iopub.execute_input":"2023-11-09T19:59:47.619480Z","iopub.status.idle":"2023-11-09T19:59:51.851273Z","shell.execute_reply.started":"2023-11-09T19:59:47.619451Z","shell.execute_reply":"2023-11-09T19:59:51.850033Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Helper functions\ndef display_df(df, name):\n    '''Display df shape and first row '''\n    PrintColor(text = f'{name} data has {df.shape[0]} rows and {df.shape[1]} columns. \\n ===> First row:')\n    display(df.head(1))\n\n# Color printing\ndef PrintColor(text:str, color = Fore.BLUE, style = Style.BRIGHT):\n    '''Prints color outputs using colorama of a text string'''\n    print(style + color + text + Style.RESET_ALL);","metadata":{"id":"GC5yQhVMi7yG","execution":{"iopub.status.busy":"2023-11-09T19:59:51.853100Z","iopub.execute_input":"2023-11-09T19:59:51.853577Z","iopub.status.idle":"2023-11-09T19:59:51.860012Z","shell.execute_reply.started":"2023-11-09T19:59:51.853547Z","shell.execute_reply":"2023-11-09T19:59:51.858891Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# CFG\nDEBUG = False\nCSV_DIR = \"/kaggle/input/blood-vessel-segmentation/\"\nSCAN_DIR_TRAIN = \"/kaggle/input/blood-vessel-segmentation/train/\"\nSCAN_DIR_TEST = \"/kaggle/input/blood-vessel-segmentation/test/\"","metadata":{"id":"P7LejQ7Ki7yH","execution":{"iopub.status.busy":"2023-11-09T19:59:51.861478Z","iopub.execute_input":"2023-11-09T19:59:51.861900Z","iopub.status.idle":"2023-11-09T19:59:51.871514Z","shell.execute_reply.started":"2023-11-09T19:59:51.861859Z","shell.execute_reply":"2023-11-09T19:59:51.870419Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Read CSVs and parse relevant date columns\ntrain = pd.read_csv(CSV_DIR + \"train_rles.csv\")\nsample_sub = pd.read_csv(CSV_DIR + \"sample_submission.csv\")\n\n# Display\ndisplay_df(train, 'train_rles')\ndisplay_df(sample_sub, 'sample_submission')","metadata":{"id":"nOO7ArAji7yI","outputId":"e4ea8584-346b-496d-dcf1-4538d28d5f62","execution":{"iopub.status.busy":"2023-11-09T19:59:51.874330Z","iopub.execute_input":"2023-11-09T19:59:51.874810Z","iopub.status.idle":"2023-11-09T19:59:52.963351Z","shell.execute_reply.started":"2023-11-09T19:59:51.874779Z","shell.execute_reply":"2023-11-09T19:59:52.962263Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Exploratory data analysis","metadata":{"id":"Vy2l-8Rhi7yJ"}},{"cell_type":"code","source":"# Helper functions\ndef create_animation_with_mask(kidney_scan, slice_interval=50, figsize = (8, 4), cmap_mask='jet', alpha_mask=0.5):\n    '''📹 Create animation for the kidney scan w/ mask, given slice_interval, cmap_mask and alpha 📹'''\n    \n    # Get the image ids\n    scan_path = os.path.join(SCAN_DIR_TRAIN, kidney_scan)\n    img_ids = os.listdir(os.path.join(scan_path, 'labels'))\n    img_ids = sorted(img_ids) # sorts by frame\n\n#     # Print number of images\n#     PrintColor(f'There are {len(img_ids)} images for {kidney_scan}')\n\n    # Select every slice_interval\n    select_ids = img_ids[::slice_interval]\n\n    # Loop over the images and their mask and collect them\n    artists = []\n    fig, ax = plt.subplots(figsize=figsize)\n    for select_id in select_ids:\n        label_file = os.path.join(scan_path, 'labels', select_id) \n        img_file = os.path.join(scan_path, 'images', select_id)\n        \n        # Special case: kidney_3_dense has no images\n        if kidney_scan == 'kidney_3_dense':\n            img_file = img_file.replace('kidney_3_dense', 'kidney_3_sparse') \n\n        img = Image.open(img_file)\n        mask = Image.open(label_file)\n\n        title_str = f\"{kidney_scan} with mask │ scan {select_id.replace('.tif', '')}\"\n\n        title = ax.text(x = 0.5,\n                        y = 1.02,\n                        s = title_str,\n                        size=12,\n                        ha=\"center\",\n                        transform=ax.transAxes\n                       )\n\n        im1 = ax.imshow(img, cmap='gray',animated=True)\n        im2 = ax.imshow(mask, cmap=cmap_mask, alpha=alpha_mask, animated=True)\n        plt.axis(\"off\")\n        plt.close()\n        artists.append([im1, im2, title])\n\n    # Create the animation\n    ani = animation.ArtistAnimation(fig = fig,\n                                    artists = artists,\n                                    interval=200,\n                                    blit = True,\n                                    repeat = True)\n\n    display(HTML(ani.to_jshtml()))\n    \ndef create_subplot_test_scans(kidney_scan, figsize = (16, 16)):\n    '''📈 Create subplot for the few images in test 📈'''\n\n    # Get the image ids\n    scan_path = os.path.join(SCAN_DIR_TEST, kidney_scan)\n    img_ids = os.listdir(os.path.join(scan_path, 'images'))\n    img_ids = sorted(img_ids) # sorts by frame\n    \n    col = 0\n    fig, ax = plt.subplots(1, 3, figsize=figsize)\n    for img_id in img_ids:\n        # Open image\n        img_file = os.path.join(scan_path, 'images', img_id)\n        img = Image.open(img_file)\n        \n        # Plot\n        title_str = f\"{kidney_scan} │ scan {img_id.replace('.tif', '')}\"\n        ax[col].imshow(img, cmap='gray')\n        ax[col].set_title(title_str)\n        ax[col].axis(\"off\")\n        col += 1\n    plt.show()","metadata":{"id":"aBDIg215i7yK","execution":{"iopub.status.busy":"2023-11-09T19:59:52.965171Z","iopub.execute_input":"2023-11-09T19:59:52.965602Z","iopub.status.idle":"2023-11-09T19:59:52.979903Z","shell.execute_reply.started":"2023-11-09T19:59:52.965561Z","shell.execute_reply":"2023-11-09T19:59:52.978672Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## Processing","metadata":{}},{"cell_type":"code","source":"train['scan'] = train['id'].apply(lambda x: x[:-5])\ntrain['mask_is_empty'] = train['id'].apply(lambda x: x[:-5])","metadata":{"execution":{"iopub.status.busy":"2023-11-09T19:59:52.981716Z","iopub.execute_input":"2023-11-09T19:59:52.982172Z","iopub.status.idle":"2023-11-09T19:59:53.001536Z","shell.execute_reply.started":"2023-11-09T19:59:52.982131Z","shell.execute_reply":"2023-11-09T19:59:53.000402Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## Number of images per scan","metadata":{}},{"cell_type":"code","source":"train['scan'] = train['id'].apply(lambda x: x[:-5])\ntrain['mask_is_empty'] = train['rle']=='1 0'","metadata":{"execution":{"iopub.status.busy":"2023-11-09T19:59:53.002770Z","iopub.execute_input":"2023-11-09T19:59:53.003318Z","iopub.status.idle":"2023-11-09T19:59:53.013899Z","shell.execute_reply.started":"2023-11-09T19:59:53.003286Z","shell.execute_reply":"2023-11-09T19:59:53.012856Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"fig, ax = plt.subplots(figsize=(8,4))\nsns.countplot(\n              data = train, \n              y = 'scan',  \n              palette=\"hls\",\n              ax = ax\n             )\nplt.title('Number of images per scan')\n\n# Add percentage and absolute value to plot\npatches = ax.patches\nfor patch in patches:\n    # Variables\n    height = patch.get_height() \n    width = patch.get_width()\n    percent = 100*width/len(train)\n    y_pos = patch.get_y()\n    \n    # Text\n    text = f'{int(width)} ({percent:.1f}%)'\n    ax.text(width, y_pos + height/2, text)\nplt.show()","metadata":{"execution":{"iopub.status.busy":"2023-11-09T19:59:53.015382Z","iopub.execute_input":"2023-11-09T19:59:53.015709Z","iopub.status.idle":"2023-11-09T19:59:53.317011Z","shell.execute_reply.started":"2023-11-09T19:59:53.015681Z","shell.execute_reply":"2023-11-09T19:59:53.316201Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## Masks empty vs. non-empty","metadata":{}},{"cell_type":"code","source":"x = 'scan'\ny = 'mask_is_empty'\n\ngb = train.groupby(x)[y].value_counts(normalize=True)\ngb = gb.round(3)*100\ngb = gb.rename('percent').reset_index()\n\n# Plot \nfig, ax = plt.subplots(figsize=(8,4))\n\nsns.barplot(data = gb,\n            x='percent',\n            y=x, #'percent',\n            hue=y,\n            ax = ax,\n            palette = 'hls'\n           )\nplt.title('Mask is empty percentage')\nax.legend(bbox_to_anchor=(1, 1))\nfor patch in ax.patches:\n    height = patch.get_height() \n    y_loc = patch.get_y()\n    width = patch.get_width()\n    if np.isnan(width):\n        width = 0\n    ax.text(width, y_loc + height/2, f'{width}%')","metadata":{"execution":{"iopub.status.busy":"2023-11-09T20:05:36.520877Z","iopub.execute_input":"2023-11-09T20:05:36.521286Z","iopub.status.idle":"2023-11-09T20:05:36.892645Z","shell.execute_reply.started":"2023-11-09T20:05:36.521255Z","shell.execute_reply":"2023-11-09T20:05:36.891093Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## kidney_1_dense scans\n<div style=\"border-radius:10px; border: #babab5 solid; padding: 15px; background-color: #e6f9ff; font-size:100%; \">\nThe whole of a right kidney at 50um resolution. The entire 3D arterial vascular tree has been densely segmented, down to two generations from the glomeruli (i.e. the capillary bed). Uses beamline BM05.","metadata":{"id":"qOt7qn0Ai7yL"}},{"cell_type":"code","source":"%%time\nif not DEBUG:\n    create_animation_with_mask(kidney_scan = 'kidney_1_dense',\n                                figsize = (8, 8),\n                                slice_interval=50,\n                                cmap_mask='jet',\n                                alpha_mask=0.5)","metadata":{"id":"vy7X1iMwi7yL","execution":{"iopub.status.busy":"2023-11-09T19:59:54.090940Z","iopub.status.idle":"2023-11-09T19:59:54.091349Z","shell.execute_reply.started":"2023-11-09T19:59:54.091165Z","shell.execute_reply":"2023-11-09T19:59:54.091184Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## kidney_1_voi scans\n<div style=\"border-radius:10px; border: #babab5 solid; padding: 15px; background-color: #e6f9ff; font-size:100%; \">\n    \nA high-resolution subset of `kidney_1`, at 5.2um resolution","metadata":{"id":"ja2WG8gki7yM"}},{"cell_type":"code","source":"%%time\nif not DEBUG:\n    create_animation_with_mask(kidney_scan = 'kidney_1_voi',\n                                figsize = (8, 8),\n                                slice_interval=50,\n                                cmap_mask='jet',\n                                alpha_mask=0.5)","metadata":{"id":"E2p1GXAmi7yM","execution":{"iopub.status.busy":"2023-11-09T19:59:54.093064Z","iopub.status.idle":"2023-11-09T19:59:54.093558Z","shell.execute_reply.started":"2023-11-09T19:59:54.093340Z","shell.execute_reply":"2023-11-09T19:59:54.093367Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## kidney_2 scans\n<div style=\"border-radius:10px; border: #babab5 solid; padding: 15px; background-color: #e6f9ff; font-size:100%; \">\n    \nThe whole of a kidney from another donor, at 50um resolution. Sparsely segmented (about 65%)","metadata":{"id":"ZAf8QcC2i7yN"}},{"cell_type":"code","source":"%%time\nif not DEBUG:\n    create_animation_with_mask(kidney_scan = 'kidney_2',\n                                figsize = (8, 8),\n                                slice_interval=50,\n                                cmap_mask='jet',\n                                alpha_mask=0.5)","metadata":{"id":"B7TvJJYYi7yN","execution":{"iopub.status.busy":"2023-11-09T19:59:54.095161Z","iopub.status.idle":"2023-11-09T19:59:54.095693Z","shell.execute_reply.started":"2023-11-09T19:59:54.095431Z","shell.execute_reply":"2023-11-09T19:59:54.095456Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## kidney_3_dense scans\n<div style=\"border-radius:10px; border: #babab5 solid; padding: 15px; background-color: #e6f9ff; font-size:100%; \">\n    \nA portion (500 slices) of a kidney at 50.16um resolution using BM05. Densely segmented. Note that we provide all of the images for `kidney_3` in the `kidney_3_sparse/images` folder. This dataset accordingly has only a `labels` folder.","metadata":{"id":"oIghhlpBi7yN"}},{"cell_type":"code","source":"%%time\nif not DEBUG:\n    create_animation_with_mask(kidney_scan = 'kidney_3_dense',\n                                figsize = (8, 8),\n                                slice_interval=50,\n                                cmap_mask='jet',\n                                alpha_mask=0.5)","metadata":{"id":"YiqaUZ_Oi7yN","execution":{"iopub.status.busy":"2023-11-09T19:59:54.096811Z","iopub.status.idle":"2023-11-09T19:59:54.097363Z","shell.execute_reply.started":"2023-11-09T19:59:54.097086Z","shell.execute_reply":"2023-11-09T19:59:54.097112Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## kidney_3_sparse scans\n<div style=\"border-radius:10px; border: #babab5 solid; padding: 15px; background-color: #e6f9ff; font-size:100%; \">\n    \n The remainder of the segmentation masks for `kidney_3`. Sparsely segmented (about 85%).","metadata":{}},{"cell_type":"code","source":"%%time\nif not DEBUG:\n    create_animation_with_mask(kidney_scan = 'kidney_3_sparse',\n                                figsize = (8, 8),\n                                slice_interval=50,\n                                cmap_mask='jet',\n                                alpha_mask=0.5)","metadata":{"execution":{"iopub.status.busy":"2023-11-09T19:59:54.099772Z","iopub.status.idle":"2023-11-09T19:59:54.100202Z","shell.execute_reply.started":"2023-11-09T19:59:54.099983Z","shell.execute_reply":"2023-11-09T19:59:54.100021Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## Test scans: kidney_5 and kidney_6\n<div style=\"border-radius:10px; border: #babab5 solid; padding: 15px; background-color: #e6f9ff; font-size:100%; \">\n    \nContains the TIFF scans for the test set. These scans may or may not use a different beamline or resolution from the scans used in the training set. The names of the datasets are `kidney_5` and `kidney_6`.","metadata":{}},{"cell_type":"code","source":"create_subplot_test_scans(kidney_scan = 'kidney_5',\n                          figsize = (10,10))\n\ncreate_subplot_test_scans(kidney_scan = 'kidney_6',\n                          figsize = (10,10))","metadata":{"execution":{"iopub.status.busy":"2023-11-09T19:59:54.102343Z","iopub.status.idle":"2023-11-09T19:59:54.102891Z","shell.execute_reply.started":"2023-11-09T19:59:54.102620Z","shell.execute_reply":"2023-11-09T19:59:54.102647Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Next steps\n> 📌**Note**: <br> Thank you for reaching the end of this notebook 🙏🙏 <br> If you liked or forked this notebook, please consider upvoting ⬆️⬆️ It encourages to keep posting relevant content. <br> Feedback is always welcome!!\n\n<div style=\"border-radius:10px; border: #babab5 solid; padding: 15px; background-color: #e6f9ff; font-size:130%;\">\n\n* More exploratory data analysis\n* Starter segmentation model   ","metadata":{}}]}