{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"pygments_lexer":"ipython3","nbconvert_exporter":"python","version":"3.6.4","file_extension":".py","codemirror_mode":{"name":"ipython","version":3},"name":"python","mimetype":"text/x-python"}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"# Data and Mask Generation using DataFrames! for use with Keras ImageDataGenerator\nThis is an update for my previous notebook [Data Import and Binary Mask Generation](http://https://www.kaggle.com/code/ezmgszi/data-import-and-binary-mask-generation/notebook)\n\nI decide to use ImageDataGenerator for training my model and switched to using a dataframe to hold the data. I made this choice to avoid holding all images in memory as ImageDataGenerator will load the images as they are request by training.\n\nAnother change i made was that i am only creating masks for blood vessels rather than all three annotation types, again to save memory and to reduce run time of the notebook.\n\nI did include importing all three types of annotations just incase i decide to use them later i have them, if memory is still an issue you can change importing annotations to only importing blood vessel annotations insted. Tough i dont think it will really save that much overall.\n\nI am leaving my old notebook up and unchanged incase someone finds that version helpful, as it does include visualization of all mask types.","metadata":{}},{"cell_type":"markdown","source":"# **Imports:**","metadata":{}},{"cell_type":"code","source":"import os\nimport glob\nimport json\n\nimport cv2\nimport matplotlib.pyplot as plt\nimport numpy as np\nimport pandas as pd\nfrom PIL import Image","metadata":{"execution":{"iopub.status.busy":"2023-06-14T08:14:50.927930Z","iopub.execute_input":"2023-06-14T08:14:50.929201Z","iopub.status.idle":"2023-06-14T08:14:51.198720Z","shell.execute_reply.started":"2023-06-14T08:14:50.929155Z","shell.execute_reply":"2023-06-14T08:14:51.197333Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# **Data Import**\n* import our test and train data using glob, this imports the file paths for all files in the path.\n* import our annotations which are in the polygons.jsonl file\n    * Durring this step we import all polygons by annotation. We first import the image id, and then all three annotation types as image annotations. We then put them in our annotations{} dictionary maping them to the image ID. \n        * annotation[image_id]->[blood_vessel],[glomerulus],[unsure]\n* Map our train and test data to image maps, this will take the image id from the file path and map it to the image file path.","metadata":{}},{"cell_type":"code","source":"# Setup paths and import all the necessary data for train and test\ntrain_path = \"/kaggle/input/hubmap-hacking-the-human-vasculature/train/*\"\ntest_path = \"/kaggle/input/hubmap-hacking-the-human-vasculature/test/*\"\n\ntrain = glob.glob(train_path)\ntest = glob.glob(test_path)\n\n# Create a dictionary to hold the annotations for each image\nannotations = {}\n\n# Open the annotations file\nwith open('/kaggle/input/hubmap-hacking-the-human-vasculature/polygons.jsonl', 'r') as f:\n    # For each line in the file\n    for line in f:\n        # Parse the line as JSON\n        annotation = json.loads(line)\n\n        # Get the image ID and the list of annotations for this image\n        image_id = annotation['id']\n        image_annotations = annotation['annotations']\n\n        # Store the annotations in the dictionary\n        annotations[image_id] = image_annotations\n        \n# Create our image map for our train data and test data\nimage_map = {impath.split('/')[-1].split('.')[0]: impath for impath in train}\nimage_map_test = {impath.split('/')[-1].split('.')[0]: impath for impath in test}\n","metadata":{"execution":{"iopub.status.busy":"2023-06-14T08:14:51.201774Z","iopub.execute_input":"2023-06-14T08:14:51.202394Z","iopub.status.idle":"2023-06-14T08:14:57.008251Z","shell.execute_reply.started":"2023-06-14T08:14:51.202345Z","shell.execute_reply":"2023-06-14T08:14:57.007135Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# DataFrame and Binary Mask generation\n\nhere we create our data fram and use our convert_to_mask function to generate the truth binary mask that will be used in training","metadata":{}},{"cell_type":"code","source":"def convert_to_mask(annotations):\n    \"\"\"\n    Converts annotations to a binary mask.\n\n    Parameters:\n    - annotations: A list of annotations for a single image.\n\n    Returns:\n    - mask: The binary mask.\n    \"\"\"\n    # Set the image dimensions directly\n    image_dimensions = (512, 512)\n    \n    # Create an empty mask of the same size as the image\n    mask = np.zeros(image_dimensions, dtype=np.uint8)\n\n    # For each annotation\n    for annotation in annotations:\n        coordinates = np.array(annotation['coordinates'])\n        coordinates = coordinates.reshape(-1, 1, 2)\n        # Draw the polygon on the mask\n        cv2.fillPoly(mask, [coordinates], 255)\n\n    return mask","metadata":{"execution":{"iopub.status.busy":"2023-06-14T08:14:57.009727Z","iopub.execute_input":"2023-06-14T08:14:57.010132Z","iopub.status.idle":"2023-06-14T08:14:57.018647Z","shell.execute_reply.started":"2023-06-14T08:14:57.010096Z","shell.execute_reply":"2023-06-14T08:14:57.017300Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# create data frame, mapping our image file paths to the correct annotation\n# Initialize an empty list to hold the DataFrame data\ndf_data = []\n\nfor image_id, image_path in image_map.items():\n    # get only annotations that have a matching image id\n    annotation = annotations.get(image_id)\n    \n    if annotation:  # If annotations for this image exist\n        # Filtering out unnecessary annotation types and only keeping 'blood_vessel'\n        blood_vessel_annotations = [an for an in annotation if an['type'] == 'blood_vessel']\n\n        if blood_vessel_annotations:\n            # Convert these annotations to a binary mask\n            mask = convert_to_mask(blood_vessel_annotations)\n            df_data.append([image_id, image_path, mask])\n\n# Convert list of dictionaries to a DataFrame\ndf = pd.DataFrame(df_data, columns=['id', 'filepath', 'mask'])","metadata":{"execution":{"iopub.status.busy":"2023-06-14T08:14:57.021615Z","iopub.execute_input":"2023-06-14T08:14:57.022212Z","iopub.status.idle":"2023-06-14T08:14:59.726038Z","shell.execute_reply.started":"2023-06-14T08:14:57.022105Z","shell.execute_reply":"2023-06-14T08:14:59.724572Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# ImageDataGenerator requires a file path in order to use our binary mask, \n# so we need to save our masks as .png and replace their values with the path to the saved mask\n# Create the directory if it doesn't exist\nos.makedirs('/kaggle/working/masks', exist_ok=True)\n# Iterate through the DataFrame and save masks\nfor idx, row in df.iterrows():\n    mask_array = row['mask'] \n    mask_img = Image.fromarray((mask_array * 255).astype(np.uint8))  # convert to an image\n    filepath = f\"/kaggle/working/masks/{row['id']}.png\"  # specify file path\n    mask_img.save(filepath)\n    df.at[idx, 'mask'] = filepath  # update mask value with file path","metadata":{"execution":{"iopub.status.busy":"2023-06-14T08:14:59.727851Z","iopub.execute_input":"2023-06-14T08:14:59.728985Z","iopub.status.idle":"2023-06-14T08:15:08.860639Z","shell.execute_reply.started":"2023-06-14T08:14:59.728937Z","shell.execute_reply":"2023-06-14T08:15:08.859339Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"**Display a portion of our data frame**\n\nshow a small part of our data frame to demonstrate the structure","metadata":{}},{"cell_type":"code","source":"# Display the DataFrame\ndf.head()","metadata":{"execution":{"iopub.status.busy":"2023-06-14T08:15:08.862299Z","iopub.execute_input":"2023-06-14T08:15:08.862886Z","iopub.status.idle":"2023-06-14T08:15:08.894100Z","shell.execute_reply.started":"2023-06-14T08:15:08.862842Z","shell.execute_reply":"2023-06-14T08:15:08.891876Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"**Display Image and Masks**\n\nshow the orignal image, the provided image binary mask, and then the mask overlay for the first 5 images","metadata":{}},{"cell_type":"code","source":"# Display the first 5 images and masks\nfor i in range(5):\n    # Load image\n    img = cv2.imread(df.iloc[i]['filepath'])\n    img = cv2.cvtColor(img, cv2.COLOR_BGR2RGB)  # OpenCV uses BGR color order\n\n    # Load mask\n    mask = cv2.imread(df.iloc[i]['mask'], cv2.IMREAD_GRAYSCALE)\n\n    # Create overlay image\n    overlay = img.copy()\n    overlay[mask > 0] =  (0, 255, 0)\n\n    # Display image and mask\n    fig, ax = plt.subplots(1, 3, figsize=(10, 5))\n    ax[0].imshow(img)\n    ax[0].set_title('Image')\n    ax[0].axis('off')  # Turn off axis for image\n    ax[1].imshow(mask, cmap='gray')\n    ax[1].set_title('Blood Vessel Mask')\n    ax[1].axis('off')  # Turn off axis for mask\n    ax[2].imshow(img)\n    ax[2].imshow(overlay, alpha=0.7)\n    ax[2].set_title('Overlay')\n    ax[2].axis('off')  # Turn off axis for mask\n\nplt.show()","metadata":{"execution":{"iopub.status.busy":"2023-06-14T08:15:08.896264Z","iopub.execute_input":"2023-06-14T08:15:08.896735Z","iopub.status.idle":"2023-06-14T08:15:11.615300Z","shell.execute_reply.started":"2023-06-14T08:15:08.896691Z","shell.execute_reply":"2023-06-14T08:15:11.614094Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"**Whats Next?**\n\nThe next steps would be to setup your preprocessing function and your ImageDataGenerator, splitting the data into training and validation steps.\n\nI will upload a note book showing how to do this later!","metadata":{}}]}