{"cells":[{"metadata":{},"cell_type":"markdown","source":"# Introduction"},{"metadata":{},"cell_type":"markdown","source":"This notebook has for goal to introduce several functions to load the data and create instance masks for every cell present in training images. By generating instance segmentation masks, we will make it possible to analyze cells individually and potentially associate each of them to one or several of the labels given for the entire image."},{"metadata":{"_uuid":"8f2839f25d086af736a60e9eeb907d3b93b6e0e5","_cell_guid":"b1076dfc-b9ad-4769-8c92-a6c4dae69d19","trusted":true},"cell_type":"code","source":"import numpy as np\nimport pandas as pd\nimport os\nimport cv2\nimport matplotlib.pyplot as plt\nfrom tqdm import tqdm","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"dataset_folder = \"/kaggle/input/hpa-single-cell-image-classification/\"\ntraining_image_folder = dataset_folder+\"train/\"\ntrain_df = pd.read_csv(dataset_folder+\"train.csv\")\ntrain_df","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"# Load the images and apply a binary mask"},{"metadata":{"trusted":true},"cell_type":"code","source":"def get_binary_mask(img):\n    '''\n    Turn the RGB image into grayscale before\n    applying an Otsu threshold to obtain a\n    binary segmentation\n    '''\n    \n    blurred_img = cv2.GaussianBlur(img,(25,25),0)\n    gray_img = cv2.cvtColor(blurred_img, cv2.COLOR_RGBA2GRAY)\n    ret, otsu = cv2.threshold(gray_img, 0, 255, cv2.THRESH_BINARY+cv2.THRESH_OTSU)\n    \n    kernel = np.ones((40,40),np.uint8)\n    closed_mask = cv2.morphologyEx(otsu, cv2.MORPH_CLOSE, kernel)\n    return closed_mask","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"def load_RGBY_image(image_id_path):\n    '''\n    Load and stack the channels that are stored separately.\n    '''\n    \n    red_image = cv2.imread(image_id_path+\"_red.png\", cv2.IMREAD_UNCHANGED)\n    green_image = cv2.imread(image_id_path+\"_green.png\", cv2.IMREAD_UNCHANGED)\n    blue_image = cv2.imread(image_id_path+\"_blue.png\", cv2.IMREAD_UNCHANGED)\n    yellow_image = cv2.imread(image_id_path+\"_yellow.png\", cv2.IMREAD_UNCHANGED)\n\n    stacked_images = np.transpose(np.array([red_image, green_image, blue_image, yellow_image]), (1,2,0))\n    return stacked_images","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"image_id_path = training_image_folder+train_df.iloc[0].ID\nstacked_images = load_RGBY_image(image_id_path)\nbinary_mask = get_binary_mask(stacked_images)","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"plt.imshow(stacked_images[:,:,:3])\nplt.show()","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"plt.imshow(binary_mask)\nplt.show()","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"# Generate instance masks and convert to RLE encoding"},{"metadata":{"trusted":true},"cell_type":"code","source":"def rle_encoding(x):\n    '''\n    Turns our masks into RLE encoding to easily store them\n    and feed them into models later on\n    https://en.wikipedia.org/wiki/Run-length_encoding\n    '''\n    \n    dots = np.where(x.T.flatten() == 255)[0]\n    run_lengths = []\n    prev = -2\n    for b in dots:\n        if (b>prev+1): run_lengths.extend((b + 1, 0))\n        run_lengths[-1] += 1\n        prev = b\n        \n    return ' '.join([str(x) for x in run_lengths])","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"def get_instance_masks(binary_mask):\n    '''\n    Using a binary mask, this function filters out \n    small items and create a separate mask for each \n    blobs\n    '''\n    \n    contours= cv2.findContours(binary_mask,\n                               cv2.RETR_TREE, \n                               cv2.CHAIN_APPROX_SIMPLE)\n    instance_masks = []\n    for contour in contours[0]:\n        if cv2.contourArea(contour)>100:\n            instance_contour = np.zeros(binary_mask.shape)\n            cv2.drawContours(instance_contour,[contour], \n                             0, 255,thickness=cv2.FILLED)\n            \n            encoded_cell_mask = rle_encoding(instance_contour)\n            instance_masks.append(encoded_cell_mask)\n            \n    return instance_masks","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"# Add the RLE encoding to the existing training dataframe"},{"metadata":{"trusted":true},"cell_type":"code","source":"process_RLE_for = 20\ntrain_df[\"RLE_encoding\"] = \"\"\n\nwith tqdm(total=process_RLE_for) as pbar:\n    for idx, item in train_df[:process_RLE_for].iterrows():\n        image_id_path = training_image_folder+item.ID\n\n        stacked_images = load_RGBY_image(image_id_path)\n        binary_mask = get_binary_mask(stacked_images)\n        instance_masks = get_instance_masks(binary_mask)\n\n        train_df.at[idx, \"RLE_encoding\"] = str(instance_masks)\n        pbar.update(1)","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"train_df","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"By creating individual masks for every cells in the training images, I now have the possibility to proceed to image analysis. Below, I display every cell and the color distribution for the RGB channels. A preliminary methods to identify the cells' classes could be to cluster them based on their color distribution signature."},{"metadata":{"trusted":true},"cell_type":"code","source":"def plot_color_distribution(isolated_cell_img):\n    color = ('r','g','b','y')\n    for i,col in enumerate(color):\n        histr = cv2.calcHist([isolated_cell_img],[i],None,[256],[1,256])\n        plt.plot(histr,color = col)\n        plt.xlim([1,256])\n    plt.show()\n\ndef analyze_individual_cells(binary_mask, original_image):\n    \n    contours= cv2.findContours(binary_mask,\n                               cv2.RETR_TREE, \n                               cv2.CHAIN_APPROX_SIMPLE)\n    \n    for contour in contours[0]:\n        if cv2.contourArea(contour)>100:\n            x, y, width, height = cv2.boundingRect(contour)\n            \n            instance_contour = np.zeros(binary_mask.shape)\n            cv2.drawContours(instance_contour,[contour], \n                             0, 255, thickness=cv2.FILLED)\n\n            isolated_cell_image = np.zeros(binary_mask.shape)\n            isolated_cell_image = cv2.bitwise_and(original_image,original_image, mask = instance_contour.astype(\"uint8\"))\n    \n            plt.imshow(isolated_cell_image[y:y+height,x:x+width,:3])\n            plt.show()\n            plot_color_distribution(isolated_cell_image[y:y+height,x:x+width])","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"As seen below, when attempting to isolate individual cells using the functions previously defined, we can observe some issues when cells are too close from each other. Nonetheless, it does appear like a promising technique to generate instance masks."},{"metadata":{"trusted":true},"cell_type":"code","source":"image_id_path = training_image_folder+train_df.iloc[0].ID\nstacked_images = load_RGBY_image(image_id_path)\nbinary_mask = get_binary_mask(stacked_images)\nanalyze_individual_cells(binary_mask, stacked_images)","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"image_id_path = training_image_folder+train_df.iloc[55].ID\nstacked_images = load_RGBY_image(image_id_path)\nbinary_mask = get_binary_mask(stacked_images)\nanalyze_individual_cells(binary_mask, stacked_images)","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"## Thanks for reading this notebook! If you found this notebook helpful, please give it an upvote. It is always greatly appreciated!"}],"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"pygments_lexer":"ipython3","nbconvert_exporter":"python","version":"3.6.4","file_extension":".py","codemirror_mode":{"name":"ipython","version":3},"name":"python","mimetype":"text/x-python"}},"nbformat":4,"nbformat_minor":4}