{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"pygments_lexer":"ipython3","nbconvert_exporter":"python","version":"3.6.4","file_extension":".py","codemirror_mode":{"name":"ipython","version":3},"name":"python","mimetype":"text/x-python"}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"# Using Sartorious and LiveCell data\n\nWe've been given new and old data. I come at these problems from an image processing background. As such, I like to have the binary masks to work with in an image format. \n\nThe purpose of this notebook is to show how to get a binary mask generated for both the Sartorious training data and the LiveCell training data.","metadata":{"_uuid":"8f2839f25d086af736a60e9eeb907d3b93b6e0e5","_cell_guid":"b1076dfc-b9ad-4769-8c92-a6c4dae69d19","_kg_hide-input":true,"_kg_hide-output":true,"execution":{"iopub.status.busy":"2021-11-12T17:51:50.665011Z","iopub.execute_input":"2021-11-12T17:51:50.665302Z","iopub.status.idle":"2021-11-12T17:51:50.669429Z","shell.execute_reply.started":"2021-11-12T17:51:50.665270Z","shell.execute_reply":"2021-11-12T17:51:50.668558Z"}}},{"cell_type":"markdown","source":"import libraries first","metadata":{}},{"cell_type":"code","source":"#from pycocotools.coco import COCO\nimport skimage.io as io\nimport matplotlib.pyplot as plt\nfrom pathlib import Path\nfrom PIL import Image","metadata":{"execution":{"iopub.status.busy":"2021-11-12T17:51:50.672770Z","iopub.execute_input":"2021-11-12T17:51:50.673564Z","iopub.status.idle":"2021-11-12T17:51:51.367938Z","shell.execute_reply.started":"2021-11-12T17:51:50.673526Z","shell.execute_reply":"2021-11-12T17:51:51.366985Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## Sartorious data\nFirst let's handle the new competition data\n\n### Load the training data\nFirst step is loading the training data and checking out the first few lines.\n\nIt is important to recognize that none of the \"metadata\" will be provided in the test set. so... without good reason, we really shouldn't use it except perhaps to organize our thoughts about the data.\n\nwe'll import pandas to read the csv and numpy because we will need it for manipulating numerical values at some point.","metadata":{"execution":{"iopub.status.busy":"2021-10-20T14:40:49.42228Z","iopub.execute_input":"2021-10-20T14:40:49.423207Z","iopub.status.idle":"2021-10-20T14:40:49.427388Z","shell.execute_reply.started":"2021-10-20T14:40:49.423161Z","shell.execute_reply":"2021-10-20T14:40:49.42652Z"}}},{"cell_type":"code","source":"import pandas as pd\nimport numpy as np\ndataDir='../input/sartorius-cell-instance-segmentation'\n\ntraindf = pd.read_csv(dataDir+'/train.csv')\ntraindf.head()","metadata":{"execution":{"iopub.status.busy":"2021-11-12T17:51:51.370035Z","iopub.execute_input":"2021-11-12T17:51:51.370356Z","iopub.status.idle":"2021-11-12T17:51:52.014314Z","shell.execute_reply.started":"2021-11-12T17:51:51.370315Z","shell.execute_reply":"2021-11-12T17:51:52.013428Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## Annotations\nare defined in the csv as a run length encoded (RLE) mask. The following function will be used to convert that RLE annotation into a binary mask image ","metadata":{}},{"cell_type":"code","source":"# ht to https://www.kaggle.com/shivansh002/getting-started for this particular version of the function\ndef rle_decode(mask_rle, shape=(520, 704)):\n    '''\n    mask_rle: run-length as string formated (start length)\n    shape: (height,width) of array to return \n    Returns numpy array, 1 - mask, 0 - background\n\n    '''\n    s = mask_rle.split()\n    starts, lengths = [np.asarray(x, dtype=int) for x in (s[0:][::2], s[1:][::2])]\n    starts -= 1\n    ends = starts + lengths\n    img = np.zeros(shape[0]*shape[1], dtype=np.uint8)\n    for lo, hi in zip(starts, ends):\n        img[lo:hi] = 1\n    return img.reshape(shape)  # Needed to align to RLE direction","metadata":{"execution":{"iopub.status.busy":"2021-11-12T17:51:52.015498Z","iopub.execute_input":"2021-11-12T17:51:52.015751Z","iopub.status.idle":"2021-11-12T17:51:52.024342Z","shell.execute_reply.started":"2021-11-12T17:51:52.015719Z","shell.execute_reply":"2021-11-12T17:51:52.023435Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## Images and masks\n\nNow we've got our function to generate the masks from the csv data, let's see if we can visualize an image and it's corresponding mask.\n\nwe can find an individual image name from the 'id' field. There are multiple entries per image, so we will just grab the first unique entry","metadata":{}},{"cell_type":"code","source":"imagename = traindf['id'].unique()[0]\nprint(imagename)","metadata":{"execution":{"iopub.status.busy":"2021-11-12T17:51:52.026427Z","iopub.execute_input":"2021-11-12T17:51:52.026690Z","iopub.status.idle":"2021-11-12T17:51:52.050996Z","shell.execute_reply.started":"2021-11-12T17:51:52.026659Z","shell.execute_reply":"2021-11-12T17:51:52.049647Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"let's load and plot that image.\n\nWe'll use skimage to read in the image and matplotlib to plot","metadata":{}},{"cell_type":"code","source":"import skimage\nfrom skimage.io import imread\nimport matplotlib.pyplot as plt\n\nimage = imread(dataDir + '/train/' + imagename + '.png')\n\n## Now plot that image (selecting for grayscale since it is not a color image)\nplt.imshow(image,cmap='gray', vmin=0, vmax=255)","metadata":{"execution":{"iopub.status.busy":"2021-11-12T17:51:52.052755Z","iopub.execute_input":"2021-11-12T17:51:52.053159Z","iopub.status.idle":"2021-11-12T17:51:52.413864Z","shell.execute_reply.started":"2021-11-12T17:51:52.053116Z","shell.execute_reply":"2021-11-12T17:51:52.412931Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Now let's plot the mask (using matplotlib again). \n\nFirst we need to make a sub dataframe that has only masks pertaining to the image we are interested in.\n\nNext, we will make one giant mask for all the images","metadata":{}},{"cell_type":"code","source":"\ndfimage=traindf[traindf['id']==imagename]\n\nprint(\"Number of masks for \" + imagename + \": \" + str(len(dfimage)))\n\nallMask = np.zeros(image.shape)\nfor i in range(len(dfimage)):\n    allMask = allMask + rle_decode(dfimage['annotation'][i])\n    \nplt.imshow(allMask==1)\n","metadata":{"execution":{"iopub.status.busy":"2021-11-12T17:51:52.415336Z","iopub.execute_input":"2021-11-12T17:51:52.415707Z","iopub.status.idle":"2021-11-12T17:51:52.914022Z","shell.execute_reply.started":"2021-11-12T17:51:52.415674Z","shell.execute_reply":"2021-11-12T17:51:52.913207Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"#and we can look at the mask overlaid on top like this:\nplt.imshow(image,cmap='gray', vmin=0, vmax=255)\nplt.imshow(allMask==1, cmap='jet', alpha=0.5) # interpolation='none'","metadata":{"execution":{"iopub.status.busy":"2021-11-12T17:51:52.915428Z","iopub.execute_input":"2021-11-12T17:51:52.915947Z","iopub.status.idle":"2021-11-12T17:51:53.260018Z","shell.execute_reply.started":"2021-11-12T17:51:52.915876Z","shell.execute_reply":"2021-11-12T17:51:53.259160Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## Livecell data\n\nLet's also load some data from the LiveCell dataset. it's in coco format. I tried using pycocotools, but got errors, so I was going to brute force it, but it turns out from much wrestling that the LiveCell data won't load using pycocotools, and even though it is in coco format, the annotations are in a polygon format, not a RLE format. \n\nSo let me walk you through how to load in a LiveCell image and it's corresponding segmentations.","metadata":{}},{"cell_type":"markdown","source":"First, we really do need pycocotools installed","metadata":{}},{"cell_type":"code","source":"!pip install pycocotools","metadata":{"execution":{"iopub.status.busy":"2021-11-12T17:51:53.261293Z","iopub.execute_input":"2021-11-12T17:51:53.261536Z","iopub.status.idle":"2021-11-12T17:52:13.107229Z","shell.execute_reply.started":"2021-11-12T17:51:53.261507Z","shell.execute_reply":"2021-11-12T17:52:13.105941Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## Loading data via json\n\nNow because of the pycocotools errors I'm getting, I can't load directly with the COCO function. So I load it first with json","metadata":{}},{"cell_type":"code","source":"import json\nfrom pycocotools.coco import COCO\nfrom pycocotools import _mask\nannFile = '../input/sartorius-cell-instance-segmentation/LIVECell_dataset_2021/annotations/LIVECell_single_cells/shsy5y/livecell_shsy5y_train.json' \n\nwith open(annFile) as f:\n    data = json.loads(f.read())\n","metadata":{"execution":{"iopub.status.busy":"2021-11-12T17:53:01.709662Z","iopub.execute_input":"2021-11-12T17:53:01.709992Z","iopub.status.idle":"2021-11-12T17:53:09.313077Z","shell.execute_reply.started":"2021-11-12T17:53:01.709946Z","shell.execute_reply":"2021-11-12T17:53:09.312158Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"And then you can load in an annotation by combining two pycocotools functions\n1. `frPoly`\n2. `decode`\n\nIf someone finds a better way to do this I'm happy to hear it, but I wasn't able to find a way to convert straight from the polygon format into the mask. But this seems to work","metadata":{}},{"cell_type":"code","source":"for key in data['annotations'].keys():\n#         print(type(data['annotations'][key]['segmentation']))\n\n    rle = _mask.frPoly(data['annotations'][key]['segmentation'],520,704 )\n    print(rle)\n    mask = _mask.decode(rle)\n    print(mask.shape)\n\n    break","metadata":{"execution":{"iopub.status.busy":"2021-11-12T17:53:10.942081Z","iopub.execute_input":"2021-11-12T17:53:10.942357Z","iopub.status.idle":"2021-11-12T17:53:10.950323Z","shell.execute_reply.started":"2021-11-12T17:53:10.942328Z","shell.execute_reply":"2021-11-12T17:53:10.949407Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"","metadata":{}},{"cell_type":"code","source":"plt.imshow(mask)","metadata":{"execution":{"iopub.status.busy":"2021-11-12T17:53:17.962584Z","iopub.execute_input":"2021-11-12T17:53:17.963202Z","iopub.status.idle":"2021-11-12T17:53:18.262716Z","shell.execute_reply.started":"2021-11-12T17:53:17.963146Z","shell.execute_reply":"2021-11-12T17:53:18.261727Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"from pycocotools import _mask\n\nfor image in data['images']:\n    print(image['file_name'] + ' ' + str(image['id']))\n    tempimg = imread('../input/sartorius-cell-instance-segmentation/LIVECell_dataset_2021/images/livecell_train_val_images/SHSY5Y/' + image['file_name'])\n\n    allMask = np.zeros(tempimg.shape)\n    \n    for key in data['annotations'].keys():\n    #         print(type(data['annotations'][key]['segmentation']))\n        if image['id']==data['annotations'][key]['image_id']:\n            rle = _mask.frPoly(data['annotations'][key]['segmentation'],520,704 )\n#             print(rle)\n            mask = _mask.decode(rle)\n#             print(mask.shape)\n            allMask = allMask + mask[:,:,0]\n\n    plt.imshow(tempimg,cmap='gray', vmin=0, vmax=255)\n    plt.imshow(allMask==1, cmap='jet', alpha=0.5) # interpolation='none'\n    break\n# for annotation in livecelldatajson['annotations'].items():\n#     print(annotation['imageid'])\n   ","metadata":{"execution":{"iopub.status.busy":"2021-11-12T17:53:18.939275Z","iopub.execute_input":"2021-11-12T17:53:18.939555Z","iopub.status.idle":"2021-11-12T17:53:19.649833Z","shell.execute_reply.started":"2021-11-12T17:53:18.939527Z","shell.execute_reply":"2021-11-12T17:53:19.649178Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"","metadata":{}},{"cell_type":"code","source":"plt.imshow(tempimg,cmap='gray', vmin=0, vmax=255)\n# plt.imshow(allMask==1, cmap='jet', alpha=0.5) # interpolation='none'","metadata":{"execution":{"iopub.status.busy":"2021-11-12T17:53:21.522549Z","iopub.execute_input":"2021-11-12T17:53:21.523033Z","iopub.status.idle":"2021-11-12T17:53:21.834439Z","shell.execute_reply.started":"2021-11-12T17:53:21.522981Z","shell.execute_reply":"2021-11-12T17:53:21.833572Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Next steps will include exporting these masks for use in building training datasets.","metadata":{}}]}