{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"pygments_lexer":"ipython3","nbconvert_exporter":"python","version":"3.6.4","file_extension":".py","codemirror_mode":{"name":"ipython","version":3},"name":"python","mimetype":"text/x-python"}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"code","source":"import os\n#for dirname, _, filenames in os.walk('/kaggle/input'):\n    #for filename in filenames:\n        #print(os.path.join(dirname, filename))","metadata":{"execution":{"iopub.status.busy":"2021-12-21T09:39:18.877443Z","iopub.execute_input":"2021-12-21T09:39:18.878238Z","iopub.status.idle":"2021-12-21T09:39:18.906532Z","shell.execute_reply.started":"2021-12-21T09:39:18.878088Z","shell.execute_reply":"2021-12-21T09:39:18.905635Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"os.listdir('/kaggle/input/sartorius-cell-instance-segmentation/')","metadata":{"execution":{"iopub.status.busy":"2021-12-21T09:39:18.908163Z","iopub.execute_input":"2021-12-21T09:39:18.908594Z","iopub.status.idle":"2021-12-21T09:39:18.927290Z","shell.execute_reply.started":"2021-12-21T09:39:18.908558Z","shell.execute_reply":"2021-12-21T09:39:18.926637Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Introduction:\n1. Neurological disorders, including neurodegenerative diseases such as Alzheimer's and brain tumors, are a leading cause of death and disability across the globe. However, it is hard to quantify how well these deadly disorders respond to treatment. \n\n2. One accepted method is to review neuronal cells via light microscopy, which is both accessible and non-invasive. Unfortunately, segmenting individual neuronal cells in microscopic images can be challenging and time-intensive. Accurate instance segmentation of these cells—with the help of computer vision—could lead to new and effective drug discoveries to treat the millions of people with these disorders.\n\n3. Current solutions have limited accuracy for neuronal cells in particular. In internal studies to develop cell instance segmentation models, the neuroblastoma cell line SH-SY5Y consistently exhibits the lowest precision scores out of eight different cancer cell types tested. This could be because neuronal cells have a very unique, irregular and concave morphology associated with them, making them challenging to segment with commonly used mask heads.\n\n4. In this competition, you’ll detect and delineate distinct objects of interest in biological images depicting neuronal cell types commonly used in the study of neurological disorders. More specifically, you'll use phase contrast microscopy images to train and test your model for instance segmentation of neuronal cells. Successful models will do this with a high level of accuracy.\n","metadata":{}},{"cell_type":"markdown","source":"# data\nIn this competition we are segmenting neuronal cells in images. The training annotations are provided as run length encoded masks, and the images are in PNG format. The number of images is small, but the number of annotated objects is quite high. The hidden test set is roughly 240 images.\n\nFiles: \nA. train.csv - IDs and masks for all training objects. None of this metadata is provided for the test set.\n1. id - unique identifier for object\n2. annotation - run length encoded pixels for the identified neuronal cell\n3. width - source image width\n4. height - source image height\n5. cell_type - the cell line\n6. plate_time - time plate was created\n\nB. sample_submission.csv - a sample submission file in the correct format\n\nC. train - train images in PNG format\n\nD. test - test images in PNG format. Only a few test set images are available for download; the remainder can only be accessed by your notebooks when you submit.\n\nE. train_semi_supervised - unlabeled images offered in case you want to use additional data for a semi-supervised approach.\n\nLIVECell_dataset_2021 - A mirror of the data from the LIVECell dataset. LIVECell is the predecessor dataset to this competition. You will find extra data for the SH-SHY5Y cell line, plus several other cell lines not covered in the competition dataset that may be of interest for transfer learning.\n\n","metadata":{}},{"cell_type":"code","source":"# import libraries:\nimport pandas as pd\nimport numpy as np\nimport cv2\nimport matplotlib.pyplot as plt\n%matplotlib inline\nimport seaborn as sns\n\nfrom tqdm.notebook import tqdm\n\n","metadata":{"execution":{"iopub.status.busy":"2021-12-21T09:52:24.150588Z","iopub.execute_input":"2021-12-21T09:52:24.150921Z","iopub.status.idle":"2021-12-21T09:52:24.389545Z","shell.execute_reply.started":"2021-12-21T09:52:24.150890Z","shell.execute_reply":"2021-12-21T09:52:24.388566Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"def getImagePaths(path):\n    \"\"\"\n    Function to Combine Directory Path with individual Image Paths\n    \n    parameters: path(string) - Path of directory\n    returns: image_names(string) - Full Image Path\n    \"\"\"\n    image_names = []\n    for dirname, _, filenames in os.walk(path):\n        for filename in tqdm(filenames):\n            fullpath = os.path.join(dirname, filename)\n            image_names.append(fullpath)\n    return image_names","metadata":{"execution":{"iopub.status.busy":"2021-12-21T09:39:20.043549Z","iopub.execute_input":"2021-12-21T09:39:20.043789Z","iopub.status.idle":"2021-12-21T09:39:20.051378Z","shell.execute_reply.started":"2021-12-21T09:39:20.043760Z","shell.execute_reply":"2021-12-21T09:39:20.050497Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"#Get complete image paths for train and test datasets\nDIRECTORY_PATH = \"../input/sartorius-cell-instance-segmentation\"\nTRAIN_CSV = DIRECTORY_PATH + \"/train.csv\"\nTRAIN_PATH = DIRECTORY_PATH + \"/train\"\nTEST_PATH = DIRECTORY_PATH + \"/test\"\nTRAIN_SEMI_SUPERVISED_PATH = DIRECTORY_PATH + \"/train_semi_supervised\"\ntrain_images_path = getImagePaths(TRAIN_PATH)\ntest_images_path = getImagePaths(TEST_PATH)\ntrain_semi_supervised_path = getImagePaths(TRAIN_SEMI_SUPERVISED_PATH)","metadata":{"execution":{"iopub.status.busy":"2021-12-21T09:39:20.052447Z","iopub.execute_input":"2021-12-21T09:39:20.052690Z","iopub.status.idle":"2021-12-21T09:39:20.777066Z","shell.execute_reply.started":"2021-12-21T09:39:20.052662Z","shell.execute_reply":"2021-12-21T09:39:20.776503Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Meta Data\nThe meta data is given only for the train data, meaning that they cannot be used to predict the test data but they should be used to construct a solid cross validation strategy. So let's start understanding the statistical properties of the meta data.\n\nThe meta data, which is given in the train.csv file, contains 7 categorical and 2 numerical features (see table below). Each row points to an image with the id column and its associated mask with the annotation column. The annotations are given in the \"run length encoded pixels\" format. Furthermore each row contains the cell type information.","metadata":{}},{"cell_type":"code","source":"df_train = pd.read_csv(TRAIN_CSV)","metadata":{"execution":{"iopub.status.busy":"2021-12-21T09:39:20.778592Z","iopub.execute_input":"2021-12-21T09:39:20.779073Z","iopub.status.idle":"2021-12-21T09:39:21.380258Z","shell.execute_reply.started":"2021-12-21T09:39:20.779040Z","shell.execute_reply":"2021-12-21T09:39:21.379341Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"display(df_train.head())\ndisplay(df_train.shape)","metadata":{"execution":{"iopub.status.busy":"2021-12-21T09:39:21.382534Z","iopub.execute_input":"2021-12-21T09:39:21.382778Z","iopub.status.idle":"2021-12-21T09:39:21.407086Z","shell.execute_reply.started":"2021-12-21T09:39:21.382751Z","shell.execute_reply":"2021-12-21T09:39:21.406514Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"df_train.info()","metadata":{"execution":{"iopub.status.busy":"2021-12-21T09:39:25.550192Z","iopub.execute_input":"2021-12-21T09:39:25.550504Z","iopub.status.idle":"2021-12-21T09:39:25.632450Z","shell.execute_reply.started":"2021-12-21T09:39:25.550469Z","shell.execute_reply":"2021-12-21T09:39:25.631852Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# column wise unique values:\ndf_train.nunique()","metadata":{"execution":{"iopub.status.busy":"2021-12-21T09:39:26.153319Z","iopub.execute_input":"2021-12-21T09:39:26.153622Z","iopub.status.idle":"2021-12-21T09:39:26.274259Z","shell.execute_reply.started":"2021-12-21T09:39:26.153588Z","shell.execute_reply":"2021-12-21T09:39:26.273337Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"#number of images in each directory:\nprint(f\"Number of train images: {len(train_images_path)}\")\nprint(f\"Number of test images:  {len(test_images_path)}\")","metadata":{"execution":{"iopub.status.busy":"2021-12-21T09:39:26.853253Z","iopub.execute_input":"2021-12-21T09:39:26.853563Z","iopub.status.idle":"2021-12-21T09:39:26.858561Z","shell.execute_reply.started":"2021-12-21T09:39:26.853526Z","shell.execute_reply":"2021-12-21T09:39:26.857914Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"There are only 606 images in the train set, which is small for training neural network models and can easily lead to an overfitting problem. However, it's well known that this problem can be easily mitigated with the use of appropriate augmentation techniques (e.g. Ronneberger et al. 2015).\n\nAll the images have the same shape: (704 x 520) px. This is nice to have in a dataset because there won't be any complications due to a variable image resolution.","metadata":{}},{"cell_type":"code","source":"print(f'Number of unique images: {df_train.id.nunique()}')\nprint(f'Do all the images have a width of 704: {(df_train[\"width\"]==704).all()}')\nprint(f'Do all the images have a height of 520: {(df_train[\"height\"]==520).all()}')","metadata":{"execution":{"iopub.status.busy":"2021-12-21T09:39:28.353486Z","iopub.execute_input":"2021-12-21T09:39:28.353978Z","iopub.status.idle":"2021-12-21T09:39:28.368371Z","shell.execute_reply.started":"2021-12-21T09:39:28.353943Z","shell.execute_reply":"2021-12-21T09:39:28.367663Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"fig, ax = plt.subplots()\n\ninstances_per_image = df_train.groupby('id').size().sort_values()\n\n#instances_per_image.index = range(606)\n#instances_per_image.median()\ninstances_per_image.plot.bar(ax=ax)\n\nax.set_xticklabels([])\nax.set_xlabel('Images')\nax.set_ylabel('Number of Instances')\nplt.show()","metadata":{"execution":{"iopub.status.busy":"2021-12-21T09:39:29.085422Z","iopub.execute_input":"2021-12-21T09:39:29.085899Z","iopub.status.idle":"2021-12-21T09:39:32.094553Z","shell.execute_reply.started":"2021-12-21T09:39:29.085865Z","shell.execute_reply":"2021-12-21T09:39:32.093733Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"The number of instances in each image is remarkably variable (see figure below). Some statistical measures are as follows:\n\n1. Most of the images have more than 47 instances annotated.\n2. The minimum number of instances is 4.\n3. The maximum number of instances is 790.\n\nThese numbers are extremely critical to train an instance segmentation model. For instance, the famous Mask RCNN model requires the information of \"maximum number of detections\".","metadata":{}},{"cell_type":"markdown","source":"Cell Types Distribution:\n\nEach image is annotated with one of the three cell types: shsy5y, asto, and cort. \n\nThe distribution of the cell types is shown in the figure below. \n\nWhile the most represented cell type is shsy5y (70%) in the train set, cell types cort and astro are annotated only ~10% each. This means that the data is biased towards cell type shsy5y.","metadata":{}},{"cell_type":"code","source":"#distribution of cell types in %:\ncell_types = df_train.cell_type.value_counts()\ncell_types = cell_types/df_train.shape[0]*100\ncell_types.plot.bar()\nplt.title('Cell Type Distribution')\nplt.ylabel(' Percentage of instances')\nplt.xlabel('Cell Type')\n","metadata":{"execution":{"iopub.status.busy":"2021-12-21T09:39:35.223692Z","iopub.execute_input":"2021-12-21T09:39:35.224582Z","iopub.status.idle":"2021-12-21T09:39:35.441275Z","shell.execute_reply.started":"2021-12-21T09:39:35.224543Z","shell.execute_reply":"2021-12-21T09:39:35.440705Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"1. The number of unique id and cell_type combinations is equal to the number of unique images. This means that each image is associated with a unique cell type! See the result below. This also explains the distribution of the image numbers in the train.csv file. The images observed abundantly are associated with the most observed cell type shsy5y.\n\n2. Since each image is associated with only 1 cell type, we can count the number of images associated with each cell type. The figure below shows that most of the images are associated with the cell type cort, which agrees with our previous findings.","metadata":{}},{"cell_type":"code","source":"# no of imges associated with each cell type:\nfig, ax = plt.subplots(1, 1)\ndf_train.groupby(['id','cell_type'])['cell_type'].first().value_counts().plot.bar(ax=ax)\nax.set_ylabel('Number of Images')\nax.set_xlabel('Cell Types')\nfig.tight_layout()\nplt.show()","metadata":{"execution":{"iopub.status.busy":"2021-12-21T09:39:36.573414Z","iopub.execute_input":"2021-12-21T09:39:36.574449Z","iopub.status.idle":"2021-12-21T09:39:36.821217Z","shell.execute_reply.started":"2021-12-21T09:39:36.574401Z","shell.execute_reply":"2021-12-21T09:39:36.820272Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"#distribution of plate time:\nplate_time = df_train.plate_time.value_counts()\nplate_time = plate_time/df_train.shape[0]*100\nplate_time.plot.bar()\nplt.title('Plate Time Distribution')\nplt.ylabel(' Percentage of instances')\nplt.xlabel('Plate Time')","metadata":{"execution":{"iopub.status.busy":"2021-12-21T09:40:55.098640Z","iopub.execute_input":"2021-12-21T09:40:55.098989Z","iopub.status.idle":"2021-12-21T09:40:55.365639Z","shell.execute_reply.started":"2021-12-21T09:40:55.098955Z","shell.execute_reply":"2021-12-21T09:40:55.364534Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"#distribution of elasped time:\nelasped_time = df_train.elapsed_timedelta.value_counts()\nelasped_time = elasped_time/df_train.shape[0]*100\nelasped_time.plot.bar()\nplt.title('Elasped Time Distribution')\nplt.ylabel(' Percentage of instances')\nplt.xlabel('Elasped Time')","metadata":{"execution":{"iopub.status.busy":"2021-12-21T09:50:21.395544Z","iopub.execute_input":"2021-12-21T09:50:21.395756Z","iopub.status.idle":"2021-12-21T09:50:21.830568Z","shell.execute_reply.started":"2021-12-21T09:50:21.395727Z","shell.execute_reply":"2021-12-21T09:50:21.829658Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Images\nLets take a look at the images.","metadata":{}},{"cell_type":"code","source":"def display_multiple_img(images_paths, rows, cols):\n    \"\"\"\n    Function to Display Images from Dataset.\n    \n    parameters: images_path(string) - Paths of Images to be displayed\n                rows(int) - No. of Rows in Output\n                cols(int) - No. of Columns in Output\n    \"\"\"\n    figure, ax = plt.subplots(nrows=rows,ncols=cols,figsize=(16,8) )\n    for ind,image_path in enumerate(images_paths):\n        image=cv2.imread(image_path)\n        image = cv2.cvtColor(image, cv2.COLOR_BGR2RGB) \n        try:\n            ax.ravel()[ind].imshow(image)\n            ax.ravel()[ind].set_axis_off()\n        except:\n            continue;\n    plt.tight_layout()\n    plt.show()","metadata":{"execution":{"iopub.status.busy":"2021-12-21T09:51:49.974718Z","iopub.execute_input":"2021-12-21T09:51:49.975052Z","iopub.status.idle":"2021-12-21T09:51:49.983419Z","shell.execute_reply.started":"2021-12-21T09:51:49.975019Z","shell.execute_reply":"2021-12-21T09:51:49.982339Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# display train images:\ndisplay_multiple_img(train_images_path[100:150], 2,2)","metadata":{"execution":{"iopub.status.busy":"2021-12-21T09:53:31.943035Z","iopub.execute_input":"2021-12-21T09:53:31.943986Z","iopub.status.idle":"2021-12-21T09:53:33.273626Z","shell.execute_reply.started":"2021-12-21T09:53:31.943935Z","shell.execute_reply":"2021-12-21T09:53:33.272776Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# display train semisupervised images:\ndisplay_multiple_img(train_semi_supervised_path[100:150], 2, 2)","metadata":{"execution":{"iopub.status.busy":"2021-12-21T09:54:44.659028Z","iopub.execute_input":"2021-12-21T09:54:44.659327Z","iopub.status.idle":"2021-12-21T09:54:45.584788Z","shell.execute_reply.started":"2021-12-21T09:54:44.659294Z","shell.execute_reply":"2021-12-21T09:54:45.583379Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"#display test images:\ndisplay_multiple_img(test_images_path, 1, 3)","metadata":{"execution":{"iopub.status.busy":"2021-12-21T09:54:38.129845Z","iopub.execute_input":"2021-12-21T09:54:38.130055Z","iopub.status.idle":"2021-12-21T09:54:38.632275Z","shell.execute_reply.started":"2021-12-21T09:54:38.130029Z","shell.execute_reply":"2021-12-21T09:54:38.631428Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"","metadata":{},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"It's time to look at the images and the masks now. The figures below shows randomly selected images corresponding to each of the three distinct cell types. Each cell type has its own unique morphological properties.\n\nastro instances are the biggest in shape. They cover a lot of space in the masks.\ncort instances are smaller than the other cell types in general and they are in circle-like shapes. They don't cover much space in the masks.\nshsy5y instances are slightly bigger, elongated and more abundant than the cort instances. They cover more space than the cort cells.","metadata":{}},{"cell_type":"code","source":"def make_mask(mask_files, image_shape=(520, 704), color=False):\n    mask = np.zeros(image_shape).ravel()\n    for i, mask_file in enumerate(mask_files):\n        couples = np.array(mask_file.split()).reshape(-1, 2).astype(int)\n        couples[:, 1] = couples[:, 0] + couples[:, 1]\n        for couple in couples:\n            if color:\n                mask[couple[0]: couple[1]] = i\n            else:\n                mask[couple[0]: couple[1]] = 1\n    mask = mask.reshape(520, 704)\n    return mask\n\ndef plot_image(image_id='0030fd0e6378'):\n    fig, ax = plt.subplots(1, 2, figsize=(14,5))\n    cell_type = df_train.loc[df_train['id'] == image_id, 'cell_type'][0:1].values\n    \n    file_name = os.path.join(\n        '../input/sartorius-cell-instance-segmentation',\n        'train', image_id + '.png')\n    image = plt.imread(file_name)\n    mask_files = df_train.loc[df_train['id'] == image_id, 'annotation']\n    mask = make_mask(mask_files)\n\n    ax[0].imshow(\n        image,\n        cmap = plt.get_cmap('winter'), \n        origin = 'upper',\n        vmax = np.quantile(image, 0.99),\n        vmin = np.quantile(image, 0.05)\n    )\n    ax[0].set_title(f'Source [{image_id}]')\n    ax[0].axis('off')\n    \n    ax[1].imshow(\n        image,\n        cmap = plt.get_cmap('winter'), \n        origin = 'upper',\n        vmax = 255,\n        vmin = 0)\n    ax[1].imshow(mask, alpha=1, cmap=plt.get_cmap('seismic'))\n    ax[1].set_title(f'Source [{image_id}] + Mask {cell_type}')\n    ax[1].axis('off')\n    plt.show()\n\nselect_image_ids = []\nselect_image_ids.append(df_train.loc[df_train['cell_type'] == 'astro', 'id'].sample(1).to_list()[0])\nselect_image_ids.append(df_train.loc[df_train['cell_type'] == 'cort', 'id'].sample(1).to_list()[0])\nselect_image_ids.append(df_train.loc[df_train['cell_type'] == 'shsy5y', 'id'].sample(1).to_list()[0])\n\nfor image_id in select_image_ids:\n    plot_image(image_id)\n","metadata":{"execution":{"iopub.status.busy":"2021-12-21T09:21:38.178016Z","iopub.execute_input":"2021-12-21T09:21:38.179051Z","iopub.status.idle":"2021-12-21T09:21:39.608541Z","shell.execute_reply.started":"2021-12-21T09:21:38.178986Z","shell.execute_reply":"2021-12-21T09:21:39.607604Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"","metadata":{},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"References:\n1. https://www.kaggle.com/ishandutta/sartorius-indepth-eda-explanation-model\n2. https://www.kaggle.com/tolgadincer/sartorius-eda-general-overview-and-outliers\n","metadata":{}},{"cell_type":"code","source":"","metadata":{},"execution_count":null,"outputs":[]}]}