{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"pygments_lexer":"ipython3","nbconvert_exporter":"python","version":"3.6.4","file_extension":".py","codemirror_mode":{"name":"ipython","version":3},"name":"python","mimetype":"text/x-python"}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"# Data Explore\nBasically, the dataset we use contains train images and the lables. Detailed imformation are shown below.","metadata":{}},{"cell_type":"markdown","source":"## Task\nOur task is to detect and delineate distinct objects of interest in biological images depicting neuronal cell types commonly used in the study of neurological disorders.","metadata":{}},{"cell_type":"code","source":"import os\nimport cv2\n\nimport numpy as np \nimport pandas as pd\n\nimport plotly.express as px\nimport seaborn as sns\nimport matplotlib.pyplot as plt\n\ndf_train = pd.read_csv('../input/sartorius-cell-instance-segmentation/train.csv')","metadata":{"execution":{"iopub.status.busy":"2021-11-29T02:09:06.225263Z","iopub.execute_input":"2021-11-29T02:09:06.225992Z","iopub.status.idle":"2021-11-29T02:09:09.354564Z","shell.execute_reply.started":"2021-11-29T02:09:06.225892Z","shell.execute_reply":"2021-11-29T02:09:09.353651Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### 1. Label\nThe meta data, i.e. lable here, is given in the `train.csv` file which containing 2 numerical features and 7 categorical. This dataset is well-designed contains no empty value.","metadata":{}},{"cell_type":"code","source":"df_train.head()","metadata":{"execution":{"iopub.status.busy":"2021-11-29T02:09:09.357315Z","iopub.execute_input":"2021-11-29T02:09:09.357574Z","iopub.status.idle":"2021-11-29T02:09:09.381288Z","shell.execute_reply.started":"2021-11-29T02:09:09.357541Z","shell.execute_reply":"2021-11-29T02:09:09.380453Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"df_train.info()","metadata":{"execution":{"iopub.status.busy":"2021-11-29T02:09:09.382587Z","iopub.execute_input":"2021-11-29T02:09:09.382837Z","iopub.status.idle":"2021-11-29T02:09:09.466328Z","shell.execute_reply.started":"2021-11-29T02:09:09.382807Z","shell.execute_reply":"2021-11-29T02:09:09.465427Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### 1.1 Image Information\nFrom the `train.csv`, all the images have the same shape *704 x 520* which contains no variable image resolution problem. However, there is only 606 images in the tarin set. The number of rows in this file is far more than 606 which indicates there are more than 1 instances in 1 image.","metadata":{}},{"cell_type":"code","source":"print(f'Number of images: {df_train.id.nunique()}')","metadata":{"execution":{"iopub.status.busy":"2021-11-29T02:09:09.468603Z","iopub.execute_input":"2021-11-29T02:09:09.468858Z","iopub.status.idle":"2021-11-29T02:09:09.480830Z","shell.execute_reply.started":"2021-11-29T02:09:09.468828Z","shell.execute_reply":"2021-11-29T02:09:09.479908Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"The number of instances in each image is different which varys from 4 to 790.","metadata":{}},{"cell_type":"code","source":"fig, ax = plt.subplots()\n\nninstances_per_image = df_train[['id']].value_counts().sort_values()\nninstances_per_image.index = range(606)\nninstances_per_image.median()\nninstances_per_image.plot.bar(ax=ax)\n\nax.set_xticklabels([])\nax.set_xlabel('Images')\nax.set_ylabel('Number of Instances')\nplt.show()","metadata":{"execution":{"iopub.status.busy":"2021-11-29T02:09:09.482144Z","iopub.execute_input":"2021-11-29T02:09:09.482609Z","iopub.status.idle":"2021-11-29T02:09:12.726068Z","shell.execute_reply.started":"2021-11-29T02:09:09.482576Z","shell.execute_reply":"2021-11-29T02:09:12.725426Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Observing this data, it is easy to find each image is associated with a unique cell type. Those types are cort (neurons), shsy5y (neuroblastoma) and astro (astrocytes).","metadata":{}},{"cell_type":"code","source":"fig, ax = plt.subplots(1, 1)\ndf_train.groupby(['id','cell_type'])['cell_type'].first().value_counts().plot.bar(ax=ax)\nax.set_ylabel('Number of Images')\nax.set_xlabel('Cell Types')\nfig.tight_layout()\nplt.show()","metadata":{"execution":{"iopub.status.busy":"2021-11-29T02:09:12.727153Z","iopub.execute_input":"2021-11-29T02:09:12.727486Z","iopub.status.idle":"2021-11-29T02:09:12.987516Z","shell.execute_reply.started":"2021-11-29T02:09:12.727457Z","shell.execute_reply":"2021-11-29T02:09:12.986636Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## 2. Image\nHere we show 3 image of each cell types.","metadata":{}},{"cell_type":"code","source":"def decode_rle_mask(rle_mask, shape):\n\n    rle_mask = rle_mask.split()\n    starts, lengths = [np.asarray(x, dtype=int) for x in (rle_mask[0:][::2], rle_mask[1:][::2])]\n    starts -= 1\n    ends = starts + lengths\n\n    mask = np.zeros((shape[0] * shape[1]), dtype=np.uint8)\n    for start, end in zip(starts, ends):\n        mask[start:end] = 1\n\n    mask = mask.reshape(shape[0], shape[1])\n    return mask\n\ndef visualize_image(df, image_id):   \n    image_path = df.loc[df['id'] == image_id, 'id'].values[0]\n    cell_type = df.loc[df['id'] == image_id, 'cell_type'].values[0]\n    plate_time = df.loc[df['id'] == image_id, 'plate_time'].values[0]\n    sample_date = df.loc[df['id'] == image_id, 'sample_date'].values[0]\n    sample_id = df.loc[df['id'] == image_id, 'sample_id'].values[0]\n\n    image = cv2.imread(f'../input/sartorius-cell-instance-segmentation/train/{image_path}.png')\n    image = cv2.cvtColor(image, cv2.COLOR_BGR2GRAY)\n\n    fig, axes = plt.subplots(figsize=(10, 10), ncols=2)\n    fig.tight_layout(pad=5.0)\n    \n    axes[0].imshow(image, cmap='gray')\n    masks = []\n    for mask in df.loc[df['id'] == image_id, 'annotation'].values:\n        decoded_mask = decode_rle_mask(rle_mask=mask, shape=image.shape)\n        masks.append(decoded_mask)\n    mask = np.stack(masks)\n    mask = np.any(mask == 1, axis=0)\n    axes[1].imshow(image, cmap='gray')\n    axes[1].imshow(mask, alpha=0.4)\n\n    for i in range(2):\n        axes[i].set_xlabel('')\n        axes[i].set_ylabel('')\n        axes[i].tick_params(axis='x', labelsize=10, pad=10)\n        axes[i].tick_params(axis='y', labelsize=10, pad=10)\n        \n    axes[0].set_title(f'{image_path} - {cell_type} Annotations\\n{plate_time} - {sample_date} - {sample_id}', fontsize=10, pad=12)\n    axes[1].set_title('Segmentation Mask', fontsize=10, pad=12)\n    plt.show()\n    plt.close(fig)","metadata":{"execution":{"iopub.status.busy":"2021-11-29T02:09:12.988750Z","iopub.execute_input":"2021-11-29T02:09:12.988995Z","iopub.status.idle":"2021-11-29T02:09:13.007944Z","shell.execute_reply.started":"2021-11-29T02:09:12.988964Z","shell.execute_reply":"2021-11-29T02:09:13.007055Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"#### astro","metadata":{}},{"cell_type":"code","source":"select_image_ids = []\nselect_image_ids.append(df_train.loc[df_train['cell_type'] == 'astro', 'id'].sample(1).to_list()[0])\nselect_image_ids.append(df_train.loc[df_train['cell_type'] == 'astro', 'id'].sample(2).to_list()[0])\nselect_image_ids.append(df_train.loc[df_train['cell_type'] == 'astro', 'id'].sample(3).to_list()[0])\n\nfor image_id in select_image_ids:\n     visualize_image(df=df_train, image_id=image_id)","metadata":{"execution":{"iopub.status.busy":"2021-11-29T02:09:13.009600Z","iopub.execute_input":"2021-11-29T02:09:13.010638Z","iopub.status.idle":"2021-11-29T02:09:15.400899Z","shell.execute_reply.started":"2021-11-29T02:09:13.010587Z","shell.execute_reply":"2021-11-29T02:09:15.399971Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"#### cort","metadata":{}},{"cell_type":"code","source":"select_image_ids = []\nselect_image_ids.append(df_train.loc[df_train['cell_type'] == 'cort', 'id'].sample(1).to_list()[0])\nselect_image_ids.append(df_train.loc[df_train['cell_type'] == 'cort', 'id'].sample(2).to_list()[0])\nselect_image_ids.append(df_train.loc[df_train['cell_type'] == 'cort', 'id'].sample(3).to_list()[0])\n\nfor image_id in select_image_ids:\n     visualize_image(df=df_train, image_id=image_id)","metadata":{"execution":{"iopub.status.busy":"2021-11-29T02:09:15.402142Z","iopub.execute_input":"2021-11-29T02:09:15.402404Z","iopub.status.idle":"2021-11-29T02:09:17.435926Z","shell.execute_reply.started":"2021-11-29T02:09:15.402373Z","shell.execute_reply":"2021-11-29T02:09:17.435038Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"#### shsy5y","metadata":{}},{"cell_type":"code","source":"select_image_ids = []\nselect_image_ids.append(df_train.loc[df_train['cell_type'] == 'shsy5y', 'id'].sample(1).to_list()[0])\nselect_image_ids.append(df_train.loc[df_train['cell_type'] == 'shsy5y', 'id'].sample(2).to_list()[0])\nselect_image_ids.append(df_train.loc[df_train['cell_type'] == 'shsy5y', 'id'].sample(3).to_list()[0])\n\nfor image_id in select_image_ids:\n     visualize_image(df=df_train, image_id=image_id)","metadata":{"execution":{"iopub.status.busy":"2021-11-29T02:09:17.438200Z","iopub.execute_input":"2021-11-29T02:09:17.438429Z","iopub.status.idle":"2021-11-29T02:09:19.935481Z","shell.execute_reply.started":"2021-11-29T02:09:17.438395Z","shell.execute_reply":"2021-11-29T02:09:19.934490Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"* `astro` instances are the biggest in shape\n* `cort` instances are smaller than the other cell and circle-like\n* `shsy5y` are slightly bigger and more abundant than cort","metadata":{}}]}