{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"pygments_lexer":"ipython3","nbconvert_exporter":"python","version":"3.6.4","file_extension":".py","codemirror_mode":{"name":"ipython","version":3},"name":"python","mimetype":"text/x-python"}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"# Object 目的\n- Understanding overview of the competition.<br>\nコンペの概要を把握する。\n- Understaning the train data and perform a simple visualization.<br>\n学習データを理解して、簡易的なビジュアライゼーションを行う。","metadata":{}},{"cell_type":"markdown","source":"# 1. Overview 概要\n- **the problem we'd like to solve 解きたい問題**<br>\nyou'll help accelerate the world’s understanding of the relationships between cell and tissue organization.<br>\n細胞や組織の構造に関する理解を加速させ、人間の健康、長寿に繋げたい。\n- **the problem to be solved 解くべき問題**<br>\nyou’ll identify and segment functional tissue units (FTUs) across five human organs. You'll build your model using a dataset of tissue section images, with the best submissions segmenting FTUs as accurately as possible.<br>\n人間の組織切片画像のデータセットを使用してモデルを構築し、functional tissue units (FTUs) セグメンテーションする。\n- **train data 学習データ**<br>\nbiopsy slides from several different organs<br>\nThe training dataset consists of data from public HPA data, the public test set is a combination of private HPA data and HuBMAP data, and the private test set contains only HuBMAP data.<br>\n複数の異なる臓器の生検スライド<br>\n学習データセットは公開されているHPAのデータ<br>\n公開テストセットは非公開のHPAデータとHuBMAPデータの組み合わせ<br>\n非公開テストセットはHuBMAPデータのみ\n- **Supervised ML Evaluation 評価手法**<br>\nmean Dice coefficient is familiar in image segmentation.<br>\nThe Dice coefficient can be used to compare the pixel-wise agreement between a predicted segmentation and its corresponding ground truth.<br>\n画像セグメンテーションではおなじみのDice係数<br>\n2つの集合の平均要素数と共通要素数の割合で、1に近づくほど高スコア<br>\n$$\n\\frac{2*|X\\cap Y|}{|X|+|Y|}\\\\\n$$\n$X$:predicted set of pixels 予測したピクセルの集合<br>\n$Y$:ground truth 正解ピクセルの集合<br>","metadata":{}},{"cell_type":"markdown","source":"# 2. Understanding Train Data 学習データの理解","metadata":{}},{"cell_type":"code","source":"import os\nimport cv2\nimport tifffile\nimport numpy as np\nimport pandas as pd\nimport matplotlib.pyplot as plt\nimport seaborn as sns","metadata":{"execution":{"iopub.status.busy":"2022-08-12T11:40:19.387994Z","iopub.execute_input":"2022-08-12T11:40:19.389477Z","iopub.status.idle":"2022-08-12T11:40:20.531749Z","shell.execute_reply.started":"2022-08-12T11:40:19.389312Z","shell.execute_reply":"2022-08-12T11:40:20.530510Z"},"_kg_hide-input":true,"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### 2-1. Data Summary データの外観","metadata":{}},{"cell_type":"code","source":"df_train = pd.read_csv('../input/hubmap-organ-segmentation/train.csv')\ndf_test = pd.read_csv('../input/hubmap-organ-segmentation/test.csv')\ndf_train.sample(5)","metadata":{"execution":{"iopub.status.busy":"2022-08-12T11:40:22.868910Z","iopub.execute_input":"2022-08-12T11:40:22.869810Z","iopub.status.idle":"2022-08-12T11:40:23.283656Z","shell.execute_reply.started":"2022-08-12T11:40:22.869752Z","shell.execute_reply":"2022-08-12T11:40:23.282225Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"The \"organ\" contains the category of the organ and the \"rle\" contains the mask rle encoding.<br>\norganに臓器のカテゴリ、rleにマスク情報が入っている。","metadata":{}},{"cell_type":"code","source":"df_train.describe(include='all')","metadata":{"execution":{"iopub.status.busy":"2022-08-12T11:40:25.231097Z","iopub.execute_input":"2022-08-12T11:40:25.231813Z","iopub.status.idle":"2022-08-12T11:40:25.303467Z","shell.execute_reply.started":"2022-08-12T11:40:25.231771Z","shell.execute_reply":"2022-08-12T11:40:25.302332Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"df_train.isnull().sum()","metadata":{"execution":{"iopub.status.busy":"2022-08-12T11:47:17.787464Z","iopub.execute_input":"2022-08-12T11:47:17.788235Z","iopub.status.idle":"2022-08-12T11:47:17.799589Z","shell.execute_reply.started":"2022-08-12T11:47:17.788193Z","shell.execute_reply":"2022-08-12T11:47:17.798529Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Deficiencies don't exist.<br>\n欠損は存在しない。","metadata":{}},{"cell_type":"code","source":"df_test","metadata":{"execution":{"iopub.status.busy":"2022-08-12T11:41:44.665277Z","iopub.execute_input":"2022-08-12T11:41:44.665720Z","iopub.status.idle":"2022-08-12T11:41:44.680224Z","shell.execute_reply.started":"2022-08-12T11:41:44.665684Z","shell.execute_reply":"2022-08-12T11:41:44.678793Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Hubmap has only one public test data set.<br>\nHubmapの公開テストデータは1つだけ。","metadata":{}},{"cell_type":"markdown","source":"### 2-2. Sex 性別","metadata":{}},{"cell_type":"code","source":"sns.countplot(x='sex',data=df_train)","metadata":{"execution":{"iopub.status.busy":"2022-08-12T11:41:58.978530Z","iopub.execute_input":"2022-08-12T11:41:58.979073Z","iopub.status.idle":"2022-08-12T11:41:59.202048Z","shell.execute_reply.started":"2022-08-12T11:41:58.979022Z","shell.execute_reply":"2022-08-12T11:41:59.201027Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Male have about twice as much data.<br>\n男性の方が約2倍データが多い。","metadata":{}},{"cell_type":"markdown","source":"### 2-3. Organ 臓器","metadata":{}},{"cell_type":"code","source":"sns.countplot(x='organ', hue='sex',data=df_train)","metadata":{"execution":{"iopub.status.busy":"2022-08-12T11:58:18.884481Z","iopub.execute_input":"2022-08-12T11:58:18.884934Z","iopub.status.idle":"2022-08-12T11:58:19.110187Z","shell.execute_reply.started":"2022-08-12T11:58:18.884890Z","shell.execute_reply":"2022-08-12T11:58:19.109161Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"- The prostate is in male data only.<br>\n前立腺は男性のみ存在する臓器\n- The kidney common in male data.<br>\n肝臓は男性データに多い\n- Almost the same number of the spleen, the lung organ and the largeintestine for male and female.<br>\n脾臓、肺臓、大腸は、男女でだいたい同じ数。","metadata":{}},{"cell_type":"markdown","source":"### 2-4. Data Source データソース","metadata":{}},{"cell_type":"code","source":"fig = plt.figure(figsize=(8,3))\nax1 = fig.add_subplot(1, 2, 1)\nax1 = df_train['data_source'].value_counts().plot.bar()\nax1.set_title('train')\nax2 = fig.add_subplot(1, 2, 2)\nax2 = df_test['data_source'].value_counts().plot.bar()\nax2.set_title('test')\nfig.show()","metadata":{"execution":{"iopub.status.busy":"2022-08-12T11:59:07.967023Z","iopub.execute_input":"2022-08-12T11:59:07.967486Z","iopub.status.idle":"2022-08-12T11:59:08.249468Z","shell.execute_reply.started":"2022-08-12T11:59:07.967435Z","shell.execute_reply":"2022-08-12T11:59:08.248169Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"All train data are HPA.<br>\n学習データは全てHPA。","metadata":{}},{"cell_type":"markdown","source":"### 2-5. Image Size 画像サイズ","metadata":{"execution":{"iopub.status.busy":"2022-08-12T12:00:19.517609Z","iopub.execute_input":"2022-08-12T12:00:19.518091Z","iopub.status.idle":"2022-08-12T12:00:19.523452Z","shell.execute_reply.started":"2022-08-12T12:00:19.518053Z","shell.execute_reply":"2022-08-12T12:00:19.522479Z"}}},{"cell_type":"code","source":"sns.jointplot(x='img_width',y='img_height',data=df_train)\nplt.show()","metadata":{"execution":{"iopub.status.busy":"2022-08-12T12:00:23.470644Z","iopub.execute_input":"2022-08-12T12:00:23.471109Z","iopub.status.idle":"2022-08-12T12:00:24.057567Z","shell.execute_reply.started":"2022-08-12T12:00:23.471070Z","shell.execute_reply":"2022-08-12T12:00:24.056679Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"The train image is mostly square images.<br>\nほとんど正方形の画像、ほとんどが3000×3000の画像サイズだが、例外もある。","metadata":{}},{"cell_type":"markdown","source":"### 2-6. Pixel Size ピクセルサイズ","metadata":{}},{"cell_type":"code","source":"sns.countplot(x='pixel_size',data=df_train)\nplt.show()","metadata":{"execution":{"iopub.status.busy":"2022-08-12T12:01:24.110313Z","iopub.execute_input":"2022-08-12T12:01:24.110811Z","iopub.status.idle":"2022-08-12T12:01:24.281968Z","shell.execute_reply.started":"2022-08-12T12:01:24.110774Z","shell.execute_reply":"2022-08-12T12:01:24.280977Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"All HPA images have a pixel size of 0.4 µm.<br>\nFor HuBMAP imagery the pixel size is 0.5 µm for kidney, 0.2290 µm for large intestine, 0.7562 µm for lung, 0.4945 µm for spleen, and 6.263 µm for prostate.<br>\n学習データのHPAデータは全て0.4µm。<br>\nHuBMAP画像は、腎臓が0.5μm、大腸が0.2290μm、肺が0.7562μm、脾臓が0.4945μm、前立腺が6.263μm。","metadata":{}},{"cell_type":"markdown","source":"### 2-7. Tissue Thickness 生検サンプルの厚み","metadata":{"execution":{"iopub.status.busy":"2022-08-12T12:05:19.086887Z","iopub.execute_input":"2022-08-12T12:05:19.087384Z","iopub.status.idle":"2022-08-12T12:05:19.093052Z","shell.execute_reply.started":"2022-08-12T12:05:19.087348Z","shell.execute_reply":"2022-08-12T12:05:19.091755Z"}}},{"cell_type":"code","source":"sns.countplot(x='tissue_thickness',data=df_train)\nplt.show()","metadata":{"execution":{"iopub.status.busy":"2022-08-12T12:04:33.229481Z","iopub.execute_input":"2022-08-12T12:04:33.230038Z","iopub.status.idle":"2022-08-12T12:04:33.384907Z","shell.execute_reply.started":"2022-08-12T12:04:33.229996Z","shell.execute_reply":"2022-08-12T12:04:33.383781Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"All HPA images have a thickness of 4 µm.<br>\nThe HuBMAP samples have tissue slice thicknesses 10 µm for kidney, 8 µm for large intestine, 4 µm for spleen, 5 µm for lung, and 5 µm for prostate.<br>\n学習データのHPAデータはすべて厚さ4μm。<br>\nHuBMAPサンプルは、腎臓は10μm、大腸は8μm、脾臓は4μm、肺は5μm、前立腺は5μm。","metadata":{}},{"cell_type":"markdown","source":"### 2-8. Age 年齢","metadata":{}},{"cell_type":"code","source":"(ax1, ax2), = pd.crosstab(pd.cut(df_train['age'], range(0, 91, 5), right=False), df_train['sex']).reindex(columns=['Male', 'Female']).plot.barh(subplots=True, layout=(1, 2), sharex=False)\nax1.invert_xaxis()\nax2.set_yticklabels([])\nax1.get_legend().remove()\nax2.get_legend().remove()\nplt.show()","metadata":{"execution":{"iopub.status.busy":"2022-08-12T12:07:26.029757Z","iopub.execute_input":"2022-08-12T12:07:26.030658Z","iopub.status.idle":"2022-08-12T12:07:26.334079Z","shell.execute_reply.started":"2022-08-12T12:07:26.030611Z","shell.execute_reply":"2022-08-12T12:07:26.332711Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Although there is a large amount of older data, the distribution is disjointed.<br>\n高齢のデータが多いものの、分布はばらばら。","metadata":{}},{"cell_type":"markdown","source":"### 2-9. Mask Information マスク情報","metadata":{}},{"cell_type":"code","source":"df_train['rle']","metadata":{"execution":{"iopub.status.busy":"2022-06-28T09:40:23.510297Z","iopub.execute_input":"2022-06-28T09:40:23.510967Z","iopub.status.idle":"2022-06-28T09:40:23.520086Z","shell.execute_reply.started":"2022-06-28T09:40:23.510929Z","shell.execute_reply":"2022-06-28T09:40:23.518996Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"This encoding method is often used in image segmentation.<br>\nFor example, \"1 3 10 5\" means that pixels 1, 2, 3, 10, 11, 12, 13, and 14 are included in the mask.<br>\nThe same encoding should be used for submission.<br>\nマスク情報はrleエンコーディングされている。画像セグメンテーションでは、よく用いられるエンコーディング方式。<br>\n例えば、「1 3 10 5」は、ピクセル1、2、3、10、11、12、13、14がマスクに含まれることを意味する。<br>\n提出時も同様にエンコーディングする必要がある。","metadata":{}},{"cell_type":"markdown","source":"# 3. Visualization 可視化","metadata":{}},{"cell_type":"code","source":"#https://www.kaggle.com/code/pestipeti/decoding-rle-masks/notebook\ndef mask2rle(img):\n    '''\n    img: numpy array, 1 - mask, 0 - background\n    Returns run length as string formated\n    '''\n    pixels= img.T.flatten()\n    pixels = np.concatenate([[0], pixels, [0]])\n    runs = np.where(pixels[1:] != pixels[:-1])[0] + 1\n    runs[1::2] -= runs[::2]\n    return ' '.join(str(x) for x in runs)\n\n\ndef rle2mask(mask_rle, shape=(3000,3000)):\n    '''\n    mask_rle: run-length as string formated (start length)\n    shape: (width,height) of array to return \n    Returns numpy array, 1 - mask, 0 - background\n\n    '''\n    s = mask_rle.split()\n    starts, lengths = [np.asarray(x, dtype=int) for x in (s[0::2], s[1::2])]\n    starts -= 1\n    ends = starts + lengths\n    img = np.zeros(shape[0]*shape[1], dtype=np.uint8)\n    for lo, hi in zip(starts, ends):\n        img[lo:hi] = 1\n    return img.reshape(shape).T","metadata":{"execution":{"iopub.status.busy":"2022-08-12T12:09:55.431755Z","iopub.execute_input":"2022-08-12T12:09:55.432697Z","iopub.status.idle":"2022-08-12T12:09:55.443509Z","shell.execute_reply.started":"2022-08-12T12:09:55.432648Z","shell.execute_reply":"2022-08-12T12:09:55.442523Z"},"_kg_hide-input":true,"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# 各organごとに、ランダムに1枚選んで表示\norgans = np.unique(df_train['organ'])\nfig, ax = plt.subplots(3,len(organs),figsize=(20,12))\nfor i in range(len(organs)):\n    df_organ = df_train[df_train['organ']==organs[i]].reset_index()\n    idx = np.random.randint(df_organ.shape[0])\n    image = plt.imread(f\"../input/hubmap-organ-segmentation/train_images/{df_organ.id[idx]}.tiff\")\n    mask = rle2mask(df_organ.rle[idx],shape=(df_organ.img_height[idx],df_organ.img_width[idx]))\n    ax[0,i].imshow(image)\n    ax[0,i].set_title(df_organ.organ[idx])\n    ax[0,i].axis(\"off\")\n    ax[1,i].imshow(mask,alpha=0.3,cmap='gray')\n    ax[1,i].axis(\"off\")\n    ax[2,i].imshow(image)\n    ax[2,i].imshow(mask,alpha=0.3,cmap='gray')\n    ax[2,i].axis(\"off\")","metadata":{"execution":{"iopub.status.busy":"2022-08-12T12:09:58.567565Z","iopub.execute_input":"2022-08-12T12:09:58.568393Z","iopub.status.idle":"2022-08-12T12:10:21.895854Z","shell.execute_reply.started":"2022-08-12T12:09:58.568350Z","shell.execute_reply":"2022-08-12T12:10:21.894543Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"The large intestine has a characteristic arrangement, with a relatively high number of FTUs.<br>\n大腸は特徴的な配置をしており、FTUsが比較的多い。","metadata":{}},{"cell_type":"markdown","source":"# 4. Submittion File 提出ファイル","metadata":{}},{"cell_type":"code","source":"sub_df=pd.read_csv('../input/hubmap-organ-segmentation/sample_submission.csv')\nsub_df","metadata":{"execution":{"iopub.status.busy":"2022-06-28T10:12:47.459578Z","iopub.execute_input":"2022-06-28T10:12:47.459918Z","iopub.status.idle":"2022-06-28T10:12:47.476756Z","shell.execute_reply.started":"2022-06-28T10:12:47.459890Z","shell.execute_reply":"2022-06-28T10:12:47.475630Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"The submit csv include the image id and the rle-encoded segmentation.<br>\n画像idとrleエンコーディングしたセグメンテーション情報を提出","metadata":{}},{"cell_type":"code","source":"# 提出する際は、indexをつけないように出力\nsub_df.to_csv('submission.csv',index=False)","metadata":{"execution":{"iopub.status.busy":"2022-06-28T10:14:19.696129Z","iopub.execute_input":"2022-06-28T10:14:19.696513Z","iopub.status.idle":"2022-06-28T10:14:19.707324Z","shell.execute_reply.started":"2022-06-28T10:14:19.696479Z","shell.execute_reply":"2022-06-28T10:14:19.705985Z"},"trusted":true},"execution_count":null,"outputs":[]}]}