{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"pygments_lexer":"ipython3","nbconvert_exporter":"python","version":"3.6.4","file_extension":".py","codemirror_mode":{"name":"ipython","version":3},"name":"python","mimetype":"text/x-python"}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"code","source":"# This Python 3 environment comes with many helpful analytics libraries installed\n# It is defined by the kaggle/python Docker image: https://github.com/kaggle/docker-python\n# For example, here's several helpful packages to load\n\nimport numpy as np # linear algebra\nimport pandas as pd # data processing, CSV file I/O (e.g. pd.read_csv)\n\n# Input data files are available in the read-only \"../input/\" directory\n# For example, running this (by clicking run or pressing Shift+Enter) will list all files under the input directory\n\nimport os\n# for dirname, _, filenames in os.walk('/kaggle/input'):\n#     for filename in filenames:\n#         print(os.path.join(dirname, filename))\n\n# You can write up to 20GB to the current directory (/kaggle/working/) that gets preserved as output when you create a version using \"Save & Run All\" \n# You can also write temporary files to /kaggle/temp/, but they won't be saved outside of the current session","metadata":{"_kg_hide-input":true,"execution":{"iopub.status.busy":"2022-09-30T15:47:09.921577Z","iopub.execute_input":"2022-09-30T15:47:09.922010Z","iopub.status.idle":"2022-09-30T15:47:09.960383Z","shell.execute_reply.started":"2022-09-30T15:47:09.921921Z","shell.execute_reply":"2022-09-30T15:47:09.958864Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## HuBMAP + HPA - Hacking the Human body\n**- Segment multi-organ functional tissue units**","metadata":{}},{"cell_type":"markdown","source":"The goal of this competition is to segment functional tissue units (FTUs) across five human organs. The competition provides data from two different sources, Human Protein Atlas(HPA) and Human BioMolecular Atlas Program(HuBMAP). The participant uses only public HPA data for training, and predicts private HuBMAP data. Because data originates from two different sources, their sizes and stain methods are dissimilar. Hence, the pariticipant needs to manage the domain difference as well as the segmentation of FTUs. ","metadata":{}},{"cell_type":"markdown","source":"### 1. Managing the domain difference","metadata":{}},{"cell_type":"markdown","source":"Augmenting data in train process significantly reduced the gap between HPA and HuBMAP data. I created and used the second data for training. I resized HPA data in pixel size and tissue thickness in order to match the shape of HuBMAP data. And then I stain-normalized the modified data based on the given image of 10078.tiff. The examples are shown below. Every original image was copied and modified in the above way, and 351 of modified images were created. So, 702 of the images were used for training: 351 original images and 351 modified images. Training models with the new data set as well as standard augmentation techniques successfully played the role of eliminating the gap between Cross Validation score(CV) and LeaderBoard score(LB), which was the domain difference.","metadata":{}},{"cell_type":"code","source":"import matplotlib.pyplot as plt\nimport cv2\n\ndf = pd.read_csv('../input/hubmap-organ-segmentation/train.csv')\nprint_list = ['prostate', 'largeintestine', 'spleen', 'lung', 'kidney']\nplt.figure(figsize=(12,7))\nfor _, r in df.iterrows():\n    if r.organ == 'prostate':\n        ori = plt.imread('../input/hub512-6ps/images/' + str(r.id) + '_0000.png')\n        sam = plt.imread('../input/hub512-6ps/images/' + str(r.id) + '_0001.png')\n        plt.subplot(1,2,1); plt.imshow(ori); plt.axis('off'); plt.title('original')\n        plt.subplot(1,2,2); plt.imshow(sam); plt.axis('off'); plt.title('modified')\n    break\nplt.show()","metadata":{"_kg_hide-input":true,"execution":{"iopub.status.busy":"2022-09-30T15:47:09.977134Z","iopub.execute_input":"2022-09-30T15:47:09.977491Z","iopub.status.idle":"2022-09-30T15:47:11.093711Z","shell.execute_reply.started":"2022-09-30T15:47:09.977464Z","shell.execute_reply":"2022-09-30T15:47:11.092655Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### 2. Segmenting functional tissue units","metadata":{}},{"cell_type":"markdown","source":"Transformer models empirically achieved the segmentation task better than CNN-based models. For example, vanila transformer models achieved average 0.65 of dice coefficient score while simple unet with CNN encoder did 0.55 on average. Transformer models observed images with bigger receptive field than CNN-based models, and it was seemingly suitable for medical image segmentation task that required to be analyzed in global context. A single transformer model was picked and optimized for this task: Co-Scale Conv-Attentional Image Transformers(CoaT). \n\n    1. Model\n        - Encoder: Co-Scale Conv-Attentional Image Transformers(CoaT)\n            - added one serial block and one parallel block\n        - Decoder: 3x3 convolution layer\n        - Aux decoder\n        \n    2. Training\n        - Data: mixed dataset (351 original + 351 modified)\n            - Image size: 1024 x 1024\n        - Loss: binary cross entropy + aux loss from each parallel block (1:0.8) \n        - Optimizer: AdamW with warmup start and gradual decaying\n        - Augmentations: crop, flip, random rotate, hue saturation value, random brightness\n\n    3. Inference\n        - Stochastic weight averaging\n        - Single model by ensembling each single fold model\n        - TTA: flip (horizontal, vertical, and diagonal), stain-normalization, luminosityStandardizer\n        - Thresholds for each organ\n            - Kidney: 0.45\n            - Large intestine: 0.3\n            - Lung: 0.1\n            - Prostate: 0.25\n            - Spleen: 0.35\n\nThe following table shows the model performance:\n\n|                 | Public  | Private |\n|-----------------|---------|---------|\n| Kidney          | 0.12079 | 0.16734 |\n| Large intestine | 0.05776 | 0.08930 |\n| Lung            | 0.09863 | 0.18821 |\n| Prostate        | 0.14672 | 0.17763 |\n| Spleen          | 0.17602 | 0.18324 |\n| Total           | 0.59992 | 0.80583 |\n","metadata":{}},{"cell_type":"markdown","source":"### 3. Q & A","metadata":{}},{"cell_type":"markdown","source":"1. Are the statistical and modeling methods used to identify FTUs appropriate for the task?\n    - Transformer model performs way better than expected. The model performance other than on lung part was outstanding compared to other models within top 30. Thus, transformer model is appropriate to identify FTUs.\n\n2. Are confidence scores and other metrics provided that help interpret the results achieved by the segmentation methods?\n    - Dice coefficient explicitly helps interpret the results of the segmentation methods. However, dice coefficient weighs twice more on true positive than false positive. This leads a model to preferably predict pixels as true, because predicting true positive returns bigger than reducing false positive. For example, a model predicts two pixels: one is true and the other is false. The model is likely to predict both pixels as true when it is not confident. The model still gets the score, because predicting correct true positive is twice returning than predicting correct false positive. The model maybe becomes biased towards true. So, other metrics such Intersection over Union (IoU) or confidence scores such as organ thresholds are required to balance out.\n\n3. Is the presented characterization of FTUs useful for understanding individual differences, e.g., the impact of donor sex, age on the size, shape, or spatial distribution of FTUs?\n    - Sex seems a useful indicator to understand individual differences while age does not show any impact on distribution of FTUs.\n    \n4. Is it possible to predict FTU area size distribution, given age and sex info across all organs?\n    - Yes, it seems possible because FTU area size in each organ is differently distributed based on sex, except lung. For example, female generally has bigger FTU area size than male's FTU area size in kidney, large intestine, and spleen. Also, female does not have prostate so prostate size is always zero. However, predicting FTU area size distribution based on age does not show any regular pattern so age is not a good indicator to predict.\n    - But, this result is based on the small sample size so it can be different in practice.","metadata":{}},{"cell_type":"code","source":"import plotly.graph_objects as go\nimport plotly.express as px\n\ndef rgb2hex(rgb_tuple):\n    r,g,b = rgb_tuple\n    def clamp(x): \n        return max(0, min(x, 255))\n    return \"#{0:02x}{1:02x}{2:02x}\".format(clamp(r), clamp(g), clamp(b))\nORGANS = ['kidney', 'largeintestine', 'lung', 'prostate', 'spleen']\n_COLOURS = [(230, 0, 73), (11, 180, 255), (80, 233, 145), (230, 216, 0), (155, 25, 245)]\nO2C_MAP = {_o:_c for _o,_c in zip(ORGANS, _COLOURS)}\nO2C_HEX_MAP = {_o:rgb2hex(_c) for _o,_c in O2C_MAP.items()}\n\nfig = px.histogram(df, df[\"sex\"], color_discrete_map=O2C_HEX_MAP, color=\"organ\", \n                   title=\"<b>Examples by Gender (0=Male, 1=Female) and Organ Type</b>\",\n                   labels={\"organ\":\"<b>Organ Legend</b>\", \n                         \"count\":\"<b>Number Of Observations</b>\",\n                         \"sex\":\"<b>Sex</b>\"},text_auto=True, barmode=\"group\",\n                  )\nfig.show()","metadata":{"_kg_hide-input":true,"execution":{"iopub.status.busy":"2022-09-30T15:47:11.095216Z","iopub.execute_input":"2022-09-30T15:47:11.095623Z","iopub.status.idle":"2022-09-30T15:47:14.905128Z","shell.execute_reply.started":"2022-09-30T15:47:11.095586Z","shell.execute_reply":"2022-09-30T15:47:14.903487Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"def enc2size(enc):\n    s = enc.split()    \n    area = 0\n    for i in range(len(s)//2):\n        area += int(s[2*i+1])\n    return area\n\ndef sizeCal(df, dic_area):\n    dic_area = dic_area\n    for i, r in df.iterrows():        \n        area = 0\n        area = enc2size(r.rle)\n        for k, v in dic_area.items():\n            if k == r.organ:\n                v += area\n                dic_area[k] = v\n    for k, v in dic_area.items():\n        v /= len(df)\n        dic_area[k] = round(v, 2)\n    return dic_area\n\n\nfem_df = df[df.loc[:,'sex'] == 'Female'].reset_index(drop=True)\nm_df = df[df.loc[:,'sex'] == 'Male'].reset_index(drop=True)\n\nm_area = {'prostate': 0., 'largeintestine': 0., 'lung': 0., 'spleen': 0., 'kidney': 0.}\nfem_area = {'prostate': 0., 'largeintestine': 0., 'lung': 0., 'spleen': 0., 'kidney': 0.}\n\na = sizeCal(m_df, m_area)\nb = sizeCal(fem_df, fem_area)\n# print(a)\n# print(b)\nX_axis = np.arange(len(a))\nplt.figure(figsize=(12,8))\nplt.xticks(X_axis, a.keys())\nplt.bar(X_axis - 0.2, a.values(), 0.4, label = 'Male')\nplt.bar(X_axis + 0.2, b.values(), 0.4, label = 'Female')\nplt.title('Organ size')\nplt.legend()\nplt.show()","metadata":{"_kg_hide-input":true,"execution":{"iopub.status.busy":"2022-09-30T16:01:20.603726Z","iopub.execute_input":"2022-09-30T16:01:20.604113Z","iopub.status.idle":"2022-09-30T16:01:21.502994Z","shell.execute_reply.started":"2022-09-30T16:01:20.604086Z","shell.execute_reply":"2022-09-30T16:01:21.501263Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"m_df = df[df.loc[:,'sex'] == 'Male'].reset_index(drop=True)\n\nm_dic = m_df.age.unique().tolist()\nm_dic.sort()\nm_dic = list(map(int, m_dic))\nm_dic = list(map(str, m_dic))\nm_dic = dict.fromkeys(m_dic, 0.)\n            \ndef sizeCal2(df, dic, org):\n    count = 0.\n    for i, r in df.iterrows():\n        area = 0.        \n        if r.organ == org:\n            count += 1\n            area = enc2size(r.rle)\n            for k, v in dic.items():\n                if float(k) == r.age:\n                    v += area\n                    dic[k] = v\n    for k, v in dic.items():\n        if count != 0.:\n            v /= count\n        dic[k] = round(v, 2)\n    return dic\n\nplt.rcParams.update({'font.size': 25})\norg_list = ['kidney','largeintestine','lung','prostate', 'spleen']\nplt.figure(figsize=(30,20))\nfor i, org in enumerate(org_list):\n    temp_dic = m_dic.copy()\n    a = sizeCal2(m_df, temp_dic, org)\n    a = {k:v for k,v in a.items() if v != 0.0}\n#     print(a)\n    plt.subplot(2,3,i+1)\n    plt.plot(a.keys(), a.values()); plt.title(f'Male {org}');plt.yticks([]); plt.xlabel('age'); plt.ylabel('area')\n\n\n# plt.show()","metadata":{"_kg_hide-input":true,"execution":{"iopub.status.busy":"2022-09-30T16:01:29.181404Z","iopub.execute_input":"2022-09-30T16:01:29.181799Z","iopub.status.idle":"2022-09-30T16:01:30.458637Z","shell.execute_reply.started":"2022-09-30T16:01:29.181770Z","shell.execute_reply":"2022-09-30T16:01:30.457179Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"plt.rcParams.update({'font.size': 25})\n\nfem_df = df[df.loc[:,'sex'] == 'Female'].reset_index(drop=True)\n\nfem_dic = fem_df.age.unique().tolist()\nfem_dic.sort()\nfem_dic = list(map(int, fem_dic))\nfem_dic = list(map(str, fem_dic))\nfem_dic = dict.fromkeys(fem_dic, 0.)\n\norg_list = ['kidney','largeintestine','lung','prostate', 'spleen']\nplt.figure(figsize=(30,20))\nfor i, org in enumerate(org_list):\n    temp_dic = fem_dic.copy()\n    a = sizeCal2(fem_df, temp_dic, org)\n    a = {k:v for k,v in a.items() if v != 0.0}\n    plt.subplot(2,3,i+1)\n    plt.plot(a.keys(), a.values()); plt.title(f'Female {org}');plt.yticks([]); plt.xlabel('age'); plt.ylabel('area')\n\n# plt.show()","metadata":{"_kg_hide-input":true,"execution":{"iopub.status.busy":"2022-09-30T16:03:11.468768Z","iopub.execute_input":"2022-09-30T16:03:11.469083Z","iopub.status.idle":"2022-09-30T16:03:12.190064Z","shell.execute_reply.started":"2022-09-30T16:03:11.469059Z","shell.execute_reply":"2022-09-30T16:03:12.188198Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"5. Did the team validate their methods and algorithm implementations and provide information on algorithm performance and limitations?\n    - Yes, I validated all the methods and algorithm implementation that I tested. For example, I checked cross validation score and leaderboard score whenever I modified my implementation. And I visualized trained and tested images in order to verify my methods. Every corresponding change was documented.\n\n6. Did the team document their method and code appropriately?\n    - Yes, every method and code that I experimented were accordingly documented.    \n    \n7. Did the team develop a creative or novel method to segment FTUs?\n    - No, I developed one of classic segmentation methods such as transformer models, and finetuned the model to be suitable for this task. \n\n8. Did the team provide insights that would be useful for generating reference FTUs for inclusion into a Human Reference Atlas?\n    - Yes, I did.\n    \n9. Does the team embrace diversity and equity, welcoming team members of different ages, genders, ethnicities, and with multiple backgrounds and perspectives?\n    - I soloed this competition, and so there was no diversity and equity, although I always respect the aspects.\n    \n10. Did the authors effectively communicate the details of their method for segmenting FTUs, and the quality and limitations of their results? For example, did they use data visualizations to present algorithm setup, run, results and/or to provide insight into the comparative performance of different methods? Were these visualizations effective at communicating insights about their approach and results to experts and novice users?\n    - Yes, I did.\n    \n11. Are the important results easily understood by the average person?\n    - Yes, the results are numerically documented and visualized so that the average person would have no trouble in understanding.    ","metadata":{}},{"cell_type":"markdown","source":"### 4. Conclusion","metadata":{}},{"cell_type":"markdown","source":"The results should not be considered as always true, because they were inferred from small size of data. The reader must be aware of different results that may be led by these methods in practice.","metadata":{}}]}