{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"pygments_lexer":"ipython3","nbconvert_exporter":"python","version":"3.6.4","file_extension":".py","codemirror_mode":{"name":"ipython","version":3},"name":"python","mimetype":"text/x-python"}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"<h1 style=\"text-align: center;\"><b>EDA Simplified:<span style=\"color:\n    #d13bff;\"> HuBMAP Hacking the Human Vasculature\n</span> (English/Mandarin Version)</b></h1>\n\n<h1 style=\"text-align: center;\"><b>简化探索性数据分析：<span style=\"color:\n    #d13bff;\"> HuBMAP 破解人体脉管系统\n</span> (中英文版)</b></h1>","metadata":{"_uuid":"8f2839f25d086af736a60e9eeb907d3b93b6e0e5","_cell_guid":"b1076dfc-b9ad-4769-8c92-a6c4dae69d19"}},{"cell_type":"markdown","source":"<h2 style=\"background-color: #d13bff; color: white; padding-right: 100vw; background-size:cover; text-align: center; padding: 10px; border-radius: 15px;\">Introduction (介绍)</h2>\n\nPreviously, we covered up all of the data we visualized from HuBMAP's last competition, which is segmenting the organs around the human body such as the spleen, lung, or prostate, as we probed the train_df dataframe we created as well as studying the segmented annotations of each cell in five organs. Right now, we landed into another competition, and it's all about segmenting the vasculature organs from the human body, unlike the forgoing competition about segmenting the five organs.\n\n之前，我们覆盖了我们从 \"HuBMAP\" 的最后一场比赛中可视化的所有数据，当我们探测我们创建的 `train_df` 数据框以及研究分段注释时，它正在分割人体周围的器官，如脾脏、肺或前列腺五个器官中的每个细胞。现在，我们又进入了另一场比赛，这一切都是关于从人体中分割出脉管器官，这与前面提到的分割五个器官的比赛不同。","metadata":{}},{"cell_type":"markdown","source":"<h3 style=\"background-color: #d13bff; color: white; padding-right: 100vw; background-size:cover; text-align: center; padding: 8px; border-radius: 15px;\">Background Info (背景资料)</h3>\n\nThe vasculare system in the human body is made up of all of the vessels that carry blood and lymph fluid across the body. In other words, it's also called the circulatory system. The three main organs of the vascular system are the heart: a muscular organ that pumps blood throughout the body, the blood vessels: which includes the arteries, veins, capillaries, and the blood: which is made up of red and white blood cells as well as plasma and platelets.\n\n人体中的血管系统由将血液和淋巴液输送到全身的所有血管组成。换句话说，它也被称为循环系统。血管系统的三个主要器官是心脏：将血液泵送到全身的肌肉器官，血管：包括动脉、静脉、毛细血管和血液：由红细胞和白细胞组成以及血浆和血小板。\n\n<center>\n    <img src=\"https://cdn.britannica.com/49/115249-050-0DFBBCD3/Human-circulatory-system.jpg\" width=500>\n    <figcaption style=\"color: gray;\">Here's a diagram of the human circulatory/vascular system. Note that blood vessels are running through the body as well as noticing the heart in the middle of the chest. Image credit: Britannica.</figcaption>\n    <figcaption style=\"color: gray;\">这是人体循环/血管系统的图表。请注意，血管穿过身体并注意胸部中间的心脏。图片来源：大英百科全书。</figcaption>\n</center>","metadata":{}},{"cell_type":"markdown","source":"Back to the topic about hacking the vascular system, there's a lot of current efforts to map cells in this system with the Vasculature Common Corrdinate Framework also known as the VCCF, which uses the blood vasculature in the human body. But, the hurdles in what researchers know about the vascular system amplifies another hurdles in the VCCF. On the bright side, if we can use machine learning to segment microvascular arrangements, then researchers could observe the real-world tissue data to begin breaking barriers and then map out the whole circulatory system. And just like what we did in the previous HuBMAP competition, let's engage in another data analysis in this competition on segmenting the instances of circulatory structures from healthy human kidney tissue slides!\n\n回到关于侵入血管系统的话题，目前有很多工作使用血管系统共同坐标框架（脉管系统共同坐标框架，也称为 \"VCCF\"）绘制该系统中的细胞，它使用人体中的血管系统。但是，研究人员对血管系统了解的障碍放大了 \"VCCF\" 的另一个障碍。从好的方面来说，如果我们可以使用机器学习来分割微血管排列，那么研究人员就可以观察真实世界的组织数据，开始打破障碍，然后绘制出整个循环系统。就像我们在之前的 \"HuBMAP\" 比赛中所做的那样，让我们在本次比赛中进行另一项数据分析，从健康的人体肾组织切片中分割循环结构的实例！","metadata":{}},{"cell_type":"markdown","source":"<h2 style=\"background-color: #d13bff; color: white; padding-right: 100vw; background-size:cover; text-align: center; padding: 10px; border-radius: 15px;\">Imports and Data Setup (导入和数据设置)</h2>\n\nTo get started, we load the pandas module as pd followed by loading the numpy module as np for using data science and possibly linear algebra stuff in the data analysis notebook. Next, we import the altair module as alt followed by the plotly module's express attribute as px, graph_objects attribute as go, and the make_subplots function from the plotly module's subplots attribute for plotting interactively on the dataframes created by the pandas module. Lastly, we import the tifffile module for reading out tif or tiff files, since there are medical images of the kidney tissue slides.\n\n首先，我们将 `pandas` 模块加载为 `pd`，然后将 `numpy` 模块加载为 `np`，以便在数据分析笔记本中使用数据科学和可能的线性代数内容。接下来，我们导入 `altair` 模块作为 `alt`，然后导入 `plotly` 模块的 `express` 属性作为 `px`，`graph_objects` 属性作为 `go`，以及来自 `plotly` 模块的 `subplots` 属性的 `make_subplots` 函数，用于在 `pandas` 模块创建的数据帧上进行交互式绘图。最后，我们导入 `tifffile` 模块以读取 *tif* 或 *tiff* 文件，因为有肾组织切片的医学图像。","metadata":{}},{"cell_type":"code","source":"# Data Science and Linear Algebra (数据科学和线性代数)\nimport pandas as pd\nimport numpy as np\n\n# Plotting Graphs (绘制图表)\nimport altair as alt\nimport plotly.express as px\nimport plotly.graph_objects as go\nfrom plotly.subplots import make_subplots\n\n# Reading tiff Files (读取 \"tiff\" 文件)\nimport tifffile","metadata":{"execution":{"iopub.status.busy":"2023-07-06T21:50:36.538567Z","iopub.execute_input":"2023-07-06T21:50:36.539026Z","iopub.status.idle":"2023-07-06T21:50:37.601201Z","shell.execute_reply.started":"2023-07-06T21:50:36.539001Z","shell.execute_reply":"2023-07-06T21:50:37.600141Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Following importing the modules needed for our exploratory data analysis, we characterize the tile_df and wsi_df variables to read out the tile_meta.csv and wsi_meta.csv files for creating the two dataframes. Once completed, we use the head function in the two newly-created dataframes for displaying the first five rows.\n\n在导入我们的探索性数据分析所需的模块之后，我们描述了 `tile_df` 和 `wsi_df` 变量以读出 \"tile_meta.csv\" 和 \"wsi_meta.csv\" 文件以创建两个数据帧。完成后，我们在两个新创建的数据框中使用 `head` 函数来显示前五行。","metadata":{}},{"cell_type":"code","source":"tile_df = pd.read_csv(\"/kaggle/input/hubmap-hacking-the-human-vasculature/tile_meta.csv\")\nwsi_df = pd.read_csv(\"/kaggle/input/hubmap-hacking-the-human-vasculature/wsi_meta.csv\")","metadata":{"execution":{"iopub.status.busy":"2023-07-06T21:50:37.602700Z","iopub.execute_input":"2023-07-06T21:50:37.602973Z","iopub.status.idle":"2023-07-06T21:50:37.640819Z","shell.execute_reply.started":"2023-07-06T21:50:37.602949Z","shell.execute_reply":"2023-07-06T21:50:37.639973Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"tile_df.head()","metadata":{"execution":{"iopub.status.busy":"2023-07-06T21:50:37.641763Z","iopub.execute_input":"2023-07-06T21:50:37.642824Z","iopub.status.idle":"2023-07-06T21:50:37.672242Z","shell.execute_reply.started":"2023-07-06T21:50:37.642796Z","shell.execute_reply":"2023-07-06T21:50:37.671268Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"wsi_df.head()","metadata":{"execution":{"iopub.status.busy":"2023-07-06T21:50:37.674217Z","iopub.execute_input":"2023-07-06T21:50:37.675161Z","iopub.status.idle":"2023-07-06T21:50:37.688721Z","shell.execute_reply.started":"2023-07-06T21:50:37.675134Z","shell.execute_reply":"2023-07-06T21:50:37.687485Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<h3 style=\"background-color: #d13bff; color: white; padding-right: 100vw; background-size:cover; text-align: center; padding: 8px; border-radius: 15px;\">Basic Analysis of the Dataframes (数据框的基本分析)</h3>\n\nNow that we created our two dataframes, let's take a quick basic look at them! First, let's find the number of data entities in the wsi_df and tile_df dataframes by using the len function.\n\n现在我们已经创建了两个数据框，让我们快速基本地看一下它们！首先，让我们使用 `len` 函数找出 `wsi_df` 和 `tile_df` 数据帧中数据实体的数量。","metadata":{}},{"cell_type":"code","source":"len(wsi_df)","metadata":{"execution":{"iopub.status.busy":"2023-07-06T21:50:37.689877Z","iopub.execute_input":"2023-07-06T21:50:37.690390Z","iopub.status.idle":"2023-07-06T21:50:37.701978Z","shell.execute_reply.started":"2023-07-06T21:50:37.690361Z","shell.execute_reply":"2023-07-06T21:50:37.700724Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"len(tile_df)","metadata":{"execution":{"iopub.status.busy":"2023-07-06T21:50:37.703770Z","iopub.execute_input":"2023-07-06T21:50:37.704716Z","iopub.status.idle":"2023-07-06T21:50:37.712726Z","shell.execute_reply.started":"2023-07-06T21:50:37.704681Z","shell.execute_reply":"2023-07-06T21:50:37.711786Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"From what we observed in the number of data entities calculated in the wsi_df and the tile_df dataframe, we counted four data entities in the wsi_df dataframe, and 7033 data entities in the tile_df dataframe. Additionally, the less number of data we noticed in the wsi_df dataframe gave us clues that the metadata in the wsi_df dataframe may be incomplete compared to the tile_df dataframe.\n\n根据我们观察到的在 `wsi_df` 和 `tile_df` 数据帧中计算的数据实体的数量，我们在 `wsi_df` 数据帧中计算了四个数据实体，在 `tile_df` 数据帧中计算了 **7033** 个数据实体。此外，我们在 `wsi_df` 数据帧中注意到的数据数量较少，这给我们提供了线索，即与 `tile_df` 数据帧相比，`wsi_df` 数据帧中的元数据可能不完整。","metadata":{}},{"cell_type":"markdown","source":"Now let's proceed to calculate the number of missing values that are present in the two dataframes! We simply use the isna function towards the wsi_df and tile_df dataframes, and then we use the sum function for calculating the total missing values present in the two specified dataframes.\n\n现在让我们继续计算两个数据框中存在的缺失值的数量！我们简单地对 `wsi_df` 和 `tile_df` 数据帧使用 `isna` 函数，然后我们使用 `sum` 函数计算两个指定数据帧中存在的总缺失值。","metadata":{}},{"cell_type":"code","source":"wsi_df.isna().sum()","metadata":{"execution":{"iopub.status.busy":"2023-07-06T21:50:37.713862Z","iopub.execute_input":"2023-07-06T21:50:37.714664Z","iopub.status.idle":"2023-07-06T21:50:37.732383Z","shell.execute_reply.started":"2023-07-06T21:50:37.714639Z","shell.execute_reply":"2023-07-06T21:50:37.731098Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"tile_df.isna().sum()","metadata":{"execution":{"iopub.status.busy":"2023-07-06T21:50:37.734510Z","iopub.execute_input":"2023-07-06T21:50:37.734927Z","iopub.status.idle":"2023-07-06T21:50:37.746482Z","shell.execute_reply.started":"2023-07-06T21:50:37.734888Z","shell.execute_reply":"2023-07-06T21:50:37.745508Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Splendid! Just like the previous HuBMAP competition, we found no missing values in both dataframes. In other words, the wsi_df and the tile_df dataframes had most of the metadata completely logged without any mistakes.\n\n灿烂！就像之前的 \"HuBMAP\" 比赛一样，我们在两个数据框中都没有发现缺失值。换句话说，`wsi_df` 和 `tile_df` 数据帧的大部分元数据都已完整记录，没有任何错误。","metadata":{}},{"cell_type":"markdown","source":"Last but not least for our basic analysis of the two dataframes, let's count the number of columns in the wsi_df and the tile_df dataframes! To do that, we plug in the shape function in the specified two dataframes to find the shape of it and then extract the last index with the slice index of 1.\n\n最后但同样重要的是，对于我们对两个数据帧的基本分析，让我们计算 `wsi_df` 和 `tile_df` 数据帧中的列数！为此，我们在指定的两个数据帧中插入形状函数以找到它的形状，然后提取切片索引为 **1** 的最后一个索引。","metadata":{}},{"cell_type":"code","source":"wsi_df.shape[1]","metadata":{"execution":{"iopub.status.busy":"2023-07-06T21:50:37.747556Z","iopub.execute_input":"2023-07-06T21:50:37.747854Z","iopub.status.idle":"2023-07-06T21:50:37.756119Z","shell.execute_reply.started":"2023-07-06T21:50:37.747826Z","shell.execute_reply":"2023-07-06T21:50:37.755223Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"tile_df.shape[1]","metadata":{"execution":{"iopub.status.busy":"2023-07-06T21:50:37.760439Z","iopub.execute_input":"2023-07-06T21:50:37.760963Z","iopub.status.idle":"2023-07-06T21:50:37.769655Z","shell.execute_reply.started":"2023-07-06T21:50:37.760926Z","shell.execute_reply":"2023-07-06T21:50:37.768793Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"As we can see, there are seven columns in the wsi_df dataframe, while there are five columns in the tile_df dataframe. In other way of explaination, the seven columns we counted from the wsi_df dataframe hinted to us that the wsi_df dataframe contained all of the general metadata about the whole slide images, while tile_df has few columns for storing the metadata for each image, as there are few characteristics in each tif file specified in the competition data.\n\n如我们所见，`wsi_df` 数据帧中有七列，而 `tile_df` 数据帧中有五列。换句话说，我们从 `wsi_df` 数据帧中计算出的七列暗示 `wsi_df` 数据帧包含有关整个幻灯片图像的所有一般元数据，而 `tile_df` 用于存储每个图像的元数据的列很少，因为有比赛数据中指定的每个 \"tif\" 文件中的几个特征。","metadata":{}},{"cell_type":"markdown","source":"As we finish our basic dataframe analysis to the two dataframes: wsi_df and tile_df, let's proceed into graphing the data of them into interactive charts for our next part of our data visualization journey! When we look forward to analyzing the wsi_df dataframe, it will go to be a quick plotting analysis, since it has a few data entities and columns.\n\n当我们完成对两个数据帧的基本数据帧分析时：`wsi_df` 和 `tile_df`，让我们继续将它们的数据绘制成交互式图表，以进行我们数据可视化之旅的下一部分！当我们期待分析 `wsi_df` 数据帧时，它将是一个快速的绘图分析，因为它有一些数据实体和列。","metadata":{}},{"cell_type":"markdown","source":"<h2 style=\"background-color: #d13bff; color: white; padding-right: 100vw; background-size:cover; text-align: center; padding: 10px; border-radius: 15px;\">Chapter 1: wsi_df (第一章：wsi_df)</h2>\n\nWhen we first come into the section about visualizing the wsi_df dataframe, we knew that it contained the metadata for the whole slide images in which the tile were extracted from. However, it contained a few data entities since it possibly had incomplete or insufficient data. Aside from the overview of the wsi_df dataframe, let's give our short explanation of the columns in the wsi_df dataframe columns!\n\n当我们第一次进入有关可视化 `wsi_df` 数据框的部分时，我们知道它包含从中提取图块的整个幻灯片图像的元数据。但是，它包含一些数据实体，因为它可能包含不完整或不充分的数据。除了 `wsi_df` 数据框的概述之外，让我们对 `wsi_df` 数据框列中的列进行简短说明！\n* **source_wsi**: Specifies the WSI. (指定 \"WSI\"。)\n* **age**, **sex**, **race**, **height**, & **bmi**: Specifies the demographic information of the donor based on their age, sex, race, height, and BMI. (根据年龄、性别、种族、身高和 BMI 指定捐赠者的人口统计信息。)","metadata":{}},{"cell_type":"markdown","source":"First of all, let's distribute and then visualize the source_wsi data column into the histogram-box graph! To do this, we characterize the fig variable figure to create our histogram graph with the px module's histogram function, setting the wsi_df dataframe as the data for the diagram, followed by configuring the x parameter to the source_wsi column for specifying the x-axes and the marginal parameter to box for installing the box chart into the top of the histogram. Lastly, we apply the show function in the fig variable figure for displaying the graph underneath the code cell.\n\n首先，让我们将 `source_wsi` 数据列分布并可视化为直方图！为此，我们对 `fig` 变量 \"figure\" 进行表征，以使用 px 模块的直方图函数创建直方图，将 `wsi_df` 数据帧设置为图表的数据，然后将 `x` 参数配置到 `source_wsi` 列以指定 \"x\" 轴和用于将箱形图安装到直方图顶部的框的边际参数。最后，我们在 `fig` 变量 \"figure\" 中应用 `show` 函数来显示代码单元格下方的图形。","metadata":{}},{"cell_type":"code","source":"fig = px.histogram(wsi_df, x=\"source_wsi\", marginal=\"box\")\nfig.show()","metadata":{"execution":{"iopub.status.busy":"2023-07-06T21:50:37.770999Z","iopub.execute_input":"2023-07-06T21:50:37.771527Z","iopub.status.idle":"2023-07-06T21:50:39.625649Z","shell.execute_reply.started":"2023-07-06T21:50:37.771496Z","shell.execute_reply":"2023-07-06T21:50:39.624495Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"As we can see from the histogram-box chart based on the source_wsi column, we found out that there are three big bins shown in the diagram as the middle one is taller than the ones on the left and right of the graph, making this distribution type of the source_wsi column symmetrical. In other ways of explaining this, the highest data counted in the source_wsi column is between 2 and 3, with 2 entities, while the two ranges from 0 to 1 and 4 to 5 are calculated once, making them have the least data. Meanwhile in the box plot, we noticed that there are weird box shapes on the left and right corners of the box plot, as the minimum and maximum of the source_wsi column is equal to the whisker boundaries of the box diagram. Specifically, the first and third quartiles in the source_wsi column is 1.5 and 3.5, the median is 2.5, and the interquartile range is 2. Overall, the symmetrical data we glimpsed from the source_wsi column gave us clues that there are a few whole slide images that the tiles were extracted from.\n\n正如我们从基于 `source_wsi` 列的直方图箱图表中看到的那样，我们发现图中显示了三个大箱子，因为中间的箱子比图表左右两侧的箱子高，使得这种分布`source_wsi` 列对称的类型。换句话说，`source_wsi` 列中统计的最高数据在 **2** 到 **3** 之间，有 **2** 个实体，而 **0** 到 **1** 和 **4** 到 **5** 这两个范围被计算一次，使它们具有最少的数据。同时在箱形图中，我们注意到箱形图的左右角有奇怪的箱形，因为 `source_wsi` 列的最小值和最大值等于箱形图的晶须边界。具体来说，`source_wsi` 列中的第一和第三四分位数为 **1.5** 和 **3.5**，中位数为 **2.5**，四分位数间距为 **2**。总体而言，我们从 `source_wsi` 列中瞥见的对称数据为我们提供了一些线索，即有几张完整的幻灯片图像瓷砖是从中提取的。","metadata":{}},{"cell_type":"markdown","source":"Let's then distribute the data from the age column into another histogram/box graph! Once again, we define the fig variable figure to create our histogram graph with the px module's histogram function, setting the wsi_df dataframe as the data needed for the graph, along with the x parameter to the age column for configuring the x-axes of the diagram, and the marginal parameter to box for plotting a box graph on the top of the histogram. And as always, we exhibit our graph below the code cell by applying the show function in the fig variable figure.\n\n然后让我们将年龄列中的数据分布到另一个直方图/箱形图中！我们再次定义 `fig` 变量 \"figure\" 以使用 `px` 模块的 `histogram` 函数创建我们的直方图，将 `wsi_df` 数据帧设置为图形所需的数据，以及用于配置 `x` 轴的 `age` 列的 `x` 参数\"diagram\" 和 `box` 的边际参数，用于在直方图的顶部绘制箱线图。和往常一样，我们通过在 `fig` 变量 \"figure\" 中应用 `show` 函数来在代码单元格下方展示我们的图形。","metadata":{}},{"cell_type":"code","source":"fig = px.histogram(wsi_df, x=\"age\", marginal=\"box\")\nfig.show()","metadata":{"execution":{"iopub.status.busy":"2023-07-06T21:50:39.627647Z","iopub.execute_input":"2023-07-06T21:50:39.628519Z","iopub.status.idle":"2023-07-06T21:50:39.739751Z","shell.execute_reply.started":"2023-07-06T21:50:39.628486Z","shell.execute_reply":"2023-07-06T21:50:39.738342Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"After compiling the above code, we noticed that there are two big bins in the histogram graph as the left one is taller than the right one, making the distribution of the age column right-skewed. To be specific, the highest data tallied is in the range between 50 to 59, with 3 entities, while there is one entity in the range from 70 to 79, making it the lowest data counted. And from the box-plot, we see the deformed part of the left side of the box plot, as the minimum of the age column is same as the left-whisker in the box diagram. Besides, the first and third quartiles of the age column is 54.5 and 65.5, the median is 57, and the interquartile range is 11. Overall, the right-skewed distribution we spotted in the histogram hinted to us that most people were aged between 50 to 59 according to the wsi_df dataframe.\n\n编译上述代码后，我们注意到直方图中有两个大箱子，左边一个比右边一个高，使年龄列的分布呈右偏态。具体来说，在**50**到**59**之间统计的数据最多，有**3**个实体，而在**70**到**79**之间有**1**个实体，统计的数据最少。从箱线图中，我们看到箱线图左侧的变形部分，因为年龄列的最小值与箱线图中的左胡须相同。此外，年龄列的第一和第三四分位数是 **54.5** 和 **65.5**，中位数是 **57**，四分位数间距是 **11**。总的来说，我们在直方图中发现的右偏分布暗示我们大多数人的年龄在 **50** 岁之间根据 `wsi_df` 数据帧到 **59**。","metadata":{}},{"cell_type":"markdown","source":"Following that, let's then analyze the \"sex\" column in the pie chart! Before we begin distributing this data, we characterize another dataframe, gender_df, to count the values in the \"sex\" column from the wsi_df dataframe with the value_counts function. Thenceforth, we use the fig variable figure to create our pie chart with the pie function from the px module's pie function, setting the gender_df dataframe as the data for plotting the pie chart, followed by configuring the names parameter to the indexes specified from the index attribute that was plugged into the gender_df dataframe and the values parameter to the values extracted by the value attribute in the gender_df dataframe. With that completed, we use the show function in the fig variable figure for displaying our graph under the code cell.\n\n接下来，我们再分析一下饼图中的`“性别”`栏！在我们开始分发此数据之前，我们描述了另一个数据框 `gender_df`，以使用 `value_counts`函数计算 `wsi_df` 数据框的`“性别”列`中的值。此后，我们使用 `fig` 变量 \"figure\" 创建我们的饼图，并使用 `px` 模块的 `pie` 函数中的 \"pie\" 函数，将 `gender_df` 数据帧设置为绘制饼图的数据，然后将 `names` 参数配置为从索引指定的索引插入到 gender_df 数据帧中的属性和 `values` 参数到 `gender_df` 数据帧中的 `value` 属性提取的值。完成后，我们使用 `fig` 变量 \"figure\" 中的 `show` 函数在代码单元格下显示我们的图形。","metadata":{}},{"cell_type":"code","source":"gender_df = wsi_df[\"sex\"].value_counts()\n\nfig = px.pie(gender_df, names=gender_df.index, values=gender_df.values)\nfig.show()","metadata":{"execution":{"iopub.status.busy":"2023-07-06T21:50:39.741099Z","iopub.execute_input":"2023-07-06T21:50:39.741458Z","iopub.status.idle":"2023-07-06T21:50:39.815658Z","shell.execute_reply.started":"2023-07-06T21:50:39.741427Z","shell.execute_reply":"2023-07-06T21:50:39.814785Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"From the pie chart graphed from above, we found out that 75% of the data from the wsi_df dataframe's \"sex\" column is labeled as \"F\", while the other 25% of the data is marked as \"M\". In other words, most of the tissue donors were females than males.\n\n从上面绘制的饼图中，我们发现 `wsi_df` 数据框的`“性别”`列中 **75%** 的数据被标记为“F”，而另外 **25%** 的数据被标记为“M”。换句话说，大多数组织捐献者是女性而不是男性。","metadata":{}},{"cell_type":"markdown","source":"Let's then probe the data from the race column and then gather it to the pie chart! Once again, we inherit the race_df dataframe to calculate the values from the wsi_df dataframe's race column with the value_counts function. Next, we generate our fig variable figure to create another pie chart provided by the px module's pie function, setting the race_df dataframe for the pie chart's data, along with arranging the names parameter to the names gathered from the race_df dataframe that has the index attribute and the values parameter to the values extracted from the values attribute that is plugged to the race_df dataframe. Finally, we display our graph by using the show function in the fig variable figure.\n\n然后让我们从 `race` 列中探测数据，然后将其收集到饼图中！再一次，我们继承了 `race_df` 数据框，使用 `value_counts` 函数从 `wsi_df` 数据框的 `race` 列中计算值。接下来，我们生成我们的 `fig` 变量 \"\"figure\" 来创建 `px` 模块的 `pie` 函数提供的另一个饼图，为饼图数据设置 `race_df` 数据框，同时将 `names` 参数设置为从具有 `index` 属性的 `race_df` 数据框收集的名称以及从插入到 `race_df` 数据帧的 `values` 属性中提取的值的 `values` 参数。最后，我们使用 `fig` 变量 \"figure\" 中的 `show` 函数显示我们的图表。","metadata":{}},{"cell_type":"code","source":"race_df = wsi_df[\"race\"].value_counts()\n\nfig = px.pie(race_df, names=race_df.index, values=race_df.values)\nfig.show()","metadata":{"execution":{"iopub.status.busy":"2023-07-06T21:50:39.816794Z","iopub.execute_input":"2023-07-06T21:50:39.819655Z","iopub.status.idle":"2023-07-06T21:50:39.863969Z","shell.execute_reply.started":"2023-07-06T21:50:39.819624Z","shell.execute_reply":"2023-07-06T21:50:39.862821Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"As we can see in the pie chart above, we spotted that 75% of the data in the wsi_df dataframe's race column were labeled as \"W\", while the rest of it was labeled as \"B\". Specifically, the 75% of the data in the race column gave us clues that most of the tissue donor's race is white.\n\n正如我们在上面的饼图中看到的那样，我们发现 `wsi_df` 数据框的种族列中 **75%** 的数据被标记为“W”，而其余数据被标记为“B”。具体来说，种族列中 **75%** 的数据为我们提供了线索，即大多数组织捐献者的种族是白人。","metadata":{}},{"cell_type":"markdown","source":"Lastly for our data visualization in the wsi_df dataframe, let's proceed to distribute the data from the height, weight, and bmi columns into three histogram charts in a subplot graph! Before we start plotting the three histograms in a subplot, we import two additional modules for creating our subplot chart: the make_subplots function from the plotly module's subplots attribute and the plotly's graph_objects attribute as go. Now that we imported the two additional modules for creating a subplot graph, we build the fig variable figure and assign it to generate our subplot chart with the make_subplots function, setting the rows parameter to 1 and the cols parameter to 3 for making our subplot have a row and three columns. Thenceforth, we use the add_trace function three times to the fig variable figure to add our histogram graph into the subplot, placing the go module's Histogram function that has the x parameter to the wsi_df dataframe's height, weight, and bmi columns individually as well as having the name parameter configured to the same name of the column specified from the x parameter, followed by arranging the row parameter to 1 and the col parameter to 1, 2, and 3 separately for placing each histogram graph to each column in the subplot graph. With our graph completed, we now use the show function to the fig variable figure for displaying the subplot below.\n\n最后，对于 `wsi_df` 数据框中的数据可视化，让我们继续将身高、体重和 `bmi` 列的数据分布到子图图中的三个直方图图表中！在我们开始在子图中绘制三个直方图之前，我们导入了两个额外的模块来创建我们的子图图表：来自 `plotly` 模块的 `subplots` 属性的 `make_subplots` 函数和 `plotly` 的 `graph_objects` 属性。现在我们导入了两个用于创建子图的附加模块，我们构建 `fig` 变量 \"figure\" 并分配它以使用 `make_subplots` 函数生成我们的子图图表，将 `rows` 参数设置为 **1**，将 `cols` 参数设置为 **3** 以使我们的子图具有一行三列。此后，我们对 `fig` 变量 \"figure\" 三次使用 `add_trace` 函数，将我们的直方图添加到子图中，将具有 `x` 参数的 `go` 模块的直方图函数分别放置到 `wsi_df` 数据帧的高度、重量和 `bmi` 列，并具有`name`参数配置为与`x`参数指定的列同名，然后将`row`参数设置为1，将`col`参数分别设置为**1**、**2**、**3**，用于将每个直方图图放置到`subplot`图中的每个列。图表完成后，我们现在使用 `fig` 变量 \"figure\" 的 `show` 函数来显示下面的子图。","metadata":{}},{"cell_type":"code","source":"from plotly.subplots import make_subplots\nimport plotly.graph_objects as go\n\nfig = make_subplots(rows=1, cols=3)\nfig.add_trace(go.Histogram(x=wsi_df[\"height\"], name=\"height\"), row=1, col=1)\nfig.add_trace(go.Histogram(x=wsi_df[\"weight\"], name=\"weight\"), row=1, col=2)\nfig.add_trace(go.Histogram(x=wsi_df[\"bmi\"], name=\"bmi\"), row=1, col=3)\nfig.show()","metadata":{"execution":{"iopub.status.busy":"2023-07-06T21:50:39.865527Z","iopub.execute_input":"2023-07-06T21:50:39.865874Z","iopub.status.idle":"2023-07-06T21:50:39.895728Z","shell.execute_reply.started":"2023-07-06T21:50:39.865847Z","shell.execute_reply":"2023-07-06T21:50:39.894924Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"From the data distribution of the wsi_df dataframe's height column, we noted that the right-skew distribution is clearly shown as the first bar in the far left of the diagram is taller than the other bars. In another way of explaining this, the highest data counted is in the range from 155 to 164 with 2 entities, while the lowest data is in the range between 165 and 184 with one entity. Specifically, the right-skew distribution we noticed from the wsi_df dataframe's height column implied to us that some organ donors' height is shorter than the other anonymous organ donors specified in the wsi_df dataframe.\n\n从 `wsi_df` 数据帧的高度列的数据分布中，我们注意到右偏分布清楚地显示为图表最左侧的第一个条比其他条高。用另一种方式解释这一点，计数的最高数据在 **155** 到 **164** 的范围内有 **2** 个实体，而最低的数据在 **165** 到 **184** 的范围内有一个实体。具体来说，我们从 `wsi_df` 数据框的高度列中注意到的右偏分布向我们暗示，一些器官捐献者的身高比 `wsi_df` 数据框中指定的其他匿名器官捐献者矮。\n\nMeanwhile, from the wsi_df dataframe's weight data distribution in the middle of the subplot chart, we envisaged on how the bar on the left of the graph is taller than the bar on the right, showing a right-skewed distribution. To be specific, the range from 50 to 90 has the highest data with 3 entities, while the other range from 100 to 140 got one entity, making them have the lowest data. Additionally, the right-skewed data distribution we observed from the middle histogram graph gave us suggestions that most of the organ donors weighed less.\n\n同时，从子图中间`wsi_df` \"dataframe\"的权重数据分布来看，我们设想了图形左边的柱子比右边的柱子高，呈右偏分布。具体来说，**50**到**90**的数据最高，有3个实体，而**100**到**140**的范围有一个实体，数据最低。此外，我们从中间直方图中观察到的右偏数据分布让我们认为大多数器官捐献者的体重较轻。\n\nLastly, for the data distribution of the bmi column from the wsi_df dataframe, we spotted the right-skew distribution when we noticed that the first bar in the far left of the diagram is taller than the last two bars on the right of the chart, likewise to the data distribution of the height column. Besides from the looks of the graph, the highest counted data is under the range from 20 to 29 with 2 entities, while the other range from 30 to 49 has the lowest counted data because of only one entity. In addition, the right-skewed distribution we envisaged in the right histogram graph implied to us that there are most organ donors that has the lowest body mass index.\n\n最后，对于 `wsi_df` 数据帧中 `bmi` 列的数据分布，当我们注意到图表最左侧的第一个柱比图表右侧的最后两个柱高时，我们发现了右偏分布，与高度列的数据分布类似。除了从图表的外观来看，最高计数的数据在 **20** 到 **29** 范围内有 **2** 个实体，而另一个范围从 **30** 到 **49** 具有最低计数数据，因为只有一个实体。此外，我们在右侧直方图中设想的右偏分布向我们暗示，大多数器官捐献者的体重指数最低。","metadata":{}},{"cell_type":"markdown","source":"As we finally visualized the weight, height, and bmi columns all at once, we completed our data analysis on the wsi_df dataframe! What's next we are going to visualize is that we will study the data from the tile_df dataframe.\n\n当我们最终同时可视化体重、身高和 `bmi` 列时，我们完成了对 `wsi_df` 数据框的数据分析！接下来我们要可视化的是我们将研究来自 `tile_df` 数据帧的数据。","metadata":{}},{"cell_type":"markdown","source":"<h2 style=\"background-color: #d13bff; color: white; padding-right: 100vw; background-size:cover; text-align: center; padding: 10px; border-radius: 15px;\">Chapter 2: tile_df (第二章：tile_df)</h2>\n\nAfter we did a quick visualization of the wsi_df dataframe because of the short metadata based on the whole slide images the tile was extracted from. Now, we come across another dataframe based on the metadata of each image which is known as the tile_df dataframe. With all of this summarized, let's detail the columns in the tile_df dataframe one at a time!\n\n在我们对 `wsi_df` 数据帧进行了快速可视化之后，因为基于从中提取图块的整个幻灯片图像的短元数据。现在，我们遇到了另一个基于每个图像元数据的数据帧，称为 `tile_df` 数据帧。总结了所有这些，让我们一次一个地详细介绍 `tile_df` 数据框中的列！\n\n* **id**: Identifier for each tile. (每个图块的标识符。)\n* **source_wsi**: Specifies the WSI. (指定 \"WSI\"。)\n* **dataset**: The dataset the tile belongs to. (图块所属的数据集。)\n* **i/j**: Specifies the location on the upper-left corner within the WSI where the tile was extracted. (指定在 WSI 中提取磁贴的左上角位置。)","metadata":{}},{"cell_type":"markdown","source":"With all the columns in the tile_df dataframe concisely detailed, let's switch to plotting with Altair on distributing the source_wsi column into a histogram and box graph in a subplot! To get started, we characterize the fig1 variable to create our histogram with the alt module's Chart function that contained the tile_df dataframe as the data for the graph followed by the mark_bar function for graphing the bars to the graph, as well as encoding the parameters in the chart with the encode function, setting the alt module's X function that has the source_wsi column and the bin parameter to True for setting the x-axes in the graph, and the y parameter to the count function encased in strings for counting the values in the x-axes to the y-axes. Next, we characterize another variable, fig2 to create another graph with the alt module's Chart function that has the tile_df dataframe as the data for the chart but this time, we use the mark_boxplot function to create our box chart to the graph, as well as encoding the parameters of the second graph with the encode function, setting the x parameter to the source_wsi column for setting the x-axes of the graph and on the outside of the encode function, we apply the properties function, setting the height parameter to 300 for adjusting the height of the second plot. Lastly, we use the concat function from the alt module to concatenate the fig1 and fig2 charts into the subplot.\n\n`tile_df` 数据框中的所有列都简明扼要，让我们切换到使用 \"Altair\" 绘图，将 `source_wsi` 列分布到子图中的直方图和箱形图中！首先，我们描述 `fig1` 变量以使用 `alt` 模块的 `Chart` 函数创建直方图，该函数包含 `tile_df` 数据帧作为图形的数据，然后是 `mark_bar` 函数，用于将条形图绘制到图形中，并将参数编码为具有编码功能的图表，将具有 `source_wsi` 列和 `bin` 参数的 `alt` 模块的 `X` 函数设置为 `True` 以设置图表中的 \"x\" 轴，并将 \"y\" 参数设置为包含在字符串中的计数函数以计算中的值\"x\" 轴到 \"y\" 轴。接下来，我们描述另一个变量 `fig2` 以使用 `alt` 模块的 `Chart` 函数创建另一个图形，该函数将 `tile_df` 数据框作为图表的数据，但这次，我们使用 `mark_boxplot` 函数创建我们的箱线图到图形，以及使用 `encode` 函数对第二个图的参数进行编码，将 `x` 参数设置为 `source_wsi` 列以设置图的 \"x\" 轴，在 `encode` 函数的外部，我们应用 `properties` 函数，将 `height` 参数设置为 **300**用于调整第二个图的高度。最后，我们使用 `alt` 模块中的 `concat` 函数将 `fig1` 和 `fig2` 图表连接到子图中。","metadata":{}},{"cell_type":"code","source":"# But before that, we use the disable_max_rows function from the alt module's data_transformers attribute so that we'll not see the MaxRowsError.\n# 但在此之前，我们使用 alt 模块的 data_transformers 属性中的 disable_max_rows 函数，这样我们就不会看到 \"MaxRowsError\"。\nalt.data_transformers.disable_max_rows()\n\nfig1 = alt.Chart(tile_df).mark_bar().encode(\n    alt.X(\"source_wsi\", bin=True),\n    y=\"count()\"\n)\n\nfig2 = alt.Chart(tile_df).mark_boxplot().encode(\n    x=\"source_wsi\"\n).properties(height=300)\n\nalt.concat(fig1, fig2)","metadata":{"execution":{"iopub.status.busy":"2023-07-06T21:50:39.896864Z","iopub.execute_input":"2023-07-06T21:50:39.897110Z","iopub.status.idle":"2023-07-06T21:50:40.027630Z","shell.execute_reply.started":"2023-07-06T21:50:39.897089Z","shell.execute_reply":"2023-07-06T21:50:40.026828Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"From the histogram consisting on the distributed data of the source_wsi column, we noted that the bins mostly ascend from the left to the far right of the diagram, making this a left-skewed distribution. Specifically, the highest data counted is in the range from 12 to 14 with 1800 entities, while the lowest data counted is in between 4 to 6, with around 230 entities. And from the box plot we graphed in the right of our subplot, we glimpsed on how the box portion is shifted to the right, as the left whisker is longer than the right whisker, as the left and right portions of the box plot are equal to each other. Additionally, the first and third quartiles of the source_wsi column are 6 to 12, the median is 9, and the interquartile range is 6. In summary of the data distribution of the source_wsi column, the left-skewed distribution we spotted in the histogram gave us some clues that there are more whole slide images in which some tiles were extracted from.\n\n从由 `source_wsi` 列的分布数据组成的直方图中，我们注意到 `bin` 主要从图表的左侧上升到最右侧，这使得这是一个左偏分布。具体来说，计数最高的数据在 **12** 到 **14** 之间，有 **1800** 个实体，而计数最低的数据在 **4** 到 **6** 之间，有大约 **230** 个实体。从我们在子图右侧绘制的箱形图中，我们瞥见了箱形部分是如何向右移动的，因为左边的胡须比右边的胡须长，因为箱形图的左右部分是相等的对彼此。此外，`source_wsi` 列的第一和第三四分位数为 **6** 到 **12**，中位数为 **9**，四分位数间距为 **6**。总结 `source_wsi` 列的数据分布，我们在直方图中发现的左偏分布给出了我们提供了一些线索，表明有更多完整的幻灯片图像，其中一些图块是从中提取的。","metadata":{}},{"cell_type":"markdown","source":"Let's then proceed towards visualizing the dataset column by distributing them to the histogram graph! We use the alt module's Chart function that has the tile_df dataframe as the data for the chart, then we use the mark_bar function to graph the bars to the chart, and then encode the graph's characteristics with the encode function, setting the alt module's X parameter that has the dataset column and the bin parameter set to True for configuring the x-axes, and the y parameter to the count function closed in strings for specifying the counted values from a specific column to the y-axes.\n\n然后让我们通过将数据集列分布到直方图来继续可视化数据集列！我们使用具有 `tile_df` 数据帧的 `alt` 模块的 Chart 函数作为图表的数据，然后我们使用 `mark_bar` 函数将条形图绘制到图表上，然后使用 `encode` 函数对图形的特征进行编码，设置 `alt` 模块的 `X` 参数其数据集列和 `bin` 参数设置为 `True` 以配置 \"x\" 轴，计数函数的 `y` 参数以字符串形式关闭以指定从特定列到 \"y\" 轴的计数值。","metadata":{}},{"cell_type":"code","source":"alt.Chart(tile_df).mark_bar().encode(\n    alt.X(\"dataset\", bin=True),\n    y=\"count()\"\n)","metadata":{"execution":{"iopub.status.busy":"2023-07-06T21:50:40.028581Z","iopub.execute_input":"2023-07-06T21:50:40.028893Z","iopub.status.idle":"2023-07-06T21:50:40.135099Z","shell.execute_reply.started":"2023-07-06T21:50:40.028864Z","shell.execute_reply":"2023-07-06T21:50:40.134095Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"As we can see from the dataset data distribution in the histogram chart above, we noted the three separate bars, since they ascend almost exponentially while maintaining distance from each other, making a left-skew distribution. Besides from the looks of the histogram distribution, the highest data counted is in the range from 2.8 to 3, with approximately 5450 entities, while the range between 1 to 1.2 has the least data with roughly 480 entities. Additionally, the left-skewed distribution we saw from the three individual bars made us infer that there are most tiles that were belonged to the high values in the dataset.\n\n从上面直方图中的数据集数据分布可以看出，我们注意到三个独立的条形图，因为它们几乎呈指数上升，同时保持彼此的距离，形成左偏分布。除了从直方图分布来看，统计到的数据最高的是**2.8**到3的范围内，大约有**5450**个实体，而**1**到**1.2**之间的数据最少，大约有**480**个实体。此外，我们从三个单独的条形图中看到的左偏分布使我们推断出大多数图块属于数据集中的高值。","metadata":{}},{"cell_type":"markdown","source":"Now here's the last part of our data analysis in the tile_df dataframe, let's distribute the i and j columns into two separate histograms in a subplot! Once again, we generate the fig1 and fig2 variables and assign them to create the two Altair graphs by using the alt module's Chart function that has the tile_df dataframe as the data for the graph we're going to create, followed by using the mark_bar function to plot the bars into two graphs as well as using the encode function for configuring the parameters in both graphs, setting the X function from the alt module that has the standalone i and j columns (since we created two graphs) along with the bin parameter set to True for specifying the x-axes of the two graphs, and the y parameter to the count function enclosed in strings for counting the values from the x-axes to the y-axes. After creating the two histogram graphs, we use the alt module's concat function to combine the fig1 and fig2 variable graphs into a subplot graph.\n\n现在这是我们在 `tile_df` 数据框中进行数据分析的最后一部分，让我们将 `i` 和 `j` 列分布到子图中的两个单独的直方图中！我们再次生成 `fig1` 和 `fig2` 变量并分配它们以使用 `alt` 模块的 `Chart` 函数创建两个 \"Altair\" 图，该函数将 `tile_df` 数据帧作为我们要创建的图的数据，然后使用 `mark_bar` 函数将条形图绘制成两个图形，并使用编码函数在两个图形中配置参数，从具有独立 `i` 和 `j` 列的 `alt` 模块设置 `X` 函数（因为我们创建了两个图形）以及 `bin` 参数设置为 `True` 用于指定两个图形的 \"x\" 轴，并将 `y` 参数设置为包含在字符串中的 `count` 函数，用于计算从 \"x\" 轴到 `y` 轴的值。创建两个直方图后，我们使用 `alt` 模块的 `concat` 函数将 `fig1` 和 `fig2` 变量图组合成子图。","metadata":{}},{"cell_type":"code","source":"fig1 = alt.Chart(tile_df).mark_bar().encode(\n    alt.X(\"i\", bin=True),\n    y=\"count()\"\n)\n\nfig2 = alt.Chart(tile_df).mark_bar().encode(\n    alt.X(\"j\", bin=True),\n    y=\"count()\"\n)\n\nalt.concat(fig1, fig2)","metadata":{"execution":{"iopub.status.busy":"2023-07-06T21:50:40.136364Z","iopub.execute_input":"2023-07-06T21:50:40.136634Z","iopub.status.idle":"2023-07-06T21:50:40.244977Z","shell.execute_reply.started":"2023-07-06T21:50:40.136604Z","shell.execute_reply":"2023-07-06T21:50:40.243401Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"From the data distribution of the i column on the left of the subplot chart, we envisaged on how there's a data peak on the left of the histogram as it shifts to the left of the diagram, making this a right-skewed distribution. Other than that, the highest data counted is in the range from 10000 to 15000, with nearly 2090 entities, while the other range between 25000 to 30000 has the lowest data counted, with around 300 entities. Funnilly enough, the right-skewed distribution we noted in the i column data distribution gave us clues that most of the whole slide images were extracted from the left of some tiles.\n\n从子图左侧第 `i` 列的数据分布来看，我们设想当直方图向图的左侧移动时，其左侧会出现一个数据峰值，从而使其成为右偏分布。除此之外，在**10000**到**15000**之间统计的数据最高，有近**2090**个实体，而在**25000**到**30000**之间统计的数据最少，大约有**300**个实体。有趣的是，我们在第 `i` 列数据分布中注意到的右偏分布为我们提供了线索，即整个幻灯片图像的大部分是从某些图块的左侧提取的。\n\nOn the right of the subplot graph, which is the j column data distribution, we noted that there's a data peak on the left portion of the histogram, as they faintly showed a right-skewed distribution since the bars' height from the right ascends exponentially. Besides from the look from the j data distribution graph, the range between 20K to 30K has the most common data, with an estimated 2550 units calculated, while the range spanning from 50K to 60K has the least common data, within the neighborhood of 90 or 100 entities calculated. In summary of the j column data distribution, the right-skewed distribution we noted gave us some cues that there were most whole slide images extracted from the upper-left portion of the tiles, likewise to the i column distribution.\n\n在子图的右侧，即第 `j` 列数据分布，我们注意到直方图的左侧部分有一个数据峰值，因为它们隐约显示出右偏分布，因为条形从右侧开始的高度呈指数上升.另外从第j个数据分布图来看，**20K**到**30K**范围内的数据最常见，估计计算了**2550**个单位，而**50K**到**60K**范围内的数据最不常见，在**90**或**90**左右计算了 **100** 个实体。在第 `j` 列数据分布的总结中，我们注意到的右偏分布给了我们一些线索，即大多数完整的幻灯片图像是从图块的左上部分提取的，类似于第 `i` 列分布。","metadata":{}},{"cell_type":"markdown","source":"With all four columns we analyzed in only three graphs, we finished our another short visualization of the tile_df dataframe! As we head into the final part of the EDA in the HuBMAP Vasculature competition, let's visualize the whole slide images with segmentation masks provided by the jsonl file!\n\n我们仅在三个图表中分析了所有四列，我们完成了 `tile_df` 数据框的另一个简短可视化！当我们进入 \"HuBMAP\" 脉管系统竞赛 \"EDA\" 的最后部分时，让我们使用 `jsonl` 文件提供的分割掩码可视化整个幻灯片图像！","metadata":{}},{"cell_type":"markdown","source":"<h2 style=\"background-color: #d13bff; color: white; padding-right: 100vw; background-size:cover; text-align: center; padding: 10px; border-radius: 15px;\">Chapter 3: Visualizing the Vasculature (第三章：血管可视化)</h2>\n\nSince we completed our data visualization on the tile_df and the wsi_df dataframes, here's the final part of our data analysis in hacking the vasculature, visualizing the tiff images of each vasculature image and the mask segments of it. Before we get started on visualizing the vasculature segmentation and images, we import two additional modules: the matplotlib module's pyplot attribute as plt, and the Image function from the PIL module.\n\n由于我们完成了对`tile_df`和`wsi_df`数据帧的数据可视化，这里是我们黑化脉管的数据分析的最后一部分，将每个脉管图像的\"tiff\"图像和它的掩膜段可视化。在我们开始可视化血管分割和图像之前，我们要导入两个额外的模块：`matplotlib`模块的`pyplot`属性为`plt`，以及`PIL`模块的`Image`函数。","metadata":{}},{"cell_type":"code","source":"import matplotlib.pyplot as plt\nfrom PIL import Image","metadata":{"execution":{"iopub.status.busy":"2023-07-06T21:50:40.246133Z","iopub.execute_input":"2023-07-06T21:50:40.246426Z","iopub.status.idle":"2023-07-06T21:50:40.251897Z","shell.execute_reply.started":"2023-07-06T21:50:40.246403Z","shell.execute_reply":"2023-07-06T21:50:40.250006Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Now that we imported the additional two modules, we generate another dataframe, poly_df, to read a jsonl file with the pd module's read_json function, since it contains the path directory to the polygons.jsonl file, followed by setting the lines parameter to True to include lines in the created dataframe. With that completed, we use the head function to the poly_df dataframe to display the first five rows of the converted jsonl dataframe.\n\n现在我们导入了额外的两个模块，我们生成另一个数据框架，`poly_df`，用`pd`模块的`read_json`函数读取`jsonl`文件，因为它包含\"polygons.jsonl\"文件的路径目录，接着将`lines`参数设置为`True`，以便在创建的数据框架中包含行。完成这些后，我们使用`poly_df`数据框架的`head`函数来显示转换后的`jsonl`数据框架的前五行。","metadata":{}},{"cell_type":"code","source":"poly_df = pd.read_json('/kaggle/input/hubmap-hacking-the-human-vasculature/polygons.jsonl', lines=True)\npoly_df.head()","metadata":{"execution":{"iopub.status.busy":"2023-07-06T21:50:40.253525Z","iopub.execute_input":"2023-07-06T21:50:40.253821Z","iopub.status.idle":"2023-07-06T21:50:43.814076Z","shell.execute_reply.started":"2023-07-06T21:50:40.253797Z","shell.execute_reply":"2023-07-06T21:50:43.812689Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Once we created the poly_df dataframe from the polygons.jsonl file, we realized that the annotations column contained a JSON array out of type and coordinates keys. Without any doubts ahead of us, we divide the data in the annotations column into separate columns by defining and then assigning the flatten_annotations_df dataframe to the explode function that contained the annotations column to the poly_df dataframe for dividing the type and coordinates keys to the dataframe columns. Next, we create the extracted_data_df dataframe and assign it to return a new object with the apply function connected to the flatten_annotations_df dataframe's annotations column, as the apply function contains the pd module's Series attribute for returning a data series. We then redefine the poly_df dataframe to combine the two objects inside the array which is the poly_df dataframe's id column and the extracted_data_df dataframe with the concat function thus configuring the axis parameter to 1 to sum the values along the id rows, then reset the indexes with the reset_index function, setting the inplace parameter to True for making changes to the original poly_df dataframe's data and the drop parameter to True for dropping unused columns in the poly_df dataframe. Finally, we display the poly_df dataframe for looking at what we changed from the annotations column.\n\n一旦我们从\"polygons.jsonl\"文件中创建了`poly_df`数据框架，我们就意识到，注释列包含了一个由类型和坐标键组成的\"JSON\"数组。在没有任何疑问的情况下，我们通过定义并将`flatten_annotations_df`数据框架分配给包含注释列的`explode`函数，将类型和坐标键划分给数据框架列，从而将注释列中的数据划分为不同的列。接下来，我们创建`extracted_data_df`数据框架，并指定它返回一个新的对象，其`apply`函数连接到`flatten_annotations_df`数据框架的注释列，因为`apply`函数包含`pd`模块的`Series`属性，用于返回一个数据系列。然后，我们重新定义`poly_df`数据框架，用`concat`函数将数组内的两个对象（即`poly_df`数据框架的id列和`extracted_data_df`数据框架）结合起来，从而将`axis`参数配置为**1**，将`id`行的值相加，然后用`reset_index`函数重置索引，将`inplace`参数设置为`True`，用于对原始`poly_df`数据框架的数据进行更改，将`drop`参数设置为`True`，用于删除`poly_df`数据框架中未使用的列。最后，我们显示`poly_df`数据框架，以查看我们从注释列中改变的内容。","metadata":{}},{"cell_type":"code","source":"flatten_annotations_df = poly_df.explode(\"annotations\")\nextracted_data_df = flatten_annotations_df[\"annotations\"].apply(pd.Series)\npoly_df = pd.concat([poly_df[\"id\"], extracted_data_df], axis=1)\npoly_df.reset_index(inplace=True, drop=True)\npoly_df","metadata":{"execution":{"iopub.status.busy":"2023-07-06T21:50:43.815270Z","iopub.execute_input":"2023-07-06T21:50:43.815561Z","iopub.status.idle":"2023-07-06T21:50:48.543294Z","shell.execute_reply.started":"2023-07-06T21:50:43.815536Z","shell.execute_reply":"2023-07-06T21:50:48.542359Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Now that we arranged the data from the annotations column to separate columns in the poly_df dataframe, let's take a peek at one of the tiff files in the Vasculature competition data! First of all, we fetch one of the tiff image files by assigning the img_id variable to the poly_df dataframe's id column that has the index of any number that is less than or equal to around 7033 counted entities in the poly_df dataframe, then we characterize the img_path variable to concatenate the file directory leading to a tiff file in the train folder from the vasculature competition data.\n\n现在，我们将注释列的数据安排在`poly_df`数据框架的不同列中，让我们来看看\"Vasculature\"竞赛数据中的一个\"tiff\"文件吧! 首先，我们通过将`img_id`变量分配给`poly_df`数据框架的id列来获取其中一个\"tiff\"图像文件，该列的索引是小于或等于`poly_df`数据框架中约**7033**个计数实体的任何数字，然后我们将`img_path`变量定性为连接到\"vasculature\"竞赛数据中`train`文件夹中的\"tiff\"文件的文件目录。","metadata":{}},{"cell_type":"code","source":"img_id = poly_df[\"id\"][40]\nimg_path = \"../input/hubmap-hacking-the-human-vasculature/train/\" + img_id + \".tif\"","metadata":{"execution":{"iopub.status.busy":"2023-07-06T21:50:48.544474Z","iopub.execute_input":"2023-07-06T21:50:48.544907Z","iopub.status.idle":"2023-07-06T21:50:48.548204Z","shell.execute_reply.started":"2023-07-06T21:50:48.544884Z","shell.execute_reply":"2023-07-06T21:50:48.547577Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Following after creating the img_path variable, we then display the specified image from the img_path variable with the open function from the Image module in the img variable, followed by displaying the image with the plt module's imshow function that has the img variable inside as well as having the graph's axis off from the plt module's axis function that has the 'off' string.\n\n在创建了`img_path`变量之后，我们用`Image`模块的`open`函数在`img`变量中显示`img_path`变量中的指定图片，然后用`plt`模块的`imshow`函数显示图片，该函数里面有`img`变量，同时用`plt`模块的`axis`函数关闭图形的轴，该函数有'off'字符串。","metadata":{}},{"cell_type":"code","source":"img = Image.open(img_path)\nplt.imshow(img)\nplt.axis('off')","metadata":{"execution":{"iopub.status.busy":"2023-07-06T21:50:48.548955Z","iopub.execute_input":"2023-07-06T21:50:48.549600Z","iopub.status.idle":"2023-07-06T21:50:48.843809Z","shell.execute_reply.started":"2023-07-06T21:50:48.549578Z","shell.execute_reply":"2023-07-06T21:50:48.842962Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"From the image shown above after the code cell was compiled, we glimpsed how there's a tissue organ seen in one of the tiff files. Not only that, we spotted several cells clumping together as there are purplish dots inside each cell, representing them as the nucleus. To be specific, most tiff files captured some cells in each segment of the vasculature organ.\n\n从上面显示的代码细胞被编译后的图像中，我们瞥见了其中一个`tiff`文件中看到的组织器官。不仅如此，我们还发现几个细胞聚集在一起，因为每个细胞内都有紫红色的小点，代表它们是细胞核。具体来说，大多数`tiff`文件在血管器官的每个部分都捕捉到一些细胞。","metadata":{}},{"cell_type":"markdown","source":"Now it's time to generate masks that segment each organ seen from each tiff image file! Initially, we build our function, mask_img, containing two parameter variables: poly_data and img_name. Inside the mask_img function, we search up the identifiers from the poly_df dataframe's id column by assigning the img_data function parameter to the poly_data parameter that has the query function, containing a string that has the poly_df dataframe's id column equaling to the img_name variable then reset the indexes of the img_data variable with the reset_index function, setting the inplace parameter to True for keeping the original data and the drop parameter to True for dropping some columns in the img_data function parameter.\n\n现在是时候生成掩码，将每个器官从每个tiff图像文件中分割出来 首先，我们建立我们的函数，`mask_img`，包含两个参数变量：``poly_data``和`img_name`。在`mask_img`函数中，我们通过将`img_data`函数参数分配给具有查询功能的`poly_data`参数，从`poly_df`数据框架的id列中搜索出标识符，其中包含一个字符串，该字符串的id列等于`img_name`变量，然后用`reset_index`函数重置`img_data`变量的索引，将`inplace`参数设置为`True`，以保持原始数据，将`drop`参数设置为`True`，以在`img_data`函数参数中放弃一些列。\n\nNext, we generate our img_path variable to a string format of a file directory that has the img_name parameter concatenated as a specific image name of one of the tiff images in the train folder then open the specified image with the open function from the Image module as it was assigned to the defined img variable. \n\n接下来，我们将`img_path`变量生成一个字符串格式的文件目录，该目录中的`img_name`参数串联成火车文件夹中的一个特定的\"tiff\"图像名称，然后用图像模块中的`open`函数打开指定的图像，因为它被分配给定义的`img`变量。\n\nIt's time to map each tiff image by vascular organ shape with color legend and coordinates! To get started, we use the for loop statement, defining and looping the idx and row variables in the img_data variable's rows iterated with the iterrows function for iterating each row in the poly_df dataframe based on the type of organ in the vasculature and the coordinates to segment the organ. Inside the for loop, an if-statement is created, determining if the type column in the row variable is equal to glomerulus, blood_vessel, or other organ in the vasculature, then it returns the color variable to each different color, mapping the organ seen in the tiff image. The coordinates are then generated when the sublist variable is assigned to the coordinates column from the row variable, as its slice index was set to zero for extracting the first object, followed by creating the x and y variables to an empty list for storing the x and y values separately. Following that, another for loop was created, looping the defined datapoint variable in the sublist variable, as it adds the first and last entities from the datapoint variable to the x and y arrays for specifying the coordinates to map each segment of the organ found in the vasculature image. Lastly, the graph was created with the scatter points generated by the scatter function that has the x and y arrays inside as well as setitng the s parameter to 0 for the marker size from the plt module, followed by applying the color legend with the plt module's fill function as it has the x, y, and color variables inside for filling inside the coordinate boundaries with a specified color alongside setting the alpha parameter to 0.5 for setting the opacity of each color legend.\n\n现在是时候通过血管器官形状与颜色图例和坐标来映射每张\"tiff\"图像了 为了开始，我们使用`for`循环语句，在`img_data`变量的行中定义并循环`idx`和行变量，用`iterrows`函数重复`poly_df`数据框中的每一行，根据血管中的器官类型和坐标来分割器官。在`for`循环中，创建了一个`if`语句，确定行变量中的类型列是否等于肾小球、血管或脉管系统中的其他器官。它将颜色变量返回到每种颜色，映射出在`tiff`图像中看到的器官。然后，当子列表变量被分配到来自行变量的坐标列时，就会产生坐标，因为其切片索引被设置为零，用于提取第一个对象，接着将x和y变量创建为空列表，用于分别存储x和y值。随后，创建了另一个`for`循环，在子列表变量中循环定义的数据点变量，因为它将数据点变量中的第一个和最后一个实体添加到`x`和`y`数组中，以指定坐标来映射血管图像中发现的器官的每个片段。最后，用散点函数生成的散点来创建图形，散点函数中包含`x`和`y`数组，并将`plt`模块中的`s`参数设置为**0**，用于标记尺寸，然后用`plt`模块的填充函数应用彩色图例，因为它包含`x`、`y`和颜色变量，用于用指定的颜色填充坐标边界，同时将`alpha`参数设置为**0.5**，用于设置每个彩色图例的不透明度。\n\nOnce we're outside of the long for-loop statement, we use the imshow function from the plt module to display one of the tiff image files specified by the img variable, and on the outside of the mask_img variable, we call the mask_img function to segment masks of each organ found in some vasculature tiff images, setting the poly_df dataframe and the img_id variable.\n\n一旦我们在长的\"for-loop\"语句之外，我们使用`plt`模块的`imshow`函数来显示`img`变量所指定的一个\"tiff\"图像文件，在`mask_img`变量之外，我们调用`mask_img`函数来分割一些脉管\"tiff\"图像中发现的每个器官的掩码，设置`poly_df`数据帧和`img_id`变量。","metadata":{}},{"cell_type":"code","source":"def mask_img(poly_data, img_name):\n    img_data = poly_data.query(\"id == @img_name\")\n    img_data.reset_index(inplace=True, drop=True)\n\n    img_path = \"../input/hubmap-hacking-the-human-vasculature/train/\" + img_name + \".tif\"\n    img = Image.open(img_path)\n    \n    for idx, row in img_data.iterrows():\n        if row[\"type\"] == \"glomerulus\":\n            color = \"yellow\"\n        elif row[\"type\"] == \"blood_vessel\":\n            color = \"red\"\n        else:\n            color = \"blue\"\n            \n        sublist = row[\"coordinates\"][0]\n        x = []\n        y = []\n        \n        for datapoint in sublist:\n            x.append(datapoint[0])\n            y.append(datapoint[1])\n            \n        plt.scatter(x, y, s=0)\n        plt.fill(x, y, color, alpha=0.5)\n        \n    plt.imshow(img)\n    \nmask_img(poly_df, img_id)","metadata":{"execution":{"iopub.status.busy":"2023-07-06T21:50:48.844898Z","iopub.execute_input":"2023-07-06T21:50:48.845573Z","iopub.status.idle":"2023-07-06T21:50:49.244545Z","shell.execute_reply.started":"2023-07-06T21:50:48.845542Z","shell.execute_reply":"2023-07-06T21:50:49.243312Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"As a result of creating the function on plotting the bounding boxes, we espied how there are red segmented masks plotted on one of the organs that were seen from the vasculature images, indicating that there are other vasculature organ types seen in one of the tiff images. \n\n由于创建了绘制边界框的功能，我们发现在其中一个器官上绘制了红色的分割掩码，这是从脉管图像中看到的，表明在其中一个`tiff`图像中看到了其他脉管器官类型。","metadata":{}},{"cell_type":"markdown","source":"Since we plotted one of the tiff images with bounding boxes, let's plot more of the tiff images! In the next three code cells below, we reassign the img_id variable to other slice indexes and then call out the mask_img function to segment the organ types seen in the vasculature image with the poly_df and img_id variables being set.\n\n既然我们绘制了一个带有边界框的`tiff`图像，我们就来绘制更多的`tiff`图像吧 在下面的三个代码单元中，我们将`img_id`变量重新分配给其他切片索引，然后呼出`mask_img`函数，在`poly_df`和`img_id`变量被设置的情况下，对血管图像中看到的器官类型进行分割。","metadata":{}},{"cell_type":"code","source":"img_id = poly_df[\"id\"][100]\nmask_img(poly_df, img_id)","metadata":{"execution":{"iopub.status.busy":"2023-07-06T21:50:49.245680Z","iopub.execute_input":"2023-07-06T21:50:49.245971Z","iopub.status.idle":"2023-07-06T21:50:49.634762Z","shell.execute_reply.started":"2023-07-06T21:50:49.245945Z","shell.execute_reply":"2023-07-06T21:50:49.633789Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"img_id = poly_df[\"id\"][5]\nmask_img(poly_df, img_id)","metadata":{"execution":{"iopub.status.busy":"2023-07-06T21:50:49.635889Z","iopub.execute_input":"2023-07-06T21:50:49.637154Z","iopub.status.idle":"2023-07-06T21:50:50.102251Z","shell.execute_reply.started":"2023-07-06T21:50:49.637105Z","shell.execute_reply":"2023-07-06T21:50:50.101298Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"img_id = poly_df[\"id\"][800]\nmask_img(poly_df, img_id)","metadata":{"execution":{"iopub.status.busy":"2023-07-06T21:50:50.103544Z","iopub.execute_input":"2023-07-06T21:50:50.104124Z","iopub.status.idle":"2023-07-06T21:50:50.702494Z","shell.execute_reply.started":"2023-07-06T21:50:50.104097Z","shell.execute_reply":"2023-07-06T21:50:50.701356Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Based on the three images we plotted, we glimpsed how there are segmented masks of the vasculature organs, like the glomerulus, blood vessels, and capillaries. Specifically, the images we graphed from the above three code cells implied to us that some organs are scattered around the vascular system, from the chest to other portions of the human body.\n\n根据我们绘制的三幅图像，我们瞥见了如何存在着血管器官的分割面具，如肾小球、血管和毛细血管。具体来说，我们从上述三个代码单元绘制的图像向我们暗示，一些器官散布在血管系统周围，从胸部到人体的其他部分。","metadata":{}},{"cell_type":"markdown","source":"Since we created the poly_df dataframe from the polygon annotations in the jsonl file and then plotted the vasculature images with segmented masks, we completed our data visualization on the vasculature images and mask annotations and reached the end of our whole data analysis on our exploratory data analysis in the vasculature competition!\n\n由于我们从\"jsonl\"文件中的多边形注释中创建了`poly_df`数据框架，然后用分割的掩码绘制脉管图像，我们完成了对脉管图像和掩码注释的数据可视化，并达到了我们在脉管竞争中探索性数据分析的终点","metadata":{}},{"cell_type":"markdown","source":"<h2 style=\"background-color: #d13bff; color: white; padding-right: 100vw; background-size:cover; text-align: center; padding: 10px; border-radius: 15px;\">Conclusion (总结)</h2>\n\nSo what are the takeaways from our EDA in the vasculature competition? We probed on the four data entities in the wsi_df dataframe since there are four organ donors included inside this dataframe, followed by observing more whole-slide images in the tile_df dataframe as well as seeing most of the whole-slide images extracted on the lower-left position and then visualizing some vasculature tiff images with segmented masks of each organ in each image. And as we foresighted the end-goal of this competition we're in, we will improve the researchers' knowledge of where the blood vessel were arranged in each human tissue, so that we'll identify and then understand how the relationships between our cells affect our health shortly.\n\n那么，我们在脉管系统竞赛中的\"EDA\"有什么收获呢？我们探究了`wsi_df`数据框中的四个数据实体，因为这个数据框中包含了四个器官捐献者，接着观察了`tile_df`数据框中更多的整幅图像，以及看到左下角位置提取的大部分整幅图像，然后可视化了一些血管\"tiff\"图像，每个图像中每个器官的分割掩码。而我们预见了我们所参加的这个比赛的最终目标，我们将提高研究人员对血管在每个人体组织中的排列位置的认识，这样我们就能在短期内识别然后了解我们的细胞之间的关系如何影响我们的健康。","metadata":{}}]}