{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"pygments_lexer":"ipython3","nbconvert_exporter":"python","version":"3.6.4","file_extension":".py","codemirror_mode":{"name":"ipython","version":3},"name":"python","mimetype":"text/x-python"}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"# **EDA Simplified: HuBMAP + HPA Organ Segmentation**","metadata":{"_uuid":"8f2839f25d086af736a60e9eeb907d3b93b6e0e5","_cell_guid":"b1076dfc-b9ad-4769-8c92-a6c4dae69d19"}},{"cell_type":"markdown","source":"## Introduction\nLast time, in the [UW-Madison GI Image Segmentation](https://www.kaggle.com/competitions/uw-madison-gi-tract-image-segmentation/overview), we covered up all of the data given in that competition, from analyzing the train.csv file with the train_df dataframe to visualizing a image mask with Weights and Biases. But now, we've warped into this competition, and this competition, is a lot **different** than the one we covered EDA on. Instead of segmenting the GI Tracts, we'll identify and segment functional tissue units (FTUs) across five human organs. It's more like image segmentation, but it's on the tissues from not just one or two, but on the five human body parts. Nevertheless, let's do EDA here!\n\nBefore we proceed, check out our previous EDA competition regardless of GI Image Segmentation:\nhttps://www.kaggle.com/code/dinowun/eda-simplified-uwm-gi-tract-segmentation-w-w-b","metadata":{}},{"cell_type":"markdown","source":"## Importing and Setup\nIn order to use EDA on this competition, we import the following modules first for data science and linear algebra, which is pandas as pd, numpy as np, and cv2 (aka OpenCV). And for plotting, we import the matplotlib module with the pyplot submodule as plt and the plotly module with the express module as px. Lastly, we import the tifffile and the os module for operating system stuff and reading tiff files since this compeitition has images that has tiff files.","metadata":{}},{"cell_type":"code","source":"import pandas as pd\nimport numpy as np\n\nimport matplotlib.pyplot as plt \nimport plotly.express as px\n\nimport os\nimport tifffile","metadata":{"execution":{"iopub.status.busy":"2022-07-14T05:36:16.353015Z","iopub.execute_input":"2022-07-14T05:36:16.353902Z","iopub.status.idle":"2022-07-14T05:36:18.497693Z","shell.execute_reply.started":"2022-07-14T05:36:16.353768Z","shell.execute_reply":"2022-07-14T05:36:18.496046Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"After we setup the necessary modules, let's load the csv files, which there is only one here in this competition, train.csv. Without further ado, we create a new dataframe, called train_df to read the csv file of train.csv file in this competition by using the pd module with the read_csv function. Finally, we display the rows of our train_df dataframe by using the head function to it.","metadata":{}},{"cell_type":"code","source":"train_df = pd.read_csv(\"../input/hubmap-organ-segmentation/train.csv\")\ntrain_df.head()","metadata":{"execution":{"iopub.status.busy":"2022-07-14T05:36:18.500712Z","iopub.execute_input":"2022-07-14T05:36:18.501235Z","iopub.status.idle":"2022-07-14T05:36:18.930649Z","shell.execute_reply.started":"2022-07-14T05:36:18.501165Z","shell.execute_reply":"2022-07-14T05:36:18.929746Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"After setting up our train_df dataframe, let's get started to EDA by EDA on every chapter!","metadata":{}},{"cell_type":"markdown","source":"## Chapter 1: The Basics of the train_df Dataframe\nHere we are, in the basics of the train_df dataframe! First we print out the number of entities of the data inside the train_df dataframe by using the len function to that.","metadata":{}},{"cell_type":"code","source":"print(\"No of data entities of train_df: \", len(train_df))","metadata":{"execution":{"iopub.status.busy":"2022-07-14T05:36:18.931831Z","iopub.execute_input":"2022-07-14T05:36:18.932375Z","iopub.status.idle":"2022-07-14T05:36:18.937864Z","shell.execute_reply.started":"2022-07-14T05:36:18.932337Z","shell.execute_reply":"2022-07-14T05:36:18.936762Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"As always, there are 351 data entities in the train_df dataframe, which means that there are NaN values prominent in this dataframe. So, let's check it out!","metadata":{}},{"cell_type":"markdown","source":"To find whether our train_df dataframe has NaN values or not, we print out the train_df dataframe with the isna function and sum them all together with the sum function.","metadata":{}},{"cell_type":"code","source":"train_df.isna().sum()","metadata":{"execution":{"iopub.status.busy":"2022-07-14T05:36:18.940543Z","iopub.execute_input":"2022-07-14T05:36:18.940869Z","iopub.status.idle":"2022-07-14T05:36:18.957758Z","shell.execute_reply.started":"2022-07-14T05:36:18.940838Z","shell.execute_reply":"2022-07-14T05:36:18.956451Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"After running this code cell, we noticed that there are no NaN values from this train_df dataframe, which is good news to us, since we are going to proceed to perform EDA on the next, next chapter!","metadata":{}},{"cell_type":"markdown","source":"## Chapter 2: Data Analysis\nNow, let's dive in to analyzing the data! First, let's analyze the id ranges with the histogram plotting with Plotly!\n\nTo do that, We first define a variable, fig, to the px module with the histogram function to create our histogram figure, setting the data_frame (dataframe input) parameter to the train_df dataframe, the x parameter (x-axis input) to the id data index from the train_df dataframe, the marginal parameter (optional small subplots) to violin, and the nbins parameter (number of bins) to 400.\n\nNext, optionally, we update our figure graph to any built-in templates by Plotly with the update_layout function to the fig variable figure, setting the template parameter for template specification to any [built-in template Plotly created](https://plotly.com/python/templates/).","metadata":{}},{"cell_type":"code","source":"fig = px.histogram(data_frame=train_df, x=\"id\", marginal=\"violin\", nbins=400)\nfig.update_layout(template=\"presentation\")\nfig.show()","metadata":{"execution":{"iopub.status.busy":"2022-07-14T05:36:18.960612Z","iopub.execute_input":"2022-07-14T05:36:18.961729Z","iopub.status.idle":"2022-07-14T05:36:20.624383Z","shell.execute_reply.started":"2022-07-14T05:36:18.961672Z","shell.execute_reply":"2022-07-14T05:36:20.623054Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"What we observed from this graph is that the ids of each image ranged from 0 to 32,999. Furthermore, we also see that the most counts of the id range in this graph is between 9400 and 9499. Without further ado, let's move on to counting the organs mentioned in this competition!","metadata":{}},{"cell_type":"markdown","source":"To plot each counts of the five organs, we define the fig variable again to the px module with the bar function to create a bar chart, setting the x (x-axis) parameter to the unique values of the train_df dataframe with the organ data index given by the np module, the y (y-axis) parameter to an array that contained a list (with the list function) of a train_df dataframe with the organ data index that was being counted (with the count function) by the i variable that looped in the unique values (listed by the unique function) of the train_df dataframe with the organ data index again, the color (color specification) parameter to same as what we did to setting up the x parameter in here, and the color_continuous_scale parameter set to any [built-in template Plotly created again](https://plotly.com/python/builtin-colorscales/). \n\nNext, we update our figure's x-axes and y-axes by using the update_xaxes and update_yaxes functions to the fig variable, setting the title (title for x and y axis) parameter to Classes (x-axis) and Number of Organs (y-axis). We then update our layout of our figure with the update_layout function, setting the showlegend (option to show the key labels of the figure) parameter to True, the title (title input and optional specification) parameter to a dictionary setup (details in the code), and the template (optional template specification) to any template Plotly gave out again (e.g. ggplot, seaborn). Finally, let's show our \"fig\" figure variable with the show function!","metadata":{}},{"cell_type":"code","source":"fig = px.bar(x=np.unique(train_df[\"organ\"]), y=[list(train_df[\"organ\"]).count(i) for i in np.unique(train_df[\"organ\"])], color=np.unique(train_df[\"organ\"]), color_continuous_scale=\"Mint\")\nfig.update_xaxes(title=\"Classes\")\nfig.update_yaxes(title=\"Number of Organs\")\nfig.update_layout(showlegend=True, \n                  title={\n                      'text': 'Organ Distribution', # text for the title\n                      'y': 0.95, # y positioning\n                      'x': 0.5, # x positioning\n                      'xanchor': 'center', # anchoring in x and y position (see below)\n                      'yanchor': 'top'}, template=\"ggplot2\")\nfig.show()","metadata":{"execution":{"iopub.status.busy":"2022-07-14T05:36:20.626309Z","iopub.execute_input":"2022-07-14T05:36:20.626670Z","iopub.status.idle":"2022-07-14T05:36:20.885724Z","shell.execute_reply.started":"2022-07-14T05:36:20.626627Z","shell.execute_reply":"2022-07-14T05:36:20.884395Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"As always, what we observed from this graph, is that the most counts of the organs mentioned in this competition is the kidney with 99 entities, while the least is the lung with 48 entities.","metadata":{}},{"cell_type":"markdown","source":"Let's move onto finding the number of entities in the data_source data index from the train_df dataframe! All we need to do is to print out the value counts of the train_df dataframe with the value_counts function.","metadata":{}},{"cell_type":"code","source":"train_df[\"data_source\"].value_counts()","metadata":{"execution":{"iopub.status.busy":"2022-07-14T05:36:20.887333Z","iopub.execute_input":"2022-07-14T05:36:20.888056Z","iopub.status.idle":"2022-07-14T05:36:20.900706Z","shell.execute_reply.started":"2022-07-14T05:36:20.888016Z","shell.execute_reply":"2022-07-14T05:36:20.899291Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Since the data_source data index of the train_df dataframe has a single unique value, HPA, there are 351 counts of this single entity.","metadata":{}},{"cell_type":"markdown","source":"And now, let's go to analyzing the image height and width! To do that, we create a histogram figure with Plotly in a straightforward way by defining the fig variable to the px module with the histogram function, inputting the train_df dataframe and setting the x (x-axis specification) parameter to img_width. Finally, we show the fig figure with the show function.","metadata":{}},{"cell_type":"code","source":"fig = px.histogram(train_df, x=\"img_width\", template=\"simple_white\") # Note: you can change templates if you want!\nfig.show()","metadata":{"execution":{"iopub.status.busy":"2022-07-14T05:36:20.902352Z","iopub.execute_input":"2022-07-14T05:36:20.902808Z","iopub.status.idle":"2022-07-14T05:36:21.013131Z","shell.execute_reply.started":"2022-07-14T05:36:20.902765Z","shell.execute_reply":"2022-07-14T05:36:21.012285Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Now let's do the same to img_height!","metadata":{}},{"cell_type":"code","source":"fig = px.histogram(train_df, x=\"img_height\", template=\"plotly_dark\") # Note: you can change templates if you want!\nfig.show()","metadata":{"execution":{"iopub.status.busy":"2022-07-14T05:36:21.014759Z","iopub.execute_input":"2022-07-14T05:36:21.016007Z","iopub.status.idle":"2022-07-14T05:36:21.129338Z","shell.execute_reply.started":"2022-07-14T05:36:21.015954Z","shell.execute_reply":"2022-07-14T05:36:21.128362Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"When we analyzed the two graphs we ran, we saw that the the image widths and heights ranging from 2300 to 2959 thus from 3060 to 3079 were counted from 1 to 3. However, we can see clearly that the image width and height's range from 3000 to 3019 were counted the most, since there are 326 entities of it.","metadata":{}},{"cell_type":"markdown","source":"Now, let's find out the data entities of age and sex! To get started, we create another histogram by defining the fig variable again to the px module with the histogram function, setting the train_df dataframe as the function's input thus setting the x (x-axis) parameter to age. After that, we display our fig figure variable with the show function.","metadata":{}},{"cell_type":"code","source":"fig = px.histogram(train_df, x=\"age\")\nfig.show()","metadata":{"execution":{"iopub.status.busy":"2022-07-14T05:36:21.132529Z","iopub.execute_input":"2022-07-14T05:36:21.133642Z","iopub.status.idle":"2022-07-14T05:36:21.193462Z","shell.execute_reply.started":"2022-07-14T05:36:21.133603Z","shell.execute_reply":"2022-07-14T05:36:21.192158Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Per the age data histogram, we can see that the most entities of this histogram is the range of 55 to 59 with counts up to 81, and the least entities of it is the range of 50 to 54 with counts up to 9. ","metadata":{}},{"cell_type":"markdown","source":"Now let's create a pie chart figure! Before we begin, let's find the value counts of the sex data index from the train_df dataframe by printing out the train_df dataframe with the sex data index along with the value_counts function!","metadata":{"execution":{"iopub.status.busy":"2022-07-10T22:29:20.587134Z","iopub.execute_input":"2022-07-10T22:29:20.587552Z","iopub.status.idle":"2022-07-10T22:29:20.597379Z","shell.execute_reply.started":"2022-07-10T22:29:20.58751Z","shell.execute_reply":"2022-07-10T22:29:20.596123Z"}}},{"cell_type":"code","source":"train_df[\"sex\"].value_counts()","metadata":{"execution":{"iopub.status.busy":"2022-07-14T05:36:21.194936Z","iopub.execute_input":"2022-07-14T05:36:21.195682Z","iopub.status.idle":"2022-07-14T05:36:21.204765Z","shell.execute_reply.started":"2022-07-14T05:36:21.195646Z","shell.execute_reply":"2022-07-14T05:36:21.203295Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"As you can see, there are 229 Male data entities and 122 Female data entities! Without further ado, let's go onto plotting them with Plotly's go.Pie!","metadata":{}},{"cell_type":"markdown","source":"To get started, we import the plotly module with the graph_objects submodule as go first, then we define the labels variable to an array with two strings, \"Male\" and \"Female\" along with the values variable to an another array with two values, 229 and 122. \n\nNext, we define the fig variable to the go module with the Figure function to create our new figure, setting the data (data specification) parameter to the go module again, but with the Pie function to create a pie chart, setting the labels (labels input) parameter to the labels variable and the values (values input) to the values variable. Finally, we show our fig figure variable by using the show function.","metadata":{}},{"cell_type":"code","source":"import plotly.graph_objects as go\n\nlabels = ['Male', 'Female']\nvalues = [229, 122]\n\nfig = go.Figure(data=[go.Pie(labels=labels, values=values)])\nfig.show()","metadata":{"execution":{"iopub.status.busy":"2022-07-14T05:36:21.206043Z","iopub.execute_input":"2022-07-14T05:36:21.206649Z","iopub.status.idle":"2022-07-14T05:36:21.228018Z","shell.execute_reply.started":"2022-07-14T05:36:21.206613Z","shell.execute_reply":"2022-07-14T05:36:21.226668Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"After running this pie chart graph, we see that the train_df dataframe's sex data covers most of Male by 65.2%, while Female covers the train_df dataframe's data by 34.8%.","metadata":{}},{"cell_type":"markdown","source":"And we cover up most of our data analysis of the train_df dataframe! Now, let's move on to analyzing the images and the masks of this data in this HuBMAP (and HPA)'s competition!","metadata":{}},{"cell_type":"markdown","source":"## Chapter 3: Image and Mask Analysis\nIn order to analyze the images and masks provided in this competition, we need to import tqdm from the tqdm module with the auto submodule first. ","metadata":{}},{"cell_type":"markdown","source":"Before we begin doing that, we're going to define two functions, converting the run-length encode to masks by defining the rle2mask function containing the mask_rle and shape variables and reading out the tiff images by defining a function called read_image, containing the image_id variable. \n\nInside of this rle2mask function, the s variable was defined to mask_rle variable input that was splitted by the split function. Next, the starts and lengths variables is also defined to an array containing the converting the x variable input that looped in the s variable with the slice index between 0 and 0 besides a seperate slice index that has step of 2 and another slice index between 1 and 0 thus having a seperate slice index that also has a step of 2 by the asarray function under the guidance of the np module, setting the dtype (datatype) parameter to integer (int). Furthermore, the starts variable's value has decreased to 1 while the ends variable is defined to the addition compromising the starts and the lengths variable values and the img variable is defined to a new array of the multiplication setting the shape variable's two slice indexes of 0 and 1 filled with zeros with the zeros function by the np module, setting the dtype (datatype) parameter to the np module's 8-bit unsigned integer, known as the uint8 attribute. After that, a for loop has been created, looping the lo and hi variables in the zip object of the starts and ends variables, containing the img variable with the slice index ranging from the two for loop variables defined to 1. Finally, the rle2mask function returns the img variable that has reshaped with the reshape function that contained the shape function, which is now being accessed by the T attribute.\n\nOn the other hand, which is the read_image function, the image variable is defined to the tifffile module with the imread function to read out the tiff images, containing a filepath format leading to the data full of images. An if statement has been initialized, making a condition whether the number of entities of the tuple from the shape attribute to the image variable is equal to 5, then the image variable is redefined to itself being squeezed out the single-dimensional entries with the squeeze function and transpose the rows of it with the transpose function, setting three values: 1, 2, and 0. The mask variable is defined to the rle2mask function call, inputting the train_df dataframe with the data index of the train_df dataframe's id data index equivalent to image_id variable input of this function along with the rle data index outside of that first one thus getting the values of it to an array with the values attribute with the slice index of 0 and the tuple containing the two shapes of the image variable with the shape attribute, with their slice indexes of 1 and 0. Finally, the read_image returns the image and mask variables.","metadata":{}},{"cell_type":"code","source":"def rle2mask(mask_rle, shape):\n    s = mask_rle.split()\n    starts, lengths = [\n        np.asarray(x, dtype=int) for x in (s[0:][::2], s[1:][::2])\n    ]\n    starts -= 1\n    ends = starts + lengths\n    img = np.zeros(shape[0] * shape[1], dtype=np.uint8)\n    \n    for lo, hi in zip(starts, ends):\n        img[lo:hi] = 1\n        \n    return img.reshape(shape).T\n\ndef read_image(image_id):\n    image = tifffile.imread(f\"../input/hubmap-organ-segmentation/train_images/{image_id}.tiff\")\n    \n    if len(image.shape) == 5:\n        image = image.squeeze().transpose(1,2,0)\n        \n    mask = rle2mask(\n        train_df[train_df[\"id\"] == image_id][\"rle\"].values[0],\n        (image.shape[1], image.shape[0])\n    )\n    \n    return image, mask","metadata":{"execution":{"iopub.status.busy":"2022-07-14T05:36:21.229805Z","iopub.execute_input":"2022-07-14T05:36:21.230890Z","iopub.status.idle":"2022-07-14T05:36:21.242429Z","shell.execute_reply.started":"2022-07-14T05:36:21.230839Z","shell.execute_reply":"2022-07-14T05:36:21.240709Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Next, we define the three variables, which is image_ids, images, masks to an array containing another empty array looping the underscore variable in the range from 0 to 3. Thus, we define item to looping in the tqdm setup function, containing the range ranging from 0 to 9, setting the total (total value) parameter to 9, the position (location) parameter to 0, and the desc (description) parameter to \"Reading out image and mask\" so that we need to indicate that we are doing something other than just only a plain, progress bar.\n\nWhilst doing that (see above), the idx variable was defined to the np module with the random attribute along with the randint function to draw out the specific number randomly, containing the number 0 and then the high (high number) parameter to measuring the number of entities with the len function towards the train_df dataframe thus the size parameter to a tuple containing two values that is 1 and no value and on the outside of that, the slice index of 0 contained. Furthermore, three more variables were defined in this for loop, in which the image_id one is defined to the integer location of the train_df dataframe with the iloc attribute with the slice index of the idx variable along with the id data attribute, the image and mask variables to the read_image function call, containing the image_id variable. After setting up the four variables, the images variable is appended to the image variable, along with the masks variable and the image_ids variable, as they're appended to the mask and image_id variables, with the append function.","metadata":{}},{"cell_type":"code","source":"from tqdm.auto import tqdm\nimage_ids, images, masks = [[] for _ in range(3)]\n\nfor item in tqdm(range(9), total=9, position=0, desc=\"Reading image and mask\"):\n    idx = np.random.randint(0, high=len(train_df), size=(1,))[0]\n    image_id = train_df.iloc[idx].id\n    image, mask = read_image(image_id)\n    \n    images.append(image)\n    masks.append(mask)\n    image_ids.append(image_id)","metadata":{"execution":{"iopub.status.busy":"2022-07-14T05:56:59.193419Z","iopub.execute_input":"2022-07-14T05:56:59.193777Z","iopub.status.idle":"2022-07-14T05:57:03.104034Z","shell.execute_reply.started":"2022-07-14T05:56:59.193747Z","shell.execute_reply":"2022-07-14T05:57:03.103123Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"After we read out the images and the masks, let's display them for our analysis, with Matplotlib! First, we create our figure by using the plt module with the figure function, setting the figsize (figure size) parameter to 16 by 16. Then, we create a for loop, looping the ind variable thus redefining image_id and image variables to enumerating the zip of the image_ids and images variables with the zip function followed by the enumerate function. \n\nInside this for loop, we create our subplots with the subplot function under the guidance of plt, setting any number that is less than 5, along with the addition of the ind variable and 1. We then display the images with the imshow function by the plt module, inputting the image varaible inside. Furthermore, we turn off the axis from our plot by using the axis function provided by the plt module, setting it to off.","metadata":{}},{"cell_type":"code","source":"plt.figure(figsize=(16,16))\nfor ind, (image_id, image) in enumerate(zip(image_ids, images)):\n    plt.subplot(4, 3, ind+1)\n    plt.imshow(image)\n    plt.axis(\"off\")","metadata":{"execution":{"iopub.status.busy":"2022-07-14T06:24:13.380545Z","iopub.execute_input":"2022-07-14T06:24:13.380943Z","iopub.status.idle":"2022-07-14T06:24:22.054894Z","shell.execute_reply.started":"2022-07-14T06:24:13.380911Z","shell.execute_reply":"2022-07-14T06:24:22.053338Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Now let's display the masked images! It's like the same as what we did for plotting out the images, but we use the imshow function from the plt module again but inputting out the mask variable, setting the cmap (colormap) parameter to hot and the alpha (transparency) to 0.5. Furthermore, the mask variable was redefined to looping the enumerated zip object of the masks variable.","metadata":{}},{"cell_type":"code","source":"plt.figure(figsize=(16,16))\nfor ind, (image_id, image, mask) in enumerate(zip(image_ids, images, masks)):\n    plt.subplot(4, 3, ind+1)\n    plt.imshow(image)\n    plt.imshow(mask, cmap=\"hot\", alpha=0.5)\n    plt.axis(\"off\")","metadata":{"execution":{"iopub.status.busy":"2022-07-14T06:30:11.674523Z","iopub.execute_input":"2022-07-14T06:30:11.675094Z","iopub.status.idle":"2022-07-14T06:30:30.942179Z","shell.execute_reply.started":"2022-07-14T06:30:11.675050Z","shell.execute_reply":"2022-07-14T06:30:30.940927Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"After running the two seperate code cells, we can see that the first plot shows the images only, while the second plot shows the masks whilst the images in this second plot were dimmed. And with all of that, we completed the data analysis of not just the images and masks, but all of the data in this EDA!","metadata":{}},{"cell_type":"markdown","source":"## Conclusion\nAfter we finished the EDA on hacking the human body, let's wrap up of what we've done! In the first chapter, we created our dataframe out of the train.csv file called train_df and analyzed the basics of it. Next, we covered most data of the train_df dataframe by plotting bars, pie charts, and histograms with violin graphs with all of Plotly. Lastly, we analyzed the images and the masks by plotting and displaying them with Matplotlib! Now that we've done analyzing and segmenting the five organs of the human body, what else are we going to segment to? Other human organs, or the hurricanes from the satellite imagery? In that case, we'll decide.","metadata":{}}]}