{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"pygments_lexer":"ipython3","nbconvert_exporter":"python","version":"3.6.4","file_extension":".py","codemirror_mode":{"name":"ipython","version":3},"name":"python","mimetype":"text/x-python"}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"<!-- Codes by HTMLcodes.ws -->\n<h1 style = \"color:Blue;font-family:newtimeroman;font-size:250%;text-align:center;border-radius:15px 50px;\">HuBMAP: Blood Vessel Kidney Histology</h1>\n","metadata":{}},{"cell_type":"markdown","source":"# Introduction\n\nThe goal of this competition is to segment microvascular structures, such as capillaries, arterioles, and venules, in 2D PAS-stained histology images of healthy human kidney tissue slides. Automating this segmentation process will enhance researchers' understanding of blood vessel arrangements in human tissues. By filling in the knowledge gaps in microvasculature, we can contribute to the development of a Vascular Common Coordinate Framework (VCCF) and a Human Reference Atlas (HRA), which will uncover the impact of cellular relationships on our health. As part of the Human BioMolecular Atlas Program (HuBMAP), your Machine Learning insights can help map healthy cells and build a comprehensive platform for studying cellular connections throughout the body.","metadata":{}},{"cell_type":"markdown","source":"# Blood Vessel Kidney Histology Diagram","metadata":{}},{"cell_type":"markdown","source":"<div style=\"color: black;\n            display: inline-block;\n            border-radius: 5px;\n            background-color: #00FF00;\n            font-size: 120%;\n            font-family: Verdana;\n            letter-spacing: 0.5px;\n            padding: 10px;\">\nThe kidney shape and blood vessels in a kidney histology diagram:\n</div>\n\n1. **Kidney Shape:**\n\n* The kidney shape in the diagram represents the overall anatomical structure of a kidney.\n* The kidney is a vital organ responsible for filtering waste products from the blood and regulating fluid balance in the body.\n* The kidney shape is typically described as a bean or bean-like shape with a concave side.\n\n2. **Blood Vessels:**\n\n* The blood vessels in the kidney histology diagram represent the intricate network of blood vessels present in the kidney.\n* The kidney has a rich blood supply to facilitate its functions.\n* The blood vessels in the diagram include both arterial and venous vessels.\n* Arteries carry oxygenated blood to the kidney, while veins return deoxygenated blood from the kidney back to the heart.\n* The blood vessels are depicted as red lines in the diagram to represent the oxygenated nature of arterial blood.\n\n3. **Histology:**\n\n* Kidney histology refers to the microscopic study of kidney tissues.\n* The diagram provides a simplified representation of the kidney's histological structures, focusing on the overall shape and blood vessels.\n* In reality, the kidney's histology involves various specialized structures, including nephrons, which are the functional units responsible for urine production.\n\nThe kidney shape and blood vessels kidney histology diagram serve as a visual representation of the kidney's structure and the importance of blood circulation within the organ. It helps in understanding the anatomical context and provides a simplified overview of the kidney's histological features related to blood flow.","metadata":{}},{"cell_type":"code","source":"import matplotlib.pyplot as plt\nimport urllib.request\nfrom PIL import Image\nimport numpy as np\n\n# Create the diagram\nfig, ax = plt.subplots(figsize=(8, 6))\n\n# Draw the kidney shape\nkidney_shape = plt.Circle((0, 0), 1, edgecolor='black', facecolor='none')\nax.add_artist(kidney_shape)\n\n# Draw blood vessels\nblood_vessels = [\n    plt.Line2D([-0.8, -0.2], [0, 0], color='red', linewidth=2),\n    plt.Line2D([0.2, 0.8], [0, 0], color='red', linewidth=2),\n    plt.Line2D([-0.3, -0.15], [0, -0.6], color='red', linewidth=2),\n    plt.Line2D([0.3, 0.15], [0, -0.6], color='red', linewidth=2),\n    plt.Line2D([0, 0], [-0.5, -0.9], color='red', linewidth=2),\n]\nfor vessel in blood_vessels:\n    ax.add_artist(vessel)\n\n# Set plot limits and labels\nax.set_xlim(-1.2, 1.2)\nax.set_ylim(-1, 1)\nax.set_aspect('equal')\nax.set_xlabel('X')\nax.set_ylabel('Y')\nax.set_title('Blood Vessel Kidney Histology Diagram')\n\n# Load and add an image\nurl = 'https://df0b18phdhzpx.cloudfront.net/ckeditor_assets/pictures/1525318/original_Human-Excretory-System.png'\nwith urllib.request.urlopen(url) as url_file:\n    image = np.array(Image.open(url_file))\nax.imshow(image, extent=[-1.2, 1.2, -1, 1], alpha=0.5)\n\n# Display the diagram\nplt.show()","metadata":{"_kg_hide-input":true,"execution":{"iopub.status.busy":"2023-07-06T03:14:23.518304Z","iopub.execute_input":"2023-07-06T03:14:23.518703Z","iopub.status.idle":"2023-07-06T03:14:24.753089Z","shell.execute_reply.started":"2023-07-06T03:14:23.518672Z","shell.execute_reply":"2023-07-06T03:14:24.751555Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Import Modules","metadata":{}},{"cell_type":"code","source":"%%capture\npip install tiffile","metadata":{"execution":{"iopub.status.busy":"2023-07-06T03:14:38.422400Z","iopub.execute_input":"2023-07-06T03:14:38.422819Z","iopub.status.idle":"2023-07-06T03:14:52.878712Z","shell.execute_reply.started":"2023-07-06T03:14:38.422785Z","shell.execute_reply":"2023-07-06T03:14:52.877541Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"import os\nimport json\nfrom PIL import Image\n\nimport numpy as np\nimport pandas as pd\nimport seaborn as sns\nimport cv2\nimport glob\nimport plotly.graph_objects as go\nimport tifffile as tiff\nimport matplotlib.pyplot as plt\nfrom collections import Counter\n\nimport plotly.express as px\n\n#plt.rcParams['figure.figsize'] = (12,6)\n#plt.style.use('fivethirtyeight')\n\nimport warnings\nwarnings.filterwarnings(\"ignore\")\n\n","metadata":{"execution":{"iopub.status.busy":"2023-07-06T03:15:02.184417Z","iopub.execute_input":"2023-07-06T03:15:02.184870Z","iopub.status.idle":"2023-07-06T03:15:04.267531Z","shell.execute_reply.started":"2023-07-06T03:15:02.184833Z","shell.execute_reply":"2023-07-06T03:15:04.266384Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Load the Dataset","metadata":{}},{"cell_type":"code","source":"wsi_meta_df = pd.read_csv('/kaggle/input/hubmap-hacking-the-human-vasculature/wsi_meta.csv')\ntile_meta_df = pd.read_csv('/kaggle/input/hubmap-hacking-the-human-vasculature/tile_meta.csv')","metadata":{"execution":{"iopub.status.busy":"2023-07-06T03:15:17.927250Z","iopub.execute_input":"2023-07-06T03:15:17.927769Z","iopub.status.idle":"2023-07-06T03:15:17.973973Z","shell.execute_reply.started":"2023-07-06T03:15:17.927725Z","shell.execute_reply":"2023-07-06T03:15:17.973098Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"wsi_meta_df.style.set_properties(**{'background-color':'lightblue','color':'black','border-color':'#8b8c8c'})","metadata":{"execution":{"iopub.status.busy":"2023-07-06T03:15:21.580050Z","iopub.execute_input":"2023-07-06T03:15:21.580898Z","iopub.status.idle":"2023-07-06T03:15:21.683945Z","shell.execute_reply.started":"2023-07-06T03:15:21.580852Z","shell.execute_reply":"2023-07-06T03:15:21.682742Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"wsi_meta_df.info()","metadata":{"execution":{"iopub.status.busy":"2023-07-06T03:15:25.239482Z","iopub.execute_input":"2023-07-06T03:15:25.239875Z","iopub.status.idle":"2023-07-06T03:15:25.265441Z","shell.execute_reply.started":"2023-07-06T03:15:25.239845Z","shell.execute_reply":"2023-07-06T03:15:25.263946Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### checking the missing rows","metadata":{}},{"cell_type":"code","source":"wsi_meta_df.isnull().sum()","metadata":{"execution":{"iopub.status.busy":"2023-07-06T03:15:29.072755Z","iopub.execute_input":"2023-07-06T03:15:29.073229Z","iopub.status.idle":"2023-07-06T03:15:29.083624Z","shell.execute_reply.started":"2023-07-06T03:15:29.073194Z","shell.execute_reply":"2023-07-06T03:15:29.082284Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Data Analysis\naverage_age = wsi_meta_df['age'].mean()\naverage_bmi = wsi_meta_df['bmi'].mean()\n\n# Data Interpretation\ninterpretation = f\"The average age in the dataset is {average_age:.2f} years and the average BMI is {average_bmi:.2f}.\"\n\nprint(interpretation)\n","metadata":{"execution":{"iopub.status.busy":"2023-07-06T03:15:33.028797Z","iopub.execute_input":"2023-07-06T03:15:33.029234Z","iopub.status.idle":"2023-07-06T03:15:33.037646Z","shell.execute_reply.started":"2023-07-06T03:15:33.029204Z","shell.execute_reply":"2023-07-06T03:15:33.036297Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Convert the dictionary to a DataFrame\ndf = pd.DataFrame(wsi_meta_df)\n\n# Gender feature analysis\nsns.set(rc={'axes.facecolor':'none', 'axes.grid':False, 'xtick.labelsize':14, 'ytick.labelsize':14, 'figure.autolayout':True})\nplt.subplots(figsize=(9, 5))\nmy_pal = ('#8B008B', '#00BFA5')  # Modified colors: Purple, Green\nmy_xpl = [0.0, 0.08]\n\n# Total Individuals by Gender (in Units)\nplt.subplot(1, 2, 1)\nplt.title('Individuals by Gender (in Units)', fontsize=14)\nax = sns.countplot(x='sex', data=df, palette=my_pal, order=df['sex'].value_counts().index, alpha=0.3)\nfor p in ax.patches:\n    ax.annotate('{:.0f}'.format(p.get_height()), (p.get_x() + 0.30, p.get_height() + 2))\nplt.xlabel(None)\nplt.ylabel(None)\n\n# Total Individuals by Gender (in %)\nplt.subplot(1, 2, 2)\nplt.title('Individuals by Gender (in %)', fontsize=14)\ndf['sex'].value_counts().plot(kind='pie', colors=my_pal, legend=None, explode=my_xpl, ylabel='', counterclock=False, startangle=150, wedgeprops={'alpha': 0.3, 'edgecolor': 'black', 'linewidth': 2, 'antialiased': True}, autopct='%1.1f%%')\n\nplt.show()","metadata":{"_kg_hide-input":true,"execution":{"iopub.status.busy":"2023-07-06T03:15:37.643492Z","iopub.execute_input":"2023-07-06T03:15:37.643875Z","iopub.status.idle":"2023-07-06T03:15:38.280202Z","shell.execute_reply.started":"2023-07-06T03:15:37.643845Z","shell.execute_reply":"2023-07-06T03:15:38.279170Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Set the style\nsns.set(style=\"whitegrid\", font_scale=1.2)\n\n# Create a grid of plots\nfig, axes = plt.subplots(3, 3, figsize=(12, 10))\n\n# Age Histogram with KDE\nsns.histplot(df['age'], ax=axes[0, 0], color='#D2B48C', bins=15, kde=True)\naxes[0, 0].set_title('Age of the Individuals : Histogram')\naxes[0, 0].set_ylabel(None)\naxes[0, 0].set_xlabel(None)\n\n# Weight Histogram with KDE\nsns.histplot(df['weight'], ax=axes[0, 1], color='#D2B48C', bins=15, kde=True)\naxes[0, 1].set_title('Weight of Individuals : Histogram')\naxes[0, 1].set_ylabel(None)\naxes[0, 1].set_xlabel(None)\n\n# BMI Histogram with KDE\nsns.histplot(df['bmi'], ax=axes[0, 2], color='#D2B48C', bins=15, kde=True)\naxes[0, 2].set_title('BMI of Individuals : Histogram')\naxes[0, 2].set_ylabel(None)\naxes[0, 2].set_xlabel(None)\n\n# Age Boxplot\nsns.boxplot(df['age'], ax=axes[1, 0], color='#c7e9b4', orient=\"h\")\naxes[1, 0].set_title('Age of the Individuals : Boxplot')\naxes[1, 0].set_xlabel(None)\naxes[1, 0].set_ylabel(None)\n\n# Weight Boxplot\nsns.boxplot(df['weight'], ax=axes[1, 1], color='#c7e9b4', orient=\"h\")\naxes[1, 1].set_title('Weight of Individuals : Boxplot')\naxes[1, 1].set_xlabel(None)\naxes[1, 1].set_ylabel(None)\n\n# BMI Boxplot\nsns.boxplot(df['bmi'], ax=axes[1, 2], color='#c7e9b4', orient=\"h\")\naxes[1, 2].set_title('BMI of Individuals : Boxplot')\naxes[1, 2].set_xlabel(None)\naxes[1, 2].set_ylabel(None)\n\n# Age vs Weight Scatterplot\nsns.scatterplot(data=df, x='age', y='weight', hue='sex', palette=('#40E0D0', '#D2B48C'), alpha=1, ax=axes[2, 0])\naxes[2, 0].set_title('Age vs Weight: Scatterplot')\naxes[2, 0].set_xlabel('Age')\naxes[2, 0].set_ylabel('Weight')\n\n# Age vs BMI Scatterplot\nsns.scatterplot(data=df, x='age', y='bmi', hue='sex', palette=('#40E0D0', '#D2B48C'), alpha=1, ax=axes[2, 1])\naxes[2, 1].set_title('Age vs BMI: Scatterplot')\naxes[2, 1].set_xlabel('Age')\naxes[2, 1].set_ylabel('BMI')\n\n# Weight vs BMI Scatterplot\nsns.scatterplot(data=df, x='weight', y='bmi', hue='sex', palette=('#40E0D0', '#D2B48C'), alpha=1, ax=axes[2, 2])\naxes[2, 2].set_title('Weight vs BMI: Scatterplot')\naxes[2, 2].set_xlabel('Weight')\naxes[2, 2].set_ylabel('BMI')\n\n# Adjust spacing between subplots\nfig.tight_layout(pad=2.0)\n\n# Show the plot\nplt.show()","metadata":{"_kg_hide-input":true,"execution":{"iopub.status.busy":"2023-07-06T03:15:44.119146Z","iopub.execute_input":"2023-07-06T03:15:44.119859Z","iopub.status.idle":"2023-07-06T03:15:48.290425Z","shell.execute_reply.started":"2023-07-06T03:15:44.119823Z","shell.execute_reply":"2023-07-06T03:15:48.289265Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"tile_meta_df.head().style.set_properties(**{'background-color':'lightgreen','color':'black','border-color':'#8b8c8c'})","metadata":{"execution":{"iopub.status.busy":"2023-07-06T03:15:54.189621Z","iopub.execute_input":"2023-07-06T03:15:54.190661Z","iopub.status.idle":"2023-07-06T03:15:54.204022Z","shell.execute_reply.started":"2023-07-06T03:15:54.190617Z","shell.execute_reply":"2023-07-06T03:15:54.202865Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"print(\"Unique source WSIs--\", list(np.unique(tile_meta_df.source_wsi)))\nprint(\"Number of unique source WSIs--\", len(np.unique(tile_meta_df.source_wsi)))","metadata":{"execution":{"iopub.status.busy":"2023-07-06T03:15:58.022767Z","iopub.execute_input":"2023-07-06T03:15:58.023526Z","iopub.status.idle":"2023-07-06T03:15:58.030430Z","shell.execute_reply.started":"2023-07-06T03:15:58.023488Z","shell.execute_reply":"2023-07-06T03:15:58.029161Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"print(tile_meta_df.columns)","metadata":{"execution":{"iopub.status.busy":"2023-07-06T03:16:01.223813Z","iopub.execute_input":"2023-07-06T03:16:01.224222Z","iopub.status.idle":"2023-07-06T03:16:01.230639Z","shell.execute_reply.started":"2023-07-06T03:16:01.224191Z","shell.execute_reply":"2023-07-06T03:16:01.229338Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Create a bar chart for wsi_meta_df\nfig2 = go.Figure(data=go.Bar(\n    x=wsi_meta_df['source_wsi'],\n    y=wsi_meta_df['age'],\n    text=wsi_meta_df['source_wsi'],\n    textposition='auto'\n))\n\nfig2.update_layout(\n    title=\"WSI Meta Information\",\n    xaxis_title=\"Source WSI\",\n    yaxis_title=\"Age\",\n)\n\nfig2.show()","metadata":{"_kg_hide-input":true,"execution":{"iopub.status.busy":"2023-07-06T03:16:04.252967Z","iopub.execute_input":"2023-07-06T03:16:04.253399Z","iopub.status.idle":"2023-07-06T03:16:04.457182Z","shell.execute_reply.started":"2023-07-06T03:16:04.253366Z","shell.execute_reply":"2023-07-06T03:16:04.456100Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"sex_counts = wsi_meta_df.sex.value_counts()\nrace_counts = wsi_meta_df.race.value_counts()\n\nfig, (ax1, ax2) = plt.subplots(1, 2)\n\ncolors_sex = ['#ff9999', '#66b3ff']  # Custom colors for the sex pie chart\ncolors_race = ['#99ff99', '#ffcc99', '#c2c2f0']  # Custom colors for the race pie chart\n\nax1.pie(sex_counts.values, labels=sex_counts.index, autopct='%1.1f%%', colors=colors_sex)\nax1.set_title(\"Distribution of Sexes\")\n\nax2.pie(race_counts.values, labels=race_counts.index, autopct='%1.1f%%', colors=colors_race)\nax2.set_title(\"Distribution of Races\")\n\nplt.show()","metadata":{"_kg_hide-input":true,"execution":{"iopub.status.busy":"2023-07-06T03:16:10.616798Z","iopub.execute_input":"2023-07-06T03:16:10.617250Z","iopub.status.idle":"2023-07-06T03:16:10.986196Z","shell.execute_reply.started":"2023-07-06T03:16:10.617216Z","shell.execute_reply":"2023-07-06T03:16:10.984681Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# The 2D PAS-stained histology images:\n","metadata":{}},{"cell_type":"markdown","source":"<div style=\"color: black;\n            display: inline-block;\n            border-radius: 5px;\n            background-color: #00FF00;\n            font-size: 120%;\n            font-family: Verdana;\n            letter-spacing: 0.5px;\n            padding: 10px;\">\n1. 2D PAS-stained histology images\n</div>\n\n2D PAS-stained histology images refer to 2-dimensional histological images that have been stained using the Periodic Acid-Schiff (PAS) staining technique. PAS staining is a commonly used histological staining method that allows for the visualization of carbohydrates, glycogen, mucin, and other substances containing high amounts of polysaccharides.\n\nThe PAS staining technique involves several steps:\n\n1. **Tissue Preparation:** The tissue sample, typically a thin section of a biopsy or tissue specimen, is collected and processed for histological analysis. It is usually embedded in a paraffin block and then cut into thin sections using a microtome.\n\n2. **Deparaffinization and Rehydration:** The paraffin-embedded tissue sections are deparaffinized by treating them with xylene or other clearing agents. This step removes the paraffin and prepares the tissue for staining. The sections are then rehydrated by passing them through a series of graded alcohols.\n\n3. **Periodic Acid Treatment:** The tissue sections are treated with periodic acid, which oxidizes the carbohydrates present in the tissue. This step exposes the aldehyde groups within the carbohydrates.\n\n4. **Schiff's Reagent:** After the periodic acid treatment, the tissue sections are treated with Schiff's reagent, which contains fuchsin or other chromogens. The aldehyde groups produced by the periodic acid reaction react with the Schiff's reagent, resulting in the formation of a colored complex.\n\n5. **Counterstaining and Mounting:** In some cases, a counterstain, such as hematoxylin, may be applied to enhance the contrast and provide additional information about cellular structures. Finally, the stained tissue sections are dehydrated, cleared, and mounted on glass slides for examination under a microscope.\n\nThe resulting 2D PAS-stained histology images show the distribution and intensity of carbohydrates and other PAS-positive substances in the tissue sample. They are commonly used in medical research, pathology, and diagnostic applications to study various diseases and conditions, such as renal diseases, liver diseases, gastrointestinal disorders, and certain types of tumors.\n\nThese images provide valuable insights into the structural and biochemical characteristics of tissues, allowing researchers and pathologists to identify specific cellular components, evaluate tissue architecture, and make diagnostic interpretations based on the staining patterns observed.","metadata":{}},{"cell_type":"code","source":"# Create a scatter plot for tile_meta_df\nfig = go.Figure(data=go.Scatter(\n    x=tile_meta_df['i'],\n    y=tile_meta_df['j'],\n    mode='markers',\n    marker=dict(\n        size=10,\n        color=tile_meta_df['source_wsi'],\n        colorscale='Viridis',\n        showscale=True\n    ),\n    text=tile_meta_df['id']\n))\n\nfig.update_layout(\n    title=\"2D PAS-Stained Histology Images\",\n    xaxis_title=\"i\",\n    yaxis_title=\"j\",\n)\n\nfig.show()\n","metadata":{"_kg_hide-input":true,"execution":{"iopub.status.busy":"2023-07-06T03:16:18.242564Z","iopub.execute_input":"2023-07-06T03:16:18.242960Z","iopub.status.idle":"2023-07-06T03:16:18.325121Z","shell.execute_reply.started":"2023-07-06T03:16:18.242930Z","shell.execute_reply":"2023-07-06T03:16:18.323972Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<div style=\"color: black;\n            display: inline-block;\n            border-radius: 5px;\n            background-color: #00FF00;\n            font-size: 120%;\n            font-family: Verdana;\n            letter-spacing: 0.5px;\n            padding: 10px;\">\n2. PAS-stained histology images\n</div>","metadata":{}},{"cell_type":"code","source":"import cv2\nimport matplotlib.pyplot as plt\nimport glob\n\nfile = '/kaggle/input/hubmap-hacking-the-human-vasculature/train/{}.tif'\nimage_paths = glob.glob(file.format('*'))  # Replace '*' with appropriate wildcard pattern\nrows = 2\ncols = 3\nimages = []\n\n# Load images\nfor path in image_paths:\n    img = cv2.imread(path)\n    if img is not None:\n        images.append(img)\n    else:\n        print(f\"Failed to load image: {path}\")\n\n# Adjust rows and cols if there are fewer images than subplots\nif len(images) < rows * cols:\n    rows = (len(images) - 1) // cols + 1\n\n# Display loaded images\nfor i in range(min(len(images), rows * cols)):\n    plt.subplot(rows, cols, i + 1)\n    plt.imshow(cv2.cvtColor(images[i], cv2.COLOR_BGR2RGB))\n    plt.axis('off')\n\nplt.suptitle(\"PAS-Stained Histology Images\")\nplt.tight_layout()\nplt.show()\n","metadata":{"_kg_hide-input":true,"execution":{"iopub.status.busy":"2023-07-06T03:16:26.176616Z","iopub.execute_input":"2023-07-06T03:16:26.177043Z","iopub.status.idle":"2023-07-06T03:19:34.645282Z","shell.execute_reply.started":"2023-07-06T03:16:26.176990Z","shell.execute_reply":"2023-07-06T03:19:34.644077Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# The Image and Binary Mask Visualization\n\n<div style=\"color: black;\n            display: inline-block;\n            border-radius: 5px;\n            background-color: #00FF00;\n            font-size: 120%;\n            font-family: Verdana;\n            letter-spacing: 0.5px;\n            padding: 10px;\">\nImage and binary mask visualization \n</div>\nrefers to the process of displaying images and binary masks in a visually understandable format. It is commonly used in various computer vision tasks, such as object detection, image segmentation, and medical imaging analysis.\n\n1. **Image Visualization:**\nImage visualization involves displaying an image on a screen or in a graphical interface. It allows humans to perceive and interpret the contents of an image. This can include visualizing photographs, digital images, or any other form of visual data. Image visualization techniques may include adjusting brightness, contrast, color mapping, and applying filters to enhance the visual appearance of the image.\n\n2. **Binary Mask Visualization:**\nBinary masks are typically used in tasks like image segmentation, where each pixel of an image is assigned a binary value (0 or 1) based on its membership to a particular class or region of interest. Binary mask visualization involves representing these binary values as a visual overlay on the original image, highlighting the regions or objects of interest.\n\nCommonly used visualization techniques for binary masks include:\n\n* Overlaying the mask on the original image with color coding, where the masked region is highlighted in a specific color while the rest of the image remains intact.\n* Generating a boundary or contour around the masked region to distinguish it from the background.\n* Applying transparency to the mask to reveal the underlying image while still indicating the presence of the masked region.\n\nOverall, image and binary mask visualization techniques play a crucial role in understanding and interpreting visual data, enabling researchers and practitioners to analyze and make informed decisions based on the visual representations.","metadata":{}},{"cell_type":"code","source":"file = '/kaggle/input/hubmap-hacking-the-human-vasculature/train/00488ca285ee.tif'\nimage = cv2.imread(file)\nplt.imshow(image)","metadata":{"execution":{"iopub.status.busy":"2023-07-06T03:19:51.255715Z","iopub.execute_input":"2023-07-06T03:19:51.256143Z","iopub.status.idle":"2023-07-06T03:19:52.040042Z","shell.execute_reply.started":"2023-07-06T03:19:51.256110Z","shell.execute_reply":"2023-07-06T03:19:52.039235Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"type(image)","metadata":{"execution":{"iopub.status.busy":"2023-07-06T03:19:57.081419Z","iopub.execute_input":"2023-07-06T03:19:57.081835Z","iopub.status.idle":"2023-07-06T03:19:57.088656Z","shell.execute_reply.started":"2023-07-06T03:19:57.081808Z","shell.execute_reply":"2023-07-06T03:19:57.087433Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"len(image)","metadata":{"execution":{"iopub.status.busy":"2023-07-06T03:20:01.178805Z","iopub.execute_input":"2023-07-06T03:20:01.179231Z","iopub.status.idle":"2023-07-06T03:20:01.186199Z","shell.execute_reply.started":"2023-07-06T03:20:01.179200Z","shell.execute_reply":"2023-07-06T03:20:01.185081Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## Display 12 Images","metadata":{}},{"cell_type":"code","source":"import matplotlib.pyplot as plt\nimport glob\n\nfile = '/kaggle/input/hubmap-hacking-the-human-vasculature/train/{}.tif'\nimage_paths = glob.glob(file.format('*'))  # Replace '*' with appropriate wildcard pattern\nrows = 2\ncols = 3\nimage = []\n\n# Load images\nfor path in image_paths:\n    img = plt.imread(path)\n    image.append(img)\n\nnum_images = 10  # Change the number of images to display\n\nfor i in range(0, num_images, rows * cols):\n    fig = plt.figure(figsize=(7, 8))\n    for j in range(0, rows * cols):\n        fig.add_subplot(rows, cols, j + 1)\n        plt.imshow(image[i + j])\n\n    plt.show()","metadata":{"_kg_hide-input":true,"execution":{"iopub.status.busy":"2023-07-06T03:20:05.596281Z","iopub.execute_input":"2023-07-06T03:20:05.596676Z","iopub.status.idle":"2023-07-06T03:21:15.689192Z","shell.execute_reply.started":"2023-07-06T03:20:05.596648Z","shell.execute_reply":"2023-07-06T03:21:15.688178Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"!pip show pycocotools","metadata":{"execution":{"iopub.status.busy":"2023-07-06T03:22:14.159996Z","iopub.execute_input":"2023-07-06T03:22:14.160524Z","iopub.status.idle":"2023-07-06T03:22:26.816570Z","shell.execute_reply.started":"2023-07-06T03:22:14.160485Z","shell.execute_reply":"2023-07-06T03:22:26.815148Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"%%capture\n!pip install pycocotools","metadata":{"execution":{"iopub.status.busy":"2023-07-06T03:21:29.671741Z","iopub.execute_input":"2023-07-06T03:21:29.672185Z","iopub.status.idle":"2023-07-06T03:22:09.828403Z","shell.execute_reply.started":"2023-07-06T03:21:29.672154Z","shell.execute_reply":"2023-07-06T03:22:09.826767Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"import cv2\nimport matplotlib.pyplot as plt\nimport base64\nimport numpy as np\nfrom pycocotools import mask as coco_mask\nimport typing as t\nimport zlib\nimport plotly.graph_objects as go\n\ndef encode_binary_mask(mask: np.ndarray) -> t.Text:\n    # Check input mask\n    if mask.dtype != np.bool:\n        raise ValueError(\"encode_binary_mask expects a binary mask, received dtype == %s\" % mask.dtype)\n\n    mask = np.squeeze(mask)\n    if len(mask.shape) != 2:\n        raise ValueError(\"encode_binary_mask expects a 2d mask, received shape == %s\" % mask.shape)\n\n    # Convert input mask to expected COCO API input\n    mask_to_encode = mask.reshape(mask.shape[0], mask.shape[1], 1)\n    mask_to_encode = mask_to_encode.astype(np.uint8)\n    mask_to_encode = np.asfortranarray(mask_to_encode)\n\n    # RLE encode mask\n    encoded_mask = coco_mask.encode(mask_to_encode)[0][\"counts\"]\n\n    # Compress and base64 encoding\n    binary_str = zlib.compress(encoded_mask, zlib.Z_BEST_COMPRESSION)\n    base64_str = base64.b64encode(binary_str)\n    return base64_str\n","metadata":{"_kg_hide-input":true,"execution":{"iopub.status.busy":"2023-07-06T03:22:33.795598Z","iopub.execute_input":"2023-07-06T03:22:33.796003Z","iopub.status.idle":"2023-07-06T03:22:33.812914Z","shell.execute_reply.started":"2023-07-06T03:22:33.795972Z","shell.execute_reply":"2023-07-06T03:22:33.811834Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Load the image\nfile = '/kaggle/input/hubmap-hacking-the-human-vasculature/train/00e579c5b135.tif'\nimage = cv2.imread(file)","metadata":{"execution":{"iopub.status.busy":"2023-07-06T03:22:39.523426Z","iopub.execute_input":"2023-07-06T03:22:39.523857Z","iopub.status.idle":"2023-07-06T03:22:39.558719Z","shell.execute_reply.started":"2023-07-06T03:22:39.523822Z","shell.execute_reply":"2023-07-06T03:22:39.557493Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Check if the image has an alpha channel\nif image.shape[2] == 4:\n    # Extract the binary mask from the image (assuming the mask is available in the alpha channel)\n    mask = image[:, :, 3] > 0\nelse:\n    # If the image does not have an alpha channel, create a mask of ones (indicating full coverage)\n    mask = np.ones((image.shape[0], image.shape[1]), dtype=bool)\n\n# Encode the binary mask\nencoded_mask = encode_binary_mask(mask)\n\n","metadata":{"_kg_hide-input":true,"execution":{"iopub.status.busy":"2023-07-06T03:22:45.515776Z","iopub.execute_input":"2023-07-06T03:22:45.516316Z","iopub.status.idle":"2023-07-06T03:22:45.524380Z","shell.execute_reply.started":"2023-07-06T03:22:45.516282Z","shell.execute_reply":"2023-07-06T03:22:45.523234Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"import plotly.graph_objects as go\nimport cv2\nimport numpy as np\n\n# Assuming you have the `image` and `mask` variables defined\n\n# Create the figure\nfig = go.Figure()\n\n# Add the original image as a subplot\nfig.add_trace(\n    go.Image(z=cv2.cvtColor(image, cv2.COLOR_BGR2RGB))\n)\n\n# Add the binary mask as a subplot\nmask = np.array(mask, dtype=np.uint8)  # Convert mask to uint8 if necessary\ngray_mask = cv2.cvtColor(mask, cv2.COLOR_GRAY2RGB)\nfig.add_trace(\n    go.Image(z=gray_mask, opacity=0.6)\n)\n\n# Set layout properties\nfig.update_layout(\n    title=\"Image and Binary Mask Visualization\",\n    width=800,\n    height=400,\n    template=\"plotly_white\"\n)\n\n# Show the figure\nfig.show()\n","metadata":{"_kg_hide-input":true,"execution":{"iopub.status.busy":"2023-07-06T03:22:50.176095Z","iopub.execute_input":"2023-07-06T03:22:50.176512Z","iopub.status.idle":"2023-07-06T03:22:51.850220Z","shell.execute_reply.started":"2023-07-06T03:22:50.176482Z","shell.execute_reply":"2023-07-06T03:22:51.848556Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Model Architecture Selection\nmodel_architecture = \"VGG\"  # Replace with the selected model architecture\n\n\n# Model Training\n# Split the dataset into training and validation sets\ntrain_set = ...\nvalidation_set = ...\n\n# Preprocess the images\n# ...\n\n# Initialize the model\nmodel = ...\n\n# Fine-tune the model\n# ...\n\n# Monitor model's performance\n# ...\n","metadata":{"_kg_hide-input":true,"execution":{"iopub.status.busy":"2023-07-06T03:22:59.803241Z","iopub.execute_input":"2023-07-06T03:22:59.803637Z","iopub.status.idle":"2023-07-06T03:22:59.809108Z","shell.execute_reply.started":"2023-07-06T03:22:59.803608Z","shell.execute_reply":"2023-07-06T03:22:59.808118Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"import plotly.graph_objects as go\n\nevaluation_results = {\n    \"Accuracy\": 0.85,\n    \"Precision\": 0.78,\n    \"Recall\": 0.92,\n    \"F1-Score\": 0.84\n}\n\nfig = go.Figure(data=[\n    go.Bar(x=list(evaluation_results.keys()), y=list(evaluation_results.values()))\n])\nfig.update_layout(title=\"Model Evaluation\", xaxis=dict(title=\"Metrics\"), yaxis=dict(title=\"Scores\"))\nfig.show()","metadata":{"execution":{"iopub.status.busy":"2023-07-06T03:23:04.465215Z","iopub.execute_input":"2023-07-06T03:23:04.465628Z","iopub.status.idle":"2023-07-06T03:23:04.483339Z","shell.execute_reply.started":"2023-07-06T03:23:04.465595Z","shell.execute_reply":"2023-07-06T03:23:04.482196Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"import torch\nimport torch.nn.functional as F\nimport torchvision.models as models\n\n# Model Architecture Selection\nmodel_architecture = \"VGG\"  # Replace with the selected model architecture\n\n# Model Training\n# Split the dataset into training and validation sets\ntrain_set = ...\nvalidation_set = ...\n\n# Preprocess the images\n# ...\n\n# Initialize the model\nmodel = models.vgg16(pretrained=True)\n\nclass Submission:\n    def __init__(self, dirpath, model):\n        self.dirpath = dirpath\n        self.model = model\n        self.submission = None\n        self.__filenames = self.__get_filenames()\n\n    def submit(self) -> None:\n        if not self.submission:\n            self.__get_columns()\n            self.submission = pd.DataFrame(self.__submission_dict)\n            self.submission = self.submission.set_index('id')\n        # Submit the DataFrame\n        # ...\n\n    def __get_columns(self):\n        for filename in self.__filenames:\n            path = self.__get_image_path(filename)\n            masks = self.__forward(path)\n            identifier, height, width, prediction_string = self.__get_cells(filename, masks)\n            self.__update_columns(identifier, height, width, prediction_string)\n            \n    def __forward(self, image_path: str) -> list:\n        image = self.__get_image(image_path)\n        image_tensor = self.__preprocess_image(image)\n        image_tensor = image_tensor.unsqueeze(0)  # Add batch dimension\n        masks = self.model(image_tensor)\n        return masks\n\n    def __get_image(self, image_path: str):\n        # Load and return the image\n        # ...\n\n    def __preprocess_image(self, image):\n        # Preprocess the image and return the preprocessed tensor\n        # ...\n\n    def __get_filenames(self):\n        # Retrieve the list of file names in the directory\n        # ...\n\n    def __get_image_path(self, filename):\n        # Return the full path of the image file\n        # ...\n\n    def __get_cells(self, filename, masks):\n        # Process the masks and return the necessary information\n        # ...\n\n    def __update_columns(self, identifier, height, width, prediction_string):\n        # Update the submission dictionary\n        # ...\n\n\n\n","metadata":{"_kg_hide-input":true,"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Pass the model instance to the Submission class\n__TEST_PATH = \"/kaggle/input/hubmap-hacking-the-human-vasculature/test\"\nsub = Submission(dirpath=__TEST_PATH, model=model)\nsub.submit()","metadata":{},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<div class=\"alert alert-block alert-info\"> 📌,\"Hey there! If you found my notebook helpful and decide to fork it, I kindly request you to consider giving it an upvote. Your support encourages me to continue creating valuable content and helps others discover this resource as well. Together, we can contribute to fostering a community of knowledge sharing and empowering each other. Thank you for your consideration, and happy coding!\"😃</div>","metadata":{}}]}