{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"pygments_lexer":"ipython3","nbconvert_exporter":"python","version":"3.6.4","file_extension":".py","codemirror_mode":{"name":"ipython","version":3},"name":"python","mimetype":"text/x-python"}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"![](https://www.fullstackpython.com/img/logos/bokeh.jpg)\n\n*Image source*: [www.fullstackpython.com](https://www.fullstackpython.com/bokeh.html)","metadata":{}},{"cell_type":"markdown","source":"## Introduction\n**Data visualization** is an interdisciplinary field that deals with the graphic representation of data [[wiki](https://en.wikipedia.org/wiki/Data_visualization)]. Data visualization can be employed at any time during a lifecycle of data science project, especially at the beginning of the project to explore the data at hand (to supplement the EDA); and at the end of the project when we want to communicate the findings of our work to the stakeholders of the project. So, understandably it is a vital tool to have in a data scientists/analysts skill set. Depending upon the programming language preference and test, one can choose from several visualization tools such as matplotlib, seaborn, plotly, **bokeh**, ggplot2, and so on. In the past I have used the first three tools from the aforementioned options. Today, using this notebook, I will explore bokeh's python data visualization library with the help of **HoloViews** when needed.\n\n**What is Bokeh ?**\n> Bokeh is a Python library for creating interactive visualizations for modern web browsers. It helps you build beautiful graphics, ranging from simple plots to complex dashboards with streaming datasets. With Bokeh, you can create JavaScript-powered visualizations without writing any JavaScript yourself [bokeh's page].\n\n**What is HoloViews ?** \n> **HoloViews** is an open-source Python library designed to make data analysis and visualization seamless and simple. With HoloViews, you can usually express what you want to do in very few lines of code, letting you focus on what you are trying to explore and convey, not on the process of plotting [[ref](https://holoviews.org/index.html)].\n\n**Objective**: \n> The **goal** of this notebook is to explore the bokeh python library for data visualizations and at the same time **customize** the plots to make them look as good as possible while conveying the message as effectively as possible.\n\n**Note**: This notebook is not an EDA on a standalone dataset, it is an exploration of a visualization tool; hence I may use several datasets for demonstration purposes.\n\n","metadata":{}},{"cell_type":"markdown","source":"<a id=\"top\"></a>\n\n<h3 class=\"list-group-item list-group-item-action active\" data-toggle=\"list\" role=\"tab\" aria-controls=\"home\">Table of Contents</h3>\n\n\n* [0. Plots overview](#0)\n* [1. Basic plots](#1)\n    * [1.1 Scatter plots](#1.1)\n    * [1.2 Line plot](#1.2)\n    * [1.3 Bar graphs](#1.3)\n    * [1.4 Histograms](#1.4)\n    * [1.5 Pie charts](#1.5)\n    * [1.6 Box plots](#1.6)\n    * [1.7 Violin plots](#1.7)\n* [2. Area plots](#2)\n* [3. Heatmap](#3)\n* [4. Jitter Scatter (categorical)](#4)\n* [5. Kde plots](#5)\n* [6. Subplots](#6)\n* [7. Range tools (useful for time series data)](#7)\n* [8. Pairplots (grouped grids)](#8)\n* [9. Closing Remarks](#9)\n* [10. References](#10)\n\n","metadata":{}},{"cell_type":"markdown","source":"## Install bokeh","metadata":{}},{"cell_type":"code","source":"!pip install bokeh","metadata":{"_kg_hide-output":true,"execution":{"iopub.status.busy":"2022-04-22T10:26:28.829784Z","iopub.execute_input":"2022-04-22T10:26:28.830168Z","iopub.status.idle":"2022-04-22T10:26:38.307986Z","shell.execute_reply.started":"2022-04-22T10:26:28.830094Z","shell.execute_reply":"2022-04-22T10:26:38.306206Z"},"jupyter":{"source_hidden":true},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## Import libraries","metadata":{}},{"cell_type":"code","source":"# all library may or may not be used \n\nimport numpy as np\nimport pandas as pd \nfrom scipy.stats.kde import gaussian_kde\n\nimport colorcet as cc\nfrom bokeh.models import BoxAnnotation, RangeTool\nfrom bokeh.models import ColumnDataSource, FixedTicker, PrintfTickFormatter\nfrom bokeh.io import output_file, show, output_notebook, curdoc\nfrom bokeh.plotting import figure\nfrom bokeh.models import ColumnDataSource, LabelSet, HoverTool, Range1d, Label\nfrom bokeh.palettes import GnBu3, OrRd3, Category20c\nfrom bokeh.transform import cumsum\nfrom bokeh.transform import jitter\nfrom bokeh.layouts import column\nfrom holoviews.operation import gridmatrix\noutput_notebook()\n\nimport holoviews as hv\nfrom holoviews import opts, dim\nhv.extension('bokeh')\n\ntheme_list = ['caliber', 'dark_minimal', 'light_minimal', \"night_sky\", \"contrast\"]\nimport warnings\nwarnings.filterwarnings('ignore')\n\nimport os\nfor dirname, _, filenames in os.walk('/kaggle/input'):\n    for filename in filenames:\n        print(os.path.join(dirname, filename))","metadata":{"_uuid":"8f2839f25d086af736a60e9eeb907d3b93b6e0e5","_cell_guid":"b1076dfc-b9ad-4769-8c92-a6c4dae69d19","_kg_hide-input":true,"_kg_hide-output":true,"execution":{"iopub.status.busy":"2022-04-22T10:26:38.310419Z","iopub.execute_input":"2022-04-22T10:26:38.310753Z","iopub.status.idle":"2022-04-22T10:26:41.428633Z","shell.execute_reply.started":"2022-04-22T10:26:38.31072Z","shell.execute_reply":"2022-04-22T10:26:41.427458Z"},"jupyter":{"source_hidden":true,"outputs_hidden":true},"collapsed":true,"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## Load data","metadata":{}},{"cell_type":"code","source":"df_iris = pd.read_csv('/kaggle/input/iris/Iris.csv')\ndf_titanic = pd.read_csv('/kaggle/input/titanic/train.csv')\ndf_rain= pd.read_csv('/kaggle/input/weather-dataset-rattle-package/weatherAUS.csv')\ndf_study = pd.read_csv('/kaggle/input/students-performance-in-exams/StudentsPerformance.csv')\ndf_heart = pd.read_csv('/kaggle/input/heart-failure-clinical-data/heart_failure_clinical_records_dataset.csv')\ndf_sale = pd.read_csv('/kaggle/input/competitive-data-science-predict-future-sales/sales_train.csv')\ndf_covid = pd.read_csv('/kaggle/input/d/gpreda/covid-world-vaccination-progress/country_vaccinations.csv')\ndf_housing = pd.read_csv('/kaggle/input/house-prices-advanced-regression-techniques/train.csv')\ndf_ts = pd.read_csv('/kaggle/input/tabular-playground-series-jul-2021/train.csv')","metadata":{"_kg_hide-input":true,"_kg_hide-output":true,"execution":{"iopub.status.busy":"2022-04-22T10:26:41.430634Z","iopub.execute_input":"2022-04-22T10:26:41.431191Z","iopub.status.idle":"2022-04-22T10:26:45.409444Z","shell.execute_reply.started":"2022-04-22T10:26:41.431151Z","shell.execute_reply":"2022-04-22T10:26:45.408319Z"},"jupyter":{"source_hidden":true},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"\n\n","metadata":{}},{"cell_type":"markdown","source":"<a id=\"0\"></a>\n<font color=\"skyblue\" size=+2.5><b>0. Plots overview</b></font>\n\n<a href=\"#top\" class=\"btn btn-primary btn-sm\" role=\"button\" aria-pressed=\"true\" style=\"color:white\" data-toggle=\"popover\">go to top</a>\n## Bokeh:\n\nAfter downloading and importing the necessary packages, creating a bokeh plot is a two-step process: First, you select from Bokeh’s building blocks to create your visualization and second, you customize these building blocks to fit your needs. The steps follows like this:\n\n* select theme : (optional)\n* get/prepare `data`\n* call `figure()` object\n* add the `renderer(s)` or the `plot(s)`\n* customize figure layout (optional)\n* `show()` the figure\n\nThemes: In Bokeh one can choose from the following themes, ['caliber', 'dark_minimal', 'light_minimal', \"night_sky\", \"contrast\"] + custom-made\n\n## HoloViews:\n\nSimilarlly, creating a holoview plot is a two-step process: First, you select from holowies’s building blocks (`elements`) to create your visualization and second, you customize the elements to fit your needs. The steps follows like this:\n\n* prepare `data`\n* call the `element` with the neccessary arguments\n* `customize` the element (optional)\n\nNote: The customization in holoviews is not as rich as Bokeh's and it has no its own theme selection as well. However, you can call bokeh's theme via `hv.renderer('bokeh').theme`= 'you choice here'\n\n## Providing Data\n\nThe basis for any data visualization is the underlying data. Below we will see the various ways to provide data to Bokeh, from passing data values directly to creating a `ColumnDataSource`\n\n**1. Python lists/arrays**: You can use standard Python lists of data to pass values directly into a plotting function. For example, you could use python lists to make a plot as follow.\n\n> x_values = [1, 2, 3, 4, 5] --> (x_valuse as python list)\n>\n> y_values = [6, 7, 2, 3, 6] --> (y_valuse as python list)\n>\n> fig = figure() --> (add the renderer)\n>\n> fig.circle(x=x_values, y=y_values) --> (pass the lists to the renderer)\n        \n\n**2. Numpy data**: Similarly to using Python lists and arrays, you can also work with NumPy data structures. Example below.\n\n> x = [1, 2, 3, 4, 5] --> (x as list)\n>\n> y = np.random.standard_normal(5) --> (y as Numpy array)\n>\n> fig = figure() --> (add the renderer)\n>\n> fig.circle(x=x, y=y) --> (pass the list and the Numpy array to the renderer)\n\n        \n**3. ColumnDataSource**: The `ColumnDataSource` is the core of most Bokeh plots. It provides the data to the glyphs of your plot. When you pass sequences like Python lists or NumPy arrays to a Bokeh renderer, Bokeh automatically creates a ColumnDataSource with this data for you. However, creating a ColumnDataSource yourself gives you access to more advanced options. Let's see example.\n\n* creating ColumnDataSource:\n>\n>          data = {'x_values': [1, 2, 3, 4, 5],\n>                 'y_values': [6, 7, 2, 3, 6]}\n>          source = ColumnDataSource(data=data)\n\n* plotting using ColumnDataSource: \n>\n>          fig = figure()\n>          fig.circle(x='x_values', y='y_values', source=source)\n>\n\n**4. Pandas dataframe**: The data parameter (in the above method, ColumnsDataSource) can also be a pandas DataFrame or GroupBy object. A simple example is given below.\n\n> source = ColumnDataSource(df) --> pass in the dataFrame to the ColumnsDataSource\n        \n \n**Remark**: Please refer to the official webpage of the Bokeh library (see in the ref. section) to read more about providing data to Bokeh.\n","metadata":{}},{"cell_type":"markdown","source":"<a id=\"1\"></a>\n<font color=\"skyblue\" size=+2.5><b>1. Basic plots</b></font>\n\n<a id=\"1.1\"></a>\n<font color=\"skyblue\" size=+2.5><b>1.1 Scatter plots</b></font>\n\n<a href=\"#top\" class=\"btn btn-primary btn-sm\" role=\"button\" aria-pressed=\"true\" style=\"color:white\" data-toggle=\"popover\">go to top</a>\n","metadata":{}},{"cell_type":"code","source":"# prepare your data to be plotted\n# I used iris dataset for this example\n\nIris_setosa = df_iris[df_iris['Species'] =='Iris-setosa']\nIris_virginica = df_iris[df_iris['Species'] =='Iris-virginica']\nIris_versicolor = df_iris[df_iris['Species'] =='Iris-versicolor']\n\n\ncurdoc().theme = 'dark_minimal'\n\n# create and define the figure object\n\nfig = figure(title='Iris flowers: sepal length vs sepal width', \n             x_axis_label='Sepal Length [cm]', \n             y_axis_label='Sepal Width [cm]',          \n             plot_width=750, plot_height=500,\n             tools= 'hover',# set this to None if you do not want data to be displayed on hover\n             toolbar_location=\"above\", \n             toolbar_sticky=False)\n\n# adding glyphs (scatter plots of your choice, i.e, circle, triangle, square)\n\nfig.circle(x=\"SepalLengthCm\",y=\"SepalWidthCm\", \n         size=12, alpha=0.5, \n         color=\"#F78888\", \n         legend_label='Setosa', \n         source=Iris_setosa),\nfig.triangle(x=\"SepalLengthCm\",y=\"SepalWidthCm\", \n         size=12, alpha=0.99, \n         color=\"#F3D250\", \n         legend_label='Virginica', \n         source=Iris_virginica),\nfig.square(x=\"SepalLengthCm\",y=\"SepalWidthCm\", \n         size=12, alpha=0.6, \n         color=\"#3AAFA9\", \n         legend_label='Versicolor', \n         source=Iris_versicolor),\n\n# layout update\n\nfig.title.text_font_size = '20pt'\nfig.title.text_font_style = 'bold'\nfig.title.text_font = 'Serif'\nfig.xaxis.axis_label_text_font_size = \"16pt\"\nfig.yaxis.axis_label_text_font_size = \"16pt\"\nfig.legend.location = 'top_left'\nfig.legend.background_fill_color = \"skyblue\"\n\nshow(fig)","metadata":{"_kg_hide-input":true,"execution":{"iopub.status.busy":"2022-04-22T10:26:45.410807Z","iopub.execute_input":"2022-04-22T10:26:45.411102Z","iopub.status.idle":"2022-04-22T10:26:45.51435Z","shell.execute_reply.started":"2022-04-22T10:26:45.411079Z","shell.execute_reply":"2022-04-22T10:26:45.513016Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# data to plot: rain in australia dataset\nsource = ColumnDataSource(df_rain)\n\n# theme style\ncurdoc().theme = 'light_minimal'\n\n# figure object\n\nfig = figure(title='Rainfall vs Sunshine', \n             x_axis_label='Sunshine', \n             y_axis_label='Rainfall [mm]',          \n             plot_width=750, plot_height=500, \n             toolbar_location=\"above\",\n             tools=\"\",\n             toolbar_sticky=False)\n\n# the scatter plot\nfig.circle(x='Sunshine',y='Rainfall', \n           source = source,\n           size=5, alpha=0.5,\n           color='#F78888',\n           )\n         \n# figure layout update\nfig.title.text_font_size = '20pt'\nfig.title.text_font_style = 'bold'\nfig.title.text_font = 'Serif'\nfig.xaxis.axis_label_text_font_size = \"16pt\"\nfig.yaxis.axis_label_text_font_size = \"16pt\"\n\nshow(fig)","metadata":{"_kg_hide-input":true,"execution":{"iopub.status.busy":"2022-04-22T10:26:45.515599Z","iopub.execute_input":"2022-04-22T10:26:45.515897Z","iopub.status.idle":"2022-04-22T10:26:48.732298Z","shell.execute_reply.started":"2022-04-22T10:26:45.515849Z","shell.execute_reply":"2022-04-22T10:26:48.731064Z"},"jupyter":{"source_hidden":true},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<a id=\"1.2\"></a>\n<font color=\"skyblue\" size=+2.5><b>1.2 Line plots</b></font>\n\n<a href=\"#top\" class=\"btn btn-primary btn-sm\" role=\"button\" aria-pressed=\"true\" style=\"color:white\" data-toggle=\"popover\">go to top</a>\n","metadata":{}},{"cell_type":"code","source":"curdoc().theme = 'night_sky'\n\nts=df_sale.groupby([\"date_block_num\"])[\"item_cnt_day\"].sum()\nts2=df_sale.groupby([\"date_block_num\"])[\"item_price\"].sum()\n\nx=list(ts.index)\ny=list(ts)\ny2=list(ts2)\n\nfig = figure(title='Sales In A Shop', \n             x_axis_label='Date block', \n             y_axis_label='Amount sold',          \n             plot_width=750, plot_height=500, \n             toolbar_location=\"above\",\n             tools=\"hover\",\n             toolbar_sticky=False,\n             #background_fill_color=\"#f4f0ec\"\n            )\n\n\n#fig.line(x, y, line_color='red', line_width=3)\nfig.line(x, y2,line_color='#5da2d5',line_width=3,line_dash=\"dashed\")\n\n# fig.add_layout(BoxAnnotation(left=10, fill_alpha=0.1, fill_color='yellow', line_color='red'))\n# fig.add_layout(BoxAnnotation(right=25, fill_alpha=0.1, fill_color='yellow', line_color='red'))\n        \n# figure layout update\nfig.title.text_font_size = '20pt'\nfig.title.text_font_style = 'bold'\nfig.title.text_font = 'Serif'\nfig.xaxis.axis_label_text_font_size = \"16pt\"\nfig.yaxis.axis_label_text_font_size = \"16pt\"\nfig.y_range.range_padding = 0.2\n\n\n\nshow(fig)","metadata":{"_kg_hide-input":true,"execution":{"iopub.status.busy":"2022-04-22T10:26:48.733536Z","iopub.execute_input":"2022-04-22T10:26:48.7339Z","iopub.status.idle":"2022-04-22T10:26:48.920157Z","shell.execute_reply.started":"2022-04-22T10:26:48.733839Z","shell.execute_reply":"2022-04-22T10:26:48.919143Z"},"jupyter":{"source_hidden":true},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<a id=\"1.3\"></a>\n<font color=\"skyblue\" size=+2.5><b>1.3 Bar graphs</b></font>\n\n<font color=\"skyblue\" size=+1.5><b>1.3.1 Vertical bars</b></font>\n\n<a href=\"#top\" class=\"btn btn-primary btn-sm\" role=\"button\" aria-pressed=\"true\" style=\"color:white\" data-toggle=\"popover\">go to top</a>\n\n","metadata":{}},{"cell_type":"code","source":"# plot theme\ncurdoc().theme = 'caliber'\n\n# data to plot\nIris_setosa = df_iris[df_iris['Species'] =='Iris-setosa']\nIris_virginica = df_iris[df_iris['Species'] =='Iris-virginica']\nIris_versicolor = df_iris[df_iris['Species'] =='Iris-versicolor']\n\nse = Iris_setosa['PetalLengthCm'].mean()\nvi = Iris_virginica['PetalLengthCm'].mean()\nve = Iris_versicolor['PetalLengthCm'].mean()\n\n\nspecies = ['setosa', 'virginica', 'versicolor']\navg_petal_length = [se, vi, ve]\nspecies_sorted = sorted(species, key=lambda x: avg_petal_length[species.index(x)])\n\ncolors = ['#F78888', '#5DA2D5', '#F3D250', '#3AAFA9']\n\n# figure object\nfig = figure(x_range=species,\n             title=\"Average Petal Length\",\n             x_axis_label='Species', \n             y_axis_label='Average petal legth [cm]',\n             plot_width=750,plot_height=400,\n             toolbar_location='above',\n             tools=\"hover\",\n             background_fill_color=\"#f4f0ec\"\n            )\n\n# the bar plot\nfig.vbar(x=species_sorted, top=avg_petal_length, width=0.75, \n         fill_color=colors[:4],\n        )\n\n\n# layout update\nfig.title.text_font_size = '20pt'\nfig.title.text_font_style = 'bold'\nfig.title.text_font = 'Serif'\nfig.xaxis.axis_label_text_font_size = \"16pt\"\nfig.yaxis.axis_label_text_font_size = \"16pt\"\n\n#optional customization\nfig.title.text_color = \"black\"\nfig.title.background_fill_color = \"#f4f0ec\"\n\n# show the figure\nshow(fig)","metadata":{"_kg_hide-input":true,"execution":{"iopub.status.busy":"2022-04-22T10:26:48.921335Z","iopub.execute_input":"2022-04-22T10:26:48.921617Z","iopub.status.idle":"2022-04-22T10:26:49.009512Z","shell.execute_reply.started":"2022-04-22T10:26:48.921587Z","shell.execute_reply":"2022-04-22T10:26:49.008293Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<a id=\"1.3. 2\"></a>\n\n<font color=\"skyblue\" size=+1.5><b>1.3.2 Horizontal bars</b></font>\n\n<a href=\"#top\" class=\"btn btn-primary btn-sm\" role=\"button\" aria-pressed=\"true\" style=\"color:white\" data-toggle=\"popover\">go to top</a>\n","metadata":{}},{"cell_type":"code","source":"# plot theme\ncurdoc().theme = 'caliber'\n\n\n# data to plot\nIris_setosa = df_iris[df_iris['Species'] =='Iris-setosa']\nIris_virginica = df_iris[df_iris['Species'] =='Iris-virginica']\nIris_versicolor = df_iris[df_iris['Species'] =='Iris-versicolor']\n\nse = Iris_setosa['SepalWidthCm'].mean()\nvi = Iris_virginica['SepalWidthCm'].mean()\nve = Iris_versicolor['SepalWidthCm'].mean()\n\nspecies = ['setosa', 'virginica', 'versicolor']\navg_petal_length = [se, vi, ve]\n\nfig = figure(y_range=species,\n             title=\"Average Sepal Width\",\n             y_axis_label='Species', \n             x_axis_label='Average petal length [cm]',\n             plot_width=750,plot_height=400,\n             toolbar_location='above',\n             background_fill_color=\"#f4f0ec\",\n             )\nfig.hbar(y=species,\n         right=avg_petal_length,\n         height=0.8, \n         left=0,\n         color=colors,\n        )\n\n# layout update\nfig.title.text_font_size = '20pt'\nfig.title.text_font_style = 'bold'\nfig.title.text_font = 'Serif'\nfig.xaxis.axis_label_text_font_size = \"16pt\"\nfig.yaxis.axis_label_text_font_size = \"16pt\"\n\n#optional customization\nfig.title.text_color = \"black\"\nfig.title.background_fill_color = \"#f4f0ec\"\n\nshow(fig)\n","metadata":{"_kg_hide-input":true,"execution":{"iopub.status.busy":"2022-04-22T10:26:49.013214Z","iopub.execute_input":"2022-04-22T10:26:49.013659Z","iopub.status.idle":"2022-04-22T10:26:49.293845Z","shell.execute_reply.started":"2022-04-22T10:26:49.013622Z","shell.execute_reply":"2022-04-22T10:26:49.292966Z"},"jupyter":{"source_hidden":true},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<a id=\"1.3.3\"></a>\n\n<font color=\"skyblue\" size=+1.5><b>1.3.3 Stacked bars</b></font>\n\n\n<a href=\"#top\" class=\"btn btn-primary btn-sm\" role=\"button\" aria-pressed=\"true\" style=\"color:white\" data-toggle=\"popover\">go to top</a>\n\n","metadata":{}},{"cell_type":"code","source":"# data to plot\nam =df_study[(df_study['race/ethnicity'] == 'group A') &  (df_study['gender'] == 'male')]\naf =df_study[(df_study['race/ethnicity'] == 'group A') &  (df_study['gender'] == 'female')]\nbm =df_study[(df_study['race/ethnicity'] == 'group B') &  (df_study['gender'] == 'male')]\nbf =df_study[(df_study['race/ethnicity'] == 'group B') &  (df_study['gender'] == 'female')]\ncm =df_study[(df_study['race/ethnicity'] == 'group C') &  (df_study['gender'] == 'male')]\ncf =df_study[(df_study['race/ethnicity'] == 'group C') &  (df_study['gender'] == 'female')]\ndm =df_study[(df_study['race/ethnicity'] == 'group D') &  (df_study['gender'] == 'male')]\ndf = df_study[(df_study['race/ethnicity'] == 'group D') &  (df_study['gender'] == 'female')]\nem =df_study[(df_study['race/ethnicity'] == 'group E') &  (df_study['gender'] == 'male')]\nef = df_study[(df_study['race/ethnicity'] == 'group E') &  (df_study['gender'] == 'female')]\n\n\n# define function\n\ndef stacked_bars_horizontal():\n    curdoc().theme = 'caliber'\n\n    groups = ['group A', 'group B','group C','group D','group E']\n    gender = ['male', 'female']\n    #colors = color#['skyblue', 'pink']\n    colors = ['#F78878', '#5DA2D5', '#F3D250', '#3AAFA9']\n\n    data = {'groups' : groups,\n            'male' : [len(am), len(bm), len(cm), len(dm), len(em)],\n\n            'female' :[len(af), len(bf), len(cf), len(df), len(ef)],\n                       }\n\n    fig = figure(y_range=groups, plot_height=550, \n               title=\"Gender distribution within race groups\",\n               x_axis_label='Students count', \n               y_axis_label='race group',\n               toolbar_location='above',\n               tools= 'wheel_zoom,box_zoom, reset, save, hover',\n               toolbar_sticky=False,\n               background_fill_color=\"#f4f0ec\", \n               border_fill_color=\"#f4f0ec\"\n              )\n\n    fig.hbar_stack(gender, y='groups',  height=0.75, color= colors[0:2], source=ColumnDataSource(data), legend_label=[\"%s\" % x for x in gender])\n\n    hover = fig.select(dict(type=HoverTool))\n    hover.tooltips = [(\"groups\", \"@groups\"), \n                      ('male', \"@male\"), \n                      (\"female\", \"@female\")]\n\n    fig.y_range.range_padding = 0.1\n    fig.ygrid.grid_line_color = None\n    fig.legend.location = \"top_right\"\n    fig.axis.minor_tick_line_color = None\n    fig.outline_line_color = None\n    fig.grid.grid_line_color = None\n    fig.axis.axis_line_color = '#f4f0ec'\n\n\n    fig.title.text_font_size = '20pt'\n    fig.title.text_font_style = 'bold'\n    fig.title.text_font = 'Serif'\n    fig.xaxis.axis_label_text_font_size = \"16pt\"\n    fig.yaxis.axis_label_text_font_size = \"16pt\"\n\n    #show(fig)\n    return (fig)\n\nshow(stacked_bars_horizontal())","metadata":{"_kg_hide-input":true,"execution":{"iopub.status.busy":"2022-04-22T10:26:49.295897Z","iopub.execute_input":"2022-04-22T10:26:49.296194Z","iopub.status.idle":"2022-04-22T10:26:49.43808Z","shell.execute_reply.started":"2022-04-22T10:26:49.296165Z","shell.execute_reply":"2022-04-22T10:26:49.435964Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"def stacked_bars_vertical():\n    curdoc().theme = 'caliber'\n\n    groups = ['group A', 'group B','group C','group D','group E']\n    gender = ['male', 'female']\n    #colors = color#['skyblue', 'pink']\n    colors = ['#F78878', '#5DA2D5', '#F3D250', '#3AAFA9']\n\n    data = {'groups' : groups,\n            'male' : [len(am), len(bm), len(cm), len(dm), len(em)],\n\n            'female' :[len(af), len(bf), len(cf), len(df), len(ef)],\n                       }\n\n    fig = figure(x_range=groups, plot_height=550, \n               title=\"Gender distribution within race groups\",\n               y_axis_label='Students count', \n               x_axis_label='race group',\n               toolbar_location='above',\n               tools= 'wheel_zoom,box_zoom, reset, save, hover',\n               toolbar_sticky=False,\n               background_fill_color=\"#f4f0ec\", \n               border_fill_color=\"#f4f0ec\"\n              )\n\n    fig.vbar_stack(gender, x='groups', width=0.8, color= colors[2:], source=ColumnDataSource(data), legend_label=[\"%s\" % x for x in gender])\n\n    hover = fig.select(dict(type=HoverTool))\n    hover.tooltips = [(\"groups\", \"@groups\"), \n                      ('male', \"@male\"), \n                      (\"female\", \"@female\")]\n\n    fig.y_range.range_padding = 0.1\n    fig.ygrid.grid_line_color = None\n    fig.legend.location = \"top_right\"\n    fig.axis.minor_tick_line_color = None\n    fig.outline_line_color = None\n    fig.grid.grid_line_color = None\n    fig.axis.axis_line_color = '#f4f0ec'\n\n\n    fig.title.text_font_size = '20pt'\n    fig.title.text_font_style = 'bold'\n    fig.title.text_font = 'Serif'\n    fig.xaxis.axis_label_text_font_size = \"16pt\"\n    fig.yaxis.axis_label_text_font_size = \"16pt\"\n\n    #show(fig)\n    return (fig)\n\nshow(stacked_bars_vertical())","metadata":{"_kg_hide-input":true,"execution":{"iopub.status.busy":"2022-04-22T10:26:49.439832Z","iopub.execute_input":"2022-04-22T10:26:49.440155Z","iopub.status.idle":"2022-04-22T10:26:49.565898Z","shell.execute_reply.started":"2022-04-22T10:26:49.440125Z","shell.execute_reply":"2022-04-22T10:26:49.564672Z"},"jupyter":{"source_hidden":true},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<a id=\"1.4\"></a>\n<font color=\"skyblue\" size=+2.5><b>1.4 Histograms </b></font>\n\n<a href=\"#top\" class=\"btn btn-primary btn-sm\" role=\"button\" aria-pressed=\"true\" style=\"color:white\" data-toggle=\"popover\">go to top</a>\n","metadata":{}},{"cell_type":"code","source":"def make_histogram(title, hist1, hist2, edges):\n    fig = figure(title=title, \n               #tools='',\n               plot_width=750, \n               plot_height=450, \n               background_fill_color=\"#f4f0ec\",\n               )\n    fig.quad(top=hist1, # histograms for survivors\n             bottom=0, \n             left=edges[:-1], \n             right=edges[1:],\n             fill_color='#F3D250', \n             line_color=\"white\", \n             #alpha=0.5, \n             legend_label=\"1\")\n    fig.quad(top=hist2, # histograms for victims\n             bottom=0, \n             left=edges[:-1], \n             right=edges[1:],\n             fill_color='#3AAFA9', \n             line_color=\"white\", \n             alpha=0.5, \n             legend_label=\"0\")\n    \n    \n    fig.y_range.start = 0\n    fig.legend.location = \"top_right\"\n    fig.legend.background_fill_color = \"#f2f0ec\"\n    fig.legend.title = 'Death Event'\n    fig.xaxis.axis_label = 'Age'\n    fig.yaxis.axis_label = 'Probability (x)'\n    fig.grid.grid_line_color=\"white\"\n    \n    fig.title.text_font_size = '16pt'\n    fig.title.text_font_style = 'bold'\n    fig.title.text_font = 'Serif'\n    fig.xaxis.axis_label_text_font_size = \"12pt\"\n    fig.yaxis.axis_label_text_font_size = \"12pt\"\n\n    return fig\n\n# data to plot: age distribution of patients (heart-failure dataset)\nsurv = df_heart[df_heart['DEATH_EVENT'] == 0]['age']\nvict = df_heart[df_heart['DEATH_EVENT'] == 1]['age']\n\nhist1, edges = np.histogram(surv, density=True, bins=10)\nhist2, edges = np.histogram(vict, density=True, bins=15)\n\nfig1 = make_histogram(\"Heart Failure: Patients Age Distribution\", hist1, hist2, edges)\n\nshow(fig1)\n","metadata":{"_kg_hide-input":true,"execution":{"iopub.status.busy":"2022-04-22T10:26:49.567303Z","iopub.execute_input":"2022-04-22T10:26:49.567612Z","iopub.status.idle":"2022-04-22T10:26:49.711507Z","shell.execute_reply.started":"2022-04-22T10:26:49.567537Z","shell.execute_reply":"2022-04-22T10:26:49.71016Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<a id=\"1.5\"></a>\n<font color=\"skyblue\" size=+2.5><b>1.5 Pie Charts</b></font>\n\n<a href=\"#top\" class=\"btn btn-primary btn-sm\" role=\"button\" aria-pressed=\"true\" style=\"color:white\" data-toggle=\"popover\">go to top</a>\n","metadata":{}},{"cell_type":"code","source":"def pie_chart():\n    curdoc().theme = 'caliber'    \n    # data: titanic Pclass       \n    x = {\n        'Pclass 1': len(df_titanic[df_titanic['Pclass'] ==1]),\n        'Pclass 2': len(df_titanic[df_titanic['Pclass'] ==2]),\n        'Pclass 3': len(df_titanic[df_titanic['Pclass'] ==3]),\n        }\n\n    data = pd.Series(x).reset_index(name='value').rename(columns={'index':'pclass'})\n    data['angle'] = data['value']/data['value'].sum()*2*np.pi\n    data['color'] = ['#F78888', '#5DA2D5', '#F3D250', '#3AAFA9'][0:len(x)]\n    data[\"value\"] = data['value'].astype(str)\n    data[\"value\"] = data[\"value\"].str.pad(35, side = \"left\")\n    source = ColumnDataSource(data)\n\n    fig = figure(plot_height=550, \n               plot_width=550,\n               title=\"Titanic: Passenger Class\", \n               toolbar_location=None,\n               tools=\"hover\", \n               tooltips=\"@pclass: @value\",\n               x_range=(-0.5, .75),                 \n               background_fill_color=\"#f4f0ec\", \n               border_fill_color=\"#f4f0ec\"\n              )\n\n    fig.wedge(x=0, y=1, radius=0.4,\n            start_angle=cumsum('angle', include_zero=True), end_angle=cumsum('angle'),\n            line_color=\"black\", fill_color='color', legend_field='pclass', source=data,)\n\n\n    labels = LabelSet(x=0, y=1, text='value',\n            angle=cumsum('angle', include_zero=True), source=source, render_mode='canvas')\n    \n    annotation = Label(x=0, y=0, x_units='screen', y_units='screen',\n                       text='More passengers in class 3! This class has lower survival rate.', \n                       text_font_size = '12pt',\n                       render_mode='css',\n                       border_line_color=None, border_line_alpha=1.0,\n                       background_fill_color='#f4f0ec', background_fill_alpha=1.0)\n\n    fig.add_layout(annotation)\n    fig.add_layout(labels)\n\n    fig.axis.axis_label=None\n    fig.axis.visible=False\n    fig.grid.grid_line_color = None\n    fig.title.text_font_size = '20pt'\n    fig.title.text_font_style = 'bold'\n\n    return (fig)\n\nshow(pie_chart())","metadata":{"_kg_hide-input":true,"execution":{"iopub.status.busy":"2022-04-22T10:26:49.714002Z","iopub.execute_input":"2022-04-22T10:26:49.714507Z","iopub.status.idle":"2022-04-22T10:26:49.853175Z","shell.execute_reply.started":"2022-04-22T10:26:49.714466Z","shell.execute_reply":"2022-04-22T10:26:49.851727Z"},"jupyter":{"source_hidden":true},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"def donut_chart():\n    curdoc().theme = 'contrast'\n    # data: titanic Pclass    \n    x = {\n        'Pclass 1': len(df_titanic[df_titanic['Pclass'] ==1]),\n        'Pclass 2': len(df_titanic[df_titanic['Pclass'] ==2]),\n        'Pclass 3': len(df_titanic[df_titanic['Pclass'] ==3]),\n        }\n\n    data = pd.Series(x).reset_index(name='value').rename(columns={'index':'pclass'})\n    data['angle'] = data['value']/data['value'].sum()*2*np.pi\n    data['color'] = ['#F78888', '#5DA2D5', '#F3D250', '#3AAFA9'][0:len(x)]\n    data[\"value\"] = data['value'].astype(str)\n    data[\"value\"] = data[\"value\"].str.pad(30, side = \"left\")\n    source = ColumnDataSource(data)\n\n    fig = figure(plot_height=550, \n               plot_width=550,\n               title=\"Titanic: Passenger Class\", \n               toolbar_location=None,\n               tools=\"hover\", \n               tooltips=\"@pclass: @value\",\n               x_range=(-0.5, .75)\n              )\n\n    fig.annular_wedge(x=0, y=1, outer_radius=0.4,inner_radius=0.2,\n            start_angle=cumsum('angle', include_zero=True), \n            end_angle=cumsum('angle'),\n            line_color=\"black\", fill_color='color', legend_field='pclass', source=data)\n\n\n    labels = LabelSet(x=0, y=1, text='value',\n            angle=cumsum('angle', include_zero=True), source=source, render_mode='canvas')\n    \n    annotation = Label(x=200, y=250, x_units='screen', y_units='screen',\n                       text='Pclass', \n                       text_font_size = '16pt',\n                       text_color='white',\n                       render_mode='css',\n                       border_line_color=None, border_line_alpha=1.0,\n                       background_fill_color=None, background_fill_alpha=1.0)\n\n    fig.add_layout(annotation)\n    fig.add_layout(labels)\n\n    fig.axis.axis_label=None\n    fig.axis.visible=False\n    fig.grid.grid_line_color = None\n    fig.title.text_font_size = '20pt'\n    fig.title.text_font_style = 'bold'\n\n    return (fig)\n\nshow(donut_chart())","metadata":{"_kg_hide-input":true,"execution":{"iopub.status.busy":"2022-04-22T10:26:49.854532Z","iopub.execute_input":"2022-04-22T10:26:49.854862Z","iopub.status.idle":"2022-04-22T10:26:49.977089Z","shell.execute_reply.started":"2022-04-22T10:26:49.854818Z","shell.execute_reply":"2022-04-22T10:26:49.975757Z"},"jupyter":{"source_hidden":true},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<a id=\"1.6\"></a>\n<font color=\"skyblue\" size=+2.5><b>1.6 Box plots</b></font>\n\n<a href=\"#top\" class=\"btn btn-primary btn-sm\" role=\"button\" aria-pressed=\"true\" style=\"color:white\" data-toggle=\"popover\">go to top</a>\n\nFunction to call:\n> boxwhisker = hv.**BoxWhisker**(data, dimensions, label)\n>\n> boxwhisker.**opts**(optional parameters)","metadata":{}},{"cell_type":"code","source":"# data: df_study\n\ntitle = \"Math test score: Effect of race and gender\"\ncolors = ['#F78888', '#5DA2D5', '#F3D250', '#3AAFA9', '#BDBDBD']\n\n\nboxwhisker = hv.BoxWhisker(df_study, \n                           ['gender', 'race/ethnicity'], \n                           'math score', \n                           label=title)\nboxwhisker.opts(show_legend=False,                \n                width=600, \n                height= 400,\n                box_fill_color=dim('race/ethnicity').str(), \n                cmap=colors,\n                show_grid=True,\n                )","metadata":{"_kg_hide-input":true,"execution":{"iopub.status.busy":"2022-04-22T10:26:49.978416Z","iopub.execute_input":"2022-04-22T10:26:49.978741Z","iopub.status.idle":"2022-04-22T10:26:50.285368Z","shell.execute_reply.started":"2022-04-22T10:26:49.978709Z","shell.execute_reply":"2022-04-22T10:26:50.28396Z"},"jupyter":{"source_hidden":true},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<a id=\"1.7\"></a>\n<font color=\"skyblue\" size=+2.5><b>1.7 Violin plots</b></font>\n\n<a href=\"#top\" class=\"btn btn-primary btn-sm\" role=\"button\" aria-pressed=\"true\" style=\"color:white\" data-toggle=\"popover\">go to top</a>\n\n\nFunction to call:\n> violin = hv.**Violin**(data, dimensions, label)\n>\n> violin.**opts**(optional parameters)","metadata":{}},{"cell_type":"code","source":"# data iris\n\ntitle=\"Iris: Variation of Petal length\"\ncolors = ['#F78888', '#5DA2D5', '#F3D250', '#3AAFA9', '#BDBDBD']\n\n# define hv element\nviolin1 = hv.Violin(df_iris, \n                   'Species', \n                   'PetalLengthCm', \n                    label=title)\n\n# optional params \nviolin1.opts(height=500, \n            width=750,\n            violin_fill_color=dim('Species').str(), \n            cmap=colors)\n\nviolin1 = hv.Violin(df_iris, \n                   'Species', \n                   'PetalLengthCm', \n                    label=title)\n\n# optional params \nviolin1.opts(height=500, \n            width=750,\n            violin_fill_color=dim('Species').str(), \n            cmap=colors)\n","metadata":{"_kg_hide-input":true,"execution":{"iopub.status.busy":"2022-04-22T10:26:50.287043Z","iopub.execute_input":"2022-04-22T10:26:50.287366Z","iopub.status.idle":"2022-04-22T10:26:50.536965Z","shell.execute_reply.started":"2022-04-22T10:26:50.287334Z","shell.execute_reply":"2022-04-22T10:26:50.535738Z"},"jupyter":{"source_hidden":true},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<a id=\"2\"></a>\n<font color=\"skyblue\" size=+2.5><b>2. Area Plots</b></font>\n\n<a href=\"#top\" class=\"btn btn-primary btn-sm\" role=\"button\" aria-pressed=\"true\" style=\"color:white\" data-toggle=\"popover\">go to top</a>\n","metadata":{}},{"cell_type":"code","source":"curdoc().theme = 'dark_minimal'\n\n# data to plot: covid vaccination USA US,and the Netherlands\n\nusa = df_covid[df_covid['country'] == 'United States']\nuk = df_covid[df_covid['country'] == 'United Kingdom']\nnl = df_covid[df_covid['country'] == 'Netherlands']\n\n\nsource = ColumnDataSource(data=dict(\n    x = pd.to_datetime(usa['date']),\n    y1=usa['daily_vaccinations_per_million'],\n    y2=uk['daily_vaccinations_per_million'],\n    y3=nl['daily_vaccinations_per_million'],\n    ))\n                          \nfig = figure(title='Daily Vaccinations Per-Million: USA, UK and Netherlands',\n             x_axis_label='Date', \n             y_axis_label='Daily vaccination', \n             plot_width=750, \n             plot_height=500, \n             x_axis_type=\"datetime\")\n\nfig.varea_stack(stackers=['y1', 'y2', 'y3'], x='x',\n              color=(['#F78888', '#5DA2D5', '#F3D250', '#3AAFA9'][0:3]), \n              alpha=0.99,\n              legend_label=['USA', 'UK', 'Netherlands'],\n              source=source)\n\nfig.title.text_font_size = '16pt'\nfig.title.text_font_style = 'bold'\nfig.title.text_font = 'Serif'\nfig.xaxis.axis_label_text_font_size = \"12pt\"\nfig.yaxis.axis_label_text_font_size = \"12pt\"\nfig.legend.location = 'top_left'\nfig.legend.background_fill_color = None\nshow(fig)\n","metadata":{"_kg_hide-input":true,"execution":{"iopub.status.busy":"2022-04-22T10:26:50.538264Z","iopub.execute_input":"2022-04-22T10:26:50.538554Z","iopub.status.idle":"2022-04-22T10:26:50.783167Z","shell.execute_reply.started":"2022-04-22T10:26:50.538525Z","shell.execute_reply":"2022-04-22T10:26:50.781751Z"},"jupyter":{"source_hidden":true},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"curdoc().theme = 'dark_minimal'\n\n# data to plot: covid vaccination USA US,and the Netherlands\n\nusa = df_covid[df_covid['country'] == 'United States']\nuk = df_covid[df_covid['country'] == 'United Kingdom']\nnl = df_covid[df_covid['country'] == 'Netherlands']\n\n\nsource = ColumnDataSource(data=dict(\n    y = pd.to_datetime(usa['date']),\n    x1=usa['daily_vaccinations_per_million'],\n    x2=uk['daily_vaccinations_per_million'],\n    x3=nl['daily_vaccinations_per_million'],\n    ))\n                          \nfig = figure(title='Daily Vaccinations Per-Million: USA, UK and Netherlands',\n             y_axis_label='Date', \n             x_axis_label='Daily vaccination', \n             plot_width=750, \n             plot_height=600, \n             y_axis_type=\"datetime\")\n\nfig.harea_stack(stackers=['x1', 'x2', 'x3'], y='y',\n              color=(['#F78888', '#5DA2D5', '#F3D250', '#3AAFA9'][1:]), \n              alpha=0.99,\n              legend_label=['USA', 'UK', 'Netherlands'],\n              source=source)\n\nfig.title.text_font_size = '16pt'\nfig.title.text_font_style = 'bold'\nfig.title.text_font = 'Serif'\nfig.xaxis.axis_label_text_font_size = \"12pt\"\nfig.yaxis.axis_label_text_font_size = \"12pt\"\nfig.legend.location = 'bottom_right'\nfig.legend.background_fill_color = None\nshow(fig)\n","metadata":{"_kg_hide-input":true,"execution":{"iopub.status.busy":"2022-04-22T10:26:50.784676Z","iopub.execute_input":"2022-04-22T10:26:50.784957Z","iopub.status.idle":"2022-04-22T10:26:51.008458Z","shell.execute_reply.started":"2022-04-22T10:26:50.784927Z","shell.execute_reply":"2022-04-22T10:26:51.006806Z"},"jupyter":{"source_hidden":true},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<a id=\"3\"></a>\n<font color=\"skyblue\" size=+2.5><b>3. Heatmap</b></font>\n\n<a href=\"#top\" class=\"btn btn-primary btn-sm\" role=\"button\" aria-pressed=\"true\" style=\"color:white\" data-toggle=\"popover\">go to top</a>\n\n\nFunction to call: \n\n> heatmap = hv.Heatmap(data, label)\n>\n> heatmap.opts(optional parameters such as plot properties)\n","metadata":{}},{"cell_type":"code","source":"# make the dataframe \nfeat= ['date_block_num', 'shop_id', 'item_cnt_day']\ndf = df_sale[feat]\n\n# color palette\ncolors = ['#F78888', '#5DA2D5', '#F3D250', '#3AAFA9', '#BDBDBD']\n\n# hv element == HeatMap\nheatmap = hv.HeatMap(df, label=\"Sales [items count per day] in a company\").aggregate(function=np.sum)\n\n# optional params to change the default setting\nheatmap.opts(width=1000, \n             height=600,\n             xrotation=0,\n             xaxis='bottom', \n             xlabel='Date block', \n             ylabel='Shop Id',  \n             tools=['hover'], \n             cmap=colors,\n            )\n","metadata":{"_kg_hide-input":true,"execution":{"iopub.status.busy":"2022-04-22T10:26:51.010213Z","iopub.execute_input":"2022-04-22T10:26:51.010579Z","iopub.status.idle":"2022-04-22T10:26:52.03386Z","shell.execute_reply.started":"2022-04-22T10:26:51.010548Z","shell.execute_reply":"2022-04-22T10:26:52.032834Z"},"jupyter":{"source_hidden":true},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<a id=\"4\"></a>\n<font color=\"skyblue\" size=+2.5><b>4. Jitter Scatter (for categorical variable)</b></font>\n\n<a href=\"#top\" class=\"btn btn-primary btn-sm\" role=\"button\" aria-pressed=\"true\" style=\"color:white\" data-toggle=\"popover\">go to top</a>\n\nWhile plotting a scatter plots of categorical variable use the `jitter()` function to avoid overlap between numerous scatter points in a single category. Jitter function gives each point a random offset.","metadata":{}},{"cell_type":"markdown","source":"### Scatter plot without jitter effect","metadata":{}},{"cell_type":"code","source":"curdoc().theme = 'caliber'\n\n# data titanic\ndata = df_titanic\nPclass = ['pclass_1', 'pclass_2', 'pclass_3']\nsource = ColumnDataSource(data)\n\nfig = figure(plot_width=800, \n           plot_height=400, \n           y_range=Pclass, \n           title=\"Titanic Passenger Age per Pclass\",\n           x_axis_label='Passenger Age',\n           y_axis_label='Passenger Class',\n           background_fill_color=\"#f4f0ec\")\n\nfig.circle(x='Age', \n           y = 'Pclass',\n           #y=jitter('Pclass', width=0.5,  range=fig.y_range), \n         source=source, \n         color='salmon',\n         alpha=0.5)\n\nfig.x_range.range_padding = 0\nfig.y_range.range_padding = 0.2\nfig.ygrid.grid_line_color = None\n\nfig.title.text_font_size = '20pt'\nfig.title.text_font_style = 'bold'\nfig.title.text_font = 'Serif'\nfig.xaxis.axis_label_text_font_size = \"16pt\"\nfig.yaxis.axis_label_text_font_size = \"16pt\"\n\nshow(fig)","metadata":{"_kg_hide-input":true,"execution":{"iopub.status.busy":"2022-04-22T10:26:52.03528Z","iopub.execute_input":"2022-04-22T10:26:52.035582Z","iopub.status.idle":"2022-04-22T10:26:52.241817Z","shell.execute_reply.started":"2022-04-22T10:26:52.035554Z","shell.execute_reply":"2022-04-22T10:26:52.240821Z"},"jupyter":{"source_hidden":true},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### Scatter plot with jitters","metadata":{}},{"cell_type":"code","source":"curdoc().theme = 'caliber'\n# data titanic\ndata = df_titanic\nPclass = ['pclass_1', 'pclass_2', 'pclass_3']\nsource = ColumnDataSource(data)\n\nfig = figure(plot_width=800, \n           plot_height=400, \n           y_range=Pclass, \n           title=\"Titanic Passenger Age per Pclass\",\n           x_axis_label='Passenger Age',\n           y_axis_label='Passenger Class',\n           background_fill_color=\"#f4f0ec\")\n\nfig.circle(x='Age', \n           y=jitter('Pclass', width=0.5,  range=fig.y_range), \n         source=source, \n         color='salmon',\n         alpha=0.5)\n\nfig.x_range.range_padding = 0\nfig.y_range.range_padding = 0.2\nfig.ygrid.grid_line_color = None\n\nfig.title.text_font_size = '20pt'\nfig.title.text_font_style = 'bold'\nfig.title.text_font = 'Serif'\nfig.xaxis.axis_label_text_font_size = \"16pt\"\nfig.yaxis.axis_label_text_font_size = \"16pt\"\n\nshow(fig)","metadata":{"_kg_hide-input":true,"execution":{"iopub.status.busy":"2022-04-22T10:26:52.243222Z","iopub.execute_input":"2022-04-22T10:26:52.243496Z","iopub.status.idle":"2022-04-22T10:26:52.409857Z","shell.execute_reply.started":"2022-04-22T10:26:52.243467Z","shell.execute_reply":"2022-04-22T10:26:52.408862Z"},"jupyter":{"source_hidden":true},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<a id=\"5\"></a>\n<font color=\"skyblue\" size=+2.5><b>5. Kde plots</b></font>\n\n<a href=\"#top\" class=\"btn btn-primary btn-sm\" role=\"button\" aria-pressed=\"true\" style=\"color:white\" data-toggle=\"popover\">go to top</a>\n\nFunction to call: \n\n> dist = hv.Distribution(data, label)\n>\n> dist.opts(optional parameters such as plot properties)\n","metadata":{}},{"cell_type":"code","source":"# Data to plot\nIris_setosa = df_iris[df_iris['Species'] =='Iris-setosa']\nIris_virginica = df_iris[df_iris['Species'] =='Iris-virginica']\nIris_versicolor = df_iris[df_iris['Species'] =='Iris-versicolor']\n\nwidth = 400\nheight = 300\n\nkde1 = (hv.Distribution(Iris_setosa.SepalWidthCm, label='Setosa')* \n        hv.Distribution(Iris_virginica.SepalWidthCm, label='Virginica')*\n        hv.Distribution(Iris_versicolor.SepalWidthCm, label='Versicolor'))                \nkde1.opts(width=width,\n         height=height,\n         title='Iris Flowers: Sepal Width Density Plot',\n         xrotation=0,\n         xaxis='bottom',\n         xlabel='Sepal Width [cm]', \n         ylabel='Density',  \n         #tools=['hover'],\n         background_fill_color=\"#f4f0ec\",\n         show_legend=False,                     \n\n         )\n\nkde2 = (hv.Distribution(Iris_setosa.PetalWidthCm, label='Setosa')* \n        hv.Distribution(Iris_virginica.PetalWidthCm, label='Virginica')*\n        hv.Distribution(Iris_versicolor.PetalWidthCm, label='Versicolor'))                \n\nkde2.opts(width=width,\n         height=height,\n         title='Iris Flowers: Petal Width Density Plot',\n         xrotation=0,\n         xaxis='bottom',\n         xlabel='Sepal Width [cm]', \n         ylabel='Density',  \n         #tools=['hover'],\n         background_fill_color=\"#f4f0ec\"\n         )\nkde3 = (hv.Distribution(Iris_setosa.SepalLengthCm, label='Setosa')* \n        hv.Distribution(Iris_virginica.SepalLengthCm, label='Virginica')*\n        hv.Distribution(Iris_versicolor.SepalLengthCm, label='Versicolor'))                \nkde3.opts(width=width,\n         height=height,\n         title='Iris Flowers: Sepal Length Density Plot',\n         xrotation=0,\n         xaxis='bottom',\n         xlabel='Sepal Width [cm]', \n         ylabel='Density',  \n         #tools=['hover'],\n         #background_fill_color=\"#f4f0ec\",\n         bgcolor ='#f4f0ec'\n         )\n\nkde4 = (hv.Distribution(Iris_setosa.PetalLengthCm, label='Setosa')* \n        hv.Distribution(Iris_virginica.PetalLengthCm, label='Virginica')*\n        hv.Distribution(Iris_versicolor.PetalLengthCm, label='Versicolor'))                \n\nkde4.opts(width=width,\n         height=height,\n         title='Iris Flowers: Petal Length Density Plot',\n         xrotation=0,\n         xaxis='bottom',\n         xlabel='Sepal Width [cm]', \n         ylabel='Density',  \n         tools=['hover'],\n         #bgcolor='#eee8e2',\n         show_legend=True,\n         show_grid=True\n         )\n\n(kde1 + kde2 + kde3 + kde4)","metadata":{"_kg_hide-input":true,"execution":{"iopub.status.busy":"2022-04-22T11:24:00.488843Z","iopub.execute_input":"2022-04-22T11:24:00.489208Z","iopub.status.idle":"2022-04-22T11:24:01.520611Z","shell.execute_reply.started":"2022-04-22T11:24:00.489179Z","shell.execute_reply":"2022-04-22T11:24:01.519074Z"},"jupyter":{"source_hidden":true},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<a id=\"6\"></a>\n<font color=\"skyblue\" size=+2.5><b>6. Subplots</b></font>\n\n<a href=\"#top\" class=\"btn btn-primary btn-sm\" role=\"button\" aria-pressed=\"true\" style=\"color:white\" data-toggle=\"popover\">go to top</a>\n\n> Making subplots in holoviews is a simple two step process; first you create the elements you want to put in your subplot, say `plot1`, `plot2`, . . . `plotn` then call `.cols(k)` methode to make a subplot of `k` columns. Below I showed an example with 4 kde's , a violin and a box plot arranged in two columns using the iris dataset.\n","metadata":{}},{"cell_type":"code","source":"##### add new boxplot\ncolors = ['#F78888', '#5DA2D5', '#F3D250', '#3AAFA9', '#BDBDBD']\nboxwhisker = hv.BoxWhisker(df_iris, \n                           ['Species'], \n                           'SepalWidthCm', \n                           label='Petal width variation')\nboxwhisker.opts(show_legend=False,                \n                width=600, \n                height= 400,\n                box_fill_color=dim('Species').str(), \n                cmap=colors,\n                show_grid=True,\n                )\n\n# make a subplots of the four kde plots + the violin plot + the new boxplot \n# note that I used already existing plots from section above\n# the four plots from section 5, the violin plot from section 1.7 and the box plot created in code this cell (top)\n\nsubplots = (kde1 + kde2 + kde3 + kde4 + \\\n            violin1.opts(width=width, height=height) + \\\n            boxwhisker.opts(width=width, height=height)).opts(width=600, height=500).cols(2) \n\nsubplots","metadata":{"_kg_hide-input":true,"execution":{"iopub.status.busy":"2022-04-22T10:26:53.869284Z","iopub.execute_input":"2022-04-22T10:26:53.869648Z","iopub.status.idle":"2022-04-22T10:26:55.833253Z","shell.execute_reply.started":"2022-04-22T10:26:53.869618Z","shell.execute_reply":"2022-04-22T10:26:55.831883Z"},"jupyter":{"source_hidden":true},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<a id=\"7\"></a>\n<font color=\"skyblue\" size=+2.5><b>7. Range tools (useful for time series data)</b></font>\n\n<a href=\"#top\" class=\"btn btn-primary btn-sm\" role=\"button\" aria-pressed=\"true\" style=\"color:white\" data-toggle=\"popover\">go to top</a>\n\n> * Can be used to link two plots together\n> * Can be useful to provide a more detailed view of a subset of the data while preserving an overview of the full data\n> * In the examples below we can highlight interesting regions of the sensor readings such as peaks and valleys or region of faulty readings using the air-pollution data","metadata":{}},{"cell_type":"code","source":"curdoc().theme = 'dark_minimal'\n\ndates = np.array(df_ts['date_time'], dtype=np.datetime64)\nsource = ColumnDataSource(data=dict(date=dates, signal=df_ts['target_carbon_monoxide']))\n\nfig = figure(title='Carbon Monoxide',plot_height=250, plot_width=800, tools=\"xpan\", toolbar_location=None,\n           x_axis_type=\"datetime\", x_axis_location=\"above\",background_fill_color=\"#efefef\", x_range=(dates[3500], dates[4500]))\n\nfig.line('date', 'signal', color='red', source=source)\nfig.yaxis.axis_label = None\n\nfig.title.text_font_size = '20pt'\nfig.title.text_font_style = 'bold'\nfig.title.text_font = 'Serif'\n\nselect = figure(title=\"Drag the middle and edges of the selection box to change the range above\",\n                plot_height=130, plot_width=800, y_range=fig.y_range,\n                x_axis_type=\"datetime\", y_axis_type=None,\n                tools=\"\", toolbar_location=None, background_fill_color=\"#efefef\")\n\nrange_tool = RangeTool(x_range=fig.x_range)\nrange_tool.overlay.fill_color = \"navy\"\nrange_tool.overlay.fill_alpha = 0.2\n\nselect.line('date', 'signal', color='red', source=source)\nselect.ygrid.grid_line_color = None\nselect.add_tools(range_tool)\nselect.toolbar.active_multi = range_tool\n\nshow(column(fig,select))","metadata":{"_kg_hide-input":true,"execution":{"iopub.status.busy":"2022-04-22T10:26:55.83651Z","iopub.execute_input":"2022-04-22T10:26:55.836823Z","iopub.status.idle":"2022-04-22T10:26:56.111668Z","shell.execute_reply.started":"2022-04-22T10:26:55.836794Z","shell.execute_reply":"2022-04-22T10:26:56.110892Z"},"jupyter":{"source_hidden":true},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"curdoc().theme = 'dark_minimal'\n\ndates = np.array(df_ts['date_time'], dtype=np.datetime64)\nsource = ColumnDataSource(data=dict(date=dates, signal=df_ts['target_benzene']))\n\nfig = figure(title='Benzene',\n    plot_height=250, plot_width=800, tools=\"xpan\", toolbar_location=None,\n           x_axis_type=\"datetime\", x_axis_location=\"above\",\n           background_fill_color=\"#efefef\", x_range=(dates[3500], dates[4500]))\n\nfig.line('date', 'signal', color='gold',source=source)\nfig.yaxis.axis_label = None\n\nfig.title.text_font_size = '20pt'\nfig.title.text_font_style = 'bold'\nfig.title.text_font = 'Serif'\n\nselect = figure(title=\"Drag the middle and edges of the selection box to change the range above\",\n                plot_height=130, plot_width=800, y_range=fig.y_range,\n                x_axis_type=\"datetime\", y_axis_type=None,\n                tools=\"\", toolbar_location=None, background_fill_color=\"#efefef\")\n\nrange_tool = RangeTool(x_range=fig.x_range)\nrange_tool.overlay.fill_color = \"salmon\"\nrange_tool.overlay.fill_alpha = 0.2\n\nselect.line('date', 'signal', color='gold', source=source)\nselect.ygrid.grid_line_color = None\nselect.add_tools(range_tool)\nselect.toolbar.active_multi = range_tool\n\nshow(column(fig, select))","metadata":{"_kg_hide-input":true,"execution":{"iopub.status.busy":"2022-04-22T10:26:56.113116Z","iopub.execute_input":"2022-04-22T10:26:56.113575Z","iopub.status.idle":"2022-04-22T10:26:56.328867Z","shell.execute_reply.started":"2022-04-22T10:26:56.113531Z","shell.execute_reply":"2022-04-22T10:26:56.327191Z"},"jupyter":{"source_hidden":true},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"curdoc().theme = 'dark_minimal'\ndates = np.array(df_ts['date_time'], dtype=np.datetime64)\nsource = ColumnDataSource(data=dict(date=dates, signal=df_ts['target_nitrogen_oxides']))\n\nfig = figure(title='Nitrogen Oxides',\n    plot_height=250, plot_width=800, tools=\"xpan\", toolbar_location=None,\n           x_axis_type=\"datetime\", x_axis_location=\"above\",\n           background_fill_color=\"#efefef\", x_range=(dates[5500], dates[6500]))\n\nfig.line('date', 'signal', color='seagreen', source=source)\nfig.yaxis.axis_label = None\n\nfig.title.text_font_size = '20pt'\nfig.title.text_font_style = 'bold'\nfig.title.text_font = 'Serif'\n\nselect = figure(title=\"Drag the middle and edges of the selection box to change the range above\",\n                plot_height=130, plot_width=800, y_range=fig.y_range,\n                x_axis_type=\"datetime\", y_axis_type=None,\n                tools=\"\", toolbar_location=None, background_fill_color=\"#efefef\")\n\nrange_tool = RangeTool(x_range=fig.x_range)\nrange_tool.overlay.fill_color = \"salmon\"\nrange_tool.overlay.fill_alpha = 0.2\n\nselect.line('date', 'signal', color='seagreen', source=source)\nselect.ygrid.grid_line_color = None\nselect.add_tools(range_tool)\nselect.toolbar.active_multi = range_tool\n\nshow(column(fig, select))","metadata":{"_kg_hide-input":true,"execution":{"iopub.status.busy":"2022-04-22T10:26:56.331129Z","iopub.execute_input":"2022-04-22T10:26:56.331653Z","iopub.status.idle":"2022-04-22T10:26:56.564358Z","shell.execute_reply.started":"2022-04-22T10:26:56.331613Z","shell.execute_reply":"2022-04-22T10:26:56.563365Z"},"jupyter":{"source_hidden":true},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<a id=\"8\"></a>\n<font color=\"skyblue\" size=+2.5><b>8. Pairplots (grouped grids)</b></font>\n","metadata":{}},{"cell_type":"code","source":"# data to use\ndf_iris.drop('Id', axis=1, inplace=True)\n\n# groupby and species and creat overlay\n_df = hv.Dataset(df_iris).groupby('Species').overlay()\n\ndensity_grid = gridmatrix(_df, diagonal_type=hv.Distribution, chart_type=hv.Bivariate)\npoint_grid = gridmatrix(_df, chart_type=hv.Points)\n\n(density_grid * point_grid).opts(\n    opts.Bivariate(bandwidth=0.5, cmap=hv.Cycle(values=['Blues', 'Oranges', 'Reds'])),\n    opts.Points(size=2, alpha=0.5),\n    opts.NdOverlay(batched=False))","metadata":{"_kg_hide-input":true,"execution":{"iopub.status.busy":"2022-04-22T10:26:56.565484Z","iopub.execute_input":"2022-04-22T10:26:56.565925Z","iopub.status.idle":"2022-04-22T10:27:15.37013Z","shell.execute_reply.started":"2022-04-22T10:26:56.565861Z","shell.execute_reply":"2022-04-22T10:27:15.369139Z"},"jupyter":{"source_hidden":true},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<a id=\"9\"></a>\n<font color=\"skyblue\" size=+2.5><b>9. Closing Remark</b></font>\n\n<font color=\"skyblue\" size=+1.5><b>In this tutorial notebook we have covered:</b></font>\n\n* What Bokeh and HoloViews data visualization libraries are.\n* How to make basic plots and also customize them.\n\nHowever, this is just a starter notebook. You can further explore the library by visiting their respective official user guides (see below). I highly recommend you to do so. There is much more to discover and have fun with visualization techniques.\n","metadata":{}},{"cell_type":"markdown","source":"<a id=\"10\"></a>\n<font color=\"skyblue\" size=+2.5><b>10. References</b></font>\n\n<a href=\"#top\" class=\"btn btn-primary btn-sm\" role=\"button\" aria-pressed=\"true\" style=\"color:white\" data-toggle=\"popover\">go to top</a>\n\n[1].  https://docs.bokeh.org/en/latest/\n\n[2].  https://holoviews.org/gallery/index.html\n\n[3].  https://malouche.github.io/notebooks/index.html","metadata":{}},{"cell_type":"markdown","source":"> ### I hope that you find this notebook interesting and most importantly `useful!` If you have any `feedback` please do not hesitate to comment.\n> ### I have made a similar notebook for `plotly python library`. If you are interested you can check it [`HERE`](https://www.kaggle.com/desalegngeb/plotly-guide-customize-for-better-visualizations).\n># <font color= 'salmon'> Thank you! </font> \n\n\n\n","metadata":{}},{"cell_type":"code","source":"","metadata":{},"execution_count":null,"outputs":[]}]}