{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"pygments_lexer":"ipython3","nbconvert_exporter":"python","version":"3.6.4","file_extension":".py","codemirror_mode":{"name":"ipython","version":3},"name":"python","mimetype":"text/x-python"}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"<div \n     style=\"padding: 20px; \n            color: black;\n            margin: 0;\n            font-size: 250%;\n            text-align: center;\n            display: fill;\n            border-radius: 5px;\n            background-color: #f79d28;\n            overflow: hidden;\n            font-weight: 700;\n            border: 5px solid black;\"\n     >\n    Plotly Tutorial\n</div>","metadata":{}},{"cell_type":"markdown","source":"We'll be taking our data visualization skills to the next level with plotly.  \nplotly is a great package that allows you to create dynamic data visualizations in Python.  \nHere's an example of what we'll be able to create by the end of this tutorial.","metadata":{}},{"cell_type":"code","source":"import numpy as np  # linear algebra\nimport pandas as pd  # data processing\nimport seaborn as sns  # datasets\nimport itertools  # iteration utils\n\nimport plotly.express as px\nimport plotly.graph_objects as go\nfrom plotly import subplots\n\n\ndf_flights = sns.load_dataset(\"flights\")\ndf_flights['passengers'] = df_flights['passengers'].map(lambda x: [i for i in range(x)])\ndf_flights = df_flights.explode('passengers')\ndf_flights['passenger_id'] = df_flights['year'].astype(str) + '-' + df_flights['month'].astype(str) + '-' + df_flights['passengers'].astype(str)\ndf_flights = df_flights.drop(columns=['passengers'], axis=1)\n\ndf_growth = df_flights['passenger_id'].groupby(df_flights['year']).count().to_frame()\ndf_growth = df_growth.reset_index(level=0)\ndf_growth = df_growth.rename(columns={\"passenger_id\": \"passengers\"})\ndf_growth['prev_passengers'] = df_growth['passengers'].shift(1)\ndf_growth['passenger_growth_yoy'] = (df_growth['passengers'] - df_growth['prev_passengers']) / df_growth['prev_passengers']\n\nfig = subplots.make_subplots(\n    rows=2, cols=1, row_heights=[0.3, 0.7],\n    subplot_titles=['YoY Growth', 'Monthly and Yearly Distribution']\n)\n\nfig.add_trace(\n    go.Bar(\n        x=df_growth['year'], y=df_growth['passenger_growth_yoy'], name='YoY Growth',\n        marker_color='green'\n    ),\n    row=1, col=1\n)\n\nfig.add_trace(\n    go.Histogram2d(\n        x=df_flights['year'], y=df_flights['month'], z=df_flights['passenger_id'],\n        name='Distribution',\n        histfunc='count', texttemplate=\"%{z}\"\n    ),\n    row=2, col=1\n)\n\nfig.update_layout(\n    title='Flight Passengers Evolution 1949-1960',\n    height=800,\n    yaxis1_tickformat='.2%',\n    xaxis2_title='Year', yaxis2_title='Month',\n    plot_bgcolor='white'\n)\n\nfig.show()","metadata":{"_kg_hide-input":true,"execution":{"iopub.status.busy":"2022-09-22T09:32:01.675703Z","iopub.execute_input":"2022-09-22T09:32:01.677288Z","iopub.status.idle":"2022-09-22T09:32:07.117241Z","shell.execute_reply.started":"2022-09-22T09:32:01.677136Z","shell.execute_reply":"2022-09-22T09:32:07.115420Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"The prerequisites for this tutorial are intermediate knowledge of Python and beginner knowledge of pandas and numpy.","metadata":{}},{"cell_type":"markdown","source":"During the course of this tutorial we'll be exploring 2 plotly modules: plotly.express and plotly.graph_objects.  \nplotly.express (px) is a fast and easy way to create dynamic data visualizations.  \nplotly.graph_objects (go) is the lower level API that grants more control over your visualizations, but is more code intensive.","metadata":{}},{"cell_type":"markdown","source":"Shoutout to Derek Banas, I created this Kaggle Notebook in part by following his [tutorial on YouTube](https://www.youtube.com/watch?v=GGL6U0k8WYA).  \n","metadata":{}},{"cell_type":"markdown","source":"If you are interested in Machine Learning, I'm developing a new series of notebooks that I hope will make you an ML Grandmaster.\n<div\n     style=\"float: left; position: relative; width: 100%;\"\n     >\n    <table\n           style=\"float: left; font-size: 16px\"\n           >\n      <tr>\n        <th><b>📈 Machine Learning Fundamentals</b></th>\n        <th><b>🌳 Decision Trees</b></th>\n      </tr>\n      <tr>\n        <td><a href=\"https://www.kaggle.com/chazzer/ml-grandmaster-linear-regression/\">Linear Regression</a></td>\n        <td><a href=\"https://www.kaggle.com/code/chazzer/ml-grandmaster-decision-tree-classifier\">Decision Tree Classification</a></td>\n      </tr>\n      <tr>\n        <td>Logistic Regression [work in progress]</td>\n        <td><a href=\"https://www.kaggle.com/code/chazzer/ml-grandmaster-decision-tree-regressor\">Decision Tree Regression</a></td>\n      </tr>\n      <tr>\n        <td></td>\n        <td><a href=\"https://www.kaggle.com/code/chazzer/ml-grandmaster-random-forest/\">Random Forest</a></td>\n      </tr>\n      <tr>\n        <td></td>\n        <td><a href=\"https://www.kaggle.com/chazzer/ml-grandmaster-gradient-boosting-xgboost/\">Gradient Boosting (XGBoost)</a></td>\n      </tr>\n    </table>\n</div>","metadata":{}},{"cell_type":"markdown","source":"# 🚚 Import","metadata":{}},{"cell_type":"code","source":"import numpy as np  # linear algebra\nimport pandas as pd  # data processing\nimport seaborn as sns  # datasets\nimport itertools  # iteration utils\nfrom scipy.interpolate import griddata  # for 3d surface plot\n\nimport plotly.express as px\nimport plotly.graph_objects as go\nfrom plotly import subplots","metadata":{"_uuid":"8f2839f25d086af736a60e9eeb907d3b93b6e0e5","_cell_guid":"b1076dfc-b9ad-4769-8c92-a6c4dae69d19","execution":{"iopub.status.busy":"2022-09-22T09:32:07.119796Z","iopub.execute_input":"2022-09-22T09:32:07.120443Z","iopub.status.idle":"2022-09-22T09:32:07.131812Z","shell.execute_reply.started":"2022-09-22T09:32:07.120391Z","shell.execute_reply":"2022-09-22T09:32:07.130077Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# 📊 Bar Chart","metadata":{}},{"cell_type":"markdown","source":"Let's start by creating a few bar charts.  \nWe'll be analizing the builtin plotly.express \"gapminder\" dataset, which contains a few metrics for each country of the world.","metadata":{}},{"cell_type":"code","source":"df_world = px.data.gapminder()\ndf_world.head()","metadata":{"execution":{"iopub.status.busy":"2022-09-22T09:32:07.135075Z","iopub.execute_input":"2022-09-22T09:32:07.136889Z","iopub.status.idle":"2022-09-22T09:32:07.192334Z","shell.execute_reply.started":"2022-09-22T09:32:07.136816Z","shell.execute_reply":"2022-09-22T09:32:07.190953Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"px.bar(df_world, x='year', y='pop', color='continent', hover_name='country', title='World Population Growth')","metadata":{"execution":{"iopub.status.busy":"2022-09-22T09:32:07.196239Z","iopub.execute_input":"2022-09-22T09:32:07.197760Z","iopub.status.idle":"2022-09-22T09:32:08.534882Z","shell.execute_reply.started":"2022-09-22T09:32:07.197681Z","shell.execute_reply":"2022-09-22T09:32:08.533875Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Next let's take a more detailed look at Europe's population distribution by country in 2007.  \nWe can query the dataset by using the query() method and by passing in a Python logic statement as a string parameter.","metadata":{}},{"cell_type":"code","source":"df_europe = px.data.gapminder().query(\"continent == 'Europe' and year == 2007 \")\ndf_europe.head()","metadata":{"execution":{"iopub.status.busy":"2022-09-22T09:32:08.536461Z","iopub.execute_input":"2022-09-22T09:32:08.537375Z","iopub.status.idle":"2022-09-22T09:32:08.569366Z","shell.execute_reply.started":"2022-09-22T09:32:08.537333Z","shell.execute_reply":"2022-09-22T09:32:08.568012Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"We can instantiate any chart/plot as an object and access its methods for more granular control over our data","metadata":{}},{"cell_type":"code","source":"fig = px.bar(df_europe, x='country', y='pop', text='pop', color='country')\nfig.update_traces(texttemplate='%{text:.2s}', textposition='outside')\nfig.show()","metadata":{"execution":{"iopub.status.busy":"2022-09-22T09:32:08.571124Z","iopub.execute_input":"2022-09-22T09:32:08.571585Z","iopub.status.idle":"2022-09-22T09:32:08.785708Z","shell.execute_reply.started":"2022-09-22T09:32:08.571548Z","shell.execute_reply":"2022-09-22T09:32:08.784381Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# 📈 Line Plot","metadata":{}},{"cell_type":"markdown","source":"Let's create a few line plots.  \nLine plots are generally used to visualize a variable y that changes along the x axis (usually time or space, but not necessarily).  \nWe'll be analizing the builtin plotly.express \"stocks\" dataset.","metadata":{}},{"cell_type":"code","source":"df_stocks = px.data.stocks()\ndf_stocks.head()","metadata":{"execution":{"iopub.status.busy":"2022-09-22T09:32:08.787447Z","iopub.execute_input":"2022-09-22T09:32:08.787815Z","iopub.status.idle":"2022-09-22T09:32:08.808708Z","shell.execute_reply.started":"2022-09-22T09:32:08.787783Z","shell.execute_reply":"2022-09-22T09:32:08.807506Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"We'll instantiate the go.Figure() class as fig, which will handle the graph layout.  \nUsing the fig.add_trace() method we'll be able to add plots to our layout.  \nThe best way to create a line chart is the go.Scatter() class and setting mode='lines'.  \nWhen using the go.Figure() class, so we pass the plot class (go.Scatter() in this case) as the first parameter of the fig.add_trace() method.","metadata":{}},{"cell_type":"code","source":"fig = go.Figure()\n# we subtract 1 to stocks price data to show performance\nfig.add_trace(go.Scatter(x=df_stocks.date, y=df_stocks.GOOG - 1, mode='lines', name='Google'))\nfig.add_trace(go.Scatter(x=df_stocks.date, y=df_stocks.AAPL - 1, mode='lines', name='Apple'))\nfig.add_trace(go.Scatter(x=df_stocks.date, y=df_stocks.AMZN - 1, mode='lines+markers', name='Amazon'))\n\nfig.update_layout(\n    title='Some Tech Stocks Performance (Jan 2018 - Jan 2020)',\n    xaxis_title='Date', yaxis_title='Price'\n)","metadata":{"execution":{"iopub.status.busy":"2022-09-22T09:32:08.810344Z","iopub.execute_input":"2022-09-22T09:32:08.810704Z","iopub.status.idle":"2022-09-22T09:32:08.841184Z","shell.execute_reply.started":"2022-09-22T09:32:08.810673Z","shell.execute_reply":"2022-09-22T09:32:08.839956Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# 🤔 Advanced Styling 1 | update_layout()","metadata":{}},{"cell_type":"markdown","source":"Now we'll be going bonkers on the the previous graph's styling by adding in a lot of parameters to the go.Figure.update_layout() method.","metadata":{}},{"cell_type":"code","source":"fig = go.Figure()\n# we subtract 1 to stocks price data to show performance\nfig.add_trace(go.Scatter(x=df_stocks.date, y=df_stocks.GOOG - 1, mode='markers', name='Google'))\nfig.add_trace(go.Scatter(x=df_stocks.date, y=df_stocks.AAPL - 1, mode='lines', name='Apple'))\nfig.add_trace(go.Scatter(x=df_stocks.date, y=df_stocks.AMZN - 1, mode='lines+markers', name='Amazon'))\n\nfig.update_layout(\n    title='Some Tech Stocks Performance (Jan 2018 - Jan 2020)',\n    xaxis=dict(\n        showline=True, showgrid=False, showticklabels=True,\n        linecolor='rgb(204, 204, 204)', linewidth=2, ticks='outside', \n        tickfont=dict(\n            family='Arial', size=12, color='rgb(82, 82, 82)'\n        )\n    ),\n    yaxis=dict(\n        showgrid=False, showline=False, tickformat='.2%'\n    ),\n    margin=dict(\n        autoexpand=False, l=100, r=20, t=110\n    ),\n    showlegend=False,\n    plot_bgcolor='white'\n)","metadata":{"execution":{"iopub.status.busy":"2022-09-22T09:32:08.842719Z","iopub.execute_input":"2022-09-22T09:32:08.843114Z","iopub.status.idle":"2022-09-22T09:32:08.896745Z","shell.execute_reply.started":"2022-09-22T09:32:08.843080Z","shell.execute_reply":"2022-09-22T09:32:08.895370Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# 🔵 Scatter Plot","metadata":{}},{"cell_type":"markdown","source":"We already introduced scatter plots in the line plot section, so let's build upon that.  \nThe main difference between a scatter plot and a line plot is that a scatter plot is used, in general, to explore correlation of 2 variables or clustering of 2 variables in discrete groups.","metadata":{}},{"cell_type":"code","source":"df_iris = px.data.iris()\ndf_iris.head()","metadata":{"execution":{"iopub.status.busy":"2022-09-22T09:32:08.902512Z","iopub.execute_input":"2022-09-22T09:32:08.903863Z","iopub.status.idle":"2022-09-22T09:32:08.924007Z","shell.execute_reply.started":"2022-09-22T09:32:08.903809Z","shell.execute_reply":"2022-09-22T09:32:08.922645Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"fig = px.scatter(\n    df_iris, x='sepal_width', y='sepal_length', \n    color='species', size='petal_width', hover_data=['petal_length']\n)\nfig.update_layout(width= 1000, height=600)\nfig.show()","metadata":{"execution":{"iopub.status.busy":"2022-09-22T09:32:08.925600Z","iopub.execute_input":"2022-09-22T09:32:08.925990Z","iopub.status.idle":"2022-09-22T09:32:09.028037Z","shell.execute_reply.started":"2022-09-22T09:32:08.925954Z","shell.execute_reply":"2022-09-22T09:32:09.026759Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"We can see there is some correlation between sepal_width and sepal_length, and the \"setosa\" species is in a cluster that's distinct from \"versicolor\" and \"virginica\".","metadata":{}},{"cell_type":"markdown","source":"For large datasets (size > 10000 rows) the Scattergl class is faster (performance-wise) than Scatter, but it has less features.","metadata":{}},{"cell_type":"code","source":"fig = go.Figure()\nfig.add_trace(\n    go.Scattergl(\n        x=np.random.randn(100000),\n        y=np.random.randn(100000),\n        mode='markers',\n        marker=dict(\n            color=np.random.randn(100000),\n            colorscale='sunsetdark',\n            line_width=1\n        )\n    )\n)\nfig.show()","metadata":{"execution":{"iopub.status.busy":"2022-09-22T09:32:09.029536Z","iopub.execute_input":"2022-09-22T09:32:09.030794Z","iopub.status.idle":"2022-09-22T09:32:09.206982Z","shell.execute_reply.started":"2022-09-22T09:32:09.030746Z","shell.execute_reply":"2022-09-22T09:32:09.204920Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# 🍩 Donut Chart (a.k.a. Pie Chart But Better)","metadata":{}},{"cell_type":"markdown","source":"Donut Charts and Pie Charts are useful when identifying which categories of data represent a bigger slice of the whole, based on a numeric variable.  \nFor example, let's see how the population of Asia (numeric variable) is distributed among Asia's countries (categoric variable).","metadata":{}},{"cell_type":"code","source":"df_asia = px.data.gapminder().query(\"continent == 'Asia' and year == 2007\")\ndf_asia.head()","metadata":{"execution":{"iopub.status.busy":"2022-09-22T09:32:09.208880Z","iopub.execute_input":"2022-09-22T09:32:09.209367Z","iopub.status.idle":"2022-09-22T09:32:09.243511Z","shell.execute_reply.started":"2022-09-22T09:32:09.209324Z","shell.execute_reply":"2022-09-22T09:32:09.241982Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"fig = px.pie(\n    df_asia, values='pop', names='country',\n    hole=0.5\n)\nfig.update_layout(height=1000, title='Population of Asia')\nfig.show()","metadata":{"execution":{"iopub.status.busy":"2022-09-22T09:32:09.245069Z","iopub.execute_input":"2022-09-22T09:32:09.245478Z","iopub.status.idle":"2022-09-22T09:32:09.323534Z","shell.execute_reply.started":"2022-09-22T09:32:09.245444Z","shell.execute_reply":"2022-09-22T09:32:09.321994Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# 🤔 Advanced Styling 2 | Data Manipulation","metadata":{}},{"cell_type":"markdown","source":"To reach our visualization goals, we'll sometimes need to manipulate the datasets we're given as input.  \nTo further style the previous chart, let's define an additional array that will specify how much we should pull each sector from the chart, based on its percentage value.  \nLet's say that our rule is that if a sector's percentage value is less than 1.5%, we pull it away from the center of the donut by 60%","metadata":{}},{"cell_type":"code","source":"tot_asia_pop = df_asia['pop'].sum()\nasia_pop_perc = df_asia['pop'] / tot_asia_pop\npull_array = asia_pop_perc.map(lambda x: 0.6 if x < 0.015 else 0).values\npull_array","metadata":{"execution":{"iopub.status.busy":"2022-09-22T09:32:09.325437Z","iopub.execute_input":"2022-09-22T09:32:09.326481Z","iopub.status.idle":"2022-09-22T09:32:09.339504Z","shell.execute_reply.started":"2022-09-22T09:32:09.326426Z","shell.execute_reply":"2022-09-22T09:32:09.337752Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"We'll also be doing some extra styling in the update_traces and update_layout methods","metadata":{}},{"cell_type":"code","source":"fig = go.Figure()\nfig.add_traces(\n    go.Pie(\n        values=df_asia['pop'], labels=df_asia['country'], hole=0.5\n    )\n)\nfig.update_traces(\n    pull=pull_array,\n    rotation=135,\n    hoverinfo='label+value+percent',\n    marker=dict(\n        colors=px.colors.sequential.Agsunset\n    )\n)\nfig.update_layout(\n    height=1000, title='Population of Asia'\n)\nfig.show()","metadata":{"execution":{"iopub.status.busy":"2022-09-22T09:32:09.341745Z","iopub.execute_input":"2022-09-22T09:32:09.342238Z","iopub.status.idle":"2022-09-22T09:32:09.374614Z","shell.execute_reply.started":"2022-09-22T09:32:09.342196Z","shell.execute_reply":"2022-09-22T09:32:09.373014Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# 🎲 Histogram","metadata":{}},{"cell_type":"markdown","source":"Let's analize a value's probability when rolling 2 6-faced dice.  \nWe'll do it with a histogram.","metadata":{}},{"cell_type":"code","source":"die1 = np.random.randint(1, 7, 100000)\ndie2 = np.random.randint(1, 7, 100000)\ndice = die1 + die2","metadata":{"execution":{"iopub.status.busy":"2022-09-22T09:32:09.377513Z","iopub.execute_input":"2022-09-22T09:32:09.378787Z","iopub.status.idle":"2022-09-22T09:32:09.389527Z","shell.execute_reply.started":"2022-09-22T09:32:09.378706Z","shell.execute_reply":"2022-09-22T09:32:09.388038Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"fig = px.histogram(\n    dice, nbins=11, labels={'value': 'Dice Roll'},\n    title='10000 Dice Roll Histogram',\n    color_discrete_sequence=['darkred']\n)\nfig.update_layout(showlegend=False)\nfig.show()","metadata":{"execution":{"iopub.status.busy":"2022-09-22T09:32:09.391462Z","iopub.execute_input":"2022-09-22T09:32:09.392070Z","iopub.status.idle":"2022-09-22T09:32:09.525411Z","shell.execute_reply.started":"2022-09-22T09:32:09.392011Z","shell.execute_reply":"2022-09-22T09:32:09.524170Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# 🤔 Advanced Styling 3 | Subplots","metadata":{}},{"cell_type":"markdown","source":"Let's say that we want to show both value counts and the probability distribution (value_count / total).  \nWe'll need to generate 2 plots. We'll do that by using subplots.  \nIn our imports we imported subplots from plotly, we'll be using that module.","metadata":{}},{"cell_type":"code","source":"fig = subplots.make_subplots(\n    rows=1, cols=2,\n    subplot_titles=['Value Counts', 'Value Probability']\n)\n\nfig.add_trace(\n    go.Histogram(\n        x=dice, nbinsx=11,    \n        marker=dict(\n            color=['darkred'] * 11\n        ),\n        hoverinfo='x+y'\n    ),\n    row=1, col=1\n)\n\nfig.add_trace(\n    go.Histogram(\n        x=dice, nbinsx=11, histnorm='probability',\n        marker=dict(\n            color=['darkblue'] * 11\n        ),\n        hoverinfo='x+y'\n    ),\n    row=1, col=2\n)\n\nfig.update_layout(\n    showlegend=False,\n    yaxis2_tickformat='.2%'\n)\n\nfig.show()","metadata":{"execution":{"iopub.status.busy":"2022-09-22T09:32:09.527137Z","iopub.execute_input":"2022-09-22T09:32:09.527614Z","iopub.status.idle":"2022-09-22T09:32:09.587345Z","shell.execute_reply.started":"2022-09-22T09:32:09.527570Z","shell.execute_reply":"2022-09-22T09:32:09.585955Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Since the 2 distributions are the same but with different y-axis scaling, we can consolidate the 2 subplots into 1 by using the secondary_y attribute.","metadata":{}},{"cell_type":"code","source":"fig = subplots.make_subplots(\n    rows=1, cols=1,\n    subplot_titles=['100000 Dice Roll Distribution'],\n    specs=[[{'secondary_y': True}]]\n)\n\nfig.add_trace(\n    go.Histogram(\n        x=dice, nbinsx=11,    \n        marker=dict(\n            color=['darkred'] * 11\n        ),\n        hoverinfo='x+y'\n    ),\n    row=1, col=1, secondary_y=False\n)\n\nfig.add_trace(\n    go.Histogram(\n        x=dice, nbinsx=11, histnorm='probability',\n        marker=dict(\n            color=['darkred'] * 11\n        ),\n        hoverinfo='x+y'\n    ),\n    row=1, col=1, secondary_y=True\n)\n\nfig.update_layout(\n    showlegend=False,\n    yaxis2_tickformat='.2%'\n)\n\nfig.show()","metadata":{"execution":{"iopub.status.busy":"2022-09-22T09:32:09.589334Z","iopub.execute_input":"2022-09-22T09:32:09.589744Z","iopub.status.idle":"2022-09-22T09:32:09.645081Z","shell.execute_reply.started":"2022-09-22T09:32:09.589709Z","shell.execute_reply":"2022-09-22T09:32:09.643857Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# 🗳️ Box Plot","metadata":{}},{"cell_type":"markdown","source":"A Box Plot is used to visualize a numeric variable's quartiles.  \nWhen you have a sorted array of numeric variables, the quartiles are defined as follows.  \nQ0 is the lowest data point, exluding outliers, so equal or close to 0% of the data.  \nQ1 is the point that splits off the bottom 25% of the data.  \nQ2 (better known as the median) is the point that splits off the bottom 50% of the data.  \nQ3 is the point that splits off the bottom 50% of the data.  \nQ4 is the highest data point, excluding outliers, so equal or close to 100% of the data.  \nIn the Box Plot, the data within the box is all the data within Q1 and Q3.  \nThe line inside the box is the median (or Q2).  \nThe bottom whisker contains the data within Q0 and Q1.  \nThe upper whisker contains the data within Q3 and Q4.  \nOutliers are displayed as data points outside the whiskers.","metadata":{}},{"cell_type":"code","source":"df_tips = px.data.tips()\ndf_tips.head()","metadata":{"execution":{"iopub.status.busy":"2022-09-22T09:32:09.646535Z","iopub.execute_input":"2022-09-22T09:32:09.647840Z","iopub.status.idle":"2022-09-22T09:32:09.670794Z","shell.execute_reply.started":"2022-09-22T09:32:09.647778Z","shell.execute_reply":"2022-09-22T09:32:09.669165Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"px.box(df_tips, x='day', y='tip', color='sex')","metadata":{"execution":{"iopub.status.busy":"2022-09-22T09:32:09.672620Z","iopub.execute_input":"2022-09-22T09:32:09.673087Z","iopub.status.idle":"2022-09-22T09:32:09.781461Z","shell.execute_reply.started":"2022-09-22T09:32:09.673048Z","shell.execute_reply":"2022-09-22T09:32:09.779838Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"col_cycle = itertools.cycle(px.colors.qualitative.Alphabet)\nfig = go.Figure()\nfor sex in df_tips['sex'].unique():\n    df_plot = df_tips.loc[df_tips['sex'] == sex]\n    color = next(col_cycle)\n    \n    fig.add_trace(\n        go.Box(\n            x=df_plot['day'], y=df_plot['tip'],\n            line=dict(\n                color=color\n            ),\n            notched=True,\n            name=sex,\n            boxmean='sd'\n        )\n    )\nfig.update_layout(\n    boxmode='group',\n    title='Tips by Day and Sex',\n    yaxis_tickformat='$'\n)\nfig.show()","metadata":{"execution":{"iopub.status.busy":"2022-09-22T09:32:09.783772Z","iopub.execute_input":"2022-09-22T09:32:09.785387Z","iopub.status.idle":"2022-09-22T09:32:09.819407Z","shell.execute_reply.started":"2022-09-22T09:32:09.785316Z","shell.execute_reply":"2022-09-22T09:32:09.818035Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# 🎻 Violin Plot","metadata":{}},{"cell_type":"markdown","source":"A Violin plot is similar to a Box Plot, but it applies a Kernel Density Estimation (KDE) to the data, thereby making the visualization smoother.","metadata":{}},{"cell_type":"code","source":"df_tips.head()","metadata":{"execution":{"iopub.status.busy":"2022-09-22T09:32:09.821560Z","iopub.execute_input":"2022-09-22T09:32:09.822462Z","iopub.status.idle":"2022-09-22T09:32:09.840416Z","shell.execute_reply.started":"2022-09-22T09:32:09.822387Z","shell.execute_reply":"2022-09-22T09:32:09.838942Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# let's compare it to the box plot by setting box=True\npx.violin(df_tips, y='total_bill', box=True, points='all')","metadata":{"execution":{"iopub.status.busy":"2022-09-22T09:32:09.842253Z","iopub.execute_input":"2022-09-22T09:32:09.842795Z","iopub.status.idle":"2022-09-22T09:32:09.941929Z","shell.execute_reply.started":"2022-09-22T09:32:09.842747Z","shell.execute_reply":"2022-09-22T09:32:09.940656Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Let's look at a 2 sided Violin Plots, a way to visualize 2 distributions of the same numeric data.","metadata":{}},{"cell_type":"code","source":"fig = go.Figure()\n\ndf_yes = df_tips.loc[df_tips['smoker'] == 'Yes']\nfig.add_trace(\n    go.Violin(\n        x=df_yes['day'], y=df_yes['total_bill'], name='Smoker',\n        legendgroup='Yes', scalegroup='Yes',\n        side='negative', line_color='red'\n    )\n)\n\ndf_no = df_tips.loc[df_tips['smoker'] == 'No']\nfig.add_trace(\n    go.Violin(\n        x=df_no['day'], y=df_no['total_bill'], name='Non Smoker',\n        legendgroup='Yes', scalegroup='Yes',\n        side='positive', line_color='green'\n    )\n)\n\nfig.update_layout(\n    title='Total Bill by Day | Smokers vs. Non Smokers'\n)\n\nfig.show()","metadata":{"execution":{"iopub.status.busy":"2022-09-22T09:32:09.943787Z","iopub.execute_input":"2022-09-22T09:32:09.944425Z","iopub.status.idle":"2022-09-22T09:32:09.979363Z","shell.execute_reply.started":"2022-09-22T09:32:09.944380Z","shell.execute_reply":"2022-09-22T09:32:09.977794Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# 🟧 Density Heatmap (or 2D Histogram)","metadata":{}},{"cell_type":"markdown","source":"A Density Heatmap is kinda like a 2D Histogram, where the x and y axis are dedicated to 2 variables and the \"z axis\" (or rather the color of the Heatmap cell) is dedicated to count/sum of occurences.","metadata":{}},{"cell_type":"code","source":"df_flights = sns.load_dataset(\"flights\")\ndf_flights.head()","metadata":{"execution":{"iopub.status.busy":"2022-09-22T09:32:09.981142Z","iopub.execute_input":"2022-09-22T09:32:09.981614Z","iopub.status.idle":"2022-09-22T09:32:10.000653Z","shell.execute_reply.started":"2022-09-22T09:32:09.981577Z","shell.execute_reply":"2022-09-22T09:32:09.999688Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"We want to use the count aggregation on our data, so we'll have to create a new row for each passenger.  \nSo the first record will be mapped to 112 records, the second one to 118 records etc.  \nWe'll be flexing our pandas skills to do this.","metadata":{}},{"cell_type":"code","source":"df_flights['passengers'] = df_flights['passengers'].map(lambda x: [i for i in range(x)])\ndf_flights","metadata":{"execution":{"iopub.status.busy":"2022-09-22T09:32:10.007349Z","iopub.execute_input":"2022-09-22T09:32:10.008228Z","iopub.status.idle":"2022-09-22T09:32:10.041540Z","shell.execute_reply.started":"2022-09-22T09:32:10.008163Z","shell.execute_reply":"2022-09-22T09:32:10.039965Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"df_flights = df_flights.explode('passengers')\ndf_flights","metadata":{"execution":{"iopub.status.busy":"2022-09-22T09:32:10.044322Z","iopub.execute_input":"2022-09-22T09:32:10.044844Z","iopub.status.idle":"2022-09-22T09:32:10.076685Z","shell.execute_reply.started":"2022-09-22T09:32:10.044798Z","shell.execute_reply":"2022-09-22T09:32:10.075243Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"df_flights['passenger_id'] = df_flights['year'].astype(str) + '-' + df_flights['month'].astype(str) + '-' + df_flights['passengers'].astype(str)\ndf_flights = df_flights.drop(columns=['passengers'], axis=1)\ndf_flights","metadata":{"execution":{"iopub.status.busy":"2022-09-22T09:32:10.079454Z","iopub.execute_input":"2022-09-22T09:32:10.079945Z","iopub.status.idle":"2022-09-22T09:32:10.170057Z","shell.execute_reply.started":"2022-09-22T09:32:10.079904Z","shell.execute_reply":"2022-09-22T09:32:10.168629Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"We'll be be showing the marginal distributions along x and y","metadata":{}},{"cell_type":"code","source":"fig = px.density_heatmap(\n    df_flights, x='year', y='month', z='passenger_id',\n    marginal_x='histogram', marginal_y='histogram', histfunc='count'\n)\nfig.show()","metadata":{"execution":{"iopub.status.busy":"2022-09-22T09:32:10.171811Z","iopub.execute_input":"2022-09-22T09:32:10.172410Z","iopub.status.idle":"2022-09-22T09:32:10.811512Z","shell.execute_reply.started":"2022-09-22T09:32:10.172351Z","shell.execute_reply":"2022-09-22T09:32:10.809835Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Next let's overlay the number of passengers in the Heatmap cells and show the Year over Year growth as a Bar Chart.  \nWe'll be doing some data transformation to get the YoY growth.","metadata":{}},{"cell_type":"code","source":"df_growth = df_flights['passenger_id'].groupby(df_flights['year']).count().to_frame()\ndf_growth = df_growth.reset_index(level=0)\ndf_growth","metadata":{"execution":{"iopub.status.busy":"2022-09-22T09:32:10.813707Z","iopub.execute_input":"2022-09-22T09:32:10.814207Z","iopub.status.idle":"2022-09-22T09:32:10.870304Z","shell.execute_reply.started":"2022-09-22T09:32:10.814156Z","shell.execute_reply":"2022-09-22T09:32:10.869198Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"df_growth = df_growth.rename(columns={\"passenger_id\": \"passengers\"})\ndf_growth['prev_passengers'] = df_growth['passengers'].shift(1)\ndf_growth['passenger_growth_yoy'] = (df_growth['passengers'] - df_growth['prev_passengers']) / df_growth['prev_passengers']\ndf_growth","metadata":{"execution":{"iopub.status.busy":"2022-09-22T09:32:10.871946Z","iopub.execute_input":"2022-09-22T09:32:10.872610Z","iopub.status.idle":"2022-09-22T09:32:10.892179Z","shell.execute_reply.started":"2022-09-22T09:32:10.872569Z","shell.execute_reply":"2022-09-22T09:32:10.890894Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"fig = subplots.make_subplots(\n    rows=2, cols=1, row_heights=[0.3, 0.7],\n    subplot_titles=['YoY Growth', 'Monthly and Yearly Distribution']\n)\n\nfig.add_trace(\n    go.Bar(\n        x=df_growth['year'], y=df_growth['passenger_growth_yoy'], name='YoY Growth',\n        marker_color='green'\n    ),\n    row=1, col=1\n)\n\nfig.add_trace(\n    go.Histogram2d(\n        x=df_flights['year'], y=df_flights['month'], z=df_flights['passenger_id'],\n        name='Distribution',\n        histfunc='count', texttemplate=\"%{z}\"\n    ),\n    row=2, col=1\n)\n\nfig.update_layout(\n    title='Flight Passengers Evolution 1949-1960',\n    height=800,\n    yaxis1_tickformat='.2%',\n    xaxis2_title='Year', yaxis2_title='Month',\n    plot_bgcolor='white'\n)\n\nfig.show()","metadata":{"execution":{"iopub.status.busy":"2022-09-22T09:32:10.894355Z","iopub.execute_input":"2022-09-22T09:32:10.894782Z","iopub.status.idle":"2022-09-22T09:32:11.259400Z","shell.execute_reply.started":"2022-09-22T09:32:10.894748Z","shell.execute_reply":"2022-09-22T09:32:11.257629Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# 🤖 3D Plots","metadata":{}},{"cell_type":"markdown","source":"3D Plots allow you to add a 3rd axis to your plots and rotate/pan/zoom on the plot with your mouse and keyboard.  \nLet's see them in action on the flights dataset.","metadata":{}},{"cell_type":"code","source":"df_flights = sns.load_dataset(\"flights\")\ndf_flights","metadata":{"execution":{"iopub.status.busy":"2022-09-22T09:32:11.261050Z","iopub.execute_input":"2022-09-22T09:32:11.262490Z","iopub.status.idle":"2022-09-22T09:32:11.295144Z","shell.execute_reply.started":"2022-09-22T09:32:11.262418Z","shell.execute_reply":"2022-09-22T09:32:11.293869Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## 3D Scatter Plot","metadata":{}},{"cell_type":"code","source":"fig = px.scatter_3d(\n    df_flights, x='year', y='month', z='passengers',\n    color='month', opacity=0.7\n)\n\nfig.show()","metadata":{"execution":{"iopub.status.busy":"2022-09-22T09:32:11.297050Z","iopub.execute_input":"2022-09-22T09:32:11.297572Z","iopub.status.idle":"2022-09-22T09:32:11.474898Z","shell.execute_reply.started":"2022-09-22T09:32:11.297523Z","shell.execute_reply":"2022-09-22T09:32:11.473521Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## 3D Line Plot","metadata":{}},{"cell_type":"code","source":"fig = px.line_3d(\n    df_flights, x='year', y='month', z='passengers',\n    color='month'\n)\n\nfig.show()","metadata":{"execution":{"iopub.status.busy":"2022-09-22T09:32:11.476711Z","iopub.execute_input":"2022-09-22T09:32:11.477140Z","iopub.status.idle":"2022-09-22T09:32:11.607824Z","shell.execute_reply.started":"2022-09-22T09:32:11.477102Z","shell.execute_reply":"2022-09-22T09:32:11.606360Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## 3D Surface Plot","metadata":{}},{"cell_type":"markdown","source":"Surface Plots require data to be formatted in a grid-like fashion. We'll be using numpy and the griddata function in scipy.interpolate for the heavy lifting.","metadata":{}},{"cell_type":"code","source":"month_to_int = {\n    'Jan': 1,\n    'Feb': 2,\n    'Mar': 3,\n    'Apr': 4,\n    'May': 5,\n    'Jun': 6,\n    'Jul': 7,\n    'Aug': 8,\n    'Sep': 9,\n    'Oct': 10,\n    'Nov': 11,\n    'Dec': 12\n}\ndf_flights['month'].replace(month_to_int)","metadata":{"execution":{"iopub.status.busy":"2022-09-22T09:32:11.609778Z","iopub.execute_input":"2022-09-22T09:32:11.610200Z","iopub.status.idle":"2022-09-22T09:32:11.626332Z","shell.execute_reply.started":"2022-09-22T09:32:11.610165Z","shell.execute_reply":"2022-09-22T09:32:11.624759Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"fig = go.Figure()\n\nxi = np.linspace(\n    min(df_flights['year']), \n    max(df_flights['year']), \n    num=12\n)\nyi = np.linspace(\n    min(df_flights['month'].replace(month_to_int)), \n    max(df_flights['month'].replace(month_to_int)), \n    num=12\n)\n\nx_grid, y_grid = np.meshgrid(xi,yi)\n\nz_grid = griddata(\n    (df_flights['year'],df_flights['month'].replace(month_to_int)), \n    df_flights['passengers'], \n    (x_grid,y_grid), \n    method='cubic'\n)\n\nfig.add_trace(\n    go.Surface(\n        x=x_grid, y=y_grid, z=z_grid,\n        colorscale='viridis'\n    )\n)\n\nfig.show()","metadata":{"execution":{"iopub.status.busy":"2022-09-22T09:32:11.628033Z","iopub.execute_input":"2022-09-22T09:32:11.628565Z","iopub.status.idle":"2022-09-22T09:32:11.687338Z","shell.execute_reply.started":"2022-09-22T09:32:11.628512Z","shell.execute_reply":"2022-09-22T09:32:11.686090Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# 🤯 Scatter Matrix","metadata":{}},{"cell_type":"markdown","source":"A scatter matrix creates multiple 2d scatter plots to explore relationships there might be between column variables.  \n$$n = \\# columns \\implies \\# plots = n^2$$","metadata":{}},{"cell_type":"code","source":"fig = px.scatter_matrix(df_flights)\nfig.show()","metadata":{"execution":{"iopub.status.busy":"2022-09-22T09:32:11.689817Z","iopub.execute_input":"2022-09-22T09:32:11.690800Z","iopub.status.idle":"2022-09-22T09:32:11.785923Z","shell.execute_reply.started":"2022-09-22T09:32:11.690741Z","shell.execute_reply":"2022-09-22T09:32:11.784412Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"You can highlight a certain variable with the color parameter.","metadata":{}},{"cell_type":"code","source":"fig = px.scatter_matrix(df_flights, color='passengers')\nfig.show()","metadata":{"execution":{"iopub.status.busy":"2022-09-22T09:32:11.788072Z","iopub.execute_input":"2022-09-22T09:32:11.788966Z","iopub.status.idle":"2022-09-22T09:32:11.855759Z","shell.execute_reply.started":"2022-09-22T09:32:11.788901Z","shell.execute_reply":"2022-09-22T09:32:11.854406Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"This tool is very useful for quickly exploring correlation between variables","metadata":{}},{"cell_type":"markdown","source":"# 🌍 Map Scatter Plot","metadata":{}},{"cell_type":"markdown","source":"The map plot is my favorite plot, because it allows us the see data on our home, Planet Earth.","metadata":{}},{"cell_type":"code","source":"df_world_now = df_world.query(\"year == 2007\")\ndf_world_now.head()","metadata":{"execution":{"iopub.status.busy":"2022-09-22T09:32:11.857895Z","iopub.execute_input":"2022-09-22T09:32:11.858752Z","iopub.status.idle":"2022-09-22T09:32:11.882407Z","shell.execute_reply.started":"2022-09-22T09:32:11.858692Z","shell.execute_reply":"2022-09-22T09:32:11.880989Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"fig = px.scatter_geo(\n    df_world_now, locations=\"iso_alpha\", color=\"continent\",\n    size=\"pop\", hover_name=\"country\",\n    projection=\"orthographic\"\n)\nfig.show()","metadata":{"execution":{"iopub.status.busy":"2022-09-22T09:32:11.884467Z","iopub.execute_input":"2022-09-22T09:32:11.884861Z","iopub.status.idle":"2022-09-22T09:32:12.012638Z","shell.execute_reply.started":"2022-09-22T09:32:11.884827Z","shell.execute_reply":"2022-09-22T09:32:12.011303Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# ⭕ Polar Charts","metadata":{}},{"cell_type":"markdown","source":"Polar Charts display data in polar coordinates.  \nAny time you have variables dependent on directions/angles, polar charts are a good idea.","metadata":{}},{"cell_type":"code","source":"df_wind = px.data.wind()\ndf_wind.head()","metadata":{"execution":{"iopub.status.busy":"2022-09-22T09:32:12.014282Z","iopub.execute_input":"2022-09-22T09:32:12.014744Z","iopub.status.idle":"2022-09-22T09:32:12.036267Z","shell.execute_reply.started":"2022-09-22T09:32:12.014705Z","shell.execute_reply":"2022-09-22T09:32:12.034600Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## Scatter Polar","metadata":{}},{"cell_type":"code","source":"px.scatter_polar(\n    df_wind, r='strength', theta='direction',\n    color='frequency', size='frequency'\n)","metadata":{"execution":{"iopub.status.busy":"2022-09-22T09:32:12.038467Z","iopub.execute_input":"2022-09-22T09:32:12.038908Z","iopub.status.idle":"2022-09-22T09:32:12.136094Z","shell.execute_reply.started":"2022-09-22T09:32:12.038870Z","shell.execute_reply":"2022-09-22T09:32:12.135131Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## Line Polar","metadata":{}},{"cell_type":"code","source":"px.line_polar(\n    df_wind, r='frequency', theta='direction',\n    color='strength', line_close=True, template='plotly_dark'\n)","metadata":{"execution":{"iopub.status.busy":"2022-09-22T09:32:12.137789Z","iopub.execute_input":"2022-09-22T09:32:12.138476Z","iopub.status.idle":"2022-09-22T09:32:12.293053Z","shell.execute_reply.started":"2022-09-22T09:32:12.138432Z","shell.execute_reply":"2022-09-22T09:32:12.291714Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# 🔺 Ternary Plot","metadata":{}},{"cell_type":"markdown","source":"Ternary Plot tries to have its cake and eat it too.  \nIt projects 3d space in 2d space by putting the x, y and z axes on the sides of an equilateral triangle.","metadata":{}},{"cell_type":"code","source":"df_exp = px.data.experiment()\ndf_exp.head()","metadata":{"execution":{"iopub.status.busy":"2022-09-22T09:32:12.294627Z","iopub.execute_input":"2022-09-22T09:32:12.295068Z","iopub.status.idle":"2022-09-22T09:32:12.315991Z","shell.execute_reply.started":"2022-09-22T09:32:12.295027Z","shell.execute_reply":"2022-09-22T09:32:12.314767Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"fig = px.scatter_ternary(\n    df_exp, a='experiment_1', b='experiment_2', c='experiment_3',\n    color='group', hover_name='gender',\n)\n\nfig.show()","metadata":{"execution":{"iopub.status.busy":"2022-09-22T09:32:12.317828Z","iopub.execute_input":"2022-09-22T09:32:12.318226Z","iopub.status.idle":"2022-09-22T09:32:12.440586Z","shell.execute_reply.started":"2022-09-22T09:32:12.318184Z","shell.execute_reply":"2022-09-22T09:32:12.439387Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# 🤔 Advanced Styling 4 | Facets","metadata":{}},{"cell_type":"markdown","source":"Facets allow us to easily create multiple versions of the same plot, depending on a certain facet variable.  \nLet's look at a simple example","metadata":{}},{"cell_type":"code","source":"df_tips.head()","metadata":{"execution":{"iopub.status.busy":"2022-09-22T09:32:12.442084Z","iopub.execute_input":"2022-09-22T09:32:12.443104Z","iopub.status.idle":"2022-09-22T09:32:12.461227Z","shell.execute_reply.started":"2022-09-22T09:32:12.443061Z","shell.execute_reply":"2022-09-22T09:32:12.459240Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"px.scatter(\n    df_tips, x='total_bill', y='tip', color='smoker',\n    facet_col= 'sex'\n)","metadata":{"execution":{"iopub.status.busy":"2022-09-22T09:32:12.463426Z","iopub.execute_input":"2022-09-22T09:32:12.464863Z","iopub.status.idle":"2022-09-22T09:32:12.563164Z","shell.execute_reply.started":"2022-09-22T09:32:12.464793Z","shell.execute_reply":"2022-09-22T09:32:12.561666Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"From a higher dimension geometry point of view, it allows us to see discrete slices along the dimension of the facet variable","metadata":{}},{"cell_type":"markdown","source":"We can define an additional facet with the facet_row parameter","metadata":{}},{"cell_type":"code","source":"px.histogram(\n    df_tips, x='total_bill', y='tip', color='sex',\n    facet_row='time', facet_col='day', \n    category_orders={\n        \"day\": [\"Thur\", \"Fri\", \"Sat\", \"Sun\"],\n        \"time\": [\"Lunch\", \"Dinner\"]\n    }\n)","metadata":{"execution":{"iopub.status.busy":"2022-09-22T09:32:12.567526Z","iopub.execute_input":"2022-09-22T09:32:12.567998Z","iopub.status.idle":"2022-09-22T09:32:12.793737Z","shell.execute_reply.started":"2022-09-22T09:32:12.567964Z","shell.execute_reply":"2022-09-22T09:32:12.792408Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Last example","metadata":{}},{"cell_type":"code","source":"df_att = sns.load_dataset(\"attention\")\ndf_att.head()","metadata":{"execution":{"iopub.status.busy":"2022-09-22T09:32:12.795632Z","iopub.execute_input":"2022-09-22T09:32:12.796221Z","iopub.status.idle":"2022-09-22T09:32:14.221538Z","shell.execute_reply.started":"2022-09-22T09:32:12.796158Z","shell.execute_reply":"2022-09-22T09:32:14.220101Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"df_att = df_att['score'].groupby([df_att['attention'], df_att['solutions'], df_att['subject']]).sum().reset_index()\ndf_att.head()","metadata":{"execution":{"iopub.status.busy":"2022-09-22T09:32:14.223670Z","iopub.execute_input":"2022-09-22T09:32:14.224174Z","iopub.status.idle":"2022-09-22T09:32:14.252581Z","shell.execute_reply.started":"2022-09-22T09:32:14.224127Z","shell.execute_reply":"2022-09-22T09:32:14.251288Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"fig = px.bar(\n    df_att, x='solutions', y='score', color='attention', facet_col='subject',\n    facet_col_wrap=5, title='Scores Based on Student Attention'\n)\nfig.show()","metadata":{"execution":{"iopub.status.busy":"2022-09-22T09:32:14.254871Z","iopub.execute_input":"2022-09-22T09:32:14.255496Z","iopub.status.idle":"2022-09-22T09:32:14.716779Z","shell.execute_reply.started":"2022-09-22T09:32:14.255439Z","shell.execute_reply":"2022-09-22T09:32:14.715845Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# 🤔 Advanced Styling 5 | Animation","metadata":{}},{"cell_type":"markdown","source":"Animation let's you animate your plots by changing a certain variable in time.  \nanimation_frame will be the variable that evolves with time","metadata":{}},{"cell_type":"code","source":"df_world.head()","metadata":{"execution":{"iopub.status.busy":"2022-09-22T09:32:14.718015Z","iopub.execute_input":"2022-09-22T09:32:14.718977Z","iopub.status.idle":"2022-09-22T09:32:14.734042Z","shell.execute_reply.started":"2022-09-22T09:32:14.718935Z","shell.execute_reply":"2022-09-22T09:32:14.732731Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"px.scatter(\n    df_world, x='gdpPercap', y='lifeExp',\n    animation_frame='year',\n    size='pop', color='continent', hover_name='country',\n    log_x=True, size_max=50,\n    range_x=[100, 100_000], range_y=[25, 90]\n)","metadata":{"execution":{"iopub.status.busy":"2022-09-22T09:32:14.735489Z","iopub.execute_input":"2022-09-22T09:32:14.735955Z","iopub.status.idle":"2022-09-22T09:32:15.162373Z","shell.execute_reply.started":"2022-09-22T09:32:14.735910Z","shell.execute_reply":"2022-09-22T09:32:15.161072Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"px.bar(\n    df_world, x='continent', y='pop', color='continent',\n    animation_frame='year',\n    range_y=[0, 4_000_000_000]\n)","metadata":{"execution":{"iopub.status.busy":"2022-09-22T09:32:15.164419Z","iopub.execute_input":"2022-09-22T09:32:15.165155Z","iopub.status.idle":"2022-09-22T09:32:15.516516Z","shell.execute_reply.started":"2022-09-22T09:32:15.165098Z","shell.execute_reply":"2022-09-22T09:32:15.514923Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# 🙇‍♂️ Conclusion","metadata":{}},{"cell_type":"markdown","source":"This marks the end of our tutorial.  \nI hope you found this notebook useful/fun!\nFor more information visit the [plotly python documentation](http://https://plotly.com/python/).","metadata":{}},{"cell_type":"markdown","source":"If you liked this tutorial, you'll love my ML Grandmaster series I'm developing!\n<div\n     style=\"float: left; position: relative; width: 100%;\"\n     >\n    <table\n           style=\"float: left; font-size: 16px\"\n           >\n      <tr>\n        <th><b>📈 Machine Learning Fundamentals</b></th>\n        <th><b>🌳 Decision Trees</b></th>\n      </tr>\n      <tr>\n        <td><a href=\"https://www.kaggle.com/chazzer/ml-grandmaster-linear-regression/\">Linear Regression</a></td>\n        <td><a href=\"https://www.kaggle.com/code/chazzer/ml-grandmaster-decision-tree-classifier\">Decision Tree Classification</a></td>\n      </tr>\n      <tr>\n        <td>Logistic Regression [work in progress]</td>\n        <td><a href=\"https://www.kaggle.com/code/chazzer/ml-grandmaster-decision-tree-regressor\">Decision Tree Regression</a></td>\n      </tr>\n      <tr>\n        <td></td>\n        <td><a href=\"https://www.kaggle.com/code/chazzer/ml-grandmaster-random-forest/\">Random Forest</a></td>\n      </tr>\n      <tr>\n        <td></td>\n        <td><a href=\"https://www.kaggle.com/chazzer/ml-grandmaster-gradient-boosting-xgboost/\">Gradient Boosting (XGBoost)</a></td>\n      </tr>\n    </table>\n</div>","metadata":{}}]}