{"cells":[{"metadata":{},"cell_type":"markdown","source":"# Introduction \n![](https://miro.medium.com/max/757/1*eYEoP5hF-IyQ4SCLEs1cbQ.png)\n### In this kernel notebook I will be focusing on initially covering the new Pandas_Bokeh Data visualisation followed by a exploratory data analysis ,a case study about Karnataka Education using Bokeh.\n\n**Pandas-Bokeh** provides a Bokeh plotting backend for Pandas, GeoPandas and Pyspark DataFrames, similar to the already existing Visualization feature of Pandas. Importing the library adds a complementary plotting method plot_bokeh() on DataFrames and Series.\n\n\nWith **Pandas-Bokeh**, creating stunning, interactive, HTML-based visualization is as easy as calling:\n\ndf.plot_bokeh()\n\n\n**Pandas-Bokeh** also provides native support as a Pandas Plotting backend for Pandas >= 0.25. When **Pandas-Bokeh** is installed, switchting the default Pandas plotting backend to Bokeh can be done via:\n\npd.set_option('plotting.backend', 'pandas_bokeh')\n\n![](https://miro.medium.com/max/1962/0*lfsR26JXj4o_QMWI.gif)\n\n\nNow its time to first install Pandas_Bokeh using PIP command.","execution_count":null},{"metadata":{"trusted":true,"_kg_hide-output":true},"cell_type":"code","source":"!pip install pandas-bokeh","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"### Import Libraries\n","execution_count":null},{"metadata":{"trusted":true},"cell_type":"code","source":"import numpy as np\nimport pandas as pd\nimport pandas_bokeh\npandas_bokeh.output_notebook()\npd.set_option('plotting.backend', 'pandas_bokeh')\n# Create Bokeh-Table with DataFrame:\nfrom bokeh.models.widgets import DataTable, TableColumn\nfrom bokeh.models import ColumnDataSource","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"# Import Data","execution_count":null},{"metadata":{},"cell_type":"markdown","source":"\n\n## Plot types\n\n\n#### Pandas & Pyspark DataFrames\n* Line plot\n* Point plot\n* Step plot\n* Scatter plot\n* Bar plot\n* Histogram\n* Area plot\n* Pie plot\n* Map plot\n\n#### Geoplots (Point, Line, Polygon) with GeoPandas\n\n\n### Lineplot\n\nThis simple lineplot in Pandas-Bokeh already contains various interactive elements:\n\n* a pannable and zoomable (zoom in plotarea and zoom on axis) plot\n* by clicking on the legend elements, one can hide and show the individual lines\n* a Hovertool for the plotted lines\n\nConsider the following simple example:\n\nWe will be importing the time series data about the power usage in various states in India. All of the values are measured in **MU(millions of units)**. **The date ranges from 28/10/2019 to 23/05/2020.**","execution_count":null},{"metadata":{"trusted":true},"cell_type":"code","source":"df = pd.read_csv('../input/state-wise-power-consumption-in-india/dataset_tk.csv')\ndf_long = pd.read_csv('../input/state-wise-power-consumption-in-india/long_data_.csv')","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"Firstly creating Date column and dropping the unwanted column and reformatting the date column","execution_count":null},{"metadata":{"trusted":true},"cell_type":"code","source":"df[\"Date\"]=df[\"Unnamed: 0\"]\ndf['Date'] = pd.to_datetime(df.Date, dayfirst=True)\ndf = df.drop([\"Unnamed: 0\"], axis = 1) ","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"df.info()","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"Now let us divide the states based on 5 regions namely \n\n1. Northern Region\n\n2. Southern Region\n\n3. Eastern Region\n\n4. Western Region\n\n5. North Eastern Region\n","execution_count":null},{"metadata":{"trusted":true},"cell_type":"code","source":"df['NR'] = df['Punjab']+ df['Haryana']+ df['Rajasthan']+ df['Delhi']+df['UP']+df['Uttarakhand']+df['HP']+df['J&K']+df['Chandigarh']\ndf['WR'] = df['Chhattisgarh']+df['Gujarat']+df['MP']+df['Maharashtra']+df['Goa']+df['DNH']\ndf['SR'] = df['Andhra Pradesh']+df['Telangana']+df['Karnataka']+df['Kerala']+df['Tamil Nadu']+df['Pondy']\ndf['ER'] = df['Bihar']+df['Jharkhand']+ df['Odisha']+df['West Bengal']+df['Sikkim']\ndf['NER'] =df['Arunachal Pradesh']+df['Assam']+df['Manipur']+df['Meghalaya']+df['Mizoram']+df['Nagaland']+df['Tripura']","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"df_line = pd.DataFrame({\"Northern Region\": df[\"NR\"].values,\n                        \"Southern Region\": df[\"SR\"].values,\n                        \"Eastern Region\": df[\"ER\"].values,\n                        \"Western Region\": df[\"WR\"].values,\n                        \"North Eastern Region\": df[\"NER\"].values},index=df.Date)\n\ndf_line.plot_bokeh(kind=\"line\",title =\"India - Power Consumption Regionwise\",\n                   figsize =(1000,800),\n                   xlabel = \"Date\",\n                   ylabel=\"MU(millions of units)\"\n                   )","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"### In the above data visualisation which is completely intereactive. You can click on any index regions and check the data .Is it an interesting data visualisation ???????\n\n##### Let us look at some other types of LinePlot\n\n#### Bar Type:\n","execution_count":null},{"metadata":{"trusted":true},"cell_type":"code","source":"df_line.plot_bokeh(kind=\"bar\",title =\"India - Power Consumption Regionwise\",figsize =(1000,800),xlabel = \"Date\",ylabel=\"MU(millions of units)\")","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"#### Point Type:","execution_count":null},{"metadata":{"trusted":true},"cell_type":"code","source":"df_line.plot_bokeh(kind=\"point\",title =\"India - Power Consumption Regionwise\",figsize =(1000,800),xlabel = \"Date\",ylabel=\"MU(millions of units)\")","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"#### Histogram Type:","execution_count":null},{"metadata":{"trusted":true},"cell_type":"code","source":"df_line.plot_bokeh(kind=\"hist\",title =\"India - Power Consumption Regionwise\",\n                   figsize =(1000,800),\n                   xlabel = \"Date\",\n                   ylabel=\"MU(millions of units)\"\n                )","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"#### Lineplot with rangetool","execution_count":null},{"metadata":{"trusted":true},"cell_type":"code","source":"df_line = pd.DataFrame({\"Northern Region\": df[\"NR\"].values,\n                        \"Southern Region\": df[\"SR\"].values,\n                        \"Eastern Region\": df[\"ER\"].values,\n                        \"Western Region\": df[\"WR\"].values,\n                        \"North Eastern Region\": df[\"NER\"].values},index=df.Date)\n\ndf_line.plot_bokeh(kind=\"line\",title =\"India - Power Consumption Regionwise\",\n                   figsize =(1000,800),\n                   xlabel = \"Date\",\n                   ylabel=\"MU(millions of units)\",rangetool=True\n                   )","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"### Pointplot\n\nIf you just wish to draw the date points for curves, the pointplot option is the right choice. It also accepts the kwargs of bokeh.plotting.figure.scatter like marker or size:","execution_count":null},{"metadata":{"trusted":true},"cell_type":"code","source":"df_line.plot_bokeh.point(\n    x=df.Date,\n    xticks=range(0,1),\n    size=5,\n    colormap=[\"#009933\", \"#ff3399\",\"#ae0399\",\"#220111\",\"#890300\"],\n    title=\" Point Plot - India Power Consumption\",\n    fontsize_title=20,\n    marker=\"x\",figsize =(1000,800))","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"### Stepplot\n\nWith a similar API as the line- & pointplots, one can generate a stepplot. Additional keyword arguments for this plot type are passes to bokeh.plotting.figure.step, e.g. mode (before, after, center), see the following example\n","execution_count":null},{"metadata":{"trusted":true},"cell_type":"code","source":"df_line.plot_bokeh.step(\n    x=df.Date,\n    xticks=range(-1, 1),\n    colormap=[\"#009933\", \"#ff3399\",\"#ae0399\",\"#220111\",\"#890300\"],\n    title=\"Step Plot - India Power Consumption\",\n    figsize=(1000,800),\n    fontsize_title=20,\n    fontsize_label=20,\n    fontsize_ticks=20,\n    fontsize_legend=8,\n    )","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"### Scatterplot\n\nA basic scatterplot can be created using the kind=\"scatter\" option. For scatterplots, the x and y parameters have to be specified and the following optional keyword argument is allowed:\n\ncategory: Determines the category column to use for coloring the scatter points\n\nkwargs**: Optional keyword arguments of bokeh.plotting.figure.scatter\n\nNote, that the pandas.DataFrame.plot_bokeh() method return per default a Bokeh figure, which can be embedded in Dashboard layouts with other figures and Bokeh objects (for more details about (sub)plot layouts and embedding the resulting Bokeh plots as HTML click here).\n\nIn the example below, we use the building grid layout support of Pandas-Bokeh to display both the DataFrame (using a Bokeh DataTable) and the resulting scatterplot:","execution_count":null},{"metadata":{"trusted":true},"cell_type":"code","source":"df = pd.read_csv(\"../input/iris/Iris.csv\")\ndf = df.sample(frac=1)","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"df.head()","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"data_table = DataTable(\n    columns=[TableColumn(field=Ci, title=Ci) for Ci in df.columns],\n    source=ColumnDataSource(df),\n    height=300,\n)\n\n# Create Scatterplot:\np_scatter = df.plot_bokeh.scatter(\n    x=\"PetalLengthCm\",\n    y=\"SepalWidthCm\",\n    category=\"Species\",\n    title=\"Iris DataSet Visualization\",\n    show_figure=False\n)\n\n# Combine Table and Scatterplot via grid layout:\npandas_bokeh.plot_grid([[data_table, p_scatter]], plot_width=400, plot_height=350)","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"### Barplot\n\nThe barplot API has no special keyword arguments, but accepts optional kwargs of bokeh.plotting.figure.vbar like alpha. It uses per default the index for the bar categories (however, also columns can be used as x-axis category using the x argument).\n\nLet us look at an example","execution_count":null},{"metadata":{"trusted":true},"cell_type":"code","source":"data = {\n    'Cars':\n    ['Maruti Suzuki', 'Honda', 'Toyota', 'Hyundai', 'Benz', 'BMW'],\n    '2018': [20000, 15722, 4340, 38000, 2890, 412],\n    '2019': [19000, 13700, 340, 31200, 290, 234],\n    '2020': [23456, 15891, 440, 36700, 890, 417]\n}\ndf = pd.DataFrame(data).set_index(\"Cars\")\n\np_bar = df.plot_bokeh.bar(\n    ylabel=\"Price per Unit\", \n    title=\"Car Units sold per Year\", \n    alpha=0.6)","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"Using the stacked keyword argument you also make stacked barplots as shown below","execution_count":null},{"metadata":{"trusted":true},"cell_type":"code","source":"stacked_bar = df.plot_bokeh.bar(\n    ylabel=\"Price per Unit\", \n    title=\"Car Units sold per Year\", \n    stacked=True,\n    alpha=0.6)","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"Also horizontal versions of the above barplot are supported with the keyword kind=\"barh\" or the accessor plot_bokeh.barh. You can still specify a column of the DataFrame as the bar category via the x argument if you do not wish to use the index.","execution_count":null},{"metadata":{"trusted":true},"cell_type":"code","source":"#Reset index, such that \"Cars\" is now a column of the DataFrame:\ndf.reset_index(inplace=True)\n\n#Create horizontal bar (via kind keyword):\np_hbar = df.plot_bokeh(\n    kind=\"barh\",\n    x=\"Cars\",\n    ylabel=\"Price per Unit\", \n    title=\"Car Units sold per Year\", \n    alpha=0.6,\n    legend = \"bottom_right\",\n    show_figure=False)\n\n#Create stacked horizontal bar (via barh accessor):\nstacked_hbar = df.plot_bokeh.barh(\n    x=\"Cars\",\n    stacked=True,\n    ylabel=\"Price per Unit\", \n    title=\"Car Units sold per Year\", \n    alpha=0.6,\n    legend = \"bottom_right\",\n    show_figure=False)\n\n#Plot all barplot examples in a grid:\npandas_bokeh.plot_grid([[p_bar, stacked_bar],\n                        [p_hbar, stacked_hbar]], \n                       plot_width=450)","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"Now let us look at a more practical example of housing prices problem to understand it better.","execution_count":null},{"metadata":{"trusted":true},"cell_type":"code","source":"df = pd.read_csv(\"../input/house-prices-advanced-regression-techniques/train.csv\",index_col='SalePrice')\nnumeric_features = df.select_dtypes(include=[np.number])\np_bar = numeric_features.plot_bokeh.bar(\n    ylabel=\"Sale Price\", \n    figsize=(1000,800),\n    title=\"Housing Prices\", \n    alpha=0.6)","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"","execution_count":null},{"metadata":{"_uuid":"2be426b8de93d22bb38be706699f014a7652f08e"},"cell_type":"markdown","source":"# Bokeh Introduction\n\nBokeh is an interactive visualization library that targets modern web browsers for presentation. Its goal is to provide elegant, concise construction of versatile graphics, and to extend this capability with high-performance interactivity over very large or streaming datasets. Bokeh can help anyone who would like to quickly and easily create interactive plots, dashboards, and data applications.\n\nFor this kernel we will taking an example of Karnataka State(India) Education dataset for our exploratory data analysis.\n\n# EDA -Bokeh - Karnataka Education\n\nAn NGO organisation takes initiatives to improve primary education in  India and want to carry out this program in Karnataka. It wants to target districts that fall behind in areas such as \n\n- Education Infrastructure\n\n- Education Awareness\n\n- Demographic features\n\nIdentify such districts that could be targeted in its first phase.\n\nThe source data for this exercise is obtained from data.gov.in\n\n## Goal :\n\nThe goal of this notebook was primarily to:\n\n1.       Explain the data, define your target and come up with features that can be used for modelling.\n\n2.       Create a model based on your features and come up with the list of target districts\n\n3.       Detailed analysis to include all components such as \n\n        - Data fetch\n        - Data cleansing\n        - Exploratory data analysis\n        - Summary and Data Visualization\n        \n","execution_count":null},{"metadata":{"_uuid":"304931f7ae97441228e7b81775398d35894b560e"},"cell_type":"markdown","source":"# Import Libraries","execution_count":null},{"metadata":{"trusted":true,"_kg_hide-output":true},"cell_type":"code","source":"!pip install pandas-bokeh","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"a8cc5bb84a8e58eb20c8270256788b793bdd08a3","_kg_hide-input":true},"cell_type":"code","source":"import pandas as pd\nimport numpy as np\nimport warnings\nwarnings.filterwarnings('ignore')\n# Import Bokeh Library for output\nfrom bokeh.io import output_notebook\noutput_notebook()\nfrom bokeh.models import ColumnDataSource\nfrom bokeh.models import HoverTool\nfrom bokeh.models import LinearInterpolator,CategoricalColorMapper\nfrom bokeh.io import show\nfrom bokeh.plotting import figure\nfrom bokeh.palettes import Spectral8","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"93d40ec7b70ccfb06cb2fb2ab3efcefc8a28586f"},"cell_type":"markdown","source":"## Data Fetching","execution_count":null},{"metadata":{"trusted":true,"_uuid":"f6da77f05738b3cf88633669a90c9beffaaf3635","_kg_hide-input":true},"cell_type":"code","source":"data = pd.read_csv('../input/karnataka-state-education/Town-wise-education - Karnataka.csv')","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"cd8c506d64c55633d80f1bd1257d733ac5102614"},"cell_type":"markdown","source":"## Exploratory Data Analysis","execution_count":null},{"metadata":{"trusted":true,"_uuid":"d969627662ac0ec4635499ced704bce912a5273e","_kg_hide-input":true},"cell_type":"code","source":"data.info()","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"2d972cce11b7bb23585dc7e6b1cb053dd7a6014a"},"cell_type":"markdown","source":"Let us have a quick glance of what the data looks like by observing the first and last few rows","execution_count":null},{"metadata":{"trusted":true,"_uuid":"cf33a8e3324adac80ee7cd592939f8a7c470bf11","_kg_hide-input":true},"cell_type":"code","source":"data.head()","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"9cc19e3daed79bbf1bd3b9c481ac2b0ae35770db","_kg_hide-input":true},"cell_type":"code","source":"data.tail()","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"c5a0369ada1b5db15ccd9086eb690cc2ef2997b9"},"cell_type":"markdown","source":"Lets us examine the shape of this dataset","execution_count":null},{"metadata":{"trusted":true,"_uuid":"c0184ceb2fb0a8f6a83070172dacdc20e3688244","_kg_hide-input":true},"cell_type":"code","source":"data.shape","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"765b5aedb1752c2b23c2c45ca9ec650532dec6df"},"cell_type":"markdown","source":"This means that we have 812 dimensions(rows) and 46 features (columns) in this dataset.\nNow let us explore the data types of the dataset","execution_count":null},{"metadata":{"trusted":true,"_uuid":"c864e8593ca34a59392b7bd7e2c0b15a3c4bdb34","_kg_hide-input":true},"cell_type":"code","source":"data.info()","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"3b6c72602189cc0044d8d654a390b7d3f5200fe9"},"cell_type":"markdown","source":"The above observation shows that there are categorical and numerical features in the dataset.Let us explore further...\n\nNow let us find out how many unique categories are available from the above categorical features.\nLet us examine if there are any nulls in the dataset","execution_count":null},{"metadata":{"trusted":true,"_uuid":"44ee92f1faf41d35130c91799a1ad9e251a7b6ab","_kg_hide-input":true},"cell_type":"code","source":"data.isnull().sum()","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"f35fe48750c7ca55ce455dda64d8c41a11b7853e"},"cell_type":"markdown","source":"Let us also look at the entire metrics including the inter quartile range,mean,standard deviation for all the features\n","execution_count":null},{"metadata":{"trusted":true,"_uuid":"846c37f73042e7130ac6b5f87f91388f5ba0bbf7","_kg_hide-input":true},"cell_type":"code","source":"data.describe(include = 'all')","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"0b7db7fbe50df670a12d8dab884e3f77c59d56ba"},"cell_type":"markdown","source":"Now let us look in detail the categorical features.For this basically extract all the categorical features into a dataframe object","execution_count":null},{"metadata":{"trusted":true,"_uuid":"219dbdc8646cc6b3b756565f0d86e46be9b2f714","_kg_hide-input":true},"cell_type":"code","source":"categorical_features = data.select_dtypes(include=[np.object])\ncategorical_features.info()","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"54f2e215cef0330666192b5e48cca7450da9b3e2"},"cell_type":"markdown","source":"#### Let us observe the unique categories for all the object varables","execution_count":null},{"metadata":{"trusted":true,"_uuid":"59e4004dd95db613efd557021686dbc2c200f5ae","_kg_hide-input":true},"cell_type":"code","source":"for column_name in data.columns:\n    if data[column_name].dtypes == 'object':\n        data[column_name] = data[column_name].fillna(data[column_name].mode().iloc[0])\n        unique_category = len(data[column_name].unique())\n        print(\"Feature '{column_name}' has '{unique_category}' unique categories\".format(column_name = column_name,\n                                                                                         unique_category=unique_category))","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"078d3bad8281eaffe0e9578f96f5d484704f1356"},"cell_type":"markdown","source":"So based on the above results it is evident that 'Table Name' and 'Total/Rural/Urban' categorical features are redundant in nature which can be eliminated as it has no significance.\n\nWe can also eliminate 'State Code' as we are dealing with only Karnataka.\n\nFrom above observations we can conclude that we do not need to do any imputation as there are no missing values.\n\nBefore we jump into data visualisations and explore further let us observe the general metrics using the describe function\n\nok now let me explore more about the data in detail and come up with some basic observations","execution_count":null},{"metadata":{"_uuid":"389b9aaf19d695495fef36971c9fd6ed7c06fe7e"},"cell_type":"markdown","source":"## Data Cleansing\n\nNow let us drop some of the columns as discussed above which has no importance in our EDA such as \n- 'Table Name'\n- 'State Code'\n- 'Total/Rural/Urban'","execution_count":null},{"metadata":{"trusted":true,"_uuid":"d362ca07f9ff595145c662fbc47858316c632a2c"},"cell_type":"code","source":"data.drop('Table Name',axis =1,inplace = True)\n\ndata.drop('State Code',axis =1,inplace = True)\n\ndata.drop('Total/ Rural/ Urban',axis =1,inplace = True)\n","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"01033c0f87a33feee969b78e99eedb7d2813c660"},"cell_type":"markdown","source":"Let us further get more insights of the data by observing first few records say for a district & town ","execution_count":null},{"metadata":{"trusted":true,"_uuid":"e3cd347d53a41c9e9efd5049a1611b712d17ef95"},"cell_type":"code","source":"data.head()","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"dcd2c8ac706b9508b32f171623583bca4e7579fc"},"cell_type":"markdown","source":"#### Few Data Observations:\n\nThe general observation observed in each of the district are as follows\n- 29 unique records for age group for each town code\n    - 'All Ages' category is a summation of all ages \n- All of the below 12 categories are depicted in the form of persons ,male and female where persons is summation of male and female\n    - Illiterate\n    - Literate\n    - Educational Level - Literate without Educational Level\n    - Educational Level - Below Primary \n    - Educational Level - Primary\n    - Educational Level - Middle\n    - Educational Level - Matric/Secondary\n    - Educational Level - Higher Secondary/Intermediate Pre-University/Senior Secondary\n    - Educational Level - Non-technical Diploma or Certificate Not Equal to Degree\n    - Educational Level - Technical Diploma or Certificate Not Equal to Degree\n    - Educational Level - Graduate & Above\n    - Unclassified\nNote : Also we have 'Total - Persons', 'Total - Male','Total - Female' which do not have any significance to our analysis as these are summation of all the above categories person/male/female wise for each of the above category\n\n#### Current Focus :\n\n - Since our main focus of this exercise is to improve the 'Primary Education'. So the analysis going forward  as per my assumption that the following categories fall in the need of primary education and rest not\n    - Illiterate \n    - Educational Level - Literate without Educational Level\n    - Educational Level - Below Primary \n    - Educational Level - Primary\n    - Unclassified\n    \n   Which means that the following features can be eliminated from the dataset\n   \n    Total - Persons                                                                              \n    Total - Males                                                                                \n    Total - Females                                                                              \n    Literate - Persons                                                                           \n    Literate - Males                                                                             \n    Literate - Females                                                                           \n    Educational Level - Middle Persons                                                           \n    Educational Level - Middle Males                                                             \n    Educational Level - Middle Females                                                           \n    Educational Level - Matric/Secondary Persons                                                 \n    Educational Level - Matric/Secondary Males                                                   \n    Educational Level - Matric/Secondary Females                                                 \n    Educational Level - Higher Secondary/Intermediate Pre-University/Senior Secondary Persons    \n    Educational Level - Higher Secondary/Intermediate Pre-University/Senior Secondary Males      \n    Educational Level - Higher Secondary/Intermediate Pre-University/Senior Secondary Females    \n    Educational Level - Non-technical Diploma or Certificate Not Equal to Degree Persons         \n    Educational Level - Non-technical Diploma or Certificate Not Equal to Degree Males           \n    Educational Level - Non-technical Diploma or Certificate Not Equal to Degree Females         \n    Educational Level - Technical Diploma or Certificate Not Equal to Degree Persons             \n    Educational Level - Technical Diploma or Certificate Not Equal to Degree Males               \n    Educational Level - Technical Diploma or Certificate Not Equal to Degree Females             \n    Educational Level - Graduate & Above Persons                                                 \n    Educational Level - Graduate & Above Males                                                   \n    Educational Level - Graduate & Above Females \n    \n \n #### Key Note:\n\nFor each district code & town code we have the total of all age groups as 'All Ages' category in 'Age Group' feature .In my opinion it is irrelevant as we need to focus on age groups which need primary education.So lets go ahead and remove these rows in the dataset","execution_count":null},{"metadata":{"trusted":true,"_uuid":"e2e492a98f0f10bfdcc9b5b0aa00cce2e7edf7a3"},"cell_type":"code","source":"data = data[data['Age-Group'] != 'All ages']","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"073368aff403a185ed304c6687c34e9f63fafbb0"},"cell_type":"markdown","source":"Now let us look in detail the numerical features that need to be dropped as they do not contribute to the primary education.\n","execution_count":null},{"metadata":{"trusted":true,"_uuid":"22bcf9c781632fd02769e1a169a6b64069440bf7"},"cell_type":"code","source":"columns = [ 'Total - Persons',\n           'Total - Males',\n           'Total - Females',\n           'Literate - Persons',\n           'Literate - Males',\n           'Literate - Females',\n           'Educational Level - Middle Persons',\n           'Educational Level - Middle Males',\n           'Educational Level - Middle Females',\n           'Educational Level - Matric/Secondary Persons',\n           'Educational Level - Matric/Secondary Males',\n           'Educational Level - Matric/Secondary Females',\n           'Educational Level - Higher Secondary/Intermediate Pre-University/Senior Secondary Persons',\n           'Educational Level - Higher Secondary/Intermediate Pre-University/Senior Secondary Males',\n           'Educational Level - Higher Secondary/Intermediate Pre-University/Senior Secondary Females',\n           'Educational Level - Non-technical Diploma or Certificate Not Equal to Degree Persons',\n           'Educational Level - Non-technical Diploma or Certificate Not Equal to Degree Males',\n           'Educational Level - Non-technical Diploma or Certificate Not Equal to Degree Females',\n           'Educational Level - Technical Diploma or Certificate Not Equal to Degree Persons',\n           'Educational Level - Technical Diploma or Certificate Not Equal to Degree Males',\n           'Educational Level - Technical Diploma or Certificate Not Equal to Degree Females',\n           'Educational Level - Graduate & Above Persons',\n           'Educational Level - Graduate & Above Males',\n           'Educational Level - Graduate & Above Females']                                                                            \n\ndata.drop(columns,axis =1,inplace = True)","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"bfbf3527d4cbeb3501baae87d871be4dcab1dff3"},"cell_type":"markdown","source":"## Exploratory Data Analysis:\n\nIn this section I am going to visualize data with respect to categories focused on improving 'Primary Education'.\n\n- Illiterate \n- Educational Level - Literate without Educational Level\n- Educational Level - Below Primary \n- Educational Level - Primary\n- Unclassified\n\n I have used Bokeh interactive data visualisation library to visualise data .Please note that the mouse hover over function is enabled to see the data visualisations for each district based on the above mentioned categories depicted below .Also please note the size of the circle depicts the size of the feature .\n\nThe visualisations depict district and age wise data representations for each of the above mentioned groups including total count as well as male and female counts.\n\n#### Illeiterate :\n\nIn this section you can observe the data visualisation of the following features with respect to district and age group \n\n- Illiterate - Persons\n- Illiterate - Males\n- Illiterate - Females","execution_count":null},{"metadata":{"trusted":true,"_uuid":"bc808ae392a93f2d836280f45b42d73164570df6","_kg_hide-input":true},"cell_type":"code","source":"source = ColumnDataSource(dict(\n    x = data['District Code'],\n    y = data['Illiterate - Persons'],\n    area = data['Area Name'],\n    illerate = data['Illiterate - Persons'],\n    illerate_male = data['Illiterate - Males'],\n    illerate_female = data['Illiterate - Females'],\n    age = data['Age-Group']\n)       \n)\n\nsize_mapper = LinearInterpolator(\n    x = [data['Illiterate - Persons'].min(),data['Illiterate - Persons'].max()],\n    y = [2,100]\n)\n\ncolor_mapper = CategoricalColorMapper(\n    factors = list(data['Area Name'].unique()),\n    palette = Spectral8\n)\n\nPLOT_OPTS = dict(height = 800,width = 800,x_range = (1,30),y_range=(10,100000))\n\np = figure(title = 'Illiteracy District/Area Wise',\n           toolbar_location = 'above',\n           tools = [HoverTool(\n               tooltips = [('Area ','@area'),\n                           ('Illerate - Total Persons ','@illerate'),\n                           ('Illerate - Total Males ','@illerate_male'),\n                           ('Illerate - Total Females ','@illerate_female'),\n                           ('Age Group ','@age'),\n                        ],show_arrow = False)],\n           x_axis_label = 'District Code',\n           y_axis_label = 'No of Illiterates',\n           **PLOT_OPTS)\n\np.circle(x='x',\n         y='y', \n         size = {'field': 'illerate','transform':size_mapper},\n         color = {'field': 'area','transform':color_mapper},\n         alpha = 0.7,\n         legend = 'area',\n         source = source)\np.legend.location = (0,-50)\np.right.append(p.legend[0])\np.legend.border_line_color = None\nshow(p,notebook_handle=True)","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"51de592a344143595c43924438d74c63a91b11e8"},"cell_type":"markdown","source":"#### Educational Level - Literate without Educational Level :\n\nIn this section you can observe the data visualisation of the following features with respect to district and age group \n- Educational Level - Literate without Educational Level - Persons\n- Educational Level - Literate without Educational Level - Males\n- Educational Level - Literate without Educational Level - Females","execution_count":null},{"metadata":{"trusted":true,"_uuid":"52a6823d01e55a7e8a31b11264be780a449e17b5","_kg_hide-input":true},"cell_type":"code","source":"source = ColumnDataSource(dict(\n    x = data['District Code'],\n    y = data['Educational Level - Literate without Educational Level Persons'],\n    area = data['Area Name'],\n    illerate = data['Educational Level - Literate without Educational Level Persons'],\n    illerate_male = data['Educational Level - Literate without Educational Level Males'],\n    illerate_female = data['Educational Level - Literate without Educational Level Females'],\n    age = data['Age-Group']\n)       \n)\n\nsize_mapper = LinearInterpolator(\n    x = [data['Educational Level - Literate without Educational Level Persons'].min(),\n         data['Educational Level - Literate without Educational Level Persons'].max()],\n    y = [2,100]\n)\n\ncolor_mapper = CategoricalColorMapper(\n    factors = list(data['Area Name'].unique()),\n    palette = Spectral8\n)\n\nPLOT_OPTS = dict(height = 800,width = 800,x_range = (1,30),y_range=(10,6000))\n\np = figure(title = 'Educational Level - Literate without Educational Level (District vs. Age Wise)',\n           toolbar_location = 'above',\n           tools = [HoverTool(\n               tooltips = [('Area ','@area'),\n                           ('Educational Level - Literate without Educational Level - Total Persons ','@illerate'),\n                           ('Educational Level - Literate without Educational Level - Total Males ','@illerate_male'),\n                           ('Educational Level - Literate without Educational Level - Total Females ','@illerate_female'),\n                           ('Age Group ','@age'),\n                        ],show_arrow = False)],\n           x_axis_label = 'District Code',\n           y_axis_label = 'No of Educational Level - Literate without Educational Level',\n           **PLOT_OPTS)\n\np.circle(x='x',\n         y='y', \n         size = {'field': 'illerate','transform':size_mapper},\n         color = {'field': 'area','transform':color_mapper},\n         alpha = 0.7,\n         legend = 'area',\n         source = source)\np.legend.location = (0,-50)\np.right.append(p.legend[0])\np.legend.border_line_color = None\nshow(p,notebook_handle=True)","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"5a9c75c950f36ce9c4f0b51cad329195c6d3613e"},"cell_type":"markdown","source":"#### Educational Level - Below Primary:\n\nIn this section you can observe the data visualisation of the following features with respect to district and age group \n- Educational Level - Below Primary - Persons\n- Educational Level - Below Primary - Males\n- Educational Level - Below Primary - Females","execution_count":null},{"metadata":{"trusted":true,"_uuid":"df3745b36963d9b32747832226ec3a7e3bd097e7","_kg_hide-input":true},"cell_type":"code","source":"source = ColumnDataSource(dict(\n    x = data['District Code'],\n    y = data['Educational Level - Below Primary Persons'],\n    area = data['Area Name'],\n    illerate = data['Educational Level - Below Primary Persons'],\n    illerate_male = data['Educational Level - Below Primary Males'],\n    illerate_female = data['Educational Level - Below Primary Females'],\n    age = data['Age-Group']\n)       \n)\n\nsize_mapper = LinearInterpolator(\n    x = [data['Educational Level - Below Primary Persons'].min(),\n         data['Educational Level - Below Primary Persons'].max()],\n    y = [2,50]\n)\n\ncolor_mapper = CategoricalColorMapper(\n    factors = list(data['Area Name'].unique()),\n    palette = Spectral8\n)\n\nPLOT_OPTS = dict(height = 800,width = 800,x_range = (1,30),y_range=(10,100000))\n\np = figure(title = 'Educational Level - Below Primary (District vs. Age Wise)',\n           toolbar_location = 'above',\n           tools = [HoverTool(\n               tooltips = [('Area ','@area'),\n                           ('Educational Level - Below Primary Total Persons ','@illerate'),\n                            ('Educational Level - Below Primary Total Males ','@illerate_male'),\n                           ('Educational Level -  Below Primary Total Females ','@illerate_female'),\n                           ('Age Group ','@age'),\n                        ],show_arrow = False)],\n           x_axis_label = 'District Code',\n           y_axis_label = 'No of Educational Level - Below Primary Persons',\n           **PLOT_OPTS)\n\np.circle(x='x',\n         y='y', \n         size = {'field': 'illerate','transform':size_mapper},\n         color = {'field': 'area','transform':color_mapper},\n         alpha = 0.7,\n         legend = 'area',\n         source = source)\np.legend.location = (0,-50)\np.right.append(p.legend[0])\np.legend.border_line_color = None\nshow(p,notebook_handle=True)","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"d02bcc1d47bf38cb96999a0bd76b8fdefa842918"},"cell_type":"markdown","source":"#### Educational Level - Primary:\n\nIn this section you can observe the data visualisation of the following features with respect to district and age group \n- Educational Level - Primary - Persons\n- Educational Level - Primary - Males\n- Educational Level - Primary - Females","execution_count":null},{"metadata":{"trusted":true,"_uuid":"2fcc5f8a61d699f9b0bb7ad1664a5d48864a09e2","_kg_hide-input":true},"cell_type":"code","source":"source = ColumnDataSource(dict(\n    x = data['District Code'],\n    y = data['Educational Level - Primary Persons'],\n    area = data['Area Name'],\n    illerate = data['Educational Level - Primary Persons'],\n    illerate_male = data['Educational Level - Primary Males'],\n    illerate_female = data['Educational Level - Primary Females'],\n    age = data['Age-Group']\n)       \n)\n\nsize_mapper = LinearInterpolator(\n    x = [data['Educational Level - Primary Persons'].min(),\n         data['Educational Level - Primary Persons'].max()],\n    y = [2,50]\n)\n\ncolor_mapper = CategoricalColorMapper(\n    factors = list(data['Area Name'].unique()),\n    palette = Spectral8\n)\n\nPLOT_OPTS = dict(height = 800,width = 800,x_range = (1,30),y_range=(10,100000))\n\np = figure(title = 'Educational Level - Primary (District vs. Age Wise)',\n           toolbar_location = 'above',\n           tools = [HoverTool(\n               tooltips = [('Area ','@area'),\n                           ('Educational Level -  Primary Total Persons ','@illerate'),\n                           ('Educational Level -  Primary Total Male ','@illerate_male'),\n                           ('Educational Level -  Primary Total Female ','@illerate_female'),\n                           ('Age Group ','@age'),\n                        ],show_arrow = False)],\n           x_axis_label = 'District Code',\n           y_axis_label = 'No of Educational Level - Primary Persons',\n           **PLOT_OPTS)\n\np.circle(x='x',\n         y='y', \n         size = {'field': 'illerate','transform':size_mapper},\n         color = {'field': 'area','transform':color_mapper},\n         alpha = 0.7,\n         legend = 'area',\n         source = source)\np.legend.location = (0,-50)\np.right.append(p.legend[0])\np.legend.border_line_color = None\nshow(p,notebook_handle=True)","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"e452f2a145a9810e4fb49972f91914a7cc328890"},"cell_type":"markdown","source":"#### Unclassified:\n\nIn this section you can observe the data visualisation of the following features with respect to district and age group \n- Unclassified - Persons\n- Unclassified - Males\n- Unclassified - Females","execution_count":null},{"metadata":{"trusted":true,"_uuid":"e4828e4412efe58d68dfe8305f4f32e193fb6d3c","_kg_hide-input":true},"cell_type":"code","source":"source = ColumnDataSource(dict(\n    x = data['District Code'],\n    y = data['Unclassified - Persons'],\n    area = data['Area Name'],\n    illerate = data['Unclassified - Persons'],\n    illerate_male = data['Unclassified - Males'],\n    illerate_female = data['Unclassified - Females'],\n    age = data['Age-Group']\n)       \n)\n\nsize_mapper = LinearInterpolator(\n    x = [data['Unclassified - Persons'].min(),\n         data['Unclassified - Persons'].max()],\n    y = [1,100]\n)\n\ncolor_mapper = CategoricalColorMapper(\n    factors = list(data['Area Name'].unique()),\n    palette = Spectral8\n)\n\nPLOT_OPTS = dict(height = 800,width = 800,x_range = (1,30),y_range=(10,400))\n\np = figure(title = 'Unclassified -  (District vs. Age Wise)',\n           toolbar_location = 'above',\n           tools = [HoverTool(\n               tooltips = [('Area ','@area'),\n                           ('Unclassified - Total Persons ','@illerate'),\n                           ('Unclassified - Total Males ','@illerate_male'),\n                           ('Unclassified - Total Females ','@illerate_female'),\n                           ('Age Group ','@age'),\n                        ],show_arrow = False)],\n           x_axis_label = 'District Code',\n           y_axis_label = 'Unclassified - Persons',\n           **PLOT_OPTS)\n\np.circle(x='x',\n         y='y', \n         size = {'field': 'illerate','transform':size_mapper},\n         color = {'field': 'area','transform':color_mapper},\n         alpha = 0.7,\n         legend = 'area',\n         source = source)\np.legend.location = (0,-50)\np.right.append(p.legend[0])\np.legend.border_line_color = None\n\nshow(p,notebook_handle=True)","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"b3769ae537e571db7309e4eb8ff77107ff5f48c5"},"cell_type":"markdown","source":"Now let us look at summary total count of all categories with respect to total persons,total males & total females for each of the current features of our focus as show below.We are going to create three new features for the same namely\n- Total\n- Total_Males\n- Total_Females","execution_count":null},{"metadata":{"trusted":true,"_uuid":"29319e4ad22845fefc73f57807eb4904cb03dfbb","_kg_hide-input":true},"cell_type":"code","source":"data['Total']=data['Illiterate - Persons']+data['Educational Level - Below Primary Persons']+data['Educational Level - Literate without Educational Level Persons']+data['Educational Level - Primary Persons']+data['Unclassified - Persons']\ndata['Total_Males']=data['Illiterate - Males']+data['Educational Level - Below Primary Males']+data['Educational Level - Literate without Educational Level Males']+data['Educational Level - Primary Males']+data['Unclassified - Males']\ndata['Total_Females']=data['Illiterate - Females']+data['Educational Level - Below Primary Females']+data['Educational Level - Literate without Educational Level Females']+data['Educational Level - Primary Females']+data['Unclassified - Females']\n","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"43e03b03ff56b64e15185e9460a699764acc0cf4"},"cell_type":"markdown","source":"Now let us visualise with the above new features created to get a summary holistic view of the entire analysis.\n\n#### Summary (Total):\n\nIn this section you can observe the data visualisation of the following features with respect to district and age group \n- Total \n- Total - Males\n- Total - Females","execution_count":null},{"metadata":{"trusted":true,"_uuid":"f5e75b83edcc770d6e690f35e2d0581d6dda0bcf","_kg_hide-input":true},"cell_type":"code","source":"source = ColumnDataSource(dict(\n    x = data['District Code'],\n    y = data['Total'],\n    area = data['Area Name'],\n    illerate = data['Total'],\n    illerate_male = data['Total_Males'],\n    illerate_female = data['Total_Females'],\n    age = data['Age-Group']\n)       \n)\n\nsize_mapper = LinearInterpolator(\n    x = [data['Total'].min(),\n         data['Total'].max()],\n    y = [5,100]\n)\n\ncolor_mapper = CategoricalColorMapper(\n    factors = list(data['Area Name'].unique()),\n    palette = Spectral8\n)\n\nPLOT_OPTS = dict(height = 800,width = 800,x_range = (1,30),y_range=(10,120000))\n\np = figure(title = 'Summary -  (District vs. Age Wise)',\n           toolbar_location = 'above',\n           tools = [HoverTool(\n               tooltips = [('Area ','@area'),\n                           ('Summary','@illerate'),\n                           ('Total Males ','@illerate_male'),\n                           ('Total Females ','@illerate_female'),\n                           ('Age Group ','@age'),\n                        ],show_arrow = False)],\n           x_axis_label = 'District Code',\n           y_axis_label = 'Total Population needing Primary Education',\n           **PLOT_OPTS)\n\np.circle(x='x',\n         y='y', \n         size = {'field': 'illerate','transform':size_mapper},\n         color = {'field': 'area','transform':color_mapper},\n         alpha = 0.7,\n         legend = 'area',\n         source = source)\np.legend.location = (0,-50)\np.right.append(p.legend[0])\np.legend.border_line_color = None\nshow(p,notebook_handle=True)","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"aa8ffab33456da74168afded6ca78c95c0663368"},"cell_type":"markdown","source":"## Conclusion:\n    \n- The above summary data visualisation depicts that the target districts that need to be focused with respect to primary education as part of the Phase 1 NGO initiative\n    - Hubli Darwad \n    - Mysore\n    - Bangalore\n    - Belguam\n    - Gulbarga\n    - Bellary\n    - Davanagiri\n    - Mangalore\n- Also it is observed that age groups of 0-6 years and 30-45 years need more attention for most of the cases\nScope of improvement :\nTo perform more detailed analysis and understand the reasons and come up with a predictive model.\nDue to time constraints of doing this exercise this part is left for further exercise .","execution_count":null},{"metadata":{"trusted":true,"collapsed":true,"_uuid":"a28cfaea97ab9f8c4ee90b7942bc76c3ed32edea"},"cell_type":"markdown","source":"## If you like this  kernel Greatly Appreciate to <font color=\"red\">UPVOTE</font> .  Thank you\n\n","execution_count":null}],"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"pygments_lexer":"ipython3","nbconvert_exporter":"python","version":"3.6.4","file_extension":".py","codemirror_mode":{"name":"ipython","version":3},"name":"python","mimetype":"text/x-python"}},"nbformat":4,"nbformat_minor":4}