{"cells":[{"metadata":{"_cell_guid":"5e38bc1c-2912-4509-af7e-c40dd4d5bf7f","_uuid":"e034f3c9a31bb82ea11422611a932a2c2bfb492f"},"cell_type":"markdown","source":"# Interactive Data Exploration, Analysis, and Reporting (IDEAR) in Python for Azure Notebooks\n\nAndrey Vykhodtsev: I decided it would be fun to run a slightly customized [TDSP IDEAR script](https://github.com/Azure/Azure-TDSP-Utilities/blob/master/DataScienceUtilities/DataReport-Utils/Python/IDEAR-Python-Instructions-JupyterNotebook.md) on this dataset. All credits go to the authors.\nNOTE: Looks like Kaggle kernels do not IPython widgets. So if you want to run interactive widgets you have to either fork and run it or  download this notebook and run it locally.\n\n- Author: Team Data Science Process Team from Microsoft \n- Date: 2017/03\n\n\nThis is the **Interactive Data Exploration, Analysis and Reporting (IDEAR)** in _**Python**_ running on [Azure Notebooks](https://notebooks.azure.com/). The data can be CSV files in Azure Blob Storage. A yaml file has to be pre-configured and saved to Azure Blob Storage before running this tool.\n\n## Running IDEAR Python using web browser in Azure Notebooks\nThis tool provides various functionalities to help users explore the data and get insights through interactive visualization and statistical testing using web browser in Azure Notebooks. You can run all the cells at once, when prompt, enter the yaml file and data file names.\n\n- [Read and Summarize the data](#read and summarize)\n\n- [Extract Descriptive Statistics of Data](#descriptive statistics)\n\n- [Explore Individual Variables](#individual variables)\n\n- [Explore Interactions between Variables](#multiple variables)\n\n    - [Rank variables](#rank variables)\n    \n    - [Interaction between two categorical variables](#two categorical)\n    \n    - [Interaction between two numerical variables](#two numerical)\n\n    - [Interaction between numerical and categorical variables](#numerical and categorical)\n\n    - [Interaction between two numerical variables and a categorical variable](#two numerical and categorical)\n\n- [Visualize High Dimensional Data via Projecting to Lower Dimension Principal Component Spaces](#pca)\n\nAfter you are done with exploring the data interactively, you can choose to [show/hide the source code](#show hide codes) to make your notebook look neater. You may also download the analysis results as different formats such as PDF, HTML, markdown, etc. \n\n**Note**:\n\n- Change the yaml file and save it to Azure blob storage before running IDEAR in Azure Notebooks."},{"metadata":{"_cell_guid":"d303ab62-05d5-48d4-afca-04990118f6e2","_uuid":"3868e3b00f7d41ca3638964840ff5aba5e4d4194"},"cell_type":"markdown","source":"## <a name=\"setup\"></a> Global configuration and set up"},{"metadata":{"_cell_guid":"227588b9-41d2-4165-830c-29aa1f1691a2","_uuid":"6f475e0844fa4e5e3f98ffbc70c29766c0574a10","_kg_hide-input":true,"trusted":true},"cell_type":"code","source":"import collections\nfrom IPython.core.display import HTML\nfrom IPython.display import display\nimport matplotlib.pyplot as plt\nimport matplotlib\nimport scipy.stats as stats\nfrom statsmodels.graphics.mosaicplot import mosaic\nimport statsmodels.api as sm\nfrom statsmodels.formula.api import ols\nimport pandas as pd\nimport numpy as np\nimport scipy\nimport matplotlib.pyplot as plt\nimport scipy.stats as stats\nfrom functools import partial\nimport IPython\nimport ipywidgets\nfrom ipywidgets import widgets\nfrom ipywidgets import interact, interactive,fixed\nimport operator\nfrom IPython.display import Javascript, display,HTML\nfrom ipywidgets import widgets, VBox\nimport seaborn as sns\nfrom collections import OrderedDict\nimport yaml\nimport warnings\nimport getpass\nimport sys\nwarnings.filterwarnings('ignore')\n\nclass ConfUtility():   \n    @staticmethod\n    def parse_yaml(input_file):\n        import yaml\n        yaml_dict = {}\n        with open (input_file,'r') as fin:\n            try:\n                yaml_dict = yaml.load(fin)\n            except Exception as ex:\n                print (ex)\n        return yaml_dict\n\n    @staticmethod\n    def dict_to_htmllist(dc, include_list=None):\n        dc2 = {}\n        output_formatting = {'Target':'Target variable is ','CategoricalColumns':'Categorical Columns are ',\n                           'NumericalColumns':'Numerical Columns are '}\n        for each in dc.keys():\n            if not include_list or each in include_list:\n                if isinstance(dc[each],  collections.Iterable) and not isinstance(dc[each], str):\n                    dc2[each] = ', \\n'.join(val for val in dc[each])\n                else:\n                    dc2[each] = dc[each]\n        html_list = \"<ul>{}</ul>\"\n        html_list_entry = \"<li>{}</li>\"\n        output3 = ''\n\n        for each in set(include_list)|set(dc2.keys()):\n            output3 += html_list_entry.format(output_formatting[each]+dc2[each])\n        html_list = html_list.format(output3)\n        return HTML(html_list)\n    \nclass InteractionAnalytics():\n    @staticmethod\n    def rank_associations(df, conf_dict, col1, col2, col3):        \n        try:\n            col2 = int(col2)\n            col3 = int(col3)\n        except:\n            pass\n        \n        # Passed Variable is Numerical\n        if (col1 in conf_dict['NumericalColumns']) :\n            fig,(ax1,ax2) = plt.subplots(1, 2)\n            if len(conf_dict['NumericalColumns'])>1:\n                \n                # Interaction with numerical variables\n                df2 = df[conf_dict['NumericalColumns']]\n                corrdf = df2.corr()\n                corrdf = abs(corrdf) \n                corrdf2 = corrdf[corrdf.index==col1].reset_index()[[each for each in corrdf.columns \\\n                                                      if col1 not in each]].unstack().sort_values(kind=\"quicksort\", \n                                                                                                  ascending=False).head(col2)\n                corrdf2 = corrdf2.reset_index()\n                corrdf2.columns = ['level0','level1','rsq']\n                corrdf2.set_index('level0', inplace=True)\n                corrdf2[['rsq']].plot(kind='bar', ax=ax1)\n                ax1.legend().set_visible(False)\n                ax1.set_xlabel('Absolute Correlation')\n                ax1.set_title('Top {} Associated Numeric Variables'.format(str(col2)))\n                # Interaction with categorical variables\n                etasquared_dict = {}\n            if len(conf_dict['CategoricalColumns']) >= 1:\n                for each in conf_dict['CategoricalColumns']:\n                    mod = ols('{} ~ C({})'.format(col1, each),data=df[[col1,each]],missing='drop').fit()\n                    aov_table = sm.stats.anova_lm(mod, typ=1)\n                    esq_sm = aov_table['sum_sq'][0]/(aov_table['sum_sq'][0]+aov_table['sum_sq'][1])\n                    etasquared_dict[each] = esq_sm\n\n                topk_esq = pd.DataFrame.from_dict(etasquared_dict, orient='index').unstack().sort_values(\\\n                    kind = 'quicksort', ascending=False).head(col3).reset_index().set_index('level_1')\n                topk_esq.columns = ['level_0', 'EtaSquared']\n                topk_esq[['EtaSquared']].plot(kind='bar',ax=ax2)\n                ax2.legend().set_visible(False)\n                ax2.set_xlabel('Eta-squared values')\n                ax2.set_title('Top {}  Associated Categoric Variables'.format(str(col2)))\n        # Passed Variable is Categorical\n        else:\n            #Interaction with numerical variables\n            fig,(ax1,ax2) = plt.subplots(1,2)\n            if len(conf_dict['NumericalColumns']) >= 1:\n                etasquared_dict = {}\n                for each in conf_dict['NumericalColumns']:\n                    mod = ols('{} ~ C({})'.format(each, col1), data = df[[col1,each]]).fit()\n                    aov_table = sm.stats.anova_lm(mod, typ=1)\n                    esq_sm = aov_table['sum_sq'][0]/(aov_table['sum_sq'][0]+aov_table['sum_sq'][1])\n                    etasquared_dict[each] = esq_sm\n\n                topk_esq = pd.DataFrame.from_dict(etasquared_dict, orient='index').unstack().sort_values(\\\n                    kind = 'quicksort', ascending=False).head(col2).reset_index().set_index('level_1')\n                topk_esq.columns = ['level_0','EtaSquared']\n                topk_esq[['EtaSquared']].plot(kind='bar',ax=ax1)\n                ax1.legend().set_visible(False)\n                ax1.set_xlabel('Eta-squared values')\n                ax1.set_title('Top {} Associated Numeric Variables'.format(str(col2)))\n\n            # Interaction with categorical variables\n            cramer_dict = {}\n            if len(conf_dict['CategoricalColumns'])>1:\n                for each in conf_dict['CategoricalColumns']:\n                    if each !=col1:\n                        tbl = pd.crosstab(df[col1], df[each])\n                        chisq = stats.chi2_contingency(tbl, correction=False)[0]\n                        try:\n                            cramer = np.sqrt(chisq/sum(tbl))\n                        except:\n                            cramer = np.sqrt(chisq/tbl.as_matrix().sum())\n                            pass\n                        cramer_dict[each] = cramer\n\n                topk_cramer = pd.DataFrame.from_dict(cramer_dict, orient='index').unstack().sort_values(\\\n                    kind = 'quicksort', ascending=False).head(col3).reset_index().set_index('level_1')\n                topk_cramer.columns = ['level_0','CramersV']\n                topk_cramer[['CramersV']].plot(kind='bar',ax=ax2)\n                ax2.legend().set_visible(False)\n                ax2.set_xlabel(\"Cramer's V\")\n                ax2.set_title('Top {} Associated Categoric Variables'.format(str(col2)))\n        \n    @staticmethod\n    def NoLabels(x):\n        return ''\n    \n    @staticmethod\n    def categorical_relations(df, col1, col2):\n        if col1 != col2:\n            df2 = df[(df[col1].isin(df[col1].value_counts().head(10).index.tolist()))&(df[col2].isin(df[col2].value_counts().head(10).index.tolist())) ]\n            df3 = pd.crosstab(df2[col1], df2[col2])\n            df3 = df3+1e-8\n        else:\n            df3 = pd.DataFrame(df[col1].value_counts())[:10]\n        fig,ax = plt.subplots()\n        fig,rects = mosaic(df3.unstack(),ax=ax, statistic=False, labelizer=InteractionAnalytics.NoLabels, label_rotation=30)\n        ax.set_ylabel(col1)\n        ax.set_xlabel(col2)\n        ax.set_title('{} vs {}'.format(col1, col2) )\n    \n    @staticmethod\n    def numerical_relations(df, col1, col2):\n        from statsmodels.nonparametric.smoothers_lowess import lowess\n        x = df[col2]\n        y = df[col1]\n        f, ax = plt.subplots(1)\n\n        # lowess\n        ax.scatter(x, y, c='g', s=6)\n        lowess_results = lowess(y, x)#[:,1]\n        xs = lowess_results[:, 0]\n        ys = lowess_results[:, 1]\n        ax.plot(xs,ys,'red',linewidth=1)\n\n        #ols\n        fit = np.polyfit(x, y, 1)\n        fit1d = np.poly1d(fit)\n        ax.plot(x, fit1d(x), '--b')\n        ax.set_xlabel(col2)\n        ax.set_ylabel(col1)\n        corr = round(scipy.stats.pearsonr(x, y)[0], 6)\n        ax.set_title('{} vs {}, Correlation {}'.format(col1, col2, corr))\n    \n    @staticmethod\n    def numerical_correlation(df, conf_dict, col1):\n        from matplotlib.pyplot import quiver, colorbar, clim,  matshow\n        df2 = df[conf_dict['NumericalColumns']].corr(method=col1)\n        col_names = list(df[conf_dict['NumericalColumns']].columns)\n        fig,ax = plt.subplots(1, 1)\n        m = ax.matshow(df2, cmap=matplotlib.pyplot.cm.coolwarm)\n        ax.grid(b=False)\n        fig.colorbar(m)\n        ax.set_xticklabels([' '] + col_names) \n        ax.set_yticklabels([' '] + col_names)\n\n    @staticmethod\n    def numerical_pca(df, conf_dict, col1, col2, col3):\n        from sklearn.decomposition import PCA\n        from sklearn.preprocessing import StandardScaler\n        num_numeric = len(conf_dict['NumericalColumns'])\n        num_pca = num_numeric\n        xticklabels = ['']\n        for i in range(1,num_pca+1):\n            xticklabels+=['Comp'+str(i)]\n            xticklabels+=['']\n        df2 = df[conf_dict['NumericalColumns']]\n        X = StandardScaler().fit_transform(df2.values)\n        pca = PCA(n_components=num_pca)\n        pca.fit(X)\n        fig, (ax1,ax2) = plt.subplots(1, 2)\n        ax1.bar(np.arange(1,(num_numeric+1),1),pca.explained_variance_ratio_ )\n        ax1.set_ylabel('% Variance Explained')\n        ax1.set_xticklabels(xticklabels)\n        x_pca_index = int(col2) - 1\n        y_pca_index = int(col3) - 1\n        Y_pca = pd.DataFrame(pca.fit_transform(X))\n        Y_pca_labels = []\n        for i in range(1,num_pca+1):\n            Y_pca_labels.append('PC'+str(i))\n        Y_pca.columns = Y_pca_labels       \n        Y_pca[col1] = df[col1]\n        colors_dict = {}\n        colors_list = ['r', 'y', 'c', 'y', 'k']\n        j = 0\n        for i in np.unique(df[col1]):\n            colors_dict[i] = colors_list[j]\n            j += 1\n            if j == len(colors_list):\n                j = 0\n        colordf = pd.DataFrame.from_dict(colors_dict, orient='index').reset_index()\n        colordf.columns = [col1, 'color']\n        merged_df = pd.merge(colordf,Y_pca)\n        grouped_df = merged_df.groupby(col1)\n        for name, group in grouped_df:\n            ax2.scatter(\n               group[Y_pca.columns[x_pca_index]], group[Y_pca.columns[y_pca_index]],label=name,  \n               c=group['color'],                            \n               marker='o',                                \n               s=6)                                       \n        ax2.set_xlabel(Y_pca.columns[x_pca_index])\n        ax2.set_ylabel(Y_pca.columns[y_pca_index])\n        ax2.legend(title=col1, fontsize=14)\n                \n    @staticmethod\n    def nc_relation(df, conf_dict, col1, col2, col3=None):\n        fig,ax = plt.subplots()\n        f = df[[col1,col2]].boxplot(by=col2, ax=ax)\n        mod = ols('{} ~ {}'.format(col1, col2), data=df[[col1, col2]]).fit()\n        aov_table = sm.stats.anova_lm(mod, typ=1)\n        p_val = round(aov_table['PR(>F)'][0], 6)\n        status = 'Passed'\n        color = 'blue'\n        if p_val < 0.05:\n            status = 'Rejected'\n            color = 'red'\n        fig.suptitle('ho {} (p_value = {})'.format( status, p_val), color=color, fontsize=10)\n    \n    @staticmethod\n    def pca_3d(df, conf_dict, col1, col2,  col3=None):\n        from sklearn.decomposition import PCA\n        from sklearn.preprocessing import StandardScaler\n        from mpl_toolkits.mplot3d import Axes3D\n        df2 = df[conf_dict['NumericalColumns']]\n        X = StandardScaler().fit_transform(df2.values)\n        pca = PCA(n_components=4)\n        pca.fit(X)\n        fig = plt.figure()\n        ax = fig.gca(projection='3d')\n        ax.view_init(elev=10, azim=int(col2))              \n        Y_pca = pd.DataFrame(pca.fit_transform(X))\n        Y_pca.columns = ['PC1','PC2','PC3','PC4']\n        Y_pca[col1] = df[col1]\n        colors_dict = {}\n        colors_list = ['r', 'y', 'c', 'y', 'k']\n        j = 0\n        for i in np.unique(df[col1]):\n            colors_dict[i] = colors_list[j]\n            j += 1\n            if j == len(colors_list):\n                j = 0\n        colordf = pd.DataFrame.from_dict(colors_dict, orient='index').reset_index()\n        colordf.columns = [col1,'color']\n        merged_df = pd.merge(colordf,Y_pca)\n        grouped_df = merged_df.groupby(col1)\n        for name, group in grouped_df:\n            ax.scatter(\n               group['PC1'], group['PC2'], group['PC3'], label=name,  \n               c = group['color'],                            \n               marker = 'o',                                \n               s=6)                                      \n        ax.set_xlabel('PC1', labelpad=18)\n        ax.set_ylabel('PC2', labelpad=18)\n        ax.set_zlabel('PC3', labelpad=18)\n        ax.legend(title=col1, fontsize=10)\n\n    @staticmethod\n    def pca_3d_new(df, conf_dict, col1, col2, col3, col4, col5):\n        from sklearn.decomposition import PCA\n        from sklearn.preprocessing import StandardScaler\n        from mpl_toolkits.mplot3d import Axes3D\n        df2 = df[conf_dict['NumericalColumns']]\n        X = StandardScaler().fit_transform(df2.values)\n        num_numeric = len(conf_dict['NumericalColumns'])\n        pca = PCA(n_components=num_numeric)\n        pca.fit(X)\n        fig = plt.figure()\n        ax = fig.gca(projection='3d')\n        ax.view_init(elev=10, azim=int(col5))                 \n        Y_pca = pd.DataFrame(pca.fit_transform(X))\n        Y_pca_names = []\n        for i in range(1, num_numeric+1):\n            Y_pca_names.append('PC'+str(i))\n        Y_pca.columns = Y_pca_names\n        Y_pca[col1] = df[col1]\n        colors_dict = {}\n        colors_list = ['r', 'y', 'c', 'y', 'k']\n        j = 0\n        for i in np.unique(df[col1]):\n            colors_dict[i] = colors_list[j]\n            j += 1\n            if j == len(colors_list):\n                j = 0\n        colordf = pd.DataFrame.from_dict(colors_dict, orient='index').reset_index()\n        colordf.columns = [col1,'color']\n        merged_df = pd.merge(colordf,Y_pca)\n        grouped_df = merged_df.groupby(col1)\n        for name, group in grouped_df:\n            ax.scatter(\n               group[Y_pca_names[int(col2)-1]], group[Y_pca_names[int(col3)-1]], group[Y_pca_names[int(col4)-1]], label=name,  \n               c = group['color'],                            \n               marker = 'o',                                \n               s=6)\n        ax.set_xlabel(Y_pca_names[int(col2)-1], labelpad=18)\n        ax.set_ylabel(Y_pca_names[int(col3)-1], labelpad=18)\n        ax.set_zlabel(Y_pca_names[int(col4)-1], labelpad=18)\n        ax.legend(title=col1, fontsize=10)\n        \n    @staticmethod\n    def nnc_relation(df, conf_dict, col1, col2, col3):\n        import itertools\n        markers = ['x', 'o', '^']\n        color = itertools.cycle(['r', 'y', 'c', 'y', 'k']) \n        groups = df[[col1, col2, col3]].groupby(col3)\n\n        # Plot\n        fig, ax = plt.subplots()\n        ax.margins(0.05) \n\n        for (name, group), marker in zip(groups, itertools.cycle(markers)):\n            ax.plot(group[col1], group[col2], marker='o', linestyle='', ms=4, label=name)\n        ax.set_xlabel(col1)\n        ax.set_ylabel(col2)\n        ax.legend(numpoints=1, loc='best', title=col3)\n        \nclass TargetAnalytics():\n    ReportedVariables = []\n    @staticmethod\n    def custom_barplot(df, col1=''):\n        f, (ax0,ax1) = plt.subplots(1, 2)\n        df[col1].value_counts().plot(ax=ax0, kind='bar')\n        ax0.set_title('Bar Plot of {}'.format(col1))\n        df[col1].value_counts().plot(ax=ax1, kind='pie')\n        ax1.set_title('Pie Chart of {}'.format(col1))\n\nclass NumericAnalytics():\n    @staticmethod\n    def shapiro_test(x):\n        p_val = round(stats.shapiro(x)[1],6)\n        status = 'passed'\n        color = 'blue'\n        if p_val < 0.05:\n            status = 'failed'\n            color = 'red'\n        return status, color, p_val\n\n    @staticmethod\n    def custom_barplot(df, col1=''):\n        fig, axes = plt.subplots(2,2)\n        axes = axes.reshape(-1)\n        df[col1].plot(ax=axes[0], kind='hist')\n        axes[0].set_title('Histogram of {}'.format(col1))\n        df[col1].plot(ax=axes[1], kind='kde')\n        axes[1].set_title('Density Plot of {}'.format(col1))\n        ax3 = plt.subplot(223)\n        stats.probplot(df[col1], plot=plt)\n        axes[2].set_title('QQ Plot of {}'.format(col1))\n        df[col1].plot(ax=axes[3], kind='box')\n        axes[3].set_title('Box Plot of {}'.format(col1))\n        status, color, p_val = NumericAnalytics.shapiro_test(df[col1])\n        fig.suptitle('Normality test for {} {} (p_value = {})'.format(col1, status, round(p_val, 6)), color=color, fontsize=12)\n    \nclass CategoricAnalytics():\n    @staticmethod\n    def custom_barplot(df, col1=''):\n        f, (ax0,ax1) = plt.subplots(1,2)\n        df[col1].value_counts().nlargest(10).plot(ax=ax0, kind='bar')\n        ax0.set_xlabel(col1)\n        ax0.set_title('Bar chart of {}'.format(col1))\n        df[col1].value_counts().nlargest(10).plot(ax=ax1, kind='pie')\n        ax1.set_title('Pie chart of {}'.format(col1))\n \n%matplotlib inline\nfont={'family':'normal','weight':'normal','size':8}\nmatplotlib.rc('font',**font)\nmatplotlib.rcParams['figure.figsize'] = (20.0, 10.0)\nmatplotlib.rc('xtick', labelsize=9) \nmatplotlib.rc('ytick', labelsize=9)\nmatplotlib.rc('axes', labelsize=10)\nmatplotlib.rc('axes', titlesize=10)\nsns.set_style('whitegrid')","execution_count":1,"outputs":[]},{"metadata":{"_cell_guid":"8bfde7a1-c8f4-4463-a32b-deb247a9ef88","_uuid":"955d0017b66b22339a2ad20531b1afb62944d8a1"},"cell_type":"markdown","source":"### Specify and load yaml file"},{"metadata":{"_cell_guid":"f07b9608-6e01-46cf-bd84-11f85ccb1183","_uuid":"a403d0b1db9856bb7dc87ee245c928ddc9fdb827","collapsed":true,"trusted":true},"cell_type":"code","source":"yaml_conf = '''\\\nTarget:\n    deal_probability \nCategoricalColums:\n    - region\n    - city\n    - parent_category_name\n    - category_name\n    - param_1\n    - param_2\n    - param_3\n    - image_top_1\n    - user_type\nNumericalColumns:\n    - price\n    - deal_probability\nColumnsToExclude:\n    - item_id\n    - user_id\n    - title\n    - description\n    - image\n    - activation_date\n'''\n\nconf_dict = yaml.load(yaml_conf)","execution_count":2,"outputs":[]},{"metadata":{"_cell_guid":"99f8a735-21ff-4ffe-8a49-61a7d0c9ad10","_uuid":"7bcbc991faab0ea036b2b5378f23d0591a78f603"},"cell_type":"markdown","source":"## <a name=\"read and summarize\"></a> Read and Summarize the Data"},{"metadata":{"_cell_guid":"3de4c876-2960-409a-aa24-f61b338ca7f3","_uuid":"48a9e5ae750d1f4e8e5d75d0b8f6c1e724a824f3","collapsed":true},"cell_type":"markdown","source":"* ### Load data"},{"metadata":{"_cell_guid":"e2d0ec7d-0689-4034-835d-dbfac51c628c","_uuid":"d0661e29f0d85a272b70ef18d27de0a936442ef9","collapsed":true,"trusted":true},"cell_type":"code","source":"#Read in data as text\ndata_file = \"../input/train.csv\"\n\n#convert to df\ndf = pd.read_csv(data_file)\n\n#change numeric columns to numeric\ndf[conf_dict['NumericalColumns']] = df[conf_dict['NumericalColumns']].apply(pd.to_numeric)","execution_count":3,"outputs":[]},{"metadata":{"_cell_guid":"9ac1cafc-fb3f-4958-8013-34aa086e7c1b","_uuid":"8d81b3e9a5c02100fc75197f3ebabb2efb87256e"},"cell_type":"markdown","source":"### Read data and infer column types"},{"metadata":{"_cell_guid":"d9e9b8f9-ec47-473a-bf84-128659ee51b0","_uuid":"e9a35ba2b836807dd4775441eec907dd71fa0864","trusted":true},"cell_type":"code","source":"# Define sample size\nSample_Size = 10000\n    \n# Making sure that we are not reading any extra column\ndf = df[[each for each in df.columns if 'Unnamed' not in each]]\n\n# Sampling Data if data size is larger than 10k\ndf0 = df # df0 is the unsampled data. Will be used in data exploration and analysis where sampling is not needed\n         # However, keep in mind that your final report will always be based on the sampled data. \nif Sample_Size < df.shape[0]:\n    df = df.sample(Sample_Size)\n\n# Getting the list of categorical columns if it was not there in the yaml file\nif 'CategoricalColumns' not in conf_dict:\n    conf_dict['CategoricalColumns'] = list(set(list(df.select_dtypes(exclude=[np.number]).columns)))\n\n# Getting the list of numerical columns if it was not there in the yaml file\nif 'NumericalColumns' not in conf_dict:\n    conf_dict['NumericalColumns'] = list(df.select_dtypes(include=[np.number]).columns)    \n\n# Exclude columns that we do not need\nif 'ColumnsToExclude' in conf_dict:\n    conf_dict['CategoricalColumns'] = list(set(conf_dict['CategoricalColumns'])-set(conf_dict['ColumnsToExclude']))\n    conf_dict['NumericalColumns'] = list(set(conf_dict['NumericalColumns'])-set(conf_dict['ColumnsToExclude']))\n\n# Ordering the categorical variables according to the number of unique categories\nfiltered_cat_columns = []\ntemp_dict = {}\n\nfor cat_var in conf_dict['CategoricalColumns']:\n    temp_dict[cat_var] = df[cat_var].nunique()\n\nsorted_x = sorted(temp_dict.items(), key=operator.itemgetter(0), reverse=True)\nconf_dict['CategoricalColumns'] = [x for (x,y) in sorted_x]\nConfUtility.dict_to_htmllist(conf_dict,['Target','CategoricalColumns','NumericalColumns'])","execution_count":4,"outputs":[]},{"metadata":{"_cell_guid":"bd8706b2-7179-4a7a-82cb-5ba45c798de5","_uuid":"13e12fa101f4e3d93799d25beb9ac42e366bbc5f","collapsed":true},"cell_type":"markdown","source":"### Print the first n (n=5 by default) rows of the data"},{"metadata":{"_cell_guid":"fc375cca-59b9-4a12-ae02-c2870fddf368","_uuid":"c3d3a86339390ce28f89bd5f37f4ec21f5fc5cb2","trusted":true},"cell_type":"code","source":"def custom_head(df,NoOfRows):\n    return HTML(df.head(NoOfRows).style.set_table_attributes(\"class='table'\").render())\ni = interact(custom_head,df=fixed(df0), NoOfRows=ipywidgets.IntSlider(min=0, max=30, step=1, value=5, description='Number of Rows'))","execution_count":5,"outputs":[]},{"metadata":{"_cell_guid":"a9df6b8e-fafe-4275-9507-529ebd7351ad","_uuid":"26c05e63a2dd42af6dddd5c1a4471c1a75308b3b","collapsed":true},"cell_type":"markdown","source":"### Print the dimensions of the data (rows, columns)"},{"metadata":{"_cell_guid":"1fd0bd1c-ed4c-498c-a645-75cf0ba6fc5e","_uuid":"6663b17f240a9e20872f72428ee51ce25f637777","trusted":true},"cell_type":"code","source":"print ('The data has {} Rows and {} columns'.format(df0.shape[0],df0.shape[1]))","execution_count":6,"outputs":[]},{"metadata":{"_cell_guid":"632f96ff-d3bb-4ee3-a1ae-12170fc4a347","_uuid":"67d6ba02dd21fe63ea310074b361213cbd7e77cb","collapsed":true},"cell_type":"markdown","source":"### Print the column names of the data"},{"metadata":{"_cell_guid":"ca4cb76a-00c6-49d4-9373-96ac7ff98e87","_uuid":"62c77382071e5991976ae7258af5800d41ede881","trusted":true},"cell_type":"code","source":"col_names = ','.join(each for each in list(df.columns))\nprint(\"The column names are:\" + col_names)","execution_count":7,"outputs":[]},{"metadata":{"_cell_guid":"432c9480-74c2-43bd-bb07-5de4d7b8ed73","_uuid":"66fd9625eaf97f2de7303c53acd1b39d302b734f","collapsed":true},"cell_type":"markdown","source":"### Print the column types"},{"metadata":{"_cell_guid":"789d0e3c-8755-4737-bb53-96396c7efc75","_uuid":"45848018e066d869ed442e7773c46624d0eead10","trusted":true},"cell_type":"code","source":"print(\"The types of columns are:\")\ndf.dtypes","execution_count":8,"outputs":[]},{"metadata":{"_cell_guid":"9f5a3115-9baf-484e-b24a-7ef63e4dbf2a","_uuid":"da3f7a44176446d99a2ed48c74b5b05dbd889051","collapsed":true},"cell_type":"markdown","source":"## <a name=\"descriptive statistics\"></a>Extract Descriptive Statistics of Each Column"},{"metadata":{"_cell_guid":"012088b0-1135-4138-a3a3-37533af632f1","collapsed":true,"_uuid":"6c60e5ac7ab13887a97faf2b165d1a24d3a5f898","trusted":true},"cell_type":"code","source":"def num_missing(x):\n    return len(x.index)-x.count()\n\ndef num_unique(x):\n    return x.nunique()\n\ntemp_df = df0.describe().T\nmissing_df = pd.DataFrame(df0.apply(num_missing, axis=0)) \nmissing_df.columns = ['missing']\nunq_df = pd.DataFrame(df0.apply(num_unique, axis=0))\nunq_df.columns = ['unique']\ntypes_df = pd.DataFrame(df0.dtypes)\ntypes_df.columns = ['DataType']","execution_count":9,"outputs":[]},{"metadata":{"_cell_guid":"54ec6fe4-e888-4ad9-8f1f-06ee0aac92c4","_uuid":"e5a5c2d0b5144a608b939588af4c5e506cd719c4","collapsed":true},"cell_type":"markdown","source":"### Print the descriptive statistics of numerical columns"},{"metadata":{"_cell_guid":"449d9cf5-bd50-4509-81e2-79a3d6590df5","_uuid":"4b4c4c61c398e487f34883c1897af6fd9dad8eea","trusted":true},"cell_type":"code","source":"summary_df = temp_df.join(missing_df).join(unq_df).join(types_df)\nsummary_df","execution_count":10,"outputs":[]},{"metadata":{"_cell_guid":"4c3b16c8-feb3-498e-a788-55419ca67c3e","_uuid":"7932f0fa7e62ba2523e5a65036dc3df30fb7b959","collapsed":true},"cell_type":"markdown","source":"### Print the descriptive statistics of categorical columns"},{"metadata":{"_cell_guid":"f0b6a25d-d4b0-469f-972d-10a1f3ab286d","_uuid":"b2e2af59f205db2cbe1caf4bf6f1f7364c0038b7","trusted":true},"cell_type":"code","source":"col_names = list(types_df.index)\nnum_cols = len(col_names)\nindex = range(num_cols)\ncat_index = []\nfor i in index:\n    if col_names[i] in conf_dict['CategoricalColumns']:\n        cat_index.append(i)\nsummary_df_cat = missing_df.join(unq_df).join(types_df.iloc[cat_index], how='inner') #Only summarize categorical columns\nsummary_df_cat","execution_count":11,"outputs":[]},{"metadata":{"_cell_guid":"c3ca062d-8811-4398-b63c-cbfd4a2ad0dc","_uuid":"3a863f0c7f5063fe5b3f8382e12364d8010f7678","collapsed":true},"cell_type":"markdown","source":"## <a name=\"individual variables\"></a>Explore Individual Variables"},{"metadata":{"_cell_guid":"9140a46d-73bf-4997-ab17-35dc36ffc034","_uuid":"e35ce09ab44dc6ecf0d78b8db762a5fce11d2dba"},"cell_type":"markdown","source":"### Explore the target variable"},{"metadata":{"_cell_guid":"2f7acde7-5493-4879-92fe-ae0a9be90ee5","_uuid":"f97f2a68afa1e7b169fdeee0c40ea1efda966aae","trusted":true},"cell_type":"code","source":"if conf_dict['Target'] in conf_dict['CategoricalColumns']:\n    w1_value = ''\n    w1 = None  \n    w1 = widgets.Dropdown(\n        options=[conf_dict['Target']],\n        value=conf_dict['Target'],\n        description='Target Variable:',\n    )\n    i = interactive(TargetAnalytics.custom_barplot, df=fixed(df), col1=w1)\n    hbox = widgets.HBox(i.children[:1])\n    display(hbox)\n    hbox.on_displayed(TargetAnalytics.custom_barplot(df=df0, col1=w1.value))\nelse:\n    w1_value = ''\n    w1 = None\n    w1 = widgets.Dropdown(\n            options=[conf_dict['Target']],\n            value=conf_dict['Target'],\n            description='Target Variable:',\n        )\n    i = interactive(NumericAnalytics.custom_barplot, df=fixed(df), col1=w1)\n    hbox = widgets.HBox(i.children)\n    display(hbox)\n    hbox.on_displayed(NumericAnalytics.custom_barplot(df=df, col1=w1.value))","execution_count":12,"outputs":[]},{"metadata":{"_cell_guid":"65cf0621-af62-4759-ab83-05424ef2387a","_uuid":"c36724eaf88cbcb5d9877fc2530687796d6cac2a","collapsed":true},"cell_type":"markdown","source":"### Explore individual numeric variables and test for normality (on sampled data)"},{"metadata":{"_cell_guid":"f1835c28-ee73-4411-ba5b-081befa23144","_uuid":"2e0f816ac9afd9cca77c5d5543ee45cad616b6b6","trusted":true},"cell_type":"code","source":"w1_value = ''\nw1 = None\nw1 = widgets.Dropdown(\n        options=conf_dict['NumericalColumns'],\n        value=conf_dict['NumericalColumns'][0],\n        description='Numeric Variable:',\n    )\n\ni = interactive(NumericAnalytics.custom_barplot, df=fixed(df), col1=w1)\nhbox = widgets.HBox(i.children)\ndisplay(hbox)\nhbox.on_displayed(NumericAnalytics.custom_barplot(df=df, col1=w1.value))","execution_count":13,"outputs":[]},{"metadata":{"_cell_guid":"7c9a0510-5a20-4b6d-9c29-503dd9aa8aba","_uuid":"5d344293e51d3bc0d93ee4b46c33119f13304de8","collapsed":true},"cell_type":"markdown","source":"### Explore individual categorical variables (sorted by frequencies)"},{"metadata":{"_cell_guid":"21b35d9a-54cb-48c7-bf8b-4e224e28c22f","_uuid":"1a85fd72a8f0b6af0b12d923a00585f810041911","trusted":true},"cell_type":"code","source":"w1_value = ''\nw1 = None\n\nw1 = widgets.Dropdown(\n    options = conf_dict['CategoricalColumns'],\n    value = conf_dict['CategoricalColumns'][0],\n    description = 'Categorical Variable:',\n)\n\n\ni = interactive(CategoricAnalytics.custom_barplot, df=fixed(df), col1=w1)\nhbox = widgets.HBox(i.children)\ndisplay(hbox)\nhbox.on_displayed(CategoricAnalytics.custom_barplot(df=df0, col1=w1.value))","execution_count":14,"outputs":[]},{"metadata":{"_cell_guid":"184f6bd4-04a9-40a8-abe2-3cb10dc648c8","_uuid":"ccca2f82240d74520db063beaf09cc8e49837cfe","collapsed":true},"cell_type":"markdown","source":"## <a name=\"multiple variables\"></a>Explore Interactions Between Variables"},{"metadata":{"_cell_guid":"38d58f26-048e-4225-a9e7-6f8296f177bd","_uuid":"c169fefd4bab3e93a987c45fbe3d504bc996a8a5"},"cell_type":"markdown","source":"### <a name=\"rank variables\"></a>Rank variables based on linear relationships with reference variable (on sampled data)"},{"metadata":{"_cell_guid":"11cabce0-820d-41fd-b08c-28b70326ab35","_uuid":"dfc61bc409244d633103251d47d828f7e350076b","trusted":true},"cell_type":"code","source":"cols_list = [conf_dict['Target']] + conf_dict['NumericalColumns'] + conf_dict['CategoricalColumns'] \ncols_list = list(OrderedDict.fromkeys(cols_list)) \nw1 = widgets.Dropdown(    \n    options=cols_list,\n    value=cols_list[0],\n    description='Ref Var:'\n)\nw2 = ipywidgets.Text(value=\"5\", description='Top Num Vars:')\nw3 = ipywidgets.Text(value=\"5\", description='Top Cat Vars:')\ni = interactive(InteractionAnalytics.rank_associations, df=fixed(df),conf_dict=fixed(conf_dict), col1=w1, col2=w2, col3=w3)\nhbox = widgets.HBox(i.children)\ndisplay(hbox)\nhbox.on_displayed(InteractionAnalytics.rank_associations(df=df, conf_dict=conf_dict, col1=w1.value, col2=w2.value, col3=w3.value))","execution_count":15,"outputs":[]},{"metadata":{"_cell_guid":"9f69ab2b-964c-4682-add6-7e567b1c6787","_uuid":"8c7443c3ce13b7dc010e79d1f403acdbf5b1c347","collapsed":true},"cell_type":"markdown","source":"### <a name=\"two categorical\"></a>Explore interactions between categorical variables"},{"metadata":{"_cell_guid":"a64cd83e-6b61-4231-a5c4-45bd2d3d50c2","_uuid":"d13d2d3f60e3329a133c8e28ef08d1d56d914b1e","trusted":true},"cell_type":"code","source":"w1, w2 = None, None\n\nif conf_dict['Target'] in conf_dict['CategoricalColumns']:\n    cols_list = [conf_dict['Target']] + conf_dict['CategoricalColumns'] \n    cols_list = list(OrderedDict.fromkeys(cols_list)) \nelse:\n    cols_list = conf_dict['CategoricalColumns']\n    \nw1 = widgets.Dropdown(\n    options=cols_list,\n    value=cols_list[0],\n    description='Categorical Var 1:'\n)\nw2 = widgets.Dropdown(\n    options=cols_list,\n    value=cols_list[1],\n    description='Categorical Var 2:'\n)\n\ni = interactive(InteractionAnalytics.categorical_relations, df=fixed(df), col1=w1, col2=w2)\nhbox = widgets.HBox(i.children)\ndisplay(hbox)\nhbox.on_displayed(InteractionAnalytics.categorical_relations(df=df0, col1=w1.value, col2=w2.value))","execution_count":16,"outputs":[]},{"metadata":{"_cell_guid":"0e574aa0-9c87-42c7-954d-404678f84ad0","_uuid":"23696506a6bb2c2d4ac226478a6212419b1a1dee","collapsed":true},"cell_type":"markdown","source":"### <a name=\"two numerical\"></a>Explore interactions between numerical variables (on sampled data)"},{"metadata":{"_cell_guid":"1128f3cf-bdf0-4abb-93f0-3bf9a45edc15","_uuid":"afc3f5265d5358ceec3f17a1fdf0ef676265f98a","trusted":true},"cell_type":"code","source":"w1, w2 = None, None\n\nif conf_dict['Target'] in conf_dict['NumericalColumns']:\n    cols_list = [conf_dict['Target']] + conf_dict['NumericalColumns'] \n    cols_list = list(OrderedDict.fromkeys(cols_list)) \nelse:\n    cols_list = conf_dict['NumericalColumns']\nw1 = widgets.Dropdown(\n    options=cols_list,\n    value=cols_list[0],\n    description='Numerical Var 1:'\n)\nw2 = widgets.Dropdown(\n    options=cols_list,\n    value=cols_list[min(1, len(cols_list)-1)],\n    description='Numerical Var 2:'\n)\ni = interactive(InteractionAnalytics.numerical_relations, df=fixed(df), col1=w1, col2=w2)\nhbox = widgets.HBox(i.children)\ndisplay(hbox)\nhbox.on_displayed(InteractionAnalytics.numerical_relations(df, col1=w1.value, col2=w2.value))","execution_count":17,"outputs":[]},{"metadata":{"_cell_guid":"780b465c-5930-4da7-a364-f67bb9949cd1","_uuid":"edb22023720bf5c5e113c5ce03a1c8f19ce7c0e9","collapsed":true},"cell_type":"markdown","source":"### Explore correlation matrix between numerical variables"},{"metadata":{"_cell_guid":"a3ac869d-fc6f-4b40-9f11-2cf669adb1f2","_uuid":"115ccf9d13387637cb8819ce868e265f12459906","trusted":true},"cell_type":"code","source":"w1 = None\nw1 = widgets.Dropdown(\n    options=['pearson','kendall','spearman'],\n    value='pearson',\n    description='Correlation Method:'\n)\ni = interactive(InteractionAnalytics.numerical_correlation, df=fixed(df), conf_dict=fixed(conf_dict), col1=w1)\nhbox = widgets.HBox(i.children)\ndisplay(hbox)\nhbox.on_displayed(InteractionAnalytics.numerical_correlation(df0, conf_dict=conf_dict, col1=w1.value))","execution_count":18,"outputs":[]},{"metadata":{"_cell_guid":"a958f253-7be6-4580-98ac-64dce258a007","_uuid":"3c741583797e5cd2408ae7589606573087498c2e","collapsed":true},"cell_type":"markdown","source":"### <a name=\"numerical and categorical\"></a>Explore interactions between numerical and categorical variables"},{"metadata":{"_cell_guid":"8677cd36-629c-49d3-81f9-257b263a1853","_uuid":"ca8bb14b5b577a0b3132c816008adf34487b22c7","trusted":true},"cell_type":"code","source":"w1, w2 = None, None\n\nif conf_dict['Target'] in conf_dict['NumericalColumns']:\n    cols_list = [conf_dict['Target']] + conf_dict['NumericalColumns'] #Make target the default reference variable\n    cols_list = list(OrderedDict.fromkeys(cols_list)) #remove variables that might be duplicates with target\nelse:\n    cols_list = conf_dict['NumericalColumns']\n    \nw1 = widgets.Dropdown(\n    options=cols_list,\n    value=cols_list[0],\n    description='Numerical Variable:'\n)\n\nif conf_dict['Target'] in conf_dict['CategoricalColumns']:\n    cols_list = [conf_dict['Target']] + conf_dict['CategoricalColumns'] #Make target the default reference variable\n    cols_list = list(OrderedDict.fromkeys(cols_list)) #remove variables that might be duplicates with target\nelse:\n    cols_list = conf_dict['CategoricalColumns']\n    \nw2 = widgets.Dropdown(\n    options=cols_list,\n    value=cols_list[0],\n    description='Categorical Variable:'\n)\ni = interactive(InteractionAnalytics.nc_relation, df=fixed(df),conf_dict=fixed(conf_dict), col1=w1, col2=w2, col3=fixed(w3))\nhbox = widgets.HBox(i.children)\ndisplay( hbox )\nhbox.on_displayed(InteractionAnalytics.nc_relation(df0, conf_dict, col1=w1.value, col2=w2.value))","execution_count":19,"outputs":[]},{"metadata":{"_cell_guid":"6a540242-38a1-4768-898c-1055c17edf9e","_uuid":"aca3c8b8f4eb2f4c462caf9d0218f6b00170f1eb","collapsed":true},"cell_type":"markdown","source":"### <a name=\"two numerical and categorical\"></a>Explore interactions between two numerical variables and a categorical variable (on sampled data)"},{"metadata":{"_cell_guid":"0f27e961-adb8-474d-a1fb-b484bce0b5bd","_uuid":"ba9bfc68c64650760cee0e536adbfdb0bb99a922","trusted":true},"cell_type":"code","source":"w1, w2, w3 = None, None, None\n\nif conf_dict['Target'] in conf_dict['NumericalColumns']:\n    cols_list = [conf_dict['Target']] + conf_dict['NumericalColumns'] \n    cols_list = list(OrderedDict.fromkeys(cols_list)) \nelse:\n    cols_list = conf_dict['NumericalColumns']\n    \nw1 = widgets.Dropdown(\n    options = cols_list,\n    value = cols_list[0],\n    description = 'Numerical Var 1:'\n)\nw2 = widgets.Dropdown(\n    options = cols_list,\n    value = cols_list[min(1, len(cols_list)-1)],\n    description = 'Numerical Var 2:'\n)\n\nif conf_dict['Target'] in conf_dict['CategoricalColumns']:\n    cols_list = [conf_dict['Target']] + conf_dict['CategoricalColumns'] \n    cols_list = list(OrderedDict.fromkeys(cols_list)) \nelse:\n    cols_list = conf_dict['CategoricalColumns']\n    \nw3 = widgets.Dropdown(\n    options = cols_list,\n    value = cols_list[0],\n    description = 'Legend Cat Var:'\n)\ni = interactive(InteractionAnalytics.nnc_relation, df=fixed(df),conf_dict=fixed(conf_dict), col1=w1, col2=w2, col3=w3)\nhbox = widgets.HBox(i.children)\ndisplay(hbox)\nhbox.on_displayed(InteractionAnalytics.nnc_relation(df, conf_dict, col1=w1.value,col2=w2.value, col3=w3.value))","execution_count":21,"outputs":[]},{"metadata":{"_cell_guid":"0498effb-adef-4d95-9148-1ee2b7fbbed9","_uuid":"b8f5b47c36e84e5b3afd3718b7f28022a091385a","collapsed":true},"cell_type":"markdown","source":"## <a name=\"pca\"></a>Visualize numerical data by projecting to principal component spaces (on sampled data)"},{"metadata":{"_cell_guid":"1bc71517-7d51-4807-b3f8-134bce393e5d","_uuid":"cebce9f3ebce12ade38d8ca083b5d2395f5e6136"},"cell_type":"markdown","source":"### Project data to 2-D principal component space (on sampled data)"},{"metadata":{"_cell_guid":"727e13ce-6b96-4468-b4d3-74f8a4e6f9d5","_uuid":"7385bcec3a2811663ce5dd653b4faa055aaa5ba9","collapsed":true,"trusted":true},"cell_type":"code","source":"num_numeric = len(conf_dict['NumericalColumns'])\nif  num_numeric > 3:\n    \n    w1, w2, w3 = None, None, None\n    if conf_dict['Target'] in conf_dict['CategoricalColumns']:\n        cols_list = [conf_dict['Target']] + conf_dict['CategoricalColumns'] \n        cols_list = list(OrderedDict.fromkeys(cols_list)) \n    else:\n        cols_list = conf_dict['CategoricalColumns']\n    w1 = widgets.Dropdown(\n        options = cols_list,\n        value = cols_list[0],\n        description = 'Legend Variable:',\n        width = 10\n    )\n    w2 = widgets.Dropdown(\n        options = [str(x) for x in np.arange(1,num_numeric+1)],\n        value = '1',\n        width = 1,\n        description='PC at X-Axis:'\n    )\n    w3 = widgets.Dropdown(\n        options = [str(x) for x in np.arange(1,num_numeric+1)],\n        value = '2',\n        description = 'PC at Y-Axis:'\n    )\n    i = interactive(InteractionAnalytics.numerical_pca, df=fixed(df),conf_dict=fixed(conf_dict), col1=w1, col2=w2, col3=w3)\n    hbox = widgets.HBox(i.children[:3])\n    display(hbox)\n    hbox.on_displayed(InteractionAnalytics.numerical_pca(df, conf_dict=conf_dict, col1=w1.value, col2=w2.value, col3=w3.value))","execution_count":22,"outputs":[]},{"metadata":{"_cell_guid":"38c09c84-445c-4f3a-8e48-0c9f1259c6c9","_uuid":"11d18bfb1f9235d6874b8054cd76d45331261028","collapsed":true},"cell_type":"markdown","source":"### Project data to 3-D principal component space (on sampled data)"},{"metadata":{"_cell_guid":"e5512d35-c954-4447-82bd-dd44a076f22c","_uuid":"5427d34c95b71db5981a713cc1eb2c279dd7666a","collapsed":true,"trusted":true},"cell_type":"code","source":"if len(conf_dict['NumericalColumns']) > 3:\n    if conf_dict['Target'] in conf_dict['CategoricalColumns']:\n        cols_list = [conf_dict['Target']] + conf_dict['CategoricalColumns'] \n        cols_list = list(OrderedDict.fromkeys(cols_list)) \n    else:\n        cols_list = conf_dict['CategoricalColumns']\n    w1, w2 = None, None\n    w1 = widgets.Dropdown(\n        options=cols_list,\n        value=cols_list[0],\n        description='Legend Variable:'\n    )\n    w2 = ipywidgets.IntSlider(min=-180, max=180, step=5, value=30, description='Angle')\n    i = interactive(InteractionAnalytics.pca_3d, df=fixed(df), conf_dict=fixed(conf_dict),col1=w1, col2=w2, col3=fixed(w3))\n    hbox=widgets.HBox(i.children[:2])\n    display(hbox)\n    hbox.on_displayed(InteractionAnalytics.pca_3d(df,conf_dict,col1=w1.value,col2=w2.value))","execution_count":23,"outputs":[]},{"metadata":{"_cell_guid":"b1176f64-7288-49a9-bcbb-3e8d00230ab0","_uuid":"acb7a6f7d422a86dd39846fbdf440e032a6b4760","collapsed":true},"cell_type":"markdown","source":"## <a name=\"show hide codes\"></a>Show/Hide the Source Codes"},{"metadata":{"_cell_guid":"e878a980-eb83-4adb-be3a-efcb7250b480","_uuid":"6db57bbd770f965f8444edf93f310cf47173eff2","trusted":true},"cell_type":"code","source":"display(HTML('''<style>\n    .widget-label { min-width: 20ex !important; }\n    .widget-text { min-width: 60ex !important; }\n</style>'''))\n\n#Toggle Code\nHTML('''<script>\ncode_show=true; \nfunction code_toggle() {\n if (code_show){\n $('div.input').hide();\n\n } else {\n $('div.input').show();\n\n }\n code_show = !code_show\n} \n//$( document ).ready(code_toggle);//commenting code disabling by default\n</script>\n<form action = \"javascript:code_toggle()\"><input type=\"submit\" value=\"Toggle Raw Code\"></form>''')","execution_count":24,"outputs":[]},{"metadata":{"trusted":true,"collapsed":true,"_uuid":"e04a55de73aca505af59261c0e152f28af3815f2"},"cell_type":"code","source":"","execution_count":null,"outputs":[]}],"metadata":{"widgets":{"state":{"d0d887fc32c940e88721b75f39ec03c5":{"views":[{"cell_index":13}]},"0506975d69c14717be17695d1a6c7132":{"views":[{"cell_index":50}]},"04f0d91d56f342708c52e57fbe9bcd06":{"views":[{"cell_index":45}]},"0dc41da3c3f445d690100fd99558f244":{"views":[{"cell_index":30}]},"e96f8216a4b94c3aa524bce1034708ef":{"views":[{"cell_index":32}]},"4a5ca2989b4d456d8d019f38514c3485":{"views":[{"cell_index":35}]},"d0c17a981a554bcc89cba56b4a1edccd":{"views":[{"cell_index":39}]},"a416687c67684df8818f966704d9f073":{"views":[{"cell_index":43}]},"9cc891c4a4dc4ca3bc4801bf7c38e19d":{"views":[{"cell_index":41}]},"56d3942cb1684869808846d861e7b0b9":{"views":[{"cell_index":48}]},"642af708746142059091fb9c21fb4612":{"views":[{"cell_index":28}]},"2196ee8000864f23bbfa4689bfaffc7d":{"views":[{"cell_index":37}]}},"version":"1.2.0"},"celltoolbar":"Raw Cell Format","language_info":{"name":"python","version":"3.6.5","mimetype":"text/x-python","codemirror_mode":{"name":"ipython","version":3},"pygments_lexer":"ipython3","nbconvert_exporter":"python","file_extension":".py"},"kernelspec":{"display_name":"Python 3","language":"python","name":"python3"}},"nbformat":4,"nbformat_minor":1}