{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"pygments_lexer":"ipython3","nbconvert_exporter":"python","version":"3.6.4","file_extension":".py","codemirror_mode":{"name":"ipython","version":3},"name":"python","mimetype":"text/x-python"}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"# Exploratory data analysis\n> The easy way","metadata":{}},{"cell_type":"code","source":"#installing the library\n! pip install sweetviz","metadata":{"execution":{"iopub.status.busy":"2022-08-05T20:59:22.395862Z","iopub.execute_input":"2022-08-05T20:59:22.396907Z","iopub.status.idle":"2022-08-05T20:59:30.900570Z","shell.execute_reply.started":"2022-08-05T20:59:22.396860Z","shell.execute_reply":"2022-08-05T20:59:30.899459Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"import pandas as pd\nimport sweetviz as sv\nimport IPython","metadata":{"execution":{"iopub.status.busy":"2022-08-05T21:01:26.781241Z","iopub.execute_input":"2022-08-05T21:01:26.781575Z","iopub.status.idle":"2022-08-05T21:01:26.786488Z","shell.execute_reply.started":"2022-08-05T21:01:26.781551Z","shell.execute_reply":"2022-08-05T21:01:26.785473Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"data = pd.read_csv('../input/tabular-playground-series-aug-2022/train.csv')\ntest = pd.read_csv('../input/tabular-playground-series-aug-2022/test.csv')","metadata":{"execution":{"iopub.status.busy":"2022-08-05T22:22:10.152484Z","iopub.execute_input":"2022-08-05T22:22:10.152849Z","iopub.status.idle":"2022-08-05T22:22:10.311339Z","shell.execute_reply.started":"2022-08-05T22:22:10.152824Z","shell.execute_reply":"2022-08-05T22:22:10.310454Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"data.columns","metadata":{"execution":{"iopub.status.busy":"2022-08-05T21:14:38.765613Z","iopub.execute_input":"2022-08-05T21:14:38.765900Z","iopub.status.idle":"2022-08-05T21:14:38.773709Z","shell.execute_reply.started":"2022-08-05T21:14:38.765875Z","shell.execute_reply":"2022-08-05T21:14:38.772487Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## Creating a quick and interactive EDA dashboard on the data\n\n[Fast & powerful EDA in notebooks using Sweetviz 2.0 - Colaboratory](https://colab.research.google.com/drive/1-md6YEwcVGWVnQWTBirQSYQYgdNoeSWg?usp=sharing#scrollTo=XdxySa-7QxJS)","metadata":{}},{"cell_type":"markdown","source":"Both Notebook and HTML reports can use a vertical or widescreen layout. The same information is displayed, but in the vertical layout users must click on a specific feature to see the detail, whereas in widescreen mode, data is displayed as soon as the mouse goes over any feature and this view can be locked-in by clicking. I will go for vertical layout","metadata":{}},{"cell_type":"code","source":"## creating a EDA report\nfeature_config = sv.FeatureConfig(skip=\"id\") #  removing ID column from EDA\nanalyze_report = sv.analyze(data, target_feat = 'failure', feat_cfg = feature_config)\n# IMPORTANT: only numerical and boolean features can be targets currently.\nanalyze_report.show_notebook(layout='vertical', w=800, h=300, scale=0.7)\n# analyze_report.show_notebook(layout='widescreen', w=1500, h=300, scale=0.7) will use the vertical instead","metadata":{"execution":{"iopub.status.busy":"2022-08-05T21:41:55.655961Z","iopub.execute_input":"2022-08-05T21:41:55.656275Z","iopub.status.idle":"2022-08-05T21:42:13.028459Z","shell.execute_reply.started":"2022-08-05T21:41:55.656250Z","shell.execute_reply":"2022-08-05T21:42:13.027356Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## Let's illustrate our findings ","metadata":{"execution":{"iopub.status.busy":"2022-08-05T21:11:49.180664Z","iopub.execute_input":"2022-08-05T21:11:49.180987Z","iopub.status.idle":"2022-08-05T21:11:49.185802Z","shell.execute_reply.started":"2022-08-05T21:11:49.180962Z","shell.execute_reply":"2022-08-05T21:11:49.184864Z"}}},{"cell_type":"markdown","source":"26570 rows with 0 duplicates. No text features; only numerical and categorical. \nNo missing values in target column. An imbalance (~ 80:20) is present","metadata":{}},{"cell_type":"markdown","source":"### Features and their relation with target values ","metadata":{}},{"cell_type":"markdown","source":"#### Product code\n5 categories. Although A is lowest category, it has the highest percentage of failures (23%). ","metadata":{}},{"cell_type":"markdown","source":"#### Loading \nseems to be associated with failure","metadata":{}},{"cell_type":"markdown","source":"#### measurement 0,1,2 \nintegers(few zeros - no missing).","metadata":{}},{"cell_type":"markdown","source":"#### measurement 3 - 8\nreal numbers (missng <5% - no zeros)","metadata":{}},{"cell_type":"markdown","source":"#### measurement 9 - 17\nreal numbers (missng 5 - 9% - no zeros)","metadata":{"execution":{"iopub.status.busy":"2022-08-05T22:05:38.692356Z","iopub.execute_input":"2022-08-05T22:05:38.692721Z","iopub.status.idle":"2022-08-05T22:05:38.700217Z","shell.execute_reply.started":"2022-08-05T22:05:38.692693Z","shell.execute_reply":"2022-08-05T22:05:38.698710Z"}}},{"cell_type":"markdown","source":"#### measurement 17 \nrange is higher than all other columns ","metadata":{}},{"cell_type":"markdown","source":"## Correlations and associations\n> [Powerful EDA (Exploratory Data Analysis) in just two lines of code using Sweetviz | by Francois Bertrand | Towards Data Science](https://towardsdatascience.com/powerful-eda-exploratory-data-analysis-in-just-two-lines-of-code-using-sweetviz-6c943d32f34)","metadata":{"execution":{"iopub.status.busy":"2022-08-05T22:37:06.204740Z","iopub.execute_input":"2022-08-05T22:37:06.205147Z","iopub.status.idle":"2022-08-05T22:37:06.209677Z","shell.execute_reply.started":"2022-08-05T22:37:06.205121Z","shell.execute_reply":"2022-08-05T22:37:06.208581Z"}}},{"cell_type":"markdown","source":"Basically, in addition to showing the traditional numerical correlations, it unifies in a single graph both numerical correlation but also the uncertainty coefficient (for categorical-categorical) and correlation ratio (for categorical-numerical). Squares represent categorical-featured-related variables and circles represent numerical-numerical correlations. Note that the trivial diagonal is left empty, for clarity.","metadata":{}},{"cell_type":"markdown","source":"## Associtations\n","metadata":{}},{"cell_type":"markdown","source":"From the associtation report, we can see that product code is highly associated with all attributes. Attributes two and three are mostly associated with each others. \nIMPORTANT: categorical-categorical associations (provided by the uncertainty coefficient) are ASSYMMETRICAL, meaning that each row represents how much the row title (on the left) gives information on each column.\nSince the association is asymettrical, it seems that product code influences all attributes but is influenced mainly by attributes 2 and 3. ","metadata":{}},{"cell_type":"markdown","source":"## Correlations ","metadata":{"execution":{"iopub.status.busy":"2022-08-05T22:14:10.412814Z","iopub.execute_input":"2022-08-05T22:14:10.413178Z","iopub.status.idle":"2022-08-05T22:14:10.417919Z","shell.execute_reply.started":"2022-08-05T22:14:10.413153Z","shell.execute_reply":"2022-08-05T22:14:10.417033Z"}}},{"cell_type":"markdown","source":"Measurements 0 and 1 are inversely correlated. Measurements 5-8 are positively correlated with measurement 17","metadata":{}},{"cell_type":"markdown","source":"Finally, it is worth noting these correlation/association methods shouldn’t be taken as gospel as they make some assumptions on the underlying distribution of data and relationships. However they can be a very useful starting point.","metadata":{}},{"cell_type":"markdown","source":"## Comparison report\n> we need to compare training and test sets to make sure the distribution os not quite different. ","metadata":{}},{"cell_type":"code","source":"comparison_report = sv.compare(data, test, target_feat='failure')\n\ncomparison_report.show_notebook()","metadata":{"execution":{"iopub.status.busy":"2022-08-05T22:23:11.319342Z","iopub.execute_input":"2022-08-05T22:23:11.319708Z","iopub.status.idle":"2022-08-05T22:23:35.015233Z","shell.execute_reply.started":"2022-08-05T22:23:11.319682Z","shell.execute_reply":"2022-08-05T22:23:35.014288Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"In fact, it seems quite different. categories within the categorical variables are quite different. Some levels are absent in test set. Other levels appear only in test set. Correlations are quite similar between train and test set. However, correlation between measurement 0 and 1 is lost in test set. ","metadata":{}},{"cell_type":"code","source":"","metadata":{},"execution_count":null,"outputs":[]}]}