{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"pygments_lexer":"ipython3","nbconvert_exporter":"python","version":"3.6.4","file_extension":".py","codemirror_mode":{"name":"ipython","version":3},"name":"python","mimetype":"text/x-python"}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"![](https://raw.githubusercontent.com/morganmcg1/images/main/Screenshot%202022-07-14%20at%2009.35.15.png)","metadata":{}},{"cell_type":"markdown","source":"# Deeply understand the dataset with Pandas Profiling\n**[Pandas-Profiling](https://github.com/ydataai/pandas-profiling) is an incredibly valuable library for data exploration, in their own words:**\n> pandas-profiling generates profile reports from a pandas DataFrame. The pandas df.describe() function is handy yet a little basic for exploratory data analysis. pandas-profiling extends pandas DataFrame with df.profile_report(), which automatically generates a standardized univariate and multivariate report for data understanding.\n\nAfter running this notebook you will have your own Pandas Profile of the Kaggle Tabular Playground Series July 2022 dataset embedded in an **online, interactive Report**: \n\n### 👉 **[Just like this!](https://wandb.ai/morgan/kaggle-tps-july-unsupervised/reports/Pandas-Profiling-Report-Kaggle-TPS-July-2022--VmlldzoyMjk3OTIy?accessToken=1l3veyt2br2u5seky5u622i3431wkdz4xsicxzqwr0irwskrejhtq8wz9xaqk22p)**","metadata":{}},{"cell_type":"markdown","source":"![](https://raw.githubusercontent.com/morganmcg1/images/main/Screenshot%202022-07-14%20at%2010.14.49.png)","metadata":{}},{"cell_type":"code","source":"!pip install -qqq -U pandas-profiling[notebook] wandb","metadata":{"execution":{"iopub.status.busy":"2022-07-14T08:44:49.131159Z","iopub.execute_input":"2022-07-14T08:44:49.131716Z","iopub.status.idle":"2022-07-14T08:45:02.485905Z","shell.execute_reply.started":"2022-07-14T08:44:49.131671Z","shell.execute_reply":"2022-07-14T08:45:02.483820Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# 1. Generate your data Profile with Pandas Profiling","metadata":{}},{"cell_type":"code","source":"import pandas as pd\nimport wandb\nfrom pandas_profiling import ProfileReport","metadata":{"editable":false,"execution":{"iopub.status.busy":"2022-07-14T08:45:54.497183Z","iopub.execute_input":"2022-07-14T08:45:54.497696Z","iopub.status.idle":"2022-07-14T08:45:57.717769Z","shell.execute_reply.started":"2022-07-14T08:45:54.497627Z","shell.execute_reply":"2022-07-14T08:45:57.715169Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Pandas `df.describe()` gives some useful information about the data, but we really need a deeper look to understand how best to work with this dataset","metadata":{"execution":{"iopub.status.busy":"2022-07-14T08:47:32.490085Z","iopub.execute_input":"2022-07-14T08:47:32.490554Z","iopub.status.idle":"2022-07-14T08:47:32.499724Z","shell.execute_reply.started":"2022-07-14T08:47:32.490515Z","shell.execute_reply":"2022-07-14T08:47:32.497856Z"}}},{"cell_type":"code","source":"df = pd.read_csv('../input/tabular-playground-series-jul-2022/data.csv')\ndf.describe()","metadata":{"execution":{"iopub.status.busy":"2022-07-14T08:46:25.335141Z","iopub.execute_input":"2022-07-14T08:46:25.335565Z","iopub.status.idle":"2022-07-14T08:46:27.023675Z","shell.execute_reply.started":"2022-07-14T08:46:25.335533Z","shell.execute_reply":"2022-07-14T08:46:27.022488Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Instead, lets generate a Pandas Profile","metadata":{}},{"cell_type":"code","source":"profile = ProfileReport(df, title=\"Pandas Profiling Report\")","metadata":{"editable":false,"execution":{"iopub.status.busy":"2022-07-14T08:50:33.162987Z","iopub.execute_input":"2022-07-14T08:50:33.163375Z","iopub.status.idle":"2022-07-14T08:50:33.174816Z","shell.execute_reply.started":"2022-07-14T08:50:33.163343Z","shell.execute_reply":"2022-07-14T08:50:33.173592Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"You can display the Profile Report in your notebook. In this case we'll instead save it as a html file and upload it to **[Weights & Biases](https://wandb.ai/site)** where we can add it to a W&B Report and it can persist there permanently. That way we can come back to it whenever we need it","metadata":{}},{"cell_type":"code","source":"# Uncomment if you'd like to view the Profile Report in the notebook\n#profile.to_widgets()","metadata":{"execution":{"iopub.status.busy":"2022-07-14T08:50:30.212962Z","iopub.execute_input":"2022-07-14T08:50:30.213411Z","iopub.status.idle":"2022-07-14T08:50:30.220186Z","shell.execute_reply.started":"2022-07-14T08:50:30.213379Z","shell.execute_reply":"2022-07-14T08:50:30.218532Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Save the Profile Report to HTML\nhtml_fn = \"tabular_playground_july_profile.html\"\nprofile.to_file(html_fn)","metadata":{"execution":{"iopub.status.busy":"2022-07-14T09:14:03.887471Z","iopub.execute_input":"2022-07-14T09:14:03.888560Z","iopub.status.idle":"2022-07-14T09:14:03.894049Z","shell.execute_reply.started":"2022-07-14T09:14:03.888508Z","shell.execute_reply":"2022-07-14T09:14:03.892714Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# 2. Upload your Pandas Profile to Weights & Biases\nWith just 3 lines of code we can upload the HTML file to Weights & Biases and have a persisent copy of the Profile Report available online. To do this you will need a Weights & Biases api key, you  can generate this once you **[sign up for a free Weigths & Biases account](https://wandb.ai/site)** - all W&B personal accounts are **free forever**","metadata":{}},{"cell_type":"code","source":"wandb.init(project='kaggle-tps-july-unsupervised', job_type='data_upload')\nwandb.log({\"Pandas_Profile\": wandb.Html(open(html_fn), inject=False)})\nwandb.finish()","metadata":{"execution":{"iopub.status.busy":"2022-07-14T09:00:12.967348Z","iopub.execute_input":"2022-07-14T09:00:12.967816Z","iopub.status.idle":"2022-07-14T09:00:12.973254Z","shell.execute_reply.started":"2022-07-14T09:00:12.967780Z","shell.execute_reply":"2022-07-14T09:00:12.972062Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# 3. View your Pandas-Profiling Report in a W&B Report\n\nOnce you upload your Profile to W&B you can then add it to a new W&B Report. \n\na. Click on the wandb run link in the cell above to go to your wandb run page\n\nb. Click the top right corner of your Pandas_Profile media panel and select \"Add to report\". This will open a new W&B Report.\n\nc. In this W&B you can resize your Profile and add any additional notes, comments, images and more to your exploratory data analysis\n\n**[Here is an example](https://wandb.ai/morgan/kaggle-tps-july-unsupervised/reports/Pandas-Profiling-Report-Kaggle-TPS-July-2022--VmlldzoyMjk3OTIy?accessToken=1l3veyt2br2u5seky5u622i3431wkdz4xsicxzqwr0irwskrejhtq8wz9xaqk22p)** of a Pandas Profile in a W&B Report","metadata":{}},{"cell_type":"markdown","source":"![](https://raw.githubusercontent.com/morganmcg1/images/main/Screenshot%202022-07-14%20at%2010.04.16.png)","metadata":{}},{"cell_type":"markdown","source":"# A Pandas-Profiling Report in W&B 🎉\n**[This is an example](https://wandb.ai/morgan/kaggle-tps-july-unsupervised/reports/Pandas-Profiling-Report-Kaggle-TPS-July-2022--VmlldzoyMjk3OTIy?accessToken=1l3veyt2br2u5seky5u622i3431wkdz4xsicxzqwr0irwskrejhtq8wz9xaqk22p)** of a Pandas Profile in a W&B Report:\n","metadata":{}},{"cell_type":"markdown","source":"![](https://raw.githubusercontent.com/morganmcg1/images/main/Screenshot%202022-07-14%20at%2010.14.49.png)","metadata":{}},{"cell_type":"markdown","source":"# What is Weights & Biases?","metadata":{}},{"cell_type":"markdown","source":"<center><img src=\"https://camo.githubusercontent.com/dd842f7b0be57140e68b2ab9cb007992acd131c48284eaf6b1aca758bfea358b/68747470733a2f2f692e696d6775722e636f6d2f52557469567a482e706e67\"></center>","metadata":{"editable":false}},{"cell_type":"markdown","source":"Track everything you need to make your models reproducible with [Weights & Biases](https://wandb.ai/site) — from hyperparameters and code to model weights and dataset versions.\n\nWeights & Biases helps your ML team unlock their productivity by optimizing, visualizing, collaborating on, and standardizing their model and data pipelines – regardless of framework, environment, or workflow.\n\nUsed by ML engineers at OpenAI, Lyft, Pfizer, Qualcomm, NVIDIA, Toyota, GitHub, and MILA, W&B is part of the new standard of best practices for machine learning. W&B is free for personal use and academic projects, and it's easy to get started.","metadata":{}}]}