{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"pygments_lexer":"ipython3","nbconvert_exporter":"python","version":"3.6.4","file_extension":".py","codemirror_mode":{"name":"ipython","version":3},"name":"python","mimetype":"text/x-python"}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"# Statistics 101 - A simple approach :-)\n\n![](https://www.meme-arsenal.com/memes/8b1a2e3b8e9598c9040140364ae04847.jpg)\n\n## Objective:\n\nThe purpose of this notebook is to present simply some important statistics concepts.\n\n## Why statistics is the study of statistics so important?\n\nSome of the reasons are:\n\n* proper methods to collect the data\n* employ the correct analyses\n* effectively present the results (storytelling)\n\n\n## What is statistics:\n\nStatistics is the study of how to **collect**, **organize**, **analyze**, and **interpret** information and data!\n\nStatistics is used to help us **make decisions**.\n\n## Drew Conway - Venn Diagram\n\n![](https://i.pinimg.com/736x/69/72/56/6972562e67fe098a954eb362136a5d0b--data-data-big-data.jpg)\n\n### Proper methods to collect the data\n\n**Always care about data collection!**\n\n![](https://i0.wp.com/marketbusinessnews.com/wp-content/uploads/2017/11/GIGO-garbage-in-garbage-out-definiion-and-illustration.jpg?fit=904%2C679&ssl=1)\n\n\n### Employ the correct analyses\n\n**Beware of hasty conclusions**\n\n![](https://imgs.xkcd.com/comics/extrapolating.png)\n\n**Book suggestion:**\n\n![](https://m.media-amazon.com/images/I/41shZGS-G+L.jpg)\n\n**Some useful statistical tools (dataviz)**\n\n![](https://fiverr-res.cloudinary.com/images/t_main1,q_auto,f_auto,q_auto,f_auto/gigs/129050648/original/b9c7eb074ab7fe69b8670fde27300752a561af0b/tasks-related-to-machine-learning-and-data-analysis-in-r-or-python.png)\n\n\n## Effectively present the results (storytelling)\n\n![](https://www.ecommercebrasil.com.br/wp-content/uploads/2020/09/storytelling-no-e-commerce.jpg)\n\n**Documentary suggestion:**\n\nA documentary about Fangio deals with statistics...\n\n![](https://m.media-amazon.com/images/M/MV5BYWViZWJiZTEtNzE5OC00OTRiLWI0MTktOGE5NzM3MjI1ODBjXkEyXkFqcGdeQXVyMzkyOTg1MzE@._V1_.jpg)","metadata":{}},{"cell_type":"markdown","source":"# 1. Basic concepts \n\n## Population and sample\n\n![](https://miro.medium.com/max/1400/1*83WLnE2QTeOHSjbPJDVKfw.png)\n\n\n## Population Parameters vs. Sample Statistics\n\n![](https://sphweb.bumc.bu.edu/otlt/MPH-Modules/BS/BS704_BiostatisticsBasics/Sampling3.jpg)\n\n## Descriptive vs inferential statistics\n\n![](https://datatab.net/assets/tutorial/inferenz_deskriptiv_en.png)\n\n\n### Practical example - Net Promoter Score \n\n* Frederick F. Reichheld\n* The One Number You Need to Grow (Harvard Business Review)\n\n![](https://resultadosdigitais.com.br/files/2018/10/nps-net-promoter-score.jpg)\n\n## Statistical Pitfalls \n\n![](https://substackcdn.com/image/fetch/w_1200,h_600,c_limit,f_jpg,q_auto:good,fl_progressive:steep/https%3A%2F%2Fbucketeer-e05bbc84-baa3-437e-9518-adb32be77984.s3.amazonaws.com%2Fpublic%2Fimages%2F381dd65c-06fc-4331-92bf-063bbe1d8958_1834x1417.png)\n\n## It's sad, but it's true:\n\n![](https://pbs.twimg.com/media/EVBeFOoWAAAkbsY.jpg)\n\n[George E. P. Box](https://en.wikipedia.org/wiki/George_E._P._Box)\n\n## Data Types:\n\n![](https://thebiologynotes.com/wp-content/uploads/2019/08/Data-and-its-types.jpg)\n\n\n## Most commons measures of central tendency\n\n![](https://www.basic-mathematics.com/images/Measures-of-central-tendency.png)\n\n## Most commons dispersion measures\n![](https://protonstalk.com/wp-content/uploads/2021/02/measures-of-dispersion.jpg)","metadata":{}},{"cell_type":"markdown","source":"# 2. Hands on","metadata":{}},{"cell_type":"code","source":"!pip install autoviz","metadata":{"execution":{"iopub.status.busy":"2022-07-23T17:32:20.256105Z","iopub.execute_input":"2022-07-23T17:32:20.256477Z","iopub.status.idle":"2022-07-23T17:32:32.257195Z","shell.execute_reply.started":"2022-07-23T17:32:20.256447Z","shell.execute_reply":"2022-07-23T17:32:32.255752Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"import numpy as np\nimport pandas as pd\nfrom autoviz.AutoViz_Class import AutoViz_Class\nimport plotly.express as px\nimport seaborn as sns\n\n\nfrom sklearn.metrics import accuracy_score\nimport scikitplot as skplt\n\nimport os\nfor dirname, _, filenames in os.walk('/kaggle/input'):\n    for filename in filenames:\n        print(os.path.join(dirname, filename))","metadata":{"_uuid":"8f2839f25d086af736a60e9eeb907d3b93b6e0e5","_cell_guid":"b1076dfc-b9ad-4769-8c92-a6c4dae69d19","execution":{"iopub.status.busy":"2022-07-23T18:44:31.621709Z","iopub.execute_input":"2022-07-23T18:44:31.622110Z","iopub.status.idle":"2022-07-23T18:44:31.635954Z","shell.execute_reply.started":"2022-07-23T18:44:31.622067Z","shell.execute_reply":"2022-07-23T18:44:31.634926Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## 2.1 Univariate analysis (diabetes)","metadata":{}},{"cell_type":"code","source":"diabetes = pd.read_csv('/kaggle/input/pima-indians-diabetes-database/diabetes.csv')\nprint(diabetes.shape)\ndiabetes.head(10)","metadata":{"execution":{"iopub.status.busy":"2022-07-23T17:32:32.279460Z","iopub.execute_input":"2022-07-23T17:32:32.279902Z","iopub.status.idle":"2022-07-23T17:32:32.306838Z","shell.execute_reply.started":"2022-07-23T17:32:32.279857Z","shell.execute_reply":"2022-07-23T17:32:32.305445Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### how many pregnancies? (mean and std)","metadata":{}},{"cell_type":"code","source":"mean = np.round(diabetes['Pregnancies'].mean(), 2)\nstd = np.round(diabetes['Pregnancies'].std(), 2)\n\nprint(f'Mean (std) = {mean} ({std})')","metadata":{"execution":{"iopub.status.busy":"2022-07-23T17:32:32.310236Z","iopub.execute_input":"2022-07-23T17:32:32.310818Z","iopub.status.idle":"2022-07-23T17:32:32.317660Z","shell.execute_reply.started":"2022-07-23T17:32:32.310785Z","shell.execute_reply":"2022-07-23T17:32:32.316665Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### how many pregnancies? (q1=0.25, q2=0.5, q3=0.75)","metadata":{}},{"cell_type":"code","source":"q1 = np.round(diabetes['Pregnancies'].quantile(0.25), 2)\nq2 = np.round(diabetes['Pregnancies'].quantile(0.5), 2)\nq3 = np.round(diabetes['Pregnancies'].quantile(0.75), 2)\n\nprint(f'q1= {q1} / q2={q2} / q3={q3}')","metadata":{"execution":{"iopub.status.busy":"2022-07-23T17:32:32.318986Z","iopub.execute_input":"2022-07-23T17:32:32.319696Z","iopub.status.idle":"2022-07-23T17:32:32.334969Z","shell.execute_reply.started":"2022-07-23T17:32:32.319660Z","shell.execute_reply":"2022-07-23T17:32:32.333844Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### pregnancies range","metadata":{}},{"cell_type":"code","source":"ma = diabetes['Pregnancies'].max()\nmi = diabetes['Pregnancies'].min()\n\nprint(f'Min = {mi} / Max ={ma}')","metadata":{"execution":{"iopub.status.busy":"2022-07-23T17:32:32.336543Z","iopub.execute_input":"2022-07-23T17:32:32.337727Z","iopub.status.idle":"2022-07-23T17:32:32.345337Z","shell.execute_reply.started":"2022-07-23T17:32:32.337527Z","shell.execute_reply":"2022-07-23T17:32:32.344083Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### dataviz (histogram)","metadata":{}},{"cell_type":"code","source":"fig = px.histogram(diabetes, x='Pregnancies')\nfig.show()","metadata":{"execution":{"iopub.status.busy":"2022-07-23T17:32:32.347128Z","iopub.execute_input":"2022-07-23T17:32:32.347454Z","iopub.status.idle":"2022-07-23T17:32:32.421145Z","shell.execute_reply.started":"2022-07-23T17:32:32.347425Z","shell.execute_reply":"2022-07-23T17:32:32.420059Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## dataviz (boxplot)","metadata":{}},{"cell_type":"code","source":"fig = px.box(diabetes, y='Pregnancies')\nfig.show()","metadata":{"execution":{"iopub.status.busy":"2022-07-23T17:32:32.422754Z","iopub.execute_input":"2022-07-23T17:32:32.423074Z","iopub.status.idle":"2022-07-23T17:32:32.480031Z","shell.execute_reply.started":"2022-07-23T17:32:32.423046Z","shell.execute_reply":"2022-07-23T17:32:32.479052Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## dataviz (boxplot + histogram)","metadata":{}},{"cell_type":"code","source":"fig = px.histogram(diabetes, x='Pregnancies', marginal=\"box\")\nfig.show()","metadata":{"execution":{"iopub.status.busy":"2022-07-23T17:32:32.481449Z","iopub.execute_input":"2022-07-23T17:32:32.481743Z","iopub.status.idle":"2022-07-23T17:32:32.569534Z","shell.execute_reply.started":"2022-07-23T17:32:32.481716Z","shell.execute_reply":"2022-07-23T17:32:32.568454Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### how about glucose? (mean and std)","metadata":{}},{"cell_type":"code","source":"mean = np.round(diabetes['Glucose'].mean(), 2)\nstd = np.round(diabetes['Glucose'].std(), 2)\n\nprint(f'Mean (std) = {mean} ({std})')","metadata":{"execution":{"iopub.status.busy":"2022-07-23T17:32:32.573138Z","iopub.execute_input":"2022-07-23T17:32:32.573509Z","iopub.status.idle":"2022-07-23T17:32:32.581033Z","shell.execute_reply.started":"2022-07-23T17:32:32.573475Z","shell.execute_reply":"2022-07-23T17:32:32.579794Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### glucose quantiles (q1=0.25, q2=0.5, q3=0.75)","metadata":{}},{"cell_type":"code","source":"q1 = np.round(diabetes['Glucose'].quantile(0.25), 2)\nq2 = np.round(diabetes['Glucose'].quantile(0.5), 2)\nq3 = np.round(diabetes['Glucose'].quantile(0.75), 2)\n\nprint(f'q1= {q1} / q2={q2} / q3={q3}')","metadata":{"execution":{"iopub.status.busy":"2022-07-23T17:32:32.582758Z","iopub.execute_input":"2022-07-23T17:32:32.583456Z","iopub.status.idle":"2022-07-23T17:32:32.598841Z","shell.execute_reply.started":"2022-07-23T17:32:32.583418Z","shell.execute_reply":"2022-07-23T17:32:32.597701Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### glucose range","metadata":{}},{"cell_type":"code","source":"ma = diabetes['Glucose'].max()\nmi = diabetes['Glucose'].min()\n\nprint(f'Min = {mi} / Max ={ma}')","metadata":{"execution":{"iopub.status.busy":"2022-07-23T17:32:32.600572Z","iopub.execute_input":"2022-07-23T17:32:32.600988Z","iopub.status.idle":"2022-07-23T17:32:32.611798Z","shell.execute_reply.started":"2022-07-23T17:32:32.600945Z","shell.execute_reply":"2022-07-23T17:32:32.610843Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### dataviz (histogram)","metadata":{}},{"cell_type":"code","source":"fig = px.histogram(diabetes, x='Glucose')\nfig.show()","metadata":{"execution":{"iopub.status.busy":"2022-07-23T17:32:32.612944Z","iopub.execute_input":"2022-07-23T17:32:32.613339Z","iopub.status.idle":"2022-07-23T17:32:32.674679Z","shell.execute_reply.started":"2022-07-23T17:32:32.613306Z","shell.execute_reply":"2022-07-23T17:32:32.673488Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### dataviz (boxplot)","metadata":{}},{"cell_type":"code","source":"fig = px.box(diabetes, y='Glucose')\nfig.show()","metadata":{"execution":{"iopub.status.busy":"2022-07-23T17:32:32.676290Z","iopub.execute_input":"2022-07-23T17:32:32.676897Z","iopub.status.idle":"2022-07-23T17:32:32.734510Z","shell.execute_reply.started":"2022-07-23T17:32:32.676859Z","shell.execute_reply":"2022-07-23T17:32:32.733295Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### how about glucose without 0s?","metadata":{}},{"cell_type":"code","source":"diabetes2 = diabetes[diabetes['Glucose'] > 0]\n\nmean = np.round(diabetes2['Glucose'].mean(), 2)\nstd = np.round(diabetes2['Glucose'].std(), 2)\n\nprint(f'Mean (std) = {mean} ({std})')\n\nq1 = np.round(diabetes2['Glucose'].quantile(0.25), 2)\nq2 = np.round(diabetes2['Glucose'].quantile(0.5), 2)\nq3 = np.round(diabetes2['Glucose'].quantile(0.75), 2)\n\nprint(f'q1= {q1} / q2={q2} / q3={q3}')\n\n\n\nfig = px.histogram(diabetes2, x='Glucose')\nfig.show()\n\nfig = px.box(diabetes2, y='Glucose')\nfig.show()","metadata":{"execution":{"iopub.status.busy":"2022-07-23T17:32:32.736682Z","iopub.execute_input":"2022-07-23T17:32:32.737151Z","iopub.status.idle":"2022-07-23T17:32:32.865594Z","shell.execute_reply.started":"2022-07-23T17:32:32.737104Z","shell.execute_reply":"2022-07-23T17:32:32.862672Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### glucose stratified by outcome (without 0s)","metadata":{}},{"cell_type":"code","source":"fig = px.box(diabetes2, y='Glucose', x='Outcome')\nfig.show()","metadata":{"execution":{"iopub.status.busy":"2022-07-23T17:32:32.867399Z","iopub.execute_input":"2022-07-23T17:32:32.867978Z","iopub.status.idle":"2022-07-23T17:32:32.927241Z","shell.execute_reply.started":"2022-07-23T17:32:32.867930Z","shell.execute_reply":"2022-07-23T17:32:32.925955Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"fig = px.histogram(diabetes2, x='Glucose', color=\"Outcome\", marginal=\"box\")\nfig.show()","metadata":{"execution":{"iopub.status.busy":"2022-07-23T17:32:32.928939Z","iopub.execute_input":"2022-07-23T17:32:32.929987Z","iopub.status.idle":"2022-07-23T17:32:33.022986Z","shell.execute_reply.started":"2022-07-23T17:32:32.929951Z","shell.execute_reply":"2022-07-23T17:32:33.022028Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"fig = px.histogram(diabetes2, x='Glucose', color=\"Outcome\", marginal=\"box\", histnorm='probability density')\nfig.show()","metadata":{"execution":{"iopub.status.busy":"2022-07-23T17:32:33.024425Z","iopub.execute_input":"2022-07-23T17:32:33.024725Z","iopub.status.idle":"2022-07-23T17:32:33.116396Z","shell.execute_reply.started":"2022-07-23T17:32:33.024697Z","shell.execute_reply":"2022-07-23T17:32:33.114948Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## Bivariate analysis (Wine quality)","metadata":{}},{"cell_type":"code","source":"wine = pd.read_csv('/kaggle/input/red-wine-quality-cortez-et-al-2009/winequality-red.csv')\nprint(wine.shape)\nwine.head(10)","metadata":{"execution":{"iopub.status.busy":"2022-07-23T17:32:33.117754Z","iopub.execute_input":"2022-07-23T17:32:33.118117Z","iopub.status.idle":"2022-07-23T17:32:33.150987Z","shell.execute_reply.started":"2022-07-23T17:32:33.118085Z","shell.execute_reply":"2022-07-23T17:32:33.149788Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"%matplotlib inline\n\nAV = AutoViz_Class()\ndft = AV.AutoViz(filename=\"\", dfte=wine, chart_format='png')","metadata":{"execution":{"iopub.status.busy":"2022-07-23T17:32:33.152476Z","iopub.execute_input":"2022-07-23T17:32:33.152770Z","iopub.status.idle":"2022-07-23T17:32:57.956120Z","shell.execute_reply.started":"2022-07-23T17:32:33.152744Z","shell.execute_reply":"2022-07-23T17:32:57.954786Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## Bivariate analysis (density and fixed acidity)","metadata":{}},{"cell_type":"code","source":"fig = px.scatter(wine, x=\"density\", y=\"fixed acidity\", trendline=\"ols\")\nfig.show()","metadata":{"execution":{"iopub.status.busy":"2022-07-23T17:32:57.957785Z","iopub.execute_input":"2022-07-23T17:32:57.958160Z","iopub.status.idle":"2022-07-23T17:32:58.712866Z","shell.execute_reply.started":"2022-07-23T17:32:57.958125Z","shell.execute_reply":"2022-07-23T17:32:58.711790Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## Bivariate analysis (pH and fixed acidity)","metadata":{}},{"cell_type":"code","source":"fig = px.scatter(wine, x=\"fixed acidity\", y=\"pH\", trendline=\"ols\")\nfig.show()","metadata":{"execution":{"iopub.status.busy":"2022-07-23T17:41:12.689264Z","iopub.execute_input":"2022-07-23T17:41:12.689651Z","iopub.status.idle":"2022-07-23T17:41:12.757104Z","shell.execute_reply.started":"2022-07-23T17:41:12.689620Z","shell.execute_reply":"2022-07-23T17:41:12.756337Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### (Anscombe’s quartet)","metadata":{}},{"cell_type":"code","source":"df = sns.load_dataset(\"anscombe\")\ndf","metadata":{"execution":{"iopub.status.busy":"2022-07-23T18:45:24.052917Z","iopub.execute_input":"2022-07-23T18:45:24.053301Z","iopub.status.idle":"2022-07-23T18:45:25.340936Z","shell.execute_reply.started":"2022-07-23T18:45:24.053271Z","shell.execute_reply":"2022-07-23T18:45:25.339862Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"df.groupby(['dataset']).agg(['mean', 'std'])","metadata":{"execution":{"iopub.status.busy":"2022-07-23T18:49:35.788811Z","iopub.execute_input":"2022-07-23T18:49:35.789789Z","iopub.status.idle":"2022-07-23T18:49:35.808537Z","shell.execute_reply.started":"2022-07-23T18:49:35.789729Z","shell.execute_reply":"2022-07-23T18:49:35.807398Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"fig = px.scatter(df[df['dataset'] == 'I'], x=\"x\", y=\"y\", trendline=\"ols\")\nfig.show()","metadata":{"execution":{"iopub.status.busy":"2022-07-23T18:47:58.566946Z","iopub.execute_input":"2022-07-23T18:47:58.567328Z","iopub.status.idle":"2022-07-23T18:47:58.641888Z","shell.execute_reply.started":"2022-07-23T18:47:58.567299Z","shell.execute_reply":"2022-07-23T18:47:58.640866Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"fig = px.scatter(df[df['dataset'] == 'II'], x=\"x\", y=\"y\", trendline=\"ols\")\nfig.show()","metadata":{"execution":{"iopub.status.busy":"2022-07-23T18:48:10.833726Z","iopub.execute_input":"2022-07-23T18:48:10.834117Z","iopub.status.idle":"2022-07-23T18:48:10.900635Z","shell.execute_reply.started":"2022-07-23T18:48:10.834086Z","shell.execute_reply":"2022-07-23T18:48:10.899415Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"fig = px.scatter(df[df['dataset'] == 'III'], x=\"x\", y=\"y\", trendline=\"ols\")\nfig.show()","metadata":{"execution":{"iopub.status.busy":"2022-07-23T18:48:27.034320Z","iopub.execute_input":"2022-07-23T18:48:27.035026Z","iopub.status.idle":"2022-07-23T18:48:27.101354Z","shell.execute_reply.started":"2022-07-23T18:48:27.034972Z","shell.execute_reply":"2022-07-23T18:48:27.100591Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"fig = px.scatter(df[df['dataset'] == 'IV'], x=\"x\", y=\"y\", trendline=\"ols\")\nfig.show()","metadata":{"execution":{"iopub.status.busy":"2022-07-23T18:48:45.316455Z","iopub.execute_input":"2022-07-23T18:48:45.317189Z","iopub.status.idle":"2022-07-23T18:48:45.384565Z","shell.execute_reply.started":"2022-07-23T18:48:45.317153Z","shell.execute_reply":"2022-07-23T18:48:45.383365Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## Multivariate analysis - Breast Cancer","metadata":{}},{"cell_type":"code","source":"cancer = pd.read_csv('/kaggle/input/breast-cancer-wisconsin-data/data.csv')\nprint(cancer.shape)\ncancer.head(10)","metadata":{"execution":{"iopub.status.busy":"2022-07-23T17:43:05.232773Z","iopub.execute_input":"2022-07-23T17:43:05.233199Z","iopub.status.idle":"2022-07-23T17:43:05.278639Z","shell.execute_reply.started":"2022-07-23T17:43:05.233164Z","shell.execute_reply":"2022-07-23T17:43:05.277722Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"%matplotlib inline\n\nAV = AutoViz_Class()\ndft = AV.AutoViz(filename=\"\", dfte=cancer, depVar='diagnosis', chart_format='png')","metadata":{"execution":{"iopub.status.busy":"2022-07-23T17:44:07.433496Z","iopub.execute_input":"2022-07-23T17:44:07.433867Z","iopub.status.idle":"2022-07-23T17:44:32.963774Z","shell.execute_reply.started":"2022-07-23T17:44:07.433838Z","shell.execute_reply":"2022-07-23T17:44:32.962540Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## Inferential statistics (Titanic)","metadata":{}},{"cell_type":"code","source":"titanic = pd.read_csv('/kaggle/input/titanic/train.csv')\nprint(titanic.shape)\ntitanic.head(10)","metadata":{"execution":{"iopub.status.busy":"2022-07-23T18:27:42.740477Z","iopub.execute_input":"2022-07-23T18:27:42.741065Z","iopub.status.idle":"2022-07-23T18:27:42.768880Z","shell.execute_reply.started":"2022-07-23T18:27:42.741027Z","shell.execute_reply":"2022-07-23T18:27:42.767644Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"titanic['Pclass'] = titanic['Pclass'].astype(\"str\")","metadata":{"execution":{"iopub.status.busy":"2022-07-23T18:27:44.768780Z","iopub.execute_input":"2022-07-23T18:27:44.770122Z","iopub.status.idle":"2022-07-23T18:27:44.777529Z","shell.execute_reply.started":"2022-07-23T18:27:44.770063Z","shell.execute_reply":"2022-07-23T18:27:44.776244Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"%matplotlib inline\n\nAV = AutoViz_Class()\ndft = AV.AutoViz(filename=\"\", dfte=titanic, depVar='Survived', chart_format='png')","metadata":{"execution":{"iopub.status.busy":"2022-07-23T18:17:04.614998Z","iopub.execute_input":"2022-07-23T18:17:04.615366Z","iopub.status.idle":"2022-07-23T18:17:12.901606Z","shell.execute_reply.started":"2022-07-23T18:17:04.615337Z","shell.execute_reply":"2022-07-23T18:17:12.899787Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### What's the rule?","metadata":{}},{"cell_type":"markdown","source":"### Rule 1 - Everybody dies","metadata":{}},{"cell_type":"code","source":"titanic['Prediction'] = 0\ntitanic","metadata":{"execution":{"iopub.status.busy":"2022-07-23T18:29:25.637639Z","iopub.execute_input":"2022-07-23T18:29:25.637999Z","iopub.status.idle":"2022-07-23T18:29:25.666312Z","shell.execute_reply.started":"2022-07-23T18:29:25.637970Z","shell.execute_reply":"2022-07-23T18:29:25.665152Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"accuracy_score(titanic['Survived'], titanic['Prediction'])","metadata":{"execution":{"iopub.status.busy":"2022-07-23T18:29:27.864743Z","iopub.execute_input":"2022-07-23T18:29:27.865123Z","iopub.status.idle":"2022-07-23T18:29:27.872979Z","shell.execute_reply.started":"2022-07-23T18:29:27.865093Z","shell.execute_reply":"2022-07-23T18:29:27.872030Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"skplt.metrics.plot_confusion_matrix(titanic['Survived'], titanic['Prediction'], normalize=True)","metadata":{"execution":{"iopub.status.busy":"2022-07-23T18:29:48.644080Z","iopub.execute_input":"2022-07-23T18:29:48.644492Z","iopub.status.idle":"2022-07-23T18:29:48.868268Z","shell.execute_reply.started":"2022-07-23T18:29:48.644457Z","shell.execute_reply":"2022-07-23T18:29:48.867170Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### Rule 2 - Every man in pclass 3 dies (otherwise everybody survives)","metadata":{}},{"cell_type":"code","source":"titanic['Prediction'] = 1\ntitanic.loc[(titanic['Pclass'] == '3') & (titanic['Sex'] == 'male'), 'Prediction'] = 0\ntitanic","metadata":{"execution":{"iopub.status.busy":"2022-07-23T18:30:03.235109Z","iopub.execute_input":"2022-07-23T18:30:03.235508Z","iopub.status.idle":"2022-07-23T18:30:03.265963Z","shell.execute_reply.started":"2022-07-23T18:30:03.235478Z","shell.execute_reply":"2022-07-23T18:30:03.264837Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"accuracy_score(titanic['Survived'], titanic['Prediction'])","metadata":{"execution":{"iopub.status.busy":"2022-07-23T18:30:06.351733Z","iopub.execute_input":"2022-07-23T18:30:06.352127Z","iopub.status.idle":"2022-07-23T18:30:06.361195Z","shell.execute_reply.started":"2022-07-23T18:30:06.352096Z","shell.execute_reply":"2022-07-23T18:30:06.360068Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"skplt.metrics.plot_confusion_matrix(titanic['Survived'], titanic['Prediction'], normalize=True)","metadata":{"execution":{"iopub.status.busy":"2022-07-23T18:30:09.015732Z","iopub.execute_input":"2022-07-23T18:30:09.016158Z","iopub.status.idle":"2022-07-23T18:30:09.247694Z","shell.execute_reply.started":"2022-07-23T18:30:09.016118Z","shell.execute_reply":"2022-07-23T18:30:09.246461Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## Exercise: House price regression","metadata":{}},{"cell_type":"code","source":"house = pd.read_csv('/kaggle/input/house-prices-advanced-regression-techniques/train.csv')\nprint(house.shape)\nhouse.head(10)","metadata":{"execution":{"iopub.status.busy":"2022-07-23T18:37:37.473154Z","iopub.execute_input":"2022-07-23T18:37:37.473543Z","iopub.status.idle":"2022-07-23T18:37:37.526516Z","shell.execute_reply.started":"2022-07-23T18:37:37.473505Z","shell.execute_reply":"2022-07-23T18:37:37.525743Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"1. What are the main statistical measures we can extract from the data?\n2. What are the main visualizations that we can present?\n3. Is it possible to create a rule for the sale price of houses? (Using only the stats seen so far)","metadata":{}},{"cell_type":"markdown","source":"![](https://www.aluralingua.com.br/artigos/assets/thank-you.jpg)","metadata":{}}]}