{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"pygments_lexer":"ipython3","nbconvert_exporter":"python","version":"3.6.4","file_extension":".py","codemirror_mode":{"name":"ipython","version":3},"name":"python","mimetype":"text/x-python"}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"# Simple linear regression with Python\n[Josh Hiller based off a notebook taken from](http://pmarcelino.com) - June/11/2022\n\nOther Kernels: [Data analysis and feature extraction with Python\n](https://www.kaggle.com/pmarcelino/data-analysis-and-feature-extraction-with-python)\n\n----------","metadata":{"_uuid":"d507f816cc74a88c9afefd02cf225d1a7dd6461f","_cell_guid":"e3bc4854-2787-eae1-950d-2742ad3d7db2"}},{"cell_type":"markdown","source":"Yesterday we looked at how to sort data, and how to extract data from a data set. Today, we are going to continue our work by examining a data set of home sales. ","metadata":{"_uuid":"014f5c099a26d9232f9d0d6ca85d5c02b812c98a","_cell_guid":"8ca352d7-08aa-36b4-fb2d-3c9854a8d86a"}},{"cell_type":"code","source":"#invite people for the Kaggle party\nimport pandas as pd\nimport matplotlib.pyplot as plt\nimport seaborn as sns\nimport numpy as np\nfrom scipy.stats import norm\nfrom sklearn.preprocessing import StandardScaler\nfrom sklearn.linear_model import LinearRegression\nfrom scipy import stats\nimport warnings\nwarnings.filterwarnings('ignore')\n%matplotlib inline","metadata":{"_uuid":"d581f6797b9fde1580271358d484df67bf6b14a1","_cell_guid":"2df621e0-e03c-7aaa-6e08-40ed1d7dfecc","_execution_state":"idle","execution":{"iopub.status.busy":"2022-07-11T20:12:26.822467Z","iopub.execute_input":"2022-07-11T20:12:26.822783Z","iopub.status.idle":"2022-07-11T20:12:26.831883Z","shell.execute_reply.started":"2022-07-11T20:12:26.822724Z","shell.execute_reply":"2022-07-11T20:12:26.830956Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"#bring in the six packs\ndf_train = pd.read_csv('../input/train.csv')","metadata":{"_uuid":"827a72128cd211cf6af16b003e7c09951e3f2b1e","_cell_guid":"d56d5e71-4277-7a74-5306-7d5af4c7f263","_execution_state":"idle","execution":{"iopub.status.busy":"2022-07-11T19:43:50.427425Z","iopub.execute_input":"2022-07-11T19:43:50.427730Z","iopub.status.idle":"2022-07-11T19:43:50.474045Z","shell.execute_reply.started":"2022-07-11T19:43:50.427666Z","shell.execute_reply":"2022-07-11T19:43:50.473017Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"#check the column names\ndf_train.columns","metadata":{"_uuid":"10814ba44786b5fea5e333324c6fe54729cabf33","_cell_guid":"02250c81-7e15-c195-2e86-5adbd15c9d30","_execution_state":"idle","execution":{"iopub.status.busy":"2022-07-11T19:43:54.526571Z","iopub.execute_input":"2022-07-11T19:43:54.526878Z","iopub.status.idle":"2022-07-11T19:43:54.535630Z","shell.execute_reply.started":"2022-07-11T19:43:54.526821Z","shell.execute_reply":"2022-07-11T19:43:54.534888Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Now let us examine the SalePrice column","metadata":{}},{"cell_type":"code","source":"#descriptive statistics summary\ndf_train['SalePrice'].describe()","metadata":{"_uuid":"5c15e1bd10b8e71c0b1d62bdb260882585a35579","_cell_guid":"54452e23-f4d3-919f-c734-80a35dc9ae08","_execution_state":"idle","execution":{"iopub.status.busy":"2022-07-11T19:44:08.991373Z","iopub.execute_input":"2022-07-11T19:44:08.992013Z","iopub.status.idle":"2022-07-11T19:44:09.006385Z","shell.execute_reply.started":"2022-07-11T19:44:08.991950Z","shell.execute_reply":"2022-07-11T19:44:09.004814Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"What do you think the above output represents? ","metadata":{"_uuid":"bb3005e8025ea75b0b4d1ef624927e7e21833ea1","_cell_guid":"6af460e5-1be2-6618-d624-2a4423bd501f"}},{"cell_type":"markdown","source":"Below is the code for a histogram. After running it, discuss the story it is telling you with your group members. ","metadata":{}},{"cell_type":"code","source":"\n#histogram\nsns.distplot(df_train['SalePrice']);","metadata":{"_uuid":"2f78c77caa7290298138caf167672e62d3bc5a67","_cell_guid":"6bbea362-77b6-5385-f0a8-fb53afd088b7","_execution_state":"idle","execution":{"iopub.status.busy":"2022-07-11T19:44:16.207635Z","iopub.execute_input":"2022-07-11T19:44:16.208274Z","iopub.status.idle":"2022-07-11T19:44:16.751710Z","shell.execute_reply.started":"2022-07-11T19:44:16.208209Z","shell.execute_reply":"2022-07-11T19:44:16.750427Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# 'SalePrice', her buddies and her interests","metadata":{"_uuid":"e0368771327b15853b3e306f39eb93d209594e75","_cell_guid":"c477edc1-b472-f55b-ba98-514159dbda2e"}},{"cell_type":"markdown","source":"We just created a histogram, but as we saw a couple of days ago when we used Google Sheets, it is often helpful to have a scatterplot. ","metadata":{}},{"cell_type":"markdown","source":"Look at the code below. What do you think is the purpose of the .concat() function? ","metadata":{}},{"cell_type":"code","source":"#scatter plot grlivarea/saleprice\nx_var = 'GrLivArea'\ndata = pd.concat([df_train['SalePrice'], df_train[x_var]], axis=1)\ndata.plot.scatter(x=x_var, y='SalePrice', ylim=(0,800000));","metadata":{"_uuid":"91160363898f5caeee965a1aa81eb3abb7dcd760","_cell_guid":"db040973-0adc-e126-e657-1d8934b5a5c8","_execution_state":"idle","execution":{"iopub.status.busy":"2022-07-11T19:54:39.876000Z","iopub.execute_input":"2022-07-11T19:54:39.876336Z","iopub.status.idle":"2022-07-11T19:54:40.244645Z","shell.execute_reply.started":"2022-07-11T19:54:39.876273Z","shell.execute_reply":"2022-07-11T19:54:40.243356Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Does it look like there is a linear relationship between the variables?*\n\n*And what about 'TotalBsmtSF'?*","metadata":{"_uuid":"6775955416e9b2d4ac43e3fec5685c8577138455","_cell_guid":"c3dccb06-206a-20c0-6060-8d5b49f4df0b"}},{"cell_type":"code","source":"#scatter plot totalbsmtsf/saleprice\nvar = 'TotalBsmtSF'\ndata = pd.concat([df_train['SalePrice'], df_train[var]], axis=1)\ndata.plot.scatter(x=var, y='SalePrice', ylim=(0,800000));","metadata":{"_uuid":"3ac3db51311338fcdc16014a7c506cf3d5315af7","_cell_guid":"353def35-0f26-998d-b9a4-7356f95e80ad","_execution_state":"idle","execution":{"iopub.status.busy":"2022-07-11T19:44:35.447437Z","iopub.execute_input":"2022-07-11T19:44:35.447879Z","iopub.status.idle":"2022-07-11T19:44:35.888035Z","shell.execute_reply.started":"2022-07-11T19:44:35.447836Z","shell.execute_reply":"2022-07-11T19:44:35.886901Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"*'TotalBsmtSF' is also a great friend of 'SalePrice' but this seems a much more emotional relationship! Everything is ok and suddenly, in a <b>strong linear (exponential?)</b> reaction, everything changes. Moreover, it's clear that sometimes 'TotalBsmtSF' closes in itself and gives zero credit to 'SalePrice'.*","metadata":{"_uuid":"4496c7d3635636e8e19296f3b5b52fdba16f7cc5","_cell_guid":"7ea69b90-de6d-c104-ff92-a49943993930"}},{"cell_type":"markdown","source":"Now let us do some linear regression.","metadata":{}},{"cell_type":"code","source":"\nlinear_regressor = LinearRegression()  # create object for the class\n\nX = data.iloc[:, 0].values.reshape(-1, 1)  # values converts it into a numpy array\nY = data.iloc[:, 1].values.reshape(-1, 1)  # -1 means that calculate the dimension of rows, but have 1 column\nlinear_regressor = LinearRegression()  # create object for the class\nlinear_regressor.fit(X, Y)  # perform linear regression\nY_pred = linear_regressor.predict(X)  # make predictions\n","metadata":{"execution":{"iopub.status.busy":"2022-07-11T20:13:50.653188Z","iopub.execute_input":"2022-07-11T20:13:50.653735Z","iopub.status.idle":"2022-07-11T20:13:50.661882Z","shell.execute_reply.started":"2022-07-11T20:13:50.653654Z","shell.execute_reply":"2022-07-11T20:13:50.660966Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Now we will plot our line of best fit. ","metadata":{}},{"cell_type":"code","source":"\nplt.scatter(X, Y)\nplt.plot(X, Y_pred, color='red')\nplt.show()","metadata":{"execution":{"iopub.status.busy":"2022-07-11T20:14:09.265315Z","iopub.execute_input":"2022-07-11T20:14:09.265894Z","iopub.status.idle":"2022-07-11T20:14:09.575975Z","shell.execute_reply.started":"2022-07-11T20:14:09.265801Z","shell.execute_reply":"2022-07-11T20:14:09.574542Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"**Transformations**","metadata":{}},{"cell_type":"markdown","source":"Recall that we learned that sometimes you have to transform your data. Below is how you would do that. ","metadata":{}},{"cell_type":"code","source":"#applying log transformation\ndf_train['SalePrice'] = np.log(df_train['SalePrice'])","metadata":{"_uuid":"f578838e98e9996b09abbec058200cb18aa38869","_cell_guid":"cced5b14-c39d-c847-6dc9-93af3f4b6e6d","_execution_state":"idle","execution":{"iopub.status.busy":"2022-07-11T19:58:31.113561Z","iopub.execute_input":"2022-07-11T19:58:31.113849Z","iopub.status.idle":"2022-07-11T19:58:31.120023Z","shell.execute_reply.started":"2022-07-11T19:58:31.113808Z","shell.execute_reply":"2022-07-11T19:58:31.119043Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"#transformed histogram and normal probability plot\nsns.distplot(df_train['SalePrice'], fit=norm);","metadata":{"_uuid":"de8366b3ad71c7cb398644766a412bee1d05642f","_cell_guid":"0e17fba2-3ff2-d6f1-841d-bc2a9746bcc2","_execution_state":"idle","execution":{"iopub.status.busy":"2022-07-11T19:58:50.482786Z","iopub.execute_input":"2022-07-11T19:58:50.483311Z","iopub.status.idle":"2022-07-11T19:58:51.022157Z","shell.execute_reply.started":"2022-07-11T19:58:50.483239Z","shell.execute_reply":"2022-07-11T19:58:51.021003Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Done! Let's check what's going on with 'GrLivArea'.","metadata":{"_uuid":"19b172dfedfef5330c53cc34837d47641a3cf135","_cell_guid":"74131161-b766-8b08-5ae3-0b70688a198a"}},{"cell_type":"code","source":"#histogram and normal probability plot\nsns.distplot(df_train['GrLivArea'], fit=norm);\n","metadata":{"_uuid":"aa29a3612ee861c42e5d6022ac3a8ad75266b0ee","_cell_guid":"bda5d77e-07ea-1d16-644c-a850487cc35d","_execution_state":"idle","execution":{"iopub.status.busy":"2022-07-11T20:00:27.819858Z","iopub.execute_input":"2022-07-11T20:00:27.820191Z","iopub.status.idle":"2022-07-11T20:00:28.347339Z","shell.execute_reply.started":"2022-07-11T20:00:27.820136Z","shell.execute_reply":"2022-07-11T20:00:28.346026Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Tastes like skewness... *Avada kedavra!*","metadata":{"_uuid":"f452585a30fe3006d4dc993941e826c793044a4b","_cell_guid":"7a8fa54e-46e2-5c09-0a43-0b200ee59874"}},{"cell_type":"code","source":"#data transformation\ndf_train['GrLivArea'] = np.log(df_train['GrLivArea'])","metadata":{"_uuid":"aa428f580e39e498dc06747423955fa714d97952","_cell_guid":"c0fe3abe-cc32-6f2e-0e79-1afd9b49c7ad","_execution_state":"idle","execution":{"iopub.status.busy":"2022-07-11T20:00:40.913088Z","iopub.execute_input":"2022-07-11T20:00:40.913440Z","iopub.status.idle":"2022-07-11T20:00:40.919470Z","shell.execute_reply.started":"2022-07-11T20:00:40.913363Z","shell.execute_reply":"2022-07-11T20:00:40.918512Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"#transformed histogram and normal probability plot\nsns.distplot(df_train['GrLivArea'], fit=norm);\n","metadata":{"_uuid":"bb004d70d1f260289f864c2ecb5c5dceedc2d6fa","_cell_guid":"0171b317-885d-7092-ade9-d2332b8ea796","_execution_state":"idle","execution":{"iopub.status.busy":"2022-07-11T20:00:55.705271Z","iopub.execute_input":"2022-07-11T20:00:55.705953Z","iopub.status.idle":"2022-07-11T20:00:56.137586Z","shell.execute_reply.started":"2022-07-11T20:00:55.705887Z","shell.execute_reply":"2022-07-11T20:00:56.136732Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"What happened when we applied the log transformation to the distribution of the data?","metadata":{"_uuid":"bdd45521276779b330af14f279bf578a9457dac0","_cell_guid":"1686254a-4391-a744-b8ae-653364e216d4"}},{"cell_type":"markdown","source":"**Use the data from the Google Sheet. Import it, run and plot the transformed log-log incidence vs time.**","metadata":{}},{"cell_type":"markdown","source":"Last week's Google Sheet is accessible [here](https://docs.google.com/spreadsheets/d/10e68ZKePEZv7YCihvjKFUDvLl5Jz7RHGDpVYz1c-I7g/edit#gid=0).","metadata":{}},{"cell_type":"markdown","source":"**When you are done, download COVID data from NY state. Plot a scatter plot and begin exploring the data.**","metadata":{}}]}