{"cells":[{"metadata":{"toc":true,"_uuid":"16cbd1780771996d688b3bad000e10465b206a18"},"cell_type":"markdown","source":"<h1>Table of Contents<span class=\"tocSkip\"></span></h1>\n<div class=\"toc\"><ul class=\"toc-item\"><li><span><a href=\"#Introduction\" data-toc-modified-id=\"Introduction-1\"><span class=\"toc-item-num\">1&nbsp;&nbsp;</span>Introduction</a></span><ul class=\"toc-item\"><li><span><a href=\"#Importing-modules\" data-toc-modified-id=\"Importing-modules-1.1\"><span class=\"toc-item-num\">1.1&nbsp;&nbsp;</span>Importing modules</a></span></li><li><span><a href=\"#Importing-data\" data-toc-modified-id=\"Importing-data-1.2\"><span class=\"toc-item-num\">1.2&nbsp;&nbsp;</span>Importing data</a></span></li><li><span><a href=\"#Mini-preprocessing\" data-toc-modified-id=\"Mini-preprocessing-1.3\"><span class=\"toc-item-num\">1.3&nbsp;&nbsp;</span>Mini-preprocessing</a></span></li></ul></li><li><span><a href=\"#Assumptions\" data-toc-modified-id=\"Assumptions-2\"><span class=\"toc-item-num\">2&nbsp;&nbsp;</span>Assumptions</a></span><ul class=\"toc-item\"><li><span><a href=\"#Numerical-variables\" data-toc-modified-id=\"Numerical-variables-2.1\"><span class=\"toc-item-num\">2.1&nbsp;&nbsp;</span>Numerical variables</a></span></li><li><span><a href=\"#No-outliers\" data-toc-modified-id=\"No-outliers-2.2\"><span class=\"toc-item-num\">2.2&nbsp;&nbsp;</span>No outliers</a></span></li><li><span><a href=\"#Normal-distribution\" data-toc-modified-id=\"Normal-distribution-2.3\"><span class=\"toc-item-num\">2.3&nbsp;&nbsp;</span>Normal distribution</a></span></li><li><span><a href=\"#Linear-relationship\" data-toc-modified-id=\"Linear-relationship-2.4\"><span class=\"toc-item-num\">2.4&nbsp;&nbsp;</span>Linear relationship</a></span></li><li><span><a href=\"#No-multicollinearity\" data-toc-modified-id=\"No-multicollinearity-2.5\"><span class=\"toc-item-num\">2.5&nbsp;&nbsp;</span>No multicollinearity</a></span><ul class=\"toc-item\"><li><span><a href=\"#VIF-to-check-multicolinearity\" data-toc-modified-id=\"VIF-to-check-multicolinearity-2.5.1\"><span class=\"toc-item-num\">2.5.1&nbsp;&nbsp;</span>VIF to check multicolinearity</a></span></li></ul></li><li><span><a href=\"#No-autocorrelation\" data-toc-modified-id=\"No-autocorrelation-2.6\"><span class=\"toc-item-num\">2.6&nbsp;&nbsp;</span>No autocorrelation</a></span></li><li><span><a href=\"#Homoskedasticity\" data-toc-modified-id=\"Homoskedasticity-2.7\"><span class=\"toc-item-num\">2.7&nbsp;&nbsp;</span>Homoskedasticity</a></span></li></ul></li><li><span><a href=\"#Fitting-&amp;-Predicting\" data-toc-modified-id=\"Fitting-&amp;-Predicting-3\"><span class=\"toc-item-num\">3&nbsp;&nbsp;</span>Fitting &amp; Predicting</a></span><ul class=\"toc-item\"><li><span><a href=\"#The-concept\" data-toc-modified-id=\"The-concept-3.1\"><span class=\"toc-item-num\">3.1&nbsp;&nbsp;</span>The concept</a></span></li><li><span><a href=\"#A-basic-algorithm-:-Least-Squares\" data-toc-modified-id=\"A-basic-algorithm-:-Least-Squares-3.2\"><span class=\"toc-item-num\">3.2&nbsp;&nbsp;</span>A basic algorithm : Least Squares</a></span></li></ul></li><li><span><a href=\"#Testing\" data-toc-modified-id=\"Testing-4\"><span class=\"toc-item-num\">4&nbsp;&nbsp;</span>Testing</a></span><ul class=\"toc-item\"><li><span><a href=\"#Train-&amp;-Test\" data-toc-modified-id=\"Train-&amp;-Test-4.1\"><span class=\"toc-item-num\">4.1&nbsp;&nbsp;</span>Train &amp; Test</a></span></li><li><span><a href=\"#Statistics\" data-toc-modified-id=\"Statistics-4.2\"><span class=\"toc-item-num\">4.2&nbsp;&nbsp;</span>Statistics</a></span></li><li><span><a href=\"#Cross-validation\" data-toc-modified-id=\"Cross-validation-4.3\"><span class=\"toc-item-num\">4.3&nbsp;&nbsp;</span>Cross-validation</a></span></li></ul></li><li><span><a href=\"#Seeking-Performance\" data-toc-modified-id=\"Seeking-Performance-5\"><span class=\"toc-item-num\">5&nbsp;&nbsp;</span>Seeking Performance</a></span><ul class=\"toc-item\"><li><span><a href=\"#Regularizing\" data-toc-modified-id=\"Regularizing-5.1\"><span class=\"toc-item-num\">5.1&nbsp;&nbsp;</span>Regularizing</a></span><ul class=\"toc-item\"><li><span><a href=\"#Overfitting-&amp;-Underfitting\" data-toc-modified-id=\"Overfitting-&amp;-Underfitting-5.1.1\"><span class=\"toc-item-num\">5.1.1&nbsp;&nbsp;</span>Overfitting &amp; Underfitting</a></span></li><li><span><a href=\"#Regularized-Models\" data-toc-modified-id=\"Regularized-Models-5.1.2\"><span class=\"toc-item-num\">5.1.2&nbsp;&nbsp;</span>Regularized Models</a></span></li></ul></li><li><span><a href=\"#Scenario-testing\" data-toc-modified-id=\"Scenario-testing-5.2\"><span class=\"toc-item-num\">5.2&nbsp;&nbsp;</span>Scenario testing</a></span></li><li><span><a href=\"#Submitting\" data-toc-modified-id=\"Submitting-5.3\"><span class=\"toc-item-num\">5.3&nbsp;&nbsp;</span>Submitting</a></span></li></ul></li></ul></div>"},{"metadata":{"_uuid":"9cf14bf40d74456e25c0a2b0d355c7f70f356bf1"},"cell_type":"markdown","source":"# Introduction"},{"metadata":{"_uuid":"4d490980e46a4e0dca1c44f6721fc053dd594f4a"},"cell_type":"markdown","source":"This tutorial intents to summarize the most essential things to know about linear regressions, how to apply them, what are the particular cases to watch out for and what to respond to them.\n\nIt is written in a didactic purpose rather than to seek performance in prediction. <br>\nTherefore we will spend minimum time on data cleaning in order to get straight to the point.\n<br><br>\n"},{"metadata":{"_uuid":"1421d42dee0bcecf6c32a9ea80f0b19baab7a356"},"cell_type":"markdown","source":"## Importing modules"},{"metadata":{"_uuid":"11646dfdd4982cb4d756339a5d69464836b70b5c","trusted":false},"cell_type":"code","source":"import warnings\nwarnings.filterwarnings('ignore')\n\nimport pandas as pd\npd.options.display.max_columns = None\npd.set_option('display.float_format', lambda x: '%.6f' % x)\nimport numpy as np\nimport seaborn as sns\nfrom matplotlib import pyplot as plt\nplt.style.use('ggplot')\nfrom sklearn import model_selection\nfrom sklearn.linear_model import LinearRegression, Ridge, Lasso, ElasticNet, RidgeCV, LassoCV, ElasticNetCV\nfrom sklearn.metrics import r2_score, mean_squared_error, mean_squared_log_error\nfrom sklearn.model_selection import train_test_split\nfrom sklearn import linear_model\nfrom scipy.stats import skew","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"1421d42dee0bcecf6c32a9ea80f0b19baab7a356"},"cell_type":"markdown","source":"## Importing data"},{"metadata":{"_uuid":"70b88b36e035af9df187c7982f9e38ae38fb8b85"},"cell_type":"markdown","source":"We have two datasets at our disposal : \n- train.csv. In this dataset, we know the values of SalePrice, and we'll use it to build our predicition models.\n- test.csv. Here the values of SalePrice are hidden, only the guys at Kaggle know them - we'll probably never get to know their exact values. However we'll use what we have learned from the train dataset to try to predict them. We'll then submit our results and Kaggle will tell us if we did good or not.<br><br>\nWe could only exercice on the train dataset, however it will be easier to merge the train and the test dataset, for two reasons :\n- The test dataset still has some precious information on independant variables (all the variables except SalePrice) that could help us improve our prediction models. The more data the better.\n- We'll need to make some transformations on the independant variables : doing them on both datasets at the same time will save us some time."},{"metadata":{"trusted":false,"_uuid":"33ee07a2175a9597781b76f402d0eeb4d01fe423"},"cell_type":"code","source":"#\"kaggle\" prefix to avoid confusion with train-test splits we'll generate later for cross-validation\nkaggle_train = pd.read_csv('../input/train.csv') \nkaggle_test = pd.read_csv('../input/test.csv')\n\n#merging the two dataframes in one\ndf = pd.concat([kaggle_train, kaggle_test]).reset_index()","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"dc251058ab207a1af258b37db26631f5a23b341f"},"cell_type":"markdown","source":"## Mini-preprocessing"},{"metadata":{"_uuid":"ae8946c159c44d0e0aaa6c5914693c2512338e14"},"cell_type":"markdown","source":"As we said, we'll spend minimum time on data cleaning. There are plenty of great Kernels out there that are really specific on this subject.<br> We still need to get rid of nan values though, we'll do this the easy way by replacing them with the mean of their column. Except for SalePrice of course."},{"metadata":{"trusted":false,"_uuid":"6dc1b6afe5e19d70e305242a33326f40094cbfbc"},"cell_type":"code","source":"noSalePrice = [x for x in df.columns if x!='SalePrice']\ndf[noSalePrice] = df[noSalePrice].fillna(df[noSalePrice].mean())","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"983ddd78199c82f3055009e6c1b36bf1a92dd312"},"cell_type":"markdown","source":"# Assumptions"},{"metadata":{"_uuid":"fa204be10908f2c6a09581f89b3c08ee9f0fb2b4"},"cell_type":"markdown","source":"First things first! In this section we intent to resume all the conditions needed to run a Linear Regression. <br>\nFor this section - and to be as clear as possible - we'll stick with the independant variables Overall Qual, GrLivArea and TotalBsmtSF."},{"metadata":{"_uuid":"465ee0a056390f9b79bc486e17f3d59ffb4e2ede"},"cell_type":"markdown","source":"- <b>Numerical variables</b>\n- <b>Linear relationship</b> between independent and dependent variables\n- <b>No outliers</b>\n- <b>Normal distribution</b>\n- <b>No multicollinearity</b>\n- <b>No autocorrelation</b> \n- <b>Homoskedasticity </b><br>\n\nLet's go through all the conditions one by one"},{"metadata":{"_uuid":"319a88f345d004ff983f250aaf5b7c3bcd988fa8"},"cell_type":"markdown","source":"## Numerical variables"},{"metadata":{"_uuid":"735518beab72bc0d0eb447010a093cf8ea6dc3fa"},"cell_type":"markdown","source":"It seems obvious... and it is. \nSo let's quickly jump to the next one."},{"metadata":{"_uuid":"e5c4cde47548add8492a5b6685d204aef9a0dd5e"},"cell_type":"markdown","source":"## No outliers"},{"metadata":{"_uuid":"be65cef9ed9bf8fddd8fcbcaf0392fac5ca2480e"},"cell_type":"markdown","source":"\"Outliers\" can be considered as the weird values, the observations that don't represent the general trend, the datapoints that stand alone vs. the mass of all others. We need to get rid of them because they represent an exception rather an a rule. <br>\nTo spot them, scatterplot the independant values vs. the dependant value."},{"metadata":{"scrolled":true,"trusted":false,"_uuid":"a130c86cc07643979a0ec0bfb563542c3c9390e0"},"cell_type":"code","source":"fig, ax = plt.subplots(1, 3, figsize=(17,5))\nsns.scatterplot('GrLivArea', 'SalePrice', data=df[:1460], ax=ax[0]); #after 1460 no more SalePrice values anymore\nsns.scatterplot('OverallQual', 'SalePrice', data=df[:1460], ax=ax[1]);\nsns.scatterplot('TotalBsmtSF', 'SalePrice', data=df[:1460], ax=ax[2]);\nax[1].set_title('Before Removing Outliers');\n\nfor a in ax:\n    a.yaxis.label.set_visible(False)\n    a.get_yaxis().set_visible(False)","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"7413922f14976a9b902511a9fcb9b29bf328391c"},"cell_type":"markdown","source":"From the graphs above, we see 'GrLivArea' and 'OverallQual' graphs each show two outliers on their lower right side. They actually correspond to the same observations, let's get rid of them."},{"metadata":{"trusted":false,"_uuid":"2e9960e9339c530a5d0435a388be26728a884b98"},"cell_type":"code","source":"df = df[~df['Id'].isin([524, 1299])] #getting rid of outliers, situated at Ids 524 and 1299\n\nfig, ax = plt.subplots(1, 3, figsize=(17,5))\nsns.scatterplot('GrLivArea', 'SalePrice', data=df[:1458], ax=ax[0]);\nsns.scatterplot('OverallQual', 'SalePrice', data=df[:1458], ax=ax[1]);\nsns.scatterplot('TotalBsmtSF', 'SalePrice', data=df[:1458], ax=ax[2]);\nax[1].set_title('After Removing Outliers');\n\nfor a in ax:\n    a.yaxis.label.set_visible(False)\n    a.get_yaxis().set_visible(False)","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"6c180c90920a5be013c76b448dd8ab15a410b659"},"cell_type":"markdown","source":"## Normal distribution"},{"metadata":{"_uuid":"7512e414ad14191a01e66c79ceed88ef6b64ef5e"},"cell_type":"markdown","source":"The independant variable must show a normal distribution, which is quickly checked with a histogram.\nHere on the left-hand side, we see the original SalePrice is not normally distributed as it is skewed on the left. So we fix this by running a log transformation to it, which gives us the right-hand side histogram."},{"metadata":{"_uuid":"79e71b0a18eaa83d31ac97e2c3ce740d44704044","scrolled":true,"trusted":false},"cell_type":"code","source":"fig, ax = plt.subplots(1, 2, figsize=(17,5))\nsns.distplot(df[:1458]['SalePrice'], ax=ax[0]);\nax[0].set_title('Before Log Transformation');\n\n#log transformation\ndf.loc[df.SalePrice.notnull(), 'SalePrice_LOG'] = np.log1p(df.loc[df.SalePrice.notnull(), 'SalePrice']) \n\nsns.distplot(df[:1458]['SalePrice_LOG'], ax=ax[1])\nax[1].set_title('After Log Transformation');","execution_count":null,"outputs":[]},{"metadata":{"trusted":false,"_uuid":"e397222e676634db4d7aa777ae0041c5b222fc20"},"cell_type":"code","source":"#TRANSFORMING SALEPRICE AND GETTING RID OF SALEPRICE_LOG, JUST NEEDED FOR THE GRAPHS\ndf['SalePrice'] = np.log1p(df['SalePrice'])\nif('SalePrice_LOG' in df.columns): df.drop('SalePrice_LOG', axis=1, inplace=True)","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"319a88f345d004ff983f250aaf5b7c3bcd988fa8"},"cell_type":"markdown","source":"## Linear relationship"},{"metadata":{"_uuid":"6dffa627c7bab7d5856cf7fb5c046b1d613c82aa"},"cell_type":"markdown","source":"First of all, note that we are studying Linear Relationship <i> after </i> having log-transformed SalePrice and <i> after </i> having removed outliers. That's important to know as both those methods can increase the linearity of some relationships.<br>"},{"metadata":{"_uuid":"ab1b0fe0ad8b834357adcbee949b42d2427086f1"},"cell_type":"markdown","source":"Linear relationship implies you could imagine some kind of straight line going through the points when you plot them one against the other. <br>\nOn the other hand, if you see datapoints scattered all around the place, that's not good. <br>"},{"metadata":{"_uuid":"61b2efa2eac2aec0a4163a29ef2028dbafd1610a","trusted":false},"cell_type":"code","source":"fig, ax = plt.subplots(1, 3, figsize=(17,5))\nsns.scatterplot('GrLivArea', 'SalePrice', data=df, ax=ax[0]);\nsns.scatterplot('OverallQual', 'SalePrice', data=df, ax=ax[1]);\nsns.scatterplot('TotalBsmtSF', 'SalePrice', data=df, ax=ax[2]);\nax[1].yaxis.label.set_visible(False)\nax[2].yaxis.label.set_visible(False)","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"a3061ae2c68140b4a8584b31ed7686f6b2725fa8"},"cell_type":"markdown","source":"Graphically, the chosen variables do show some linear relations.<br>\nA good rule of thumb is to chose only variables that have a correlation score (r2) of over 0.5 with the independant variable."},{"metadata":{"trusted":false,"_uuid":"d4d73ccddb83aba511e6dde8afe4c7dfb23ebf0f"},"cell_type":"code","source":"dfc = df[['SalePrice', 'GrLivArea', 'OverallQual', 'TotalBsmtSF']].corr()\ndfc[['SalePrice']][1:]","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"954a23183ff512762ffb76dc219998184b0ac017"},"cell_type":"markdown","source":"So far so good, all variables show linear relationships with 'SalePrice'"},{"metadata":{"_uuid":"21a045755c7a3b641b8165f9f076a5860e73cb7a"},"cell_type":"markdown","source":"## No multicollinearity"},{"metadata":{"_uuid":"3f9f792bad31f6eb5fb8e5292c8764aedc178aaf"},"cell_type":"markdown","source":"Concept of multicollinearity is pretty intuitive : if your independant variables show some correlations in-between themselves, then you are facing multicollinearity. <br>"},{"metadata":{"_uuid":"40539eb38c9e28c3a5a9b875c906d1ec30cbd1aa"},"cell_type":"markdown","source":"### VIF to check multicolinearity"},{"metadata":{"_uuid":"e1540018e6590e3b2d6083dfb32875dd6c345512"},"cell_type":"markdown","source":"Best method to check for multicolinearity is Variance Inflation Factor  (VIF). The analysis gives a factor for each of the dependant variables : variables with values over 5 should be dropped."},{"metadata":{"_uuid":"c99e6d71c83f9efc87c8e1dd574d48072fbaa745","scrolled":true,"trusted":false},"cell_type":"code","source":"from patsy import dmatrices\nY, X = dmatrices('SalePrice ~ OverallQual+GrLivArea+TotalBsmtSF', df, return_type='dataframe')\nfrom statsmodels.stats.outliers_influence import variance_inflation_factor\nvif = pd.DataFrame()\nvif[\"VIF Factor\"] = [variance_inflation_factor(X.values, i) for i in range(X.shape[1])]\nvif[\"features\"] = X.columns\nvif","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"5538d01a6aa3d494d1924e981252239ef79b40a9"},"cell_type":"markdown","source":"Here, no variables should be dropped as all features have small Variance Inflation Factors."},{"metadata":{"_uuid":"a2a591db773d4312054ffa6df506b24cf34317d1"},"cell_type":"markdown","source":"## No autocorrelation"},{"metadata":{"_uuid":"5b0782d7e5410cc74d3c8d1bde8a83aa8b380c10"},"cell_type":"markdown","source":"Autocorrelation is a particular case that mainly appears with time series, which is not the case in this dataset.\nA pretty intuituive way to understand autocorrelation is with temperature data over days : a temperature on one day will likely be similar to the temperature on the day before (due to seasons). Therefore this data will be autocorelated as a value on day i can help you predict the value on the day i+1."},{"metadata":{"_uuid":"0b373833f6b020d6f960379c8b34a282ae1b8e0c"},"cell_type":"markdown","source":"## Homoskedasticity"},{"metadata":{"_uuid":"fc1ae127c04b86c188c1e6a6051e586198c3968c"},"cell_type":"markdown","source":"Homoskedasticity is the fact of having equal residuals across the regression line. <br>\nHomoskedasticity and heteroskedasticity can be understood in one graphic.\n<a href=\"https://ibb.co/DWq4R9v\"><img src=\"https://i.ibb.co/3dJ1Cpj/homoskedasticity.png\" alt=\"homoskedasticity\" border=\"0\" style=\"height: 200px;\"></a>"},{"metadata":{"_uuid":"be7ea16c4167154604d1afe408aec5b146a41f29"},"cell_type":"markdown","source":"From the graphs we saw above, we can say heteroskedasticity was not too important : we don't have to remove any of the independant variables."},{"metadata":{"_uuid":"072072871390b616613c989545a385152a3b42db"},"cell_type":"markdown","source":"# Fitting & Predicting"},{"metadata":{"_uuid":"ccb27230fc5e3e63c8a38b59228332badd725930"},"cell_type":"markdown","source":"Now that we know all our assumptions are respected, we can start to predict."},{"metadata":{"_uuid":"6896167da04526b26de5897d05409cf3f26c4b8e"},"cell_type":"markdown","source":"## The concept"},{"metadata":{"_uuid":"7f6ae90e366a52d05056bd6ba7bc09b29bb175c7"},"cell_type":"markdown","source":"This is where the magic happens.<br>\n\"Fitting a model\" is like designing a machine that has one goal : make some predictions. It will transform some X values into Y values<br> \nHowever the machine cant' be designed from scratch, it first needs known X and Y values to train with. So from the moment you have some known X and Y values, you can design your very own predicting machine in two lines of codes : "},{"metadata":{"trusted":false,"_uuid":"4f23523cc16c96091efcc2cc83e5fe32e1e04625"},"cell_type":"code","source":"#training dataset\nX = [[1, 2], [3, 4]]\nY = [5, 6]\n\n#defining and training the machine in two lines of code\nmodel = linear_model.LinearRegression()\nmodel.fit(X, Y);\n\n#from this point, the machine can compute a Y value every time it is given some X values\n#here, with 5 and 6 as inputs, the machine predicts the output will be 7 \nmodel.predict([[5,6]])[0]","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"795796a663056040378f3d40a98a9c501881cb81"},"cell_type":"markdown","source":"## A basic algorithm : Least Squares "},{"metadata":{"_uuid":"3ee495aaa1174d325124e91e492d65f0bdc50dfd"},"cell_type":"markdown","source":"How is the machine designed? Many methods exist and we'll review some of them later in the Kernel, but for now we'll only cover the most commonly used method : Least Squares.<br>\nLet's apply it to the House Prices dataset, by predicting SalePrice from GrLivArea."},{"metadata":{"trusted":false,"_uuid":"5bd9b9939e44976ad8ffd915b833db0d857b133a"},"cell_type":"code","source":"X = df[:1458][['GrLivArea']] #independant variables\nY = df[:1458]['SalePrice'] #dependant variable\nmodel = linear_model.LinearRegression() #LinearRegression() = Least Square Algorithm\nmodel.fit(X, Y); #fitting, or \"training\" the model","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"f8a064d2e8f5a6f8d35e784a6e796bfe9b1bade1"},"cell_type":"markdown","source":"The Least Squares algorithm intents to find the \"Line of Best Fit\" by minimizing the sum of the distances between each of the points and the regression line. Best way to understand this is with the help of a plot.<br>\nSo let's plot the Least Squares regression line and compare it with our own custom line that comes (almost) straight out of our imagination."},{"metadata":{"trusted":false,"_uuid":"12c34cc5d3bad6493401d70a5bda3d5fc727f4e2"},"cell_type":"code","source":"plt.rcParams['figure.figsize'] = (12.0, 6.0)\nsns.scatterplot('GrLivArea', 'SalePrice', data=df); \n\nYpredicted = model.predict(X)\nsns.lineplot(X.iloc[:, 0].tolist(), Ypredicted, color='black', label='Least Squares Line'); \n\nYinvented = X.iloc[:, 0].as_matrix() * 0.00145 + 10\nsns.lineplot(X.iloc[:, 0].tolist(), Yinvented, color='blue', label='My Custom Line');","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"1e823a020f133313f3c0ca012efcdec264e6a5fe"},"cell_type":"markdown","source":"What the OLS does is it computes the regression line so that there is no other \npossible line out there that can result in a lower sum of squared residuals than the OLS line. <br>\nHere you can instictively see that the total distance \"points -> blue line\" is superior to \"points -> black line\". You can modify the blue line's equation with any intercept and coefficient you want, you'll never get a smaller distance than to the \"points -> black line\" distance. That's what OLS does.<br>"},{"metadata":{"_uuid":"e3c4f9ff1a6da42b74ff6d6c5a690952c554bbc7"},"cell_type":"markdown","source":"# Testing"},{"metadata":{"_uuid":"e48c937628a287eaf8fbd4c7f2e97a7462cec258"},"cell_type":"markdown","source":"## Train & Test"},{"metadata":{"_uuid":"0f1c7417b916c9261f3ed7a7002281969107f19d"},"cell_type":"markdown","source":"Remember the prediction machine we trained earlier?\nWell, we did something quite stupid doing that : we gave it all the SalePrice values we had at disposal. So now, there's no way of testing it on a new set of data and see if it performs well.<br>\n\nIndeed, if we want to test the accuracy of the model, we need to set a part of the dataframe aside before training it. So we'll have two sets :\n - Train : the data that will train the machine\n - Test : the \"untouched\" data that will be used to test the machine's accuracy. (Note that we can't use Kaggle's \"test\" dataset, because we don't know its SalePrice values)\n\nIn other words, once you have trained a machine with dataset \"Train\", you'll use dataset \"Test\" to put it to the test.<br>\nYou'll pretend to the machine you know nothing about the Ys from \"Test\" it is supposed to predict. So once it has predicted them, you bring in the real YTest values and compare them with the predicted values : you'll then know if it's a good machine or not.<br><br>\nLet's create our own Train and Test, thanks to the train_test_split method, which comes in quite handy.<br>\nIt will create, in one line of code, both datasets, and extract independant and dependant values from them.\n"},{"metadata":{"trusted":false,"_uuid":"bd32cdde0e5c742a2a9da5529e51a62142e0b88f"},"cell_type":"code","source":"ind_vars = ['GrLivArea', 'OverallQual', 'TotalBsmtSF']\nXtrain, Xtest, Ytrain, Ytest = train_test_split(df[:1458][ind_vars], df[:1458]['SalePrice'], test_size=0.33)","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"4e72bb96690c4314f967d6e4c1fa1b25138e4dba"},"cell_type":"markdown","source":"Now that we have our variables, let's go back to fitting & predicting"},{"metadata":{"trusted":false,"_uuid":"a1541d1ad5deda7db9c2b52e8fb44755359af5ef"},"cell_type":"code","source":"model = LinearRegression()\nmodel.fit(Xtrain, Ytrain);\nYpredicted = model.predict(Xtest)\ndff = pd.DataFrame({\"The values our model predicted\":np.expm1(Ypredicted), \n              \"The values it should have predicted, assuming it was perfect\":np.expm1(Ytest.tolist())})\ndff.astype(int).head()","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"8286a379f118456e0df620b646b90b7c3eccf0e0"},"cell_type":"markdown","source":"From the five rows above, some predicted values don't seem too far away from the actual values. But how can we quantify the overall proximity?"},{"metadata":{"_uuid":"1ecbb1b89654c1a2a9e6baa67e512045ec819d48"},"cell_type":"markdown","source":"## Statistics"},{"metadata":{"_uuid":"646e907df37c1468ec9db689f4dad5eb8efa4d69"},"cell_type":"markdown","source":"In order to evaluate a regression model's accuracy, there are <a href = \"https://scikit-learn.org/stable/modules/model_evaluation.html\"> many statistics </a> that can be tested. Let's have an overview of the main ones.<br>\n\nThe most common stat is <b>R-square</b>.\nRsquare \"is the percentage of the response variable variation that is explained by a linear model\". So if you get 0.8, you'll be able to brag claiming \"80% of SalePrice can be explained by my prediction machine\" (doesn't mean 80% of your predictions are exactly right though). It is easy to interpret as its values ranges from 0 to 1 : the closer to 1, the better the model. <br><br>\nHowever the problem with R-square is whenever you add more independant variable into the model, R-square always increases. So you can be tricked into thinking your model is getting better as its gets more complex, but it's not the case. To avoid that, you'll be better off using <b>Ajdusted R-square</b>, which takes into account the number of independant values. Sklearn does not have a built-in function to compute that, you'll find one below."},{"metadata":{"trusted":false,"_uuid":"f641b69ee8677f8f8c1c8634da15922cfca630e2"},"cell_type":"code","source":"def adjusted_r2(Xtest, r2):\n\n    p = len(Xtest.columns) #number of independant values\n    n = len(Xtest) #length of test dataset\n    adj_r2 = 1-(1-r2)*(n-1)/(n-p-1)\n    return adj_r2","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"03166eabd2c720df7d8988be3b2544059d95a397"},"cell_type":"markdown","source":"Another important stat is the <b> Mean Squared Error </b>. It is a representation of the distance between the points and the regression line. So the lower the value, the better. It has an absolute value so it's better used when comparing models in-between themselves. "},{"metadata":{"_uuid":"83d706df66ab6b04f6ff50400d1b29247345d392"},"cell_type":"markdown","source":"Let's have a shot at testing these main statistics on our predictions."},{"metadata":{"trusted":false,"_uuid":"33f19554a80b5e48117d14d5eb953c63d1e5b35e"},"cell_type":"code","source":"r2 = r2_score(Ytest, Ypredicted)\nadj_r2 = adjusted_r2(Xtest, r2)\nmse = mean_squared_error(Ytest, Ypredicted)\npd.DataFrame(data=[r2, adj_r2, mse], index=['r2', 'r2_adj', 'mse'])","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"00974c1f128b2519d3ff7f3108ccb8b0c020f095"},"cell_type":"markdown","source":"## Cross-validation"},{"metadata":{"_uuid":"8334432e4a9ce5727ccfdbf7528eaadecf989202"},"cell_type":"markdown","source":"However the best way to test you model is using sklearn's method, cross_val_score. <br>\nIndeed, this method needs one line of code to :\n- Consider the model we want to use (in this case LinearRegression())\n- Make itself a train dataset with a 20% sample of X and Y\n- Fit the model with the train dataset\n- Test the model on the other 80%\n- Give us an accuracy score (based on a given statistic)\n\nMoreover, it will do that 10 times (assuming you define cv=10) and compute the mean.<br>\nSo for calculating mean squared error, it would look like this."},{"metadata":{"trusted":false,"_uuid":"0c4929ec18df8bf4ccbee03ebab8517b0868d17f"},"cell_type":"code","source":"X = df[:1458][ind_vars]\nY = df[:1458]['SalePrice']\nmse = -model_selection.cross_val_score(model, X, Y, cv=10, scoring='neg_mean_squared_error').mean()\nround(mse, 4)","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"bcdde37903120cfb453d95d6d1a98eee6363476c"},"cell_type":"markdown","source":"# Seeking Performance"},{"metadata":{"_uuid":"efda1978a97f6e92e948996d988ad6c336127aa8"},"cell_type":"markdown","source":"## Regularizing"},{"metadata":{"_uuid":"35a17785786544957e0e1587d62581f30d897ec5"},"cell_type":"markdown","source":"### Overfitting & Underfitting"},{"metadata":{"_uuid":"3057dbb82100467f6c80ec3648ff9d6b78d61fcc"},"cell_type":"markdown","source":"Regularized models are meant to avoid Overfitting. But push them too far and you'll end up with Underfitting. So before we dive into each model indivdually, we must understand both concepts.<br>\n\nOverfitting means that your regression will probably fit well to the data you are working on, but as soon as you bring in yet unseen values, your prediction will lose accuracy. Indeed, if the dataset contains too many \"weird\" values or outliers for instance, an overfitted regression line will try to be the closest possible to these points though they are a not a good representation of the dataset.The regression will then not be able to predict new points that <i>do</i> represent the dataset. <br>\n<i> If there is noise in the training data, then the estimated coefficients won’t generalize well to the future data. </i><br>\n<br>\nBelow a schema that personnaly made me quickly understand the concepts."},{"metadata":{"_uuid":"17c148f118458f1aab967a192df13fcdfac427ef"},"cell_type":"markdown","source":"<img src=\"https://s3-ap-south-1.amazonaws.com/av-blog-media/wp-content/uploads/2017/06/05210948/overunder1.png\"/>"},{"metadata":{"_uuid":"e301bf5037afac7c8624d370e5430f0f65ec5c58"},"cell_type":"markdown","source":"Now imagine the lines below, but applied to all the variables of your model. Some variables will tend to overfit more than others.<br>\nThis is where Regularization comes in and helps improve accuracy in two ways : \n- It considers each variable and is able to lower the coefficients of variables that overfit\n- It prevents from multicollinearity by lowering influence of variables already \"represented\" by others\n\n<br>\nOn the other hand, if regularization is too severe, we'll be in a situation of Underfitting, meaning that the Regression Line will be too \"generalized\" : by trying to satisfy any possible situation, it ends up satisfying none. It takes so little risk at predicting that it becomes useless."},{"metadata":{"_uuid":"dcaaa1c333780c04006c848331f22782b657e18e"},"cell_type":"markdown","source":"### Regularized Models"},{"metadata":{"_uuid":"82364d59352f71c0ab3cb3d022ae842e4618466d"},"cell_type":"markdown","source":"There are different regularized models out there, main ones are : Ridge, Lasso and Elastic-Net.<br>\nMain differences are :\n - Ridge performs well on small datasets\n - Lasso and Elastic-Net have the capacity to bring coefficients down to 0. Meaning completely canceling the influence of some variables in the model.\n - Elastic-Net is good when you have a lot lot lot of independant variables\n \n<br>\nBecause regularized models have the ability to lower the coefficients of useless variables, you can throw all the independant variables in there, and let the models decide which ones are valuable. <br>\nMeaning that somehow, you can forget about the assumptions on multi-collinearity and linear relationships.\n<br>\nThen you can test them all the models using cross-validation and pick the one that gets the best results.\n"},{"metadata":{"trusted":false,"_uuid":"00899075db696c8f3540f071b9801e2ef4035f49"},"cell_type":"code","source":"#extracting numeric independant variables\ndfc = df[:1458].corr()[['SalePrice']].sort_values('SalePrice', ascending=False) \nnum_cols = dfc.drop(['SalePrice', 'Id']).index \n\n# df_scaled = pd.DataFrame(scaler.fit_transform(df[num_cols].astype(float)), columns=num_cols) #scaling the numeric columns\nY = df[:1458]['SalePrice']\nX = df[:1458][num_cols]\nXtrain, Xtest, Ytrain, Ytest = train_test_split(X, Y, test_size=0.33, random_state=42)","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"60c8634fc8476a120c551517a1328ee2b6323866"},"cell_type":"markdown","source":"One last thing to know before we define the models : each of them take alphas as inputs. These values define how \"intense\" are the regularizations going to be. If the value is not relevant, we can end up either with overfitting or underfitting.<br>\nUsing SkLearn's RidgeCV() instead of Ridge() allows to input a range of alphas instead of a single one; the method will then select itself the best one and can tell you which one it chose with the alpha_ attribute.<br>\nI came up with the below alphas \"manually\" : I tried a random range, ran the methods, got the alpha_ attribute. If the alpha_ was a maximum or a minimum of my range, I added values to my range (higher in case of maximum, lower in case of minimum).<br>\nThe best method though is probably the one exposed in this great [kernel](https://www.kaggle.com/apapiu/regularized-linear-models) : you loop through a range of alphas and plot a line, and pick the alpha where the line is the lowest."},{"metadata":{"trusted":false,"_uuid":"ae845023549f04d86a1749bb4fa2b7eedacd570c"},"cell_type":"code","source":"#TRAINING MODELS\nlinear_model = LinearRegression().fit(Xtrain, Ytrain)\nridge_model = RidgeCV(alphas=[0.01, 0.1, 0.5, 1, 5, 10, 50, 100, 200], normalize=True).fit(Xtrain, Ytrain)\nlasso_model = LassoCV(alphas=[0.00001, 0.0001, 0.001, 0.01], normalize=True).fit(Xtrain, Ytrain)\nelastic_model = ElasticNetCV(alphas= [0.00001, 0.0001, 0.001, 0.005, 0.01, 0.02], normalize=True).fit(Xtrain, Ytrain)\n\n#TESTING PERFORMANCES\nprint('Linear  | mse :', round(mean_squared_error(Ytest, linear_model.predict(Xtest)), 5))\nprint('Ridge   | mse :', round(mean_squared_error(Ytest, ridge_model.predict(Xtest)), 5), '| alpha :', ridge_model.alpha_)\nprint('Lasso   | mse :', round(mean_squared_error(Ytest, lasso_model.predict(Xtest)), 5), '| alpha :', lasso_model.alpha_)\nprint('Elastic | mse :', round(mean_squared_error(Ytest, elastic_model.predict(Xtest)), 5), '| alpha :', elastic_model.alpha_)","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"b78317a6b0b554dc420668efa328acc19055902f"},"cell_type":"markdown","source":"Lasso is performing slightly better than the classic Linear Regression.<br>"},{"metadata":{"_uuid":"40f8c520be38338315d41eb7acf9424b987139c6"},"cell_type":"markdown","source":"## Scenario testing"},{"metadata":{"_uuid":"d7e5d94f2c1c7d8c7878d6c6989a9ea2d79982f2"},"cell_type":"markdown","source":"There are a lot of factors influencing a regression's quality, and depending on the dataset you are working on, these factors will behave differently. One way to improve the prediction's accuracy is to test the regression by \"playing\" with these factors' parameters. Let's test a few scenarios here : \n<br>\n- Log transformation. \n    - log-transforming all ind. variables\n    - log-transforming only ind. variables showing over 50% of skewness\n    - log-transforming only ind. variables showing over 75% of skewness\n    - log-transforming none of the ind. variables\n- Correlation\n    - considering all numeric variables\n    - considering variables correlated with dependant variable at 0.3 threshold\n    - considering variables correlated with dependant variable at 0.5 threshold\n- Dummies\n    - considering continuous variables only\n    - considering continuous & categorical variables\n- Regression model\n    - Linear\n    - Ridge\n    - Lasso\n    - ElasticNet\n    \nWe'll test every possible combination. The code will be a bit harsh, but at the end we'll have a nice table presenting which combinations performed the best."},{"metadata":{"trusted":false,"_uuid":"9ab2c7a04c0d5abc3f9de70e666f9a31ec1be977"},"cell_type":"code","source":"# FUNCTION TO RETURN THE ACCURACY OF THE REGRESSION, THROUGH MEAN SQUARED ERROR\ndef test_regression(model, X, Y): \n    return round(np.sqrt(-model_selection.cross_val_score(model, X, Y, cv=5, scoring='neg_mean_squared_error')).mean(), 5)","execution_count":null,"outputs":[]},{"metadata":{"trusted":false,"_uuid":"fca4a4a0800ba6c3be8fb39068ad52e51fb219c2"},"cell_type":"code","source":"# FUNCTION TO CREATE ALL COMBINATIONS OF CORRELATION, SKEWNESS AND MODELS\n # AND STOCKING RESULTS IN LIST OF DICTS\ndef tests(X, Y, test_dicts, include_dummies):\n\n    #LOOPING THROUGH CORRELATIONS THRESHOLDS\n    for i in corr_list:\n\n        #getting only ind. variables that have over the requested correlation with dep. variable\n        not_correlated_vars = correlations[correlations.iloc[:,0]<i].index\n        cols_to_keep = [x for x in X.columns if x not in not_correlated_vars]\n        X2 = X[cols_to_keep]\n\n        #LOOPING THROUGH SKEWNESS THRESHOLDS\n        for i2 in skew_list:\n\n            X3 = X2.copy()\n            #getting only ind. variables that have over the requested skewness\n            skewed_feats = skewness[skewness.iloc[:,0]>i2].index\n            skewed_feats = [x for x in skewed_feats if x in X3.columns]\n            X3[skewed_feats] = np.log1p(X3[skewed_feats])\n            \n            #stocking relevant information into dict\n            param_dict = {'correlation_threshold':str(i), 'skewness_thresold':str(i2), 'number_of_variables':X3.shape[1], 'include_dummies':include_dummies}\n\n            #TESTING THE FOUR MODELS FOR EACH COMBINATION OF CORRELATION/SKEWNESS\n                #AND APPENDING THEIR SCORE TO LIST OF DICTS, WITH INFORMATION ASSOCIATED\n            model = LinearRegression()\n            #z in zscore only to have column at end of dataframe\n            test_dicts.append({**{'zscore':test_regression(model, X3, Y), 'model':'linear'}, **param_dict})\n\n            model = RidgeCV(alphas=[0.01, 0.1, 0.5, 1, 5, 10, 50, 100, 200], normalize=True)\n            test_dicts.append({**{'zscore':test_regression(model, X3, Y), 'model':'ridge'}, **param_dict})\n\n\n            model = LassoCV(alphas=[0.00001, 0.0001, 0.001, 0.01], tol=0.1, normalize=True)\n            test_dicts.append({**{'zscore':test_regression(model, X3, Y), 'model':'lasso'}, **param_dict})\n\n\n            model = ElasticNetCV(alphas=[0.0001, 0.001, 0.01, 0.1, 1], tol=0.1, normalize=True)\n            test_dicts.append({**{'zscore':test_regression(model, X3, Y), 'model':'elastic'}, **param_dict})","execution_count":null,"outputs":[]},{"metadata":{"trusted":false,"_uuid":"cb102225fef816b09897a0ccab0e4aba557da2f2"},"cell_type":"code","source":"test_dicts = []\n\n#defining correlation and skewness threshold\ncorr_list = [0, 0.25, 0.5]\nskew_list = [0, 0.5, 0.75, 0.9, 999] #999 threshold is same as skewing no variable\n\n#computing skewness and correlation of ind.variables\nskewness = df[num_cols].apply(lambda x: skew(x)).abs().to_frame() #skewness can be evaluated on both train and test\ncorrelations = df[:1458].corr().abs().sort_values('SalePrice', ascending=False)['SalePrice'].to_frame()[1:].drop('Id')\n\n#RUNNING THE FIRST TESTS, WITHOUT CATEGORICAL VARIABLES\ntests(X, Y, test_dicts, include_dummies='no')\n\n\n#RUNNING THE SECOND SET OF TESTS, INCLUDING CATEGORICAL VARIABLES\ndf_dummies = pd.get_dummies(df)\nX = df_dummies[:1458].drop(['SalePrice', 'Id'], axis=1)\ntests(X, Y, test_dicts, include_dummies='yes')\n\n#computing the results dataframe\nresults = pd.DataFrame(test_dicts)\nresults.sort_values('zscore').head()","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"8c6ae00b4035f2095f882ebe4cf6921351310a32"},"cell_type":"markdown","source":"Here we go! Our best result turns out to be : using Lasso on all variables, including dummies, and log-transforming only independant variables that show more than 90% of skweness.<br>"},{"metadata":{"_uuid":"8e15de4d13075499a40992b84146abeea1cd4b46"},"cell_type":"markdown","source":"## Submitting"},{"metadata":{"_uuid":"18d5eb5fd675724a0280facbeebac0c8a8ec8575"},"cell_type":"markdown","source":"Let's predict SalePrice with these conditions and submit our results.<br>\nHope you liked the kernel!"},{"metadata":{"trusted":false,"_uuid":"d2f8b1bb6f6d6fba6d7cbff530880836746d3832"},"cell_type":"code","source":"#RECREATING BEST SCENARIO\nXtrain = df_dummies[:1458].drop(['SalePrice'], axis=1) \nYtrain = df[:1458]['SalePrice']\nXtest = df_dummies[1458:].drop(['SalePrice'], axis=1) \n\nskewed_feats = skewness[skewness.iloc[:,0]>0.9].index\nskewed_feats = [x for x in skewed_feats if x in Xtest.columns]\nXtest[skewed_feats] = np.log1p(Xtest[skewed_feats])\nXtrain[skewed_feats] = np.log1p(Xtrain[skewed_feats])\n\nmodel = LassoCV(alphas=[0.00001, 0.0001, 0.001, 0.01], tol=0.1, normalize=True)\n\n#FITTING AND PREDICTING\nmodel.fit(Xtrain, Ytrain)\npredicted_prices = np.expm1(model.predict(Xtest))\n\n#SUBMITTING\nmy_submission = pd.DataFrame({'Id': Xtest.Id, 'SalePrice': predicted_prices}) \nmy_submission.to_csv('submission.csv', index=False)","execution_count":null,"outputs":[]}],"metadata":{"kernelspec":{"display_name":"Python 3","language":"python","name":"python3"},"language_info":{"codemirror_mode":{"name":"ipython","version":3},"file_extension":".py","mimetype":"text/x-python","name":"python","nbconvert_exporter":"python","pygments_lexer":"ipython3","version":"3.7.2"},"toc":{"base_numbering":1,"nav_menu":{},"number_sections":true,"sideBar":true,"skip_h1_title":false,"title_cell":"Table of Contents","title_sidebar":"Contents","toc_cell":true,"toc_position":{},"toc_section_display":true,"toc_window_display":false}},"nbformat":4,"nbformat_minor":1}