{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"pygments_lexer":"ipython3","nbconvert_exporter":"python","version":"3.6.4","file_extension":".py","codemirror_mode":{"name":"ipython","version":3},"name":"python","mimetype":"text/x-python"}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"## **Housing Prices Advanced Regression : EDA and Regression Models**\n\nIn this notebook I have performed thorough EDA on the housing dataset and tried to identify some keen underlying trends. Dtaa pre-processing is done after which different models have been applied. This kernel is an introduction to predictive modelling and demonstrates the various techniques. And it is must see if you are a beginner in regression and predictive modelling like me ;)\n\nLet's dive in !!!.","metadata":{"_uuid":"bc7c0cc80181349f53eca656f94a233ba5a24855"}},{"cell_type":"code","source":"","metadata":{"_uuid":"2242a83f7a5f547fb3f72aaae9f3f8cece4af2fa"},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## CONTENTS::","metadata":{"_uuid":"a28bac8e808e0e9a3fe1bbd2cc3503565c611179"}},{"cell_type":"markdown","source":"[ **1 ) Importing the Modules and Loading the Dataset**](#content1)","metadata":{"_uuid":"5040024e9ac845eb51583ded8a91b4bfb9baddb9"}},{"cell_type":"markdown","source":"[ **2 ) Exploratory Data Analysis (EDA)**](#content2)","metadata":{"_uuid":"29ba23b0ef9144363b46946f62376050cc9946fc"}},{"cell_type":"markdown","source":"[ **3 ) Missing Values Treatment**](#content3)","metadata":{"_uuid":"bea85cb468c8dafe8de87768e912d1e4ff15f4cb"}},{"cell_type":"markdown","source":"[ **4 ) Handling Skewness of Features**](#content4)","metadata":{"_uuid":"f8bf6178c36685604de741bb4d967a070deaf0fb"}},{"cell_type":"markdown","source":"[ **5 ) Prepare the Data**](#content5)","metadata":{"_uuid":"333d282e4c9932bfee8e381ab32fcc692ad3be13"}},{"cell_type":"markdown","source":"[ **6 ) Regression Models**](#content6)","metadata":{"_uuid":"41b5b2453784882c646e105c92ef9b11c3a32e52"}},{"cell_type":"markdown","source":"[ **7 ) Saving and Making Submission to Kaggle**](#content7)","metadata":{"_uuid":"b3d5e0080ea58dbfb425b9e405762ed5d596ae42"}},{"cell_type":"code","source":"","metadata":{"_uuid":"89f6f514db287cd496873fe14b2246ca5ac86473"},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<a id=\"content1\"></a>\n## 1) Importing the Modules and Loading the Dataset","metadata":{"_uuid":"25c8d67b312aedaf18d321eaa197b162c4b1fe28"}},{"cell_type":"code","source":"# Ignore  the warnings\nimport warnings\nwarnings.filterwarnings('always')\nwarnings.filterwarnings('ignore')\n\n# data visualisation and manipulation\nimport numpy as np\nimport pandas as pd\nimport matplotlib.pyplot as plt\nfrom matplotlib import style\nfrom matplotlib.legend_handler import HandlerBase\nimport seaborn as sns\nimport missingno as msno\n#configure\n# sets matplotlib to inline and displays graphs below the corressponding cell.\n%matplotlib inline  \nstyle.use('fivethirtyeight')\nsns.set(style='whitegrid',color_codes=True)\n\n#import the necessary modelling algos.\n\n#classifiaction.\nfrom sklearn.linear_model import LogisticRegression\nfrom sklearn.svm import LinearSVC,SVC\nfrom sklearn.neighbors import KNeighborsClassifier\nfrom sklearn.ensemble import RandomForestClassifier,GradientBoostingClassifier\nfrom sklearn.tree import DecisionTreeClassifier\nfrom sklearn.naive_bayes import GaussianNB\n\n#regression\nfrom sklearn.linear_model import LinearRegression,Ridge,Lasso,RidgeCV,ElasticNet\nfrom sklearn.ensemble import RandomForestRegressor,BaggingRegressor,GradientBoostingRegressor,AdaBoostRegressor\nfrom sklearn.svm import SVR\nfrom sklearn.neighbors import KNeighborsRegressor\n\n#model selection\nfrom sklearn.model_selection import train_test_split,cross_validate\nfrom sklearn.model_selection import KFold\nfrom sklearn.model_selection import GridSearchCV\nfrom sklearn.preprocessing import LabelEncoder\n\n#evaluation metrics\nfrom sklearn.metrics import mean_squared_log_error,mean_squared_error, r2_score,mean_absolute_error # for regression\nfrom sklearn.metrics import accuracy_score,precision_score,recall_score,f1_score  # for classification\n\nfrom scipy import stats\nfrom scipy.stats import norm, skew   # specifically for staistics","metadata":{"id":"1JI18AXTlKff","_uuid":"fe306398c6247419de28f6b8da43eb4e9adabd45","trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"train=pd.read_csv(r'../input/train.csv')\ntest=pd.read_csv(r'../input/test.csv')","metadata":{"id":"3N2G1LMylKfq","_uuid":"2889a16c07b2f71ea611b3f4fdcd52d372eb05f4","trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"train.head(10)\n#test.head(10)","metadata":{"id":"SFMv7gaXlKf0","outputId":"90f92190-0544-4bb8-cad4-904ce6b08dec","_uuid":"1483c3cb62409aad88a3fa96ff775781fcdc916a","trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"","metadata":{"_uuid":"3a03a7dcac994a50acfe14b31e799838b3f1fb35","trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<a id=\"content2\"></a>\n## 2) Exploratory Data Analysis (EDA)","metadata":{"_uuid":"516a4a16739f97c372516fd7f6f5565bf182dea8"}},{"cell_type":"markdown","source":"## 2.1 ) The Features and the 'Target' variable","metadata":{"_uuid":"2442455ae44bd2e13cbb205227fb8539b674b124"}},{"cell_type":"code","source":"df=train.copy()\n#df.head(10)\ndf.shape","metadata":{"id":"lpQgItEJlKgE","outputId":"4cb77d34-bd76-4b24-cba8-7134f12e29c7","_uuid":"d6c11eae1643c0a6fbeb01925fdad233d88e1cd9","trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"df.drop(['Id'],axis=1,inplace=True)\ntest.drop(['Id'],axis=1,inplace=True)","metadata":{"id":"N5re2y4tlKgQ","_uuid":"27935512dd78ebc5dbe37a484393b2a496f638d9","trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"We can drop the 'Id' column as the frames are already indexed.","metadata":{"id":"fYDdbeowlKgb","_uuid":"b6e3264fe83758a0d623ffce222a3707caa87d8d"}},{"cell_type":"code","source":"df.index # the indices of the rows.","metadata":{"id":"0kTwvMBBlKgu","outputId":"3f2eb31d-4586-4edd-ac24-65d586b3206d","_uuid":"1b0e63aa2e7f03e1cbe28d4d185bfbd00708d7a9","trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"df.columns ","metadata":{"id":"PcsFPq8alKg4","outputId":"1fc6f336-3d8a-4483-9726-167346bd7b09","_uuid":"4119c9aa53d917e0077e3e868205ddaf1bd03eda","trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## 2.2 ) Check for Missing Values","metadata":{"_uuid":"e7094ccff8532c0492f9f6036e24d9c0457b6119"}},{"cell_type":"code","source":"df.isnull().any()","metadata":{"id":"xf18XSaulKg-","outputId":"2fbc551f-9e9d-4139-fa1e-260fef8da746","_uuid":"1c01ec5fe018c4948b14cfb68fedc83654836a2e","trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"msno.matrix(df) # just to visulaize. ","metadata":{"id":"jUQh7aE6pmh-","outputId":"40af09c2-3807-4f81-c4ab-d23bc95b5945","_uuid":"4650f121eda3d85c0c8bf20e6afd46e5ebf26a9f","trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"* #### Many columns have missing values and that will be treated later in the notebook.","metadata":{"_uuid":"93147d06c56a4956fc90aed3b88dac4f7cd8c8bb"}},{"cell_type":"code","source":"","metadata":{"_uuid":"8d325f811864ab60d93406e0d1ee7dcf5929755c","trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## 2.3 ) Separate Dataframes (depending on data type)","metadata":{"_uuid":"ddfe2acb6f213944220033fc43a42b72ddb1ffa1"}},{"cell_type":"markdown","source":"Might be useful when we consider features of different data types.","metadata":{"_uuid":"1ea7c0b1d58d00824bf50b9854e8d5ea36ddd1a8"}},{"cell_type":"markdown","source":"#### CATEGORICAL FEATURES","metadata":{"_uuid":"5cfd9b025cdeda28ad41d344d51d74dbf517785c"}},{"cell_type":"code","source":"cat_df=df.select_dtypes(include='object')","metadata":{"id":"d5Li8tK1pxGw","_uuid":"769aacc692db88e7a756815eca1fa1da034ddab4","trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"cat_df.head(10)\ncat_df.shape","metadata":{"id":"yoJ-mKriyQ1t","outputId":"02eeb3c9-54a3-4533-9129-e2fbcab98747","_uuid":"d821136907f4eefe5ec23e0cb4f338de08d31be0","trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"cat_df.columns   # list of the categorical columns.","metadata":{"id":"4pONahVy1PP-","outputId":"cacad1ee-adfc-4f4b-a633-e0a9aa7af401","_uuid":"c8694a6566c56949dee00382990fa2170af2b70e","trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"#### NUMERIC FEATURES","metadata":{"_uuid":"61a1cc1d294a89bdb38b3de66f4a480987712073"}},{"cell_type":"code","source":"num_df=df.select_dtypes(include='number')\nnum_df.shape","metadata":{"id":"zjIhIHmNySOJ","outputId":"0993f9a2-fdeb-4e79-e9a9-f9961f08e963","_uuid":"c9207d3b2f1a0f7d301f6eeefd1647628b741d86","trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"num_df.columns # list of numeric columns.","metadata":{"id":"1wzwVwGCztOf","outputId":"b984b21f-67e2-4591-ad74-32a822b7b44b","_uuid":"0a81731c4bbce8d4b5c790b7321acb5f171926b0","trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"#### FEATURES WITH MISSING VALUES","metadata":{"_uuid":"480a5d00a08a2e82d9a374916f21bbf00c896b4f"}},{"cell_type":"code","source":"nan_df=df.loc[:, df.isna().any()]\nnan_df.shape\nnan_df.columns   # list of columns with missing values.","metadata":{"id":"zVXh9aiy1Y8b","outputId":"dc02cbdd-10fd-4a97-9118-696ae3989d69","_uuid":"4cecfcbebd1be1e7f6de7c69271a00d47a2fa0b0","trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"#### MERGING THE TRAIN & TEST SETS","metadata":{"_uuid":"8a4e196a5b539b145b4b36387468c63b64c479e5"}},{"cell_type":"code","source":"all_data=pd.concat([train,test])","metadata":{"id":"T4HlzGhB2Uvc","_uuid":"74ff3fb2a712910b84b0c266ec5beb2a5d785bc2","trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"print(all_data.shape)\nall_data = all_data.reset_index(drop=True)","metadata":{"id":"QgqobkHUMTfw","outputId":"98a32181-bc43-4e44-c05e-2f93c03c10cf","_uuid":"bc1bad854cf9c060ebdb6301a7927cb3e1587001","trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"all_data.head(10)","metadata":{"id":"k7MKA1BYOVgV","outputId":"67350384-16fa-495f-eb12-25c27a2e9539","_uuid":"10f12ddd2758d9679e60e9b8063f02b2bf09cb1b","trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"print(all_data.loc[1461:,'SalePrice'])  \n# note that it is Nan for the values in test set as expected. so we drop it here for now.\nall_data.drop(['SalePrice'],axis=1,inplace=True)\n","metadata":{"id":"sxO9eOMzPrn_","outputId":"62150687-eace-40f4-d3f3-f2cbcfabe5ff","_uuid":"93158fabaab046f39582d156d7d2374e1958f271","trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## 2.4 ) Analyzing the Target i.e. 'SalePrice'","metadata":{"id":"1H56qWlDQbOV","_uuid":"401e566786a9250050c2f36e1d663b1362ad5fb0"}},{"cell_type":"code","source":"# analyzing the target variable ie 'Saleprice'\nsns.distplot(a=df['SalePrice'],color='#ff4125',axlabel=False).set_title('Sale Price')","metadata":{"id":"eX-37QQpQeCZ","outputId":"6da9e996-4a55-4e0c-f581-d730a03de05e","_uuid":"bd4ea76e34fe7dd6668b8c93e610f2fa0e6f426b","trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"#### **The distribution of target is a bit right skewed. Hence taking the 'log transform' is a reasonable option.**","metadata":{"_uuid":"f2399d492e0f97c0cfe48e300958f473855793cb"}},{"cell_type":"markdown","source":"#### ALSO LINEAR REGRESSION IS BASED ON THE ASSUMPTION OF THE 'HOMOSCADESITY' AND HENCE TAKING LOG WILL  BE A GOOD IDEA TO ENSURE 'HOMOSCADESITY' (that the varince of errors is constant.). A bit scary but simple ;) \n\n**You can read more about this on wikipedia.**","metadata":{"_uuid":"fe31c0ccf042f5102c5f2f3a0fd44a8b0f6ae0bf"}},{"cell_type":"code","source":"#Get also the qq-plot (the quantile-quantile plot)\nfig = plt.figure()\nres = stats.probplot(train['SalePrice'], plot=plt)\nplt.show()","metadata":{"id":"Y-Jf-Vx7StmU","outputId":"82a29ea9-3bb0-4be2-9b82-f98bc65cb82a","_uuid":"f2ef731a188b7b0afebe8c32848b4e5b687f5eaf","trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"####  TAKING 'Log Transform' OF THE TARGET","metadata":{"_uuid":"6c5a76522390df690194508b89d956c2413a4b11"}},{"cell_type":"code","source":"df['SalePrice']=np.log1p(df['SalePrice']) ","metadata":{"id":"9AlxRrRSUf54","_uuid":"34dfd5216c1208204c504d96356e91b6b1eca0b9","trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# now again see the distribution.\nsns.distplot(a=df['SalePrice'],color='#ff4125',axlabel=False).set_title('log(1+SalePrice)')  # better.\n","metadata":{"id":"kclqwnBcVK91","outputId":"f56455e6-9d21-48b6-8169-20b23da60fd5","_uuid":"f9db78388901e75b0e688fd2fa331a13b44135d6","trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"","metadata":{"_uuid":"cdc6592a4cb3162bc7613603d756b4b734e64118","trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## 2.5 ) Most Related Features to the Target","metadata":{"id":"3IobanMCVdjr","_uuid":"23acf3a9794d4efe96dc777a32fc43410d00266d"}},{"cell_type":"code","source":"cor_mat= df[:].corr()\ncor_with_tar=cor_mat.sort_values(['SalePrice'],ascending=False)","metadata":{"id":"_Lk4xeTwYb55","_uuid":"a768637b63c8ea05929584393235ddcbd546f4f5","trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"print(\"The most relevant features (numeric) for the target are :\")\ncor_with_tar.SalePrice","metadata":{"id":"NBFtyDPFY7w2","outputId":"4c5ee7ae-a463-4a84-80dc-0c1ed925c06e","_uuid":"9f5f7a6893c869a5f2df206c8e5362e8a72d9d96","trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"#### INFERENCES--\n\n1. Note that some of the features have quite high corelation with the target. These features are really significant.\n\n2. Of these the features with corelation value >0.5 are really important. Some features like GrLivArea etc.. are even more important.\n\n3. We will consider these features (i.e. GrLivArea,OverallQual) etc.. in more detail in subsequent sections during univariate and bivariate analysis.","metadata":{"_uuid":"f03326abdf136153dfe6a98b949bd631d245f7fb"}},{"cell_type":"code","source":"# using a corelation map to visualize features with high corelation.\ncor_mat= df[['OverallQual','GrLivArea','GarageCars','GarageArea','TotalBsmtSF','1stFlrSF','FullBath',\n             'YearBuilt','YearRemodAdd','GarageYrBlt','TotRmsAbvGrd','SalePrice']].corr()\nmask = np.array(cor_mat)\nmask[np.tril_indices_from(mask)] = False\nfig=plt.gcf()\nfig.set_size_inches(30,12)\nsns.heatmap(data=cor_mat,mask=mask,square=True,annot=True,cbar=True)\n\n# some inference section.","metadata":{"id":"-syFY9b2x8Lk","outputId":"a3dae7c2-7eca-4964-f073-e46940d69d19","_uuid":"49701df68af095c56881ba4174e3c1f90e28f508","trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## 2.6 ) Univariate Analysis","metadata":{"id":"fw-6lK9fbPmV","_uuid":"5657caa45c1fe5be6ae0fbad2aa138696df1ad96"}},{"cell_type":"markdown","source":"In this section the univariate analysis is performed; More importantly I have considered the features that are more importanht with the 'Target' that  have high corelation with the Target.\n\nFor the numeric features I have used a 'distplot' and 'boxplot' to analyze their distribution.\n\nSimilarly for categorical features the most reasonable way to visualize the distribution is to use a 'countplot' which shows the relative counts for each category or class. Can use a pie-plot also to be a bit more fancy.","metadata":{"_uuid":"08ad35f97d0a944ca6374dd2f33beefcab91e662"}},{"cell_type":"markdown","source":"#### NUMERIC FEATURES","metadata":{"_uuid":"d74aaaf0a07f7e8e928f94f5718ec59fad3f0a57"}},{"cell_type":"code","source":"def plot_num(feature):\n    fig,axes=plt.subplots(1,2)\n    sns.boxplot(data=df,x=feature,ax=axes[0])\n    sns.distplot(a=df[feature],ax=axes[1],color='#ff4125')\n    fig.set_size_inches(15,5)","metadata":{"id":"yvgAULu33MK7","_uuid":"ef7889b886ad93a3398e609c3efdf315ade8fc3c","trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"plot_num('GrLivArea')","metadata":{"id":"644x257a3ib2","outputId":"97246f79-0104-4da3-a412-deb8c9bd32b0","_uuid":"93ea4044587ce64a2101a752136ce553b580b0d1","trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"plot_num('GarageArea')","metadata":{"id":"gARg_gB_3k6M","outputId":"4e6f0136-f36c-4483-e6f1-33f37b02ceeb","_uuid":"872fa0472467eb781e194ad4ca71e0b338370f39","trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"plot_num('TotalBsmtSF') ","metadata":{"id":"zlbLPUvq4PQE","outputId":"492ea021-e4f1-493d-9ca8-a57b44093a24","_uuid":"8fc1b228a2b1f520a08919cafc8b4f550b82c230","trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"#### Note the features are a bit right skewed. We can therefore take 'log transform' of the features or a BoXCox transformation. Both shall work well. ","metadata":{"_uuid":"697ab496c69e9f74e365358b36097caefebabb80"}},{"cell_type":"markdown","source":"#### CATEGORICAL FEATURES","metadata":{"_uuid":"583b3bed2d49332ccf340260028959b7dfad6526"}},{"cell_type":"code","source":"def plot_cat(feature):\n  sns.countplot(data=df,x=feature)\n  ax=sns.countplot(data=df,x=feature)\n   ","metadata":{"id":"w2JJvTFk7hAW","_uuid":"2943424a50e07c8214d95d7288acad4a15373444","trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"plot_cat('OverallQual')","metadata":{"id":"f9lvMcEn9MGR","outputId":"e8065eb7-c4ec-4b4a-c40f-16becf880fe8","_uuid":"84c254b4fc1622f4170b9cd048173a0fe7a72300","trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Most of them are in 'average','above average' or 'good' classes.","metadata":{"_uuid":"5f3a23820f7659a0da3f1ace34c9f7bc139d2be0"}},{"cell_type":"code","source":"plot_cat('FullBath')","metadata":{"id":"Hnvc1R9a9O6w","outputId":"21ceec02-ddc2-4549-ecd5-f8543d710fd4","_uuid":"06b707b93bbfd1f8e2270729c39c51b1c4690ec0","trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"plot_cat('YearBuilt')","metadata":{"id":"vgBTm_RlDTEi","outputId":"abb31e01-0bd9-451e-ab00-d9ae431bd4a2","_uuid":"9fb657eed568b9b9927afd3e90fc0281cce55054","trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"plot_cat('TotRmsAbvGrd') # most of the houses have 5-7 rooms above the grd floor.","metadata":{"id":"4MMOaHh_DnCV","outputId":"1a7a2923-ae72-4990-fbb6-a2ac170de662","_uuid":"9639f81e330f269e6a44b6c33d38c47e7d7f21f2","trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"#### Lastly we plot the countplot for some important features that are numerical here but are actually categorica. It seems if they have been label encoded.","metadata":{"_uuid":"39665897a56e39e2b05b8bec90424edefb438ef7"}},{"cell_type":"code","source":"plot_cat('GarageCars')","metadata":{"id":"U-ENC6Z-D1xY","outputId":"447e0dee-7bd0-4a62-fa57-58083bec0e8f","_uuid":"9b0fe7b263804a35c55070c7b353262e80ed5f27","trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"sns.factorplot(data=df,x='Neighborhood',kind='count',size=10,aspect=1.5)","metadata":{"id":"QaTXlRlGFGOh","outputId":"d3b79482-b946-43f4-c8f9-f1019b99ed64","_uuid":"f9c35b6d4dd65dd83d292bee7ee82cc55c28e7f5","trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## 2.7 ) Bivariate Analysis","metadata":{"id":"E5PRfSVBFLb_","_uuid":"f996588441684c9837607971daa7abacd52612cd"}},{"cell_type":"markdown","source":"In this section the Bivariate Analysis have been done. I have plotted various numeric as well as categorical features against the target ie 'SalePrice'.","metadata":{"id":"w4rE3R1VaxRr","_uuid":"758efcb2501350c67e8b9e1f81a3e72073395a6c"}},{"cell_type":"markdown","source":"#### NUMERIC FEATURES","metadata":{"_uuid":"894b1d488fa5128fbafbeaae435406098d0f57e8"}},{"cell_type":"code","source":"fig, ax = plt.subplots()\nax.scatter(x = df['GrLivArea'], y = df['SalePrice'])\nplt.ylabel('SalePrice')\nplt.xlabel('GrLivArea')\nplt.show()","metadata":{"id":"0nu1lW4TOeUP","outputId":"b6d6d0ca-a3ea-4bc1-a33b-af36b8481adf","_uuid":"e8c6de274505ca40ac069878efa8d2506ccea0d4","trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"#### Note that there are two outliers on the lower right hand side and can remove them.","metadata":{"_uuid":"e4d2bcf38a5c11c77998a6a7ee563fba337adb6c"}},{"cell_type":"code","source":"df = df.drop(df[(df['GrLivArea']>4000) & (df['SalePrice']<13)].index) # removing some outliers on lower right side.","metadata":{"id":"iXz5EF3SOtXD","_uuid":"7a51231a67364d069f11be8154f17f9664152549","trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# again checking\nfig, ax = plt.subplots()\nax.scatter(x = df['GrLivArea'], y = df['SalePrice'])\nplt.ylabel('SalePrice')\nplt.xlabel('GrLivArea')\nplt.show()","metadata":{"id":"8EGtwyqYPmGd","outputId":"ba081003-29d8-4bbc-9a7f-fc5ea179bd10","_uuid":"3d774da1b04f96ef21d608d2b6afe2794a57bfd5","trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# garage area\nfig, ax = plt.subplots()\nax.scatter(x =(df['GarageArea']), y = df['SalePrice'])\nplt.ylabel('SalePrice')\nplt.xlabel('GarageArea')\nplt.show()\n# can try to fremove the points with gargae rea > than 1200.","metadata":{"id":"eRNHHRAEQNsR","outputId":"89f8f78d-e8e0-47b6-f24b-608d53dd4f9b","_uuid":"7da3ba1702816ad563c2e9ec56ea3ccffc14517e","trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# basment area\nfig, ax = plt.subplots()\nax.scatter(x =(df['TotalBsmtSF']), y = df['SalePrice'])\nplt.ylabel('SalePrice')\nplt.xlabel('TotalBsmtSF')\nplt.show()   # check >3000 can leave here.","metadata":{"id":"vb0EHQcvQa9i","outputId":"22d6f17a-55b1-40ee-83ad-30332a635147","_uuid":"8ea4c670bf2c73034260d20403f5b1966ff9024a","trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"#### CATEGORICAL FEATURES","metadata":{"id":"B6RmKMy3aYjC","_uuid":"818b76257578740524b3ad214345e4d613f1da4c"}},{"cell_type":"code","source":"#overall qual\nsns.factorplot(data=df,x='OverallQual',y='SalePrice',kind='box',size=5,aspect=1.5)","metadata":{"id":"uVoxXD4qbkwg","outputId":"837cb6e2-790d-4cbc-b288-65fd94dae7de","_uuid":"5cda223fd843480ccb929353af01ed49ca0c8ec7","trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"The SalePrice increases with the overall quality as expected.","metadata":{"_uuid":"0dd421bc3b36259a10954fefe66c350b094eea35"}},{"cell_type":"markdown","source":"**Similar inferences can be drawn from other plots and graphs.**","metadata":{"_uuid":"d50d9c93a621f7872c99e9a1d241dd6cea10432d"}},{"cell_type":"code","source":"#garage cars\nsns.factorplot(data=df,x='GarageCars',y='SalePrice',kind='box',size=5,aspect=1.5)","metadata":{"id":"r07wSkIqb6qm","outputId":"acffa9e3-5905-4c86-ebad-56be27dc24a2","_uuid":"84a3341247dfb1aed27949fba1ecd04d89e57790","trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"#no of rooms\nsns.factorplot(data=df,x='TotRmsAbvGrd',y='SalePrice',kind='bar',size=5,aspect=1.5) # increasing rooms imply increasing SalePrice as expected.","metadata":{"id":"1UQLMSfoccUa","outputId":"d9decf8d-d46f-4cbe-aa1a-eeb132a428fd","_uuid":"eeb74fef8f61fa303049d0ed8a01c0b4aba731f0","trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"#neighborhood\nsns.factorplot(data=df,x='Neighborhood',y='SalePrice',kind='box',size=10,aspect=1.5)","metadata":{"id":"CLVdUMFpccid","outputId":"7bc8444a-c952-4e64-a95a-84feada252a8","_uuid":"abd9933403120828e6ee5c25ba9cdfb826689204","trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Price varies with neighborhood.More posh areas of the city will have more price as expected.","metadata":{"_uuid":"e18ca9ff95617b1f1fe4b8970ba3db83ac28b6af"}},{"cell_type":"code","source":"#sale conditioin\nsns.factorplot(data=df,x='SaleCondition',y='SalePrice',kind='box',size=10,aspect=1.5)","metadata":{"id":"Gsfb00Bec7Bh","outputId":"ef5de537-a6b6-4f94-a4ae-b544e514477d","_uuid":"6be5bee40213f868a18a4beb8dfe4b85dbba4684","trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"","metadata":{"_uuid":"78e6b63db31bcee2ddb6456ac81e726706af7273","trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<a id=\"content3\"></a>\n## 3 ) Missing Values Treatment","metadata":{"id":"B3midIVVNiwC","_uuid":"46f7f5741f5aa1707a6299f3001faf697830cbe0"}},{"cell_type":"markdown","source":"In this section of the notebook I  have handled the missing values in the columns.\n\nFirstly I have droped a couple of columns that have a really high % of missing values.\n\nFor other features I have analyzed if it that feaure is important or not and accordingly either have drooped it or imputed the values in it.\n\nFor imputation I have considered the meaning of the corressponding feature from the description. Like for a categorical feature if values are missing I have imputed \"None\" just to mark a separate category meaning absence of that thing. Similarly for a numeric feature I have imputed with 0 in case the missing value implies the 'absence' of that feature.\n\nIn all other cases I have imputed the categorical features with 'mode' i.e the most frequent class and with 'mean' for the numeric features.","metadata":{"_uuid":"f64ab433106a127cd9cad6fab11d4d3d3c0f5ad0"}},{"cell_type":"code","source":"nan_all_data = (all_data.isnull().sum())\nnan_all_data= nan_all_data.drop(nan_all_data[nan_all_data== 0].index).sort_values(ascending=False)\nnan_all_data\nmiss_df = pd.DataFrame({'Missing Ratio' :nan_all_data})\nmiss_df\n","metadata":{"id":"D9qTYDqW4vTN","outputId":"141afe87-3874-4b62-f3c8-6a7215a5b60a","_uuid":"ecf7e7ffe2d0761a1784f2ff036ba7f86442c751","trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"#delet some features withvery high number of missing values.  \nall_data.drop(['PoolQC','Alley','Fence','Id','MiscFeature'],axis=1,inplace=True)\n","metadata":{"id":"Y8uiHcDnVsFy","_uuid":"70b85d4ee64e0aa64130ee1d5f99ef596a6624de","trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"test.drop(['PoolQC','Alley','Fence','MiscFeature'],axis=1,inplace=True)\ndf.drop(['PoolQC','Alley','Fence','MiscFeature'],axis=1,inplace=True)","metadata":{"id":"EpLjs9I0V40N","_uuid":"f38666912718e049f2410d25238ef1e3cdef01a2","trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# FireplaceQu\n# it is useful but many of the values nearly half are missing makes no sense to fill half of the values. so deleting this\nall_data.drop(['FireplaceQu'],axis=1,inplace=True)\ntest.drop(['FireplaceQu'],axis=1,inplace=True)\ndf.drop(['FireplaceQu'],axis=1,inplace=True)\n","metadata":{"id":"YVAe7zo4Y9eM","_uuid":"fe2681e864972a546416d134f67f5fc62ec54266","trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"#Lot Frontage\nprint(df['LotFrontage'].dtype)\nplt.scatter(x=np.log1p(df['LotFrontage']),y=df['SalePrice'])\ncr=df.corr()\nprint(df['LotFrontage'].describe())\nprint(\"The corelation of the LotFrontage with the Target : \" , cr.loc['LotFrontage','SalePrice'])\n","metadata":{"id":"_9FkZamQeLYs","outputId":"0ccd1bc6-4873-40eb-ee3b-c1c2aa3b075d","_uuid":"a969ace59b043a085935a9f49d7f5042fbe93e74","trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"#### Above analysis shows that there is some relation of LotArea with the SalePrice both by scatter plot and also by the corelation value. Therefore instead of deleting I will impute the values with the mean for now.","metadata":{"_uuid":"0c5d8e6743d0d91a200302150885b63677b7aeec"}},{"cell_type":"code","source":"all_data['LotFrontage'].fillna(np.mean(all_data['LotFrontage']),inplace=True)\nall_data['LotFrontage'].isna().sum()","metadata":{"id":"GrM9SeT_eV51","outputId":"1238829c-dc72-47e0-e165-088f0a58ee9c","_uuid":"eacffac67f960042b1bf54edce1bf964ab5b3c34","trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"#Garage  related features.\n# these features eg like garage qual,cond,finish,type seems to be important and relevant for buying car. \n# hence I will not drop these features insted i will fill them with the 'none' for categorical and 0 for numeric as nan here implies that there is no garage.\n\nall_data['GarageYrBlt'].fillna(0,inplace=True)\nprint(all_data['GarageYrBlt'].isnull().sum())\n\nall_data['GarageArea'].fillna(0,inplace=True)\nprint(all_data['GarageArea'].isnull().sum())\n\nall_data['GarageCars'].fillna(0,inplace=True)\nprint(all_data['GarageCars'].isnull().sum())\n\nall_data['GarageQual'].fillna('None',inplace=True)   # creating a separate category 'none' which means no garage.\nprint(all_data['GarageQual'].isnull().sum())\n\nall_data['GarageFinish'].fillna('None',inplace=True)   # creating a separate category 'none' which means no garage.\nprint(all_data['GarageFinish'].isnull().sum())\n\nall_data['GarageCond'].fillna('None',inplace=True)   # creating a separate category 'none' which means no garage.\nprint(all_data['GarageCond'].isnull().sum())\n\nall_data['GarageType'].fillna('None',inplace=True)   # creating a separate category 'none' which means no garage.\nprint(all_data['GarageType'].isnull().sum())\n\n","metadata":{"id":"HVTiAhW0fdUD","outputId":"1071cff2-6a96-48f6-e167-00a9ddd0b912","_uuid":"f45752ff4b70d6fb1f7a9360ef33554e34513d7d","trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# basement related features.\n#missing values are likely zero for having no basement\n\nfor col in ('BsmtFinSF1', 'BsmtFinSF2', 'BsmtUnfSF','TotalBsmtSF', 'BsmtFullBath', 'BsmtHalfBath'):\n    all_data[col].fillna(0,inplace=True)\n    \n# for categorical features we will create a separate class 'none' as before.\n\nfor col in ('BsmtQual', 'BsmtCond', 'BsmtExposure', 'BsmtFinType1', 'BsmtFinType2'):\n    all_data[col].fillna('None',inplace=True)\n    \nprint(all_data['TotalBsmtSF'].isnull().sum())\n\n","metadata":{"id":"5FVm7T-TJcOp","outputId":"e54543e3-ef00-4437-ea38-c9c07bfc21bd","_uuid":"e6fec292343a882f280f4c911960ef4bbf1e6240","trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# MasVnrArea 0 and MasVnrType 'None'.\nall_data['MasVnrArea'].fillna(0,inplace=True)\nprint(all_data['MasVnrArea'].isnull().sum())\n\nall_data['MasVnrType'].fillna('None',inplace=True)\nprint(all_data['MasVnrType'].isnull().sum())","metadata":{"id":"9kVtbvQ9fDI-","outputId":"61669f93-0d95-48b1-924b-237ed036d3ee","_uuid":"2c2af2fb54c9298db6975568a8b4cf1fb73eefe8","trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"#MSZoning.\n# Here nan does not mean no so I will with the most common one ie the mode.\nall_data['MSZoning'].fillna(all_data['MSZoning'].mode()[0],inplace=True)\nprint(all_data['MSZoning'].isnull().sum())","metadata":{"id":"TGcPNw3eh0oM","outputId":"88de891b-bfdb-491b-d4e6-fa380b7069e3","_uuid":"76542b40795ba9a8bc8c1d012279d0051c919dd9","trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# utilities\nsns.factorplot(data=df,kind='box',x='Utilities',y='SalePrice',size=5,aspect=1.5)","metadata":{"id":"zfZzEk-visch","outputId":"fb70640c-7ca6-4482-98ca-9cba4da89054","_uuid":"e126c026c4f1b558c877bb7c3b8ebb0882874a53","trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"#### Note that training set has only 2 of the possible 4 categories (ALLPub and NoSeWa) while test set has other categories. Hence it is of no use to us.","metadata":{"_uuid":"5b1df1c9ba2942a22dd47c6a9d6ab71ccfcf3941"}},{"cell_type":"code","source":"all_data.drop(['Utilities'],axis=1,inplace=True)","metadata":{"id":"HTRIvsuSqAgw","_uuid":"3b8e0aa08dfb8458501da6210da7161373a7ae7c","trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"#functional\n# fill with mode\nall_data['Functional'].fillna(all_data['Functional'].mode()[0],inplace=True)\nprint(all_data['Functional'].isnull().sum())","metadata":{"id":"GQKczgCbqwEo","outputId":"f396e006-4445-403c-f73c-50f400abf99e","_uuid":"73aeea1cdf2bb20e9f8dd8f3323cdea87edcd964","trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# other rem columns rae all cat like kitchen qual etc.. and so filled with mode.\nfor col in ['SaleType','KitchenQual','Exterior2nd','Exterior1st','Electrical']:\n  all_data[col].fillna(all_data[col].mode()[0],inplace=True)\n  print(all_data[col].isnull().sum())","metadata":{"id":"CGfbRZL8rMeL","outputId":"902de455-a075-42de-dba6-cf43913a12cc","_uuid":"fbedd807e1c4f622da5ed03ea7df60b60a19df8d","trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"#### Lastly checking if any null value still remains.","metadata":{"_uuid":"ba514bda4d20b3f112faade1999f227e531e0b3a"}},{"cell_type":"code","source":"nan_all_data = (all_data.isnull().sum())\nnan_all_data= nan_all_data.drop(nan_all_data[nan_all_data== 0].index).sort_values(ascending=False)\nnan_all_data\nmiss_df = pd.DataFrame({'Missing Ratio' :nan_all_data})\nmiss_df\n\n","metadata":{"id":"9bgxDFPXtAKx","outputId":"f6f86f66-0227-4baf-b2fd-d894e0c0d5e8","_uuid":"e1ff655694d7491bbddc2a4b2aa8dc50b4107a8e","trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"#### Finally no null value remain now;)","metadata":{"id":"Dw_6EQsptalc","_uuid":"5ab57510ff5cdc2f0abb8a31ff8130d03a204dcf"}},{"cell_type":"code","source":"all_data.shape","metadata":{"id":"uFQLIDCztcvg","outputId":"f6876acc-5676-43dc-d99b-7047d7ac9469","_uuid":"d354f032ce5419c9273b57bf126c29b1a6206d62","trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"","metadata":{"_uuid":"69a3a9f93644be3187e861034d0f70eec2bb6103","trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<a id=\"content4\"></a>\n## 4 ) Handling Skewness","metadata":{"_uuid":"c5165aa27f636f5fe5bc3ec13bd8e0aedf6fb923"}},{"cell_type":"markdown","source":"For handling skewnesss I will take the log transform of the features with skewness > 0.5.\n\nYou can also try the BoxCox transformation as mentioned before.","metadata":{"_uuid":"feb3d9c3a1f0aac59b5fc35cad9ffc9ce10c6325"}},{"cell_type":"code","source":"#log transform skewed numeric features:\nnumeric_feats = all_data.dtypes[all_data.dtypes != \"object\"].index\n\nskewed_feats = train[numeric_feats].apply(lambda x: skew(x.dropna())) #compute skewness\nskewed_feats = skewed_feats[skewed_feats > 0.50]\nskewed_feats = skewed_feats.index\n\nall_data[skewed_feats] = np.log1p(all_data[skewed_feats])","metadata":{"id":"UTq-0hDytkOB","_uuid":"a72c4df18f969747648b36374e055fb13c2a4b30","trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"","metadata":{"_uuid":"076823eac43269683f318f05a532998f07948465","trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<a id=\"content5\"></a>\n## 5 ) Prepare the Data","metadata":{"_uuid":"d60bda46c5320fae186cbf903921a2ff87c35074"}},{"cell_type":"markdown","source":"## 5.1 ) LabelEncode the Categorical Features","metadata":{"_uuid":"479a9ee569dbb9df2c39a778dc430a8b1ea2e25c"}},{"cell_type":"code","source":"for col in all_data.columns:\n    if(all_data[col].dtype == 'object'):\n        le=LabelEncoder()\n        all_data[col]=le.fit_transform(all_data[col])","metadata":{"id":"wXsWJvV1wFih","_uuid":"312389634b0670d37cc3713d8460959ed6441630","trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## 5.2 ) Splitting into Training and Validation Sets","metadata":{"_uuid":"4c80e36bb26ef306243b2fd2d185ee8ba876d5ed"}},{"cell_type":"code","source":"train=all_data.loc[:(df.shape)[0]+2,:]\ntest=all_data.loc[(df.shape)[0]+2:,:]","metadata":{"id":"vQvA2psvuIxs","_uuid":"f74e9522e62edd161a7e6454bf8c228e0ef2cd6f","trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"train['SalePrice']=df['SalePrice']\ntrain['SalePrice'].fillna(np.mean(train['SalePrice']),inplace=True)\ntrain.shape\nprint(train['SalePrice'].isnull().sum())","metadata":{"id":"BmYWTpiIudCF","outputId":"4646f8fa-4554-4463-b6d2-28a2fc762443","_uuid":"498023c0be50cbcab456bfe829b4b5d16896f323","trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"print(train.shape)\nprint(test.shape)","metadata":{"id":"IzfFpt4TvNhP","outputId":"886588f2-00d2-4847-a51b-cdcc1bcd99ba","_uuid":"def82a6baf6ce194c02c3e5956e28af81852dd67","trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"x_train,x_test,y_train,y_test=train_test_split(train.drop(['SalePrice'],axis=1),train['SalePrice'],test_size=0.20,random_state=42)","metadata":{"id":"9nbaJ_0PvV_Z","_uuid":"e4e66c811fa201617bc95116deb7172c9afb5e92","trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"","metadata":{"_uuid":"a1f15ce7bc98bd075e42739d858150edf77b452a","trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<a id=\"content6\"></a>\n## 6 ) Regression Models","metadata":{"_uuid":"197c3170e216d8b11706fcca233189b7d3aac72a"}},{"cell_type":"markdown","source":"Lastly it is the time to apply various regression models and check how are we doing. I have used various regression models from the scikit.\n\nParameter tuning using GridSearchCV is also done to improve performance of some algos.","metadata":{"_uuid":"cb15ddccd2ebd308123bbba35dba8ce54ffdfc67"}},{"cell_type":"markdown","source":"#### The evalauton metric that I have used is the Root Mean Squared Error between the 'Log of the actual price' and 'Log of the predicted value' which is also the evaluation metric used by the kaggle.\n\n#### To get abetter idea one may also use the K-fold cross validation insteadof the normal holdout set approach to cross validation.","metadata":{"_uuid":"db6c22d8772b6349962e8542c07dd63f1686c78c"}},{"cell_type":"markdown","source":"#### LINEAR REGRESSION","metadata":{"_uuid":"7b1193e60a51ea6655f9cc85a686923084d77943"}},{"cell_type":"code","source":"reg_lin=LinearRegression()\nreg_lin.fit(x_train,y_train)\npred=reg_lin.predict(x_test)\nprint(np.sqrt(mean_squared_error(y_test,pred)))","metadata":{"id":"TS6hi7x1vq1z","outputId":"a2c203bc-b9ff-4883-dfb6-e3a93393f2e3","_uuid":"6b9382b81c30023b140bdba176c31797385cd5c9","trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"#### LASSO (and tuning with GridSearchCV)","metadata":{"_uuid":"f897b88b90a2c2af032eaf8a0d0bc7d2fb7fedba"}},{"cell_type":"code","source":"reg_lasso=Lasso()\nreg_lasso.fit(x_train,y_train)\npred=reg_lasso.predict(x_test)\nprint(np.sqrt(mean_squared_error(y_test,pred)))","metadata":{"id":"stj2l4-dMpmZ","outputId":"764f6c2c-c1c4-42ff-8f33-bcb77d2bc6fd","_uuid":"7a3c7f2318ced9b90825293c0dbef674cb7f6012","trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"params_dict={'alpha':[0.001, 0.005, 0.01,0.05,0.1,0.5,1]}\nreg_lasso_CV=GridSearchCV(estimator=Lasso(),param_grid=params_dict,scoring='neg_mean_squared_error',cv=10)\nreg_lasso_CV.fit(x_train,y_train)\npred=reg_lasso_CV.predict(x_test)\nprint(np.sqrt(mean_squared_error(y_test,pred)))","metadata":{"id":"_847HJkh7VND","outputId":"e4724acf-71b6-4a9d-ac2f-ffb209639075","_uuid":"519732e88f3adb346e5df293d3b6dca7161305eb","trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"**Note the significant decrease in the RMSE on tuning the Lasso Regression.**","metadata":{"_uuid":"3356687bb6fce59ca44671b3eab443bd7820415f"}},{"cell_type":"code","source":"reg_lasso_CV.best_params_","metadata":{"id":"Qu6QAYmeE6si","outputId":"1266dde3-b76e-4bca-9888-7e62569a7656","_uuid":"39781decd5042c3526430650d438f55ca1a5fa5d","trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"#### RIDGE (and tuning with GridSearchCV)","metadata":{"_uuid":"51cd48ecb704744b024947fea49a29a0946f9e91"}},{"cell_type":"code","source":"reg_ridge=Ridge()\nreg_ridge.fit(x_train,y_train)\npred=reg_ridge.predict(x_test)\nprint(np.sqrt(mean_squared_error(y_test,pred)))","metadata":{"id":"F9DY3TMeAOow","outputId":"ac868822-0ef7-4ce1-9de8-25990d5c63b5","_uuid":"413490aaf8f00ebbf7427ee41891a8a21eea36d7","trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"params_dict={'alpha':[0.1, 0.15, 0.20,0.25,0.30,0.35,0.4,0.45,0.50,0.55,0.60]}\nreg_ridge_CV=GridSearchCV(estimator=Ridge(),param_grid=params_dict,scoring='neg_mean_squared_error',cv=10)\nreg_ridge_CV.fit(x_train,y_train)\npred=reg_ridge_CV.predict(x_test)\nprint(np.sqrt(mean_squared_error(y_test,pred)))","metadata":{"id":"h6v9m8XvM8CE","outputId":"cf214a27-7b2d-4501-fc02-2eb151a0b273","_uuid":"27dbf2110c9d68428c5cc79f0986bd172bdffd28","trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"reg_ridge_CV.best_params_","metadata":{"id":"11hc4js1H5mn","outputId":"71a30477-c42d-4402-a01a-fd905b83494c","_uuid":"5bc5357dfd01bf0b520b3f5fdefbd70f9412e31c","trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"#### GRADIENT BOOSTING","metadata":{"_uuid":"f51822af54321c643beed1fbe7c538bba463563c"}},{"cell_type":"code","source":"#the params are tuned with grid searchCV.\n\nreg_gb=GradientBoostingRegressor(n_estimators=2000,learning_rate=0.05,max_depth=3,min_samples_split=10,max_features='sqrt',subsample=0.75 ,loss='huber')\nreg_gb.fit(x_train,y_train)\npred=reg_gb.predict(x_test)\nprint(np.sqrt(mean_squared_error(y_test,pred)))","metadata":{"id":"Ui-TWe7dBGS1","outputId":"d5cb8bbb-6d64-4542-c285-4b94eb81cb02","_uuid":"2fab514d917340d537eba886ed3d7d394bd25208","trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"#### XGBoost","metadata":{"_uuid":"e841d23995587b72e9eddc0206b016fb715feace"}},{"cell_type":"code","source":"import xgboost as xgb\nmodel_xgb = xgb.XGBRegressor(colsample_bytree=0.4603, gamma=0.0468, \n                             learning_rate=0.05, max_depth=3, \n                             min_child_weight=1.7817, n_estimators=2200,\n                             reg_alpha=0.4640, reg_lambda=0.8571,\n                             subsample=0.5213, silent=1,\n                             random_state =7, nthread = -1)\nmodel_xgb.fit(x_train,y_train)\npred=model_xgb.predict(x_test)\nprint(np.sqrt(mean_squared_error(y_test,pred)))","metadata":{"id":"MLpTvMXC9Xsf","outputId":"15f1813c-98d2-47e9-b53b-962f6605c650","_uuid":"0ea75f225380c2d7ba15993273aff081a9fea9ff","trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Note that the parameters aren't optimized. This can get a lot better tahn this for sure.","metadata":{"_uuid":"17e7d3a42e92c46a7a8d57b9fc7a3fa8007976cc","trusted":true}},{"cell_type":"markdown","source":"<a id=\"content7\"></a>\n## 7 ) Saving and Making Submission to Kaggle","metadata":{"_uuid":"827fa6202505a7a0e49cb1acf87cd1743b8a0144"}},{"cell_type":"markdown","source":"**The Gradient Boosting gives the best performance on the validation set and so I am using it to make predictions to Kaggle (on the test set).**","metadata":{"_uuid":"690a1e1b362025f3e9656f4ebb56d2dc575beec6"}},{"cell_type":"code","source":"# predictions on the test set.\n \npred=reg_gb.predict(test)\npred_act=np.exp(pred)\npred_act=pred_act-1\nlen(pred_act)","metadata":{"id":"VFZMAR5gvw4p","outputId":"ba4017c8-cdfb-4d98-c0d6-bc28696655fa","_uuid":"db7ce015cbb8273ef4cb39160e7f1c837c21273b","trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"test_id=[]\nfor i in range (1461,2920):\n    test_id.append(i)\nd={'Id':test_id,'SalePrice':pred_act}\nans_df=pd.DataFrame(d)\nans_df.head(10)","metadata":{"id":"2dIptx4g3Kbq","outputId":"90315730-998b-454a-ea9f-bdad8f7e7fde","_uuid":"efc1406eba2a50593e2c4a0a19daba2032102dfe","trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"ans_df.to_csv('answer.csv',index=False)","metadata":{"id":"-1dCzHpc3PgO","_uuid":"c9201e2c59533dba1ec59dc5509116029d639fac","trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## THE END!!!","metadata":{"id":"XzFp9Kxd7Der","_uuid":"f897fe4a62539bc8feed329891180060894e9cc8"}},{"cell_type":"markdown","source":"## [Please star/ upvote if u liked it.]","metadata":{"id":"Eh8Op4fo7S2M","_uuid":"307577ab3e583728b69dc1f4ccc0c31fd73017df"}}]}