{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"pygments_lexer":"ipython3","nbconvert_exporter":"python","version":"3.6.4","file_extension":".py","codemirror_mode":{"name":"ipython","version":3},"name":"python","mimetype":"text/x-python"}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"# Kaggle Competition: House Prices - Advanced Regression Techniques","metadata":{}},{"cell_type":"markdown","source":"This dataset was posted on Kaggle, a data science competition website, for beginners to apply their knowledge of regression analysis to solve real life problems. Given the features describing almost every aspect of the houses in Ames, Iowa, participants are required to predict the final price of each house.","metadata":{}},{"cell_type":"markdown","source":"## Look at the Big Picture","metadata":{}},{"cell_type":"markdown","source":"**Why do we want to predict housing prices?**\n\nHousing as a form of asset represents a significant portion of people's wealth.","metadata":{}},{"cell_type":"markdown","source":"**Frame the Problem**\n\nThis is a regression task since we want to estiamte the price of each house as a continuous variable. With features and labels available, the task can be conducted using supervised learning techniques. Since real time updates of the learning models using new data is not required, we can safely assume that it is an offline task.","metadata":{}},{"cell_type":"markdown","source":"**Select a Performance Metric**\n\nAs this is a regression task, we would like to determine the accuracy of our predictiosn using poplular performance metrics such as root mean square error (RMSE) or mean absolute error (MAE). RMSE is more sensitive to outliers and generally performs better than MAE. Therefore, RMSE is used as a metric to make comparison across models and assess the model performance.","metadata":{}},{"cell_type":"markdown","source":"**Check the Assumptions**\n\nNow we check the assumptions made when we were framing the problem and selecting the performance metric. Upon observing the data in the training set, we can see that the selling prices are in fact continuous, and both features and the label are available. We can now move on to obtaining the relevant data.","metadata":{}},{"cell_type":"markdown","source":"## Get the Data","metadata":{}},{"cell_type":"markdown","source":"**Find and document where you can get the data**\n\nSince this project is based on a Kaggle competition, the data is already availble on the competition page. However, under the circumstances in real life, useful data can be obtained through the following means:\n\n1. Public databases on the Internet\n2. Census and Statistica report and data from the government\n3. Data from housing agencies and real estate companies\n\nImportantly, we should always seek authorization from the owner and be wary of the legal obligations involved.","metadata":{}},{"cell_type":"markdown","source":"**Format the Data**\n\nThe separation of the data into a train set and a test set is useful for evaluating the performance of the final model. We should also format the data so that it could be easily read and manipulated. If the size of the data is comparable to the RAM storage space, we should consider pipelining the data for better efficiency during the analysis and model training steps. Again, since the formatted train and test sets are readily available on the competition website, we are prepared to begin our data exploration.","metadata":{}},{"cell_type":"markdown","source":"## Explore the Data","metadata":{}},{"cell_type":"markdown","source":"Before developing prediction models for estimating the housing prices, we should first inspect and explore the nature of the data and the relationship between the variables. Data exploration helps us gain intuition about the task, which in turn gives us insights on feaure engineering and handling missing values later on.","metadata":{}},{"cell_type":"markdown","source":"**Create a copy of the data for exploration**\n\nFirst, we import the train set and the test set. A copy of the train set is made for data exploration to avoid contamination.","metadata":{}},{"cell_type":"code","source":"import numpy as np\nimport pandas as pd\ntrain = pd.read_csv('../input/house-prices-advanced-regression-techniques/train.csv')\ntest = pd.read_csv('../input/house-prices-advanced-regression-techniques/test.csv')\ntrain_exp = train.copy()","metadata":{"execution":{"iopub.status.busy":"2022-07-11T01:35:36.550881Z","iopub.execute_input":"2022-07-11T01:35:36.551349Z","iopub.status.idle":"2022-07-11T01:35:36.657252Z","shell.execute_reply.started":"2022-07-11T01:35:36.551259Z","shell.execute_reply":"2022-07-11T01:35:36.655994Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"train_exp.head()","metadata":{"execution":{"iopub.status.busy":"2022-07-11T01:35:36.659473Z","iopub.execute_input":"2022-07-11T01:35:36.659997Z","iopub.status.idle":"2022-07-11T01:35:36.70027Z","shell.execute_reply.started":"2022-07-11T01:35:36.659951Z","shell.execute_reply":"2022-07-11T01:35:36.698955Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"**Study each attribute and its characteristics**\n\nThen we study the attributes and their characterics such as the names, types, and number of missing values.","metadata":{}},{"cell_type":"code","source":"train_exp.info()","metadata":{"execution":{"iopub.status.busy":"2022-07-11T01:35:36.701811Z","iopub.execute_input":"2022-07-11T01:35:36.702242Z","iopub.status.idle":"2022-07-11T01:35:36.739224Z","shell.execute_reply.started":"2022-07-11T01:35:36.7022Z","shell.execute_reply":"2022-07-11T01:35:36.73805Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"train_exp.describe()","metadata":{"execution":{"iopub.status.busy":"2022-07-11T01:35:36.742595Z","iopub.execute_input":"2022-07-11T01:35:36.743089Z","iopub.status.idle":"2022-07-11T01:35:36.853901Z","shell.execute_reply.started":"2022-07-11T01:35:36.743043Z","shell.execute_reply":"2022-07-11T01:35:36.852725Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"train_exp.shape","metadata":{"execution":{"iopub.status.busy":"2022-07-11T01:35:36.855753Z","iopub.execute_input":"2022-07-11T01:35:36.856613Z","iopub.status.idle":"2022-07-11T01:35:36.864126Z","shell.execute_reply.started":"2022-07-11T01:35:36.856564Z","shell.execute_reply":"2022-07-11T01:35:36.863019Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"train_exp.dtypes.value_counts()","metadata":{"execution":{"iopub.status.busy":"2022-07-11T01:35:36.865625Z","iopub.execute_input":"2022-07-11T01:35:36.866098Z","iopub.status.idle":"2022-07-11T01:35:36.878155Z","shell.execute_reply.started":"2022-07-11T01:35:36.866037Z","shell.execute_reply":"2022-07-11T01:35:36.876906Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"We can see that there is a total of 80 features, 43 of which are 'objects', 34 of which are 'int64', and the rest of which are 'float64'. According to the documentation of the competition, The variable 'SalePrice', which has a data type of 'int64', is the target or label that we are going to predict.","metadata":{}},{"cell_type":"code","source":"train_na = train_exp.isna().sum()\ntrain_na[train_na != 0]","metadata":{"execution":{"iopub.status.busy":"2022-07-11T01:35:36.879275Z","iopub.execute_input":"2022-07-11T01:35:36.880322Z","iopub.status.idle":"2022-07-11T01:35:36.899958Z","shell.execute_reply.started":"2022-07-11T01:35:36.880286Z","shell.execute_reply":"2022-07-11T01:35:36.899067Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Percentage of missing values for each variable that contains missing values\ntrain_na[train_na != 0]/1460","metadata":{"execution":{"iopub.status.busy":"2022-07-11T01:35:36.901455Z","iopub.execute_input":"2022-07-11T01:35:36.902052Z","iopub.status.idle":"2022-07-11T01:35:36.915565Z","shell.execute_reply.started":"2022-07-11T01:35:36.902021Z","shell.execute_reply":"2022-07-11T01:35:36.914186Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"len(train_na[train_na != 0])","metadata":{"execution":{"iopub.status.busy":"2022-07-11T01:35:36.917432Z","iopub.execute_input":"2022-07-11T01:35:36.918204Z","iopub.status.idle":"2022-07-11T01:35:36.927293Z","shell.execute_reply.started":"2022-07-11T01:35:36.918151Z","shell.execute_reply":"2022-07-11T01:35:36.926344Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Among the 80 features, 19 of them contain missing values. 50 of the features, namely 'Alley', 'FireplaceQu', 'PoolQC', 'Fence', and 'MiscFeature' contain more than 40% of missing values. Due to the large percentage of missing values and the difficulties in handling them, we should consider dropping these features to avoid complexities, given that they do not significantly contribute to our predictions. We could decide upon finding their correlations to the housing price and reading the relevant documentation about these variables and their relationship to the target variable.","metadata":{}},{"cell_type":"markdown","source":"After finding the missing values in the train set, we continue our exploration by visualizing the data using graphs and plots. However, before jumping into the visualization step, we should classify the variables into cateogrical and numerical data. Although in general the categorical variables are usually stored as 'objects' and numerical data are usually stored as floats and integers, there are exceptions which should be converted and preprocessed. It is useful to take a look at the documentation, which provides descriptions and hints on the variable types.","metadata":{}},{"cell_type":"code","source":"train_type = pd.read_csv('../input/house-price/house_price_variable_analysis.csv')\nvar_type = train_type.iloc[0].T\nvar_type","metadata":{"execution":{"iopub.status.busy":"2022-07-11T01:35:36.931182Z","iopub.execute_input":"2022-07-11T01:35:36.93224Z","iopub.status.idle":"2022-07-11T01:35:36.98151Z","shell.execute_reply.started":"2022-07-11T01:35:36.932193Z","shell.execute_reply":"2022-07-11T01:35:36.980361Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"var_type_concat = pd.concat([var_type, train_exp.dtypes], axis=1)\nvar_type_concat.columns = ['VarType', 'DType']\nvar_type_concat","metadata":{"execution":{"iopub.status.busy":"2022-07-11T01:35:36.982808Z","iopub.execute_input":"2022-07-11T01:35:36.983151Z","iopub.status.idle":"2022-07-11T01:35:37.001139Z","shell.execute_reply.started":"2022-07-11T01:35:36.983118Z","shell.execute_reply":"2022-07-11T01:35:36.999757Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"num_filter = ((var_type_concat['VarType'] == 'num') & ((var_type_concat['DType'] == 'int64') | (var_type_concat['DType'] == 'float64')))\ncat_filter = (((var_type_concat['VarType'] == 'cat') | (var_type_concat['VarType'] == 'cat (ordinal)')) & (var_type_concat['DType'] == 'object'))\nfilter = ~(num_filter | cat_filter)\nvar_type_concat[filter]","metadata":{"execution":{"iopub.status.busy":"2022-07-11T01:35:37.002852Z","iopub.execute_input":"2022-07-11T01:35:37.003749Z","iopub.status.idle":"2022-07-11T01:35:37.019055Z","shell.execute_reply.started":"2022-07-11T01:35:37.003697Z","shell.execute_reply":"2022-07-11T01:35:37.017755Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"var_type_concat[var_type_concat['VarType'] == 'cat (ordinal)']","metadata":{"execution":{"iopub.status.busy":"2022-07-11T01:35:37.020604Z","iopub.execute_input":"2022-07-11T01:35:37.021082Z","iopub.status.idle":"2022-07-11T01:35:37.038248Z","shell.execute_reply.started":"2022-07-11T01:35:37.021037Z","shell.execute_reply":"2022-07-11T01:35:37.036921Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"The variable types are labelled and saved using Excel Spreadsheets after looking through the data description. From the above analysis, we can see that the variable 'MSSubClass' is the only variable that is categorical but its data type is int64. We should treat it as a categorical variable at the visualization and preprocessing stage. Besides, the values of some variables actually represents ordered levels, so we might consider treating them as ordinal variables.","metadata":{}},{"cell_type":"markdown","source":"**Visualize numerical data**\n\nNow we are ready to visualize our data. First of all, we plot the histograms for each numerical variable to take look at the distributions. We want to know if the distributions are skewed or symmetric, and whether there are any ceiling/floor effects.","metadata":{}},{"cell_type":"code","source":"# We will improve this part later on\nfrom matplotlib import pyplot as plt\nplt.style.use('fivethirtyeight')\nfig = plt.figure(figsize=(12,30))\ntrain_exp_var = train_exp[list(var_type_concat[num_filter].index)]\nfor ax_num, num_var in zip(range(1,len(train_exp_var.columns.values)+1), train_exp_var.columns.values):\n    plt.subplot(13,3,ax_num)\n    plt.hist(train_exp_var[num_var])\n    plt.xlabel(num_var)\nplt.tight_layout()","metadata":{"execution":{"iopub.status.busy":"2022-07-11T01:35:37.039734Z","iopub.execute_input":"2022-07-11T01:35:37.040097Z","iopub.status.idle":"2022-07-11T01:35:42.05153Z","shell.execute_reply.started":"2022-07-11T01:35:37.040054Z","shell.execute_reply":"2022-07-11T01:35:42.05016Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Next, we want to see how these numerical variables are correlated to the target. Therefore, we create a correlation matrix between variables and plot scatterplots to visualize the relationships.","metadata":{}},{"cell_type":"code","source":"train_exp_var.corr()['SalePrice'].nlargest(10)","metadata":{"execution":{"iopub.status.busy":"2022-07-11T01:35:42.053392Z","iopub.execute_input":"2022-07-11T01:35:42.05409Z","iopub.status.idle":"2022-07-11T01:35:42.075238Z","shell.execute_reply.started":"2022-07-11T01:35:42.054047Z","shell.execute_reply":"2022-07-11T01:35:42.07416Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"fig = plt.figure(figsize=(12, 30))\n\nfor ax_num, num_var in zip(range(1,len(train_exp_var.columns.values)+1), train_exp_var.columns.values):\n    plt.subplot(13,3,ax_num)\n    plt.scatter(train_exp_var['SalePrice'], train_exp_var[num_var])\n    plt.ylabel(num_var)\n    plt.gca().set(yticklabels=[])\n    plt.gca().set(xticklabels=[])\nfig.supxlabel('Saleprice')\nplt.tight_layout()","metadata":{"execution":{"iopub.status.busy":"2022-07-11T01:35:42.076798Z","iopub.execute_input":"2022-07-11T01:35:42.077486Z","iopub.status.idle":"2022-07-11T01:35:46.033341Z","shell.execute_reply.started":"2022-07-11T01:35:42.07745Z","shell.execute_reply":"2022-07-11T01:35:46.031936Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"**Visualize Categorical Data**","metadata":{}},{"cell_type":"code","source":"train_exp_cat = train_exp[['MSSubClass'] + list(var_type_concat[cat_filter].index)]\nfig = plt.figure(figsize=(12,30))\nfor ax_num, cat_var in zip(range(1, len(['MSSubClass'] + var_type_concat[cat_filter].index)+1), ['MSSubClass'] + list(var_type_concat[cat_filter].index)):\n    plt.subplot(9,5,ax_num)\n    plt.bar(train_exp_cat[cat_var].value_counts().index, train_exp_cat[cat_var].value_counts().values)\n    plt.xlabel(cat_var)\n\nplt.tight_layout()\n","metadata":{"execution":{"iopub.status.busy":"2022-07-11T01:35:46.035346Z","iopub.execute_input":"2022-07-11T01:35:46.035836Z","iopub.status.idle":"2022-07-11T01:35:51.758621Z","shell.execute_reply.started":"2022-07-11T01:35:46.03579Z","shell.execute_reply":"2022-07-11T01:35:51.757197Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## Prepare the Data","metadata":{}},{"cell_type":"markdown","source":"Separate the features and target in the train set.","metadata":{}},{"cell_type":"code","source":"X_train = train.drop(columns='SalePrice')\ny_train = train['SalePrice']\nX_train","metadata":{"execution":{"iopub.status.busy":"2022-07-11T01:35:51.764718Z","iopub.execute_input":"2022-07-11T01:35:51.765534Z","iopub.status.idle":"2022-07-11T01:35:51.808178Z","shell.execute_reply.started":"2022-07-11T01:35:51.765484Z","shell.execute_reply":"2022-07-11T01:35:51.806825Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"**Data Cleaning**\n\nFirstly, we should decide how we are going to handle the missing data. We have found that 19 features contain missing values, five of which have more than 40% missing. We try and impute the missing data according to the strategies below:\n\n1. For numerical data, we fill in the missing data along with data transformation using the median of the variable.\n2. For categorical data, we fill in the missing data along with data transformation using the mode of the variable.","metadata":{}},{"cell_type":"markdown","source":"**Feature Selection**\n\nSince Id is not likely to provide any information on the prediction of housing prices, the variable will be dropped.","metadata":{}},{"cell_type":"markdown","source":"**Handling Categorical and Ordinal Variables**\n\n1. For categorical variables, we use one-hot encoding to encode the data.\n2. For ordinal variables, we use ordinal encoding to encode the data.","metadata":{}},{"cell_type":"markdown","source":"**Feature Scaling**\n\nAll numerical variables will be standardized for more accurate predictions and faster convergence.","metadata":{}},{"cell_type":"markdown","source":"https://scikit-learn.org/stable/modules/generated/sklearn.preprocessing.OrdinalEncoder.html#:~:text=list%20%3A%20categories%5Bi%5D%20holds%20the%20categories%20expected%20in%20the%20ith%20column.%20The%20passed%20categories%20should%20not%20mix%20strings%20and%20numeric%20values%2C%20and%20should%20be%20sorted%20in%20case%20of%20numeric%20values.","metadata":{}},{"cell_type":"markdown","source":"**Transformation Pipelines**\n\nAll things considered, we can now combine our data cleaning strategies and build a data transformation pipeline. The pipeline is built using preprocessing methods in sklearn due to the following reasons:\n\n1. It ensures the compatibility of the data with the sklearn models.\n2. It can be reused for the test set.\n3. The original datasets will not be contaminated.","metadata":{}},{"cell_type":"code","source":"# all numerical variables except Id\nnum_filter = ((var_type_concat['VarType'] == 'num') & ((var_type_concat['DType'] == 'int64') | (var_type_concat['DType'] == 'float64')))\nnum_var = list(var_type_concat[num_filter].index)\nnum_var.remove('SalePrice')\nnum_var.remove('Id')\n# all nominal variables\ncat_filter = (((var_type_concat['VarType'] == 'cat')) & (var_type_concat['DType'] == 'object'))\ncat_var = ['MSSubClass'] + list(var_type_concat[cat_filter].index)\n# all ordinal variables\nord_filter = (((var_type_concat['VarType'] == 'cat (ordinal)')) & (var_type_concat['DType'] == 'object'))\nord_var = list(var_type_concat[ord_filter].index)\nord_var","metadata":{"execution":{"iopub.status.busy":"2022-07-11T01:35:51.810139Z","iopub.execute_input":"2022-07-11T01:35:51.810934Z","iopub.status.idle":"2022-07-11T01:35:51.827887Z","shell.execute_reply.started":"2022-07-11T01:35:51.810883Z","shell.execute_reply":"2022-07-11T01:35:51.826852Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"levels = {\n    'LotShape': ['IR3', 'IR2', 'IR1', 'Reg'],\n    'LandSlope': ['Sev', 'Mod', 'Gtl'],\n    'ExterQual': ['Po', 'Fa', 'TA', 'Gd', 'Ex'],\n    'ExterCond': ['Po', 'Fa', 'TA', 'Gd', 'Ex'],\n    'BsmtQual': ['NA', 'Po', 'Fa', 'TA', 'Gd', 'Ex'],\n    'BsmtCond': ['NA', 'Po', 'Fa', 'TA', 'Gd', 'Ex'],\n    'BsmtExposure': ['NA', 'No', 'Mn', 'Av', 'Gd'],\n    'BsmtFinType1': ['NA', 'Unf', 'LwQ', 'Rec', 'BLQ', 'ALQ', 'GLQ'],\n    'BsmtFinType2': ['NA', 'Unf', 'LwQ', 'Rec', 'BLQ', 'ALQ', 'GLQ'],\n    'HeatingQC': ['Po', 'Fa', 'TA', 'Gd', 'Ex'],\n    'FireplaceQu': ['NA', 'Po', 'Fa', 'TA', 'Gd', 'Ex'],\n    'GarageQual': ['NA', 'Po', 'Fa', 'TA', 'Gd', 'Ex'],\n    'GarageCond': ['NA', 'Po', 'Fa', 'TA', 'Gd', 'Ex'],\n    'PoolQC': ['NA', 'Fa', 'TA', 'Gd', 'Ex']\n}\n\ncategories = []\nfor name, level in levels.items():\n    categories.append(level)\n\n","metadata":{"execution":{"iopub.status.busy":"2022-07-11T01:35:51.829201Z","iopub.execute_input":"2022-07-11T01:35:51.830044Z","iopub.status.idle":"2022-07-11T01:35:51.840247Z","shell.execute_reply.started":"2022-07-11T01:35:51.83001Z","shell.execute_reply":"2022-07-11T01:35:51.839193Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"from sklearn.pipeline import Pipeline\nfrom sklearn.impute import SimpleImputer\nfrom sklearn.preprocessing import StandardScaler\nfrom sklearn.compose import ColumnTransformer\nfrom sklearn.preprocessing import OrdinalEncoder\nfrom sklearn.preprocessing import OneHotEncoder\nfrom sklearn.preprocessing import MinMaxScaler\n\nnum_pipeline = Pipeline([\n    ('num_imputer', SimpleImputer(strategy='mean')),\n    ('num_std_scaler', StandardScaler())\n])\n\ncat_pipeline = Pipeline([\n    ('cat_imputer', SimpleImputer(strategy='constant')),\n    ('oh_encoder', OneHotEncoder(handle_unknown='ignore')) # ignore new value when transforming test set\n])\n\n\nord_pipeline = Pipeline([\n    ('ord_imputer', SimpleImputer(strategy='constant', fill_value='NA')),\n    ('ord_encoder', OrdinalEncoder(categories=categories, handle_unknown='use_encoded_value', unknown_value=-1)),\n    ('ord_std_scaler', StandardScaler())\n])\n\nfull_pipeline = ColumnTransformer([\n    ('num', num_pipeline, num_var),\n    ('cat', cat_pipeline, cat_var),\n    ('ord', ord_pipeline, ord_var)\n], remainder='drop') # The 'id' column will be dropped","metadata":{"execution":{"iopub.status.busy":"2022-07-11T01:35:51.842205Z","iopub.execute_input":"2022-07-11T01:35:51.842985Z","iopub.status.idle":"2022-07-11T01:35:52.662743Z","shell.execute_reply.started":"2022-07-11T01:35:51.842935Z","shell.execute_reply":"2022-07-11T01:35:52.661774Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## Shortlist Promising Models","metadata":{}},{"cell_type":"markdown","source":"We propose some commonly used regression models and compare them using a hold-out validation set. Complex models such as neural networks might not work that well given the relatively small sample size. We also compare the models with different hyperparameters.","metadata":{}},{"cell_type":"code","source":"# import all regression methods\nfrom sklearn.linear_model import LinearRegression\nfrom sklearn.preprocessing import PolynomialFeatures\nfrom sklearn.linear_model import Ridge\nfrom sklearn.linear_model import Lasso\nfrom sklearn.linear_model import ElasticNet\nfrom sklearn.svm import LinearSVR\nfrom sklearn.svm import SVR\nfrom sklearn.tree import DecisionTreeRegressor\nfrom sklearn.ensemble import GradientBoostingRegressor\nfrom xgboost import XGBRegressor\n\n# instantiate the methods\nlin_reg = LinearRegression()\nridge_reg = Ridge()\nlasso_reg = Lasso(tol = 1)\nelastic_reg = ElasticNet()\nlin_svr_reg = LinearSVR(C=5, random_state=42)\nsvr_reg = SVR()\ntree_reg = DecisionTreeRegressor()\ngbr_reg = GradientBoostingRegressor()\nxgb_reg = XGBRegressor()\n\n# store the methods in a dictionary\nmethods = {\n    'Linear Regressor': lin_reg,\n    'Ridge Regressor': ridge_reg,\n    'Lasso Regressor': lasso_reg,\n    'Elastic Net Regressor': elastic_reg,\n    'Linear SVM Regressor': lin_svr_reg,\n    'SVM Regressor': svr_reg,\n    'Decision Tree Regressor': tree_reg,\n    'Gradient Boosting Regressor': gbr_reg,\n    'XGBoost Regressor': xgb_reg\n}\n\n# Grids for grid search\nlin_grid = {\n\n}\nridge_grid = {\n    'alpha': [0.1, 0.2, 0.5, 1]\n}\nlasso_grid = {\n    'alpha': [0.1, 0.2, 0.5, 1]\n}\nelastic_grid = {\n    'alpha': [0.1, 0.2, 0.5, 1]\n}\nlin_svr_grid = {\n    'C': [10**x for x in range(-2, 3)],\n}\nsvr_grid = {\n    'C': [10**x for x in range(-2, 3)],\n    'kernel': ['rbf', 'poly'],\n    'coef0': [10**x for x in range(-2, 3)],\n    'gamma': [10**x for x in range(-2, 3)]\n}\ntree_grid = {}\ngbr_grid = {\n    'learning_rate': [0.1, 0.2, 0.5, 1],\n    'n_estimators': [50, 80, 100, 150],\n    'loss': ['squared_error', 'huber'],\n    'alpha': [0.1, 0.5, 0.9]\n}\nxgb_grid = {'n_estimators': [600,700,800,900,1000],\n        'learning_rate': [0.01,0.05,0.1],\n        'max_depth': [3,5,6],\n        'gamma': [0,0.5,1,]\n       }\n\ngrids = {\n    'Linear Regressor': lin_grid,\n    'Ridge Regressor': ridge_grid,\n    'Lasso Regressor': lasso_grid,\n    'Elastic Net Regressor': elastic_grid,\n    'Linear SVM Regressor': lin_svr_grid,\n    'SVM Regressor': svr_grid,\n    'Decision Tree Regressor': tree_grid,\n    'Gradient Boosting Regressor': gbr_grid,\n    'XGBoost Regressor': xgb_grid\n}\n\nparameters = {}\nfor name, regressor in methods.items():\n    parameters[name] = regressor.get_params()\n\nfor name, params in parameters.items():\n    print(name)\n    print(params)\n    print('\\n')","metadata":{"execution":{"iopub.status.busy":"2022-07-11T01:35:52.664717Z","iopub.execute_input":"2022-07-11T01:35:52.665519Z","iopub.status.idle":"2022-07-11T01:35:52.99288Z","shell.execute_reply.started":"2022-07-11T01:35:52.665476Z","shell.execute_reply":"2022-07-11T01:35:52.991659Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"X_train_prepared = full_pipeline.fit_transform(X_train)","metadata":{"execution":{"iopub.status.busy":"2022-07-11T01:35:52.994742Z","iopub.execute_input":"2022-07-11T01:35:52.995587Z","iopub.status.idle":"2022-07-11T01:35:53.058413Z","shell.execute_reply.started":"2022-07-11T01:35:52.995501Z","shell.execute_reply":"2022-07-11T01:35:53.05718Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"from sklearn.model_selection import GridSearchCV\n\nperformance = pd.DataFrame(\n    columns=['Regressor', 'Best CV Mean Score', 'Best Parameters'])\n\nbest_mean = 0\nfor name, method in methods.items():\n    method_cv = GridSearchCV(\n        method, param_grid=grids[name], n_jobs=-1, scoring='r2', cv=3)\n    method_cv.fit(X_train_prepared, y_train)\n    print(f'Regressor: {name}')\n    print(f'Mean Score: {method_cv.best_score_}')\n    std = method_cv.cv_results_['std_test_score'][method_cv.best_index_]\n    print(f'Standard Deviation: {std}')\n    performance = performance.append({'Regressor': name, 'Best CV Mean Score': method_cv.best_score_, 'Best CV Standard Deviation': std,\n                                     'Best Parameters': method_cv.best_params_}, ignore_index=True)\n    if method_cv.best_score_ > best_mean:\n        best_method = name\n        best_mean = method_cv.best_score_\n        best_std = std\n        best_params = method_cv.best_params_\n\n\nprint(f'Best Model: {best_method}')\nprint(f'Best Mean Score: {best_mean}')\nprint(f'Standard Deviation of the Best Mean Score: {best_std}')\nprint(f'Best Parameters: {best_params}')\nprint(performance)\n","metadata":{"execution":{"iopub.status.busy":"2022-07-11T01:35:53.060211Z","iopub.execute_input":"2022-07-11T01:35:53.060865Z","iopub.status.idle":"2022-07-11T01:37:53.823555Z","shell.execute_reply.started":"2022-07-11T01:35:53.060815Z","shell.execute_reply":"2022-07-11T01:37:53.822111Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"performance_copy = performance.copy()\nperformance_copy.loc[0, 'Best CV Mean Score'] = 0\nperformance_copy = performance_copy.sort_values(by='Best CV Mean Score', ascending=True)\nplt.barh(performance_copy['Regressor'], performance_copy['Best CV Mean Score'])","metadata":{"execution":{"iopub.status.busy":"2022-07-11T01:37:53.824815Z","iopub.status.idle":"2022-07-11T01:37:53.825806Z","shell.execute_reply.started":"2022-07-11T01:37:53.825551Z","shell.execute_reply":"2022-07-11T01:37:53.825573Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"From the above bar graph we see that the Gradient Boosting Regressor performs the best. Let's find out the best hyperparameters of the best model.","metadata":{}},{"cell_type":"code","source":"pd.set_option('display.max_colwidth', None)  # or 199\nprint(performance_copy[(performance_copy['Regressor'] == 'Gradient Boosting Regressor')]['Best Parameters'])","metadata":{"execution":{"iopub.status.busy":"2022-07-11T01:37:53.827025Z","iopub.status.idle":"2022-07-11T01:37:53.827659Z","shell.execute_reply.started":"2022-07-11T01:37:53.82745Z","shell.execute_reply":"2022-07-11T01:37:53.827471Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## Fine-tune the System","metadata":{}},{"cell_type":"markdown","source":"Knowing that the Gradient Boosting Regressor performs the best, we fine-tune the model by fine-tuning the hyperparameters using GridSearchCV once again.","metadata":{}},{"cell_type":"code","source":"gbr_params = {'learning_rate': [0.03, 0.05, 0.1, 0.13, 0.15],'n_estimators': [130, 150, 180, 200, 250]}\ngbr_cv = GridSearchCV(gbr_reg, param_grid=gbr_params, n_jobs=-1, scoring='r2', cv=5, verbose=2)\ngbr_cv.fit(X_train_prepared, y_train)  \nprint(f'Best Parameters: {gbr_cv.best_params_}')\nprint(f'Best Score: {gbr_cv.best_score_}') # mean cv score","metadata":{"execution":{"iopub.status.busy":"2022-07-11T01:37:53.828883Z","iopub.status.idle":"2022-07-11T01:37:53.829513Z","shell.execute_reply.started":"2022-07-11T01:37:53.829312Z","shell.execute_reply":"2022-07-11T01:37:53.829333Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"The best parameters are `learning_rate = 0.13` and `n_estimators = 150`. Finally, we can use these values to predict the targets in the test set.","metadata":{}},{"cell_type":"code","source":"X_test_prepared = full_pipeline.transform(test)\nbest_reg = GradientBoostingRegressor(alpha=0.13, n_estimators=150)\nbest_reg.fit(X_train_prepared, y_train)\ny_test_pred = best_reg.predict(X_test_prepared)","metadata":{"execution":{"iopub.status.busy":"2022-07-11T01:37:53.830707Z","iopub.status.idle":"2022-07-11T01:37:53.831338Z","shell.execute_reply.started":"2022-07-11T01:37:53.831137Z","shell.execute_reply":"2022-07-11T01:37:53.831158Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"y_test_pred","metadata":{"execution":{"iopub.status.busy":"2022-07-11T01:37:53.832476Z","iopub.status.idle":"2022-07-11T01:37:53.833137Z","shell.execute_reply.started":"2022-07-11T01:37:53.83289Z","shell.execute_reply":"2022-07-11T01:37:53.832922Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"final = pd.read_csv('../input/house-prices-advanced-regression-techniques/sample_submission.csv')\nfinal.iloc[:,1] = y_test_pred\nfinal.to_csv('final_submission.csv', index=False)","metadata":{"execution":{"iopub.status.busy":"2022-07-11T01:37:53.834485Z","iopub.status.idle":"2022-07-11T01:37:53.834914Z","shell.execute_reply.started":"2022-07-11T01:37:53.834717Z","shell.execute_reply":"2022-07-11T01:37:53.834737Z"},"trusted":true},"execution_count":null,"outputs":[]}]}