{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"pygments_lexer":"ipython3","nbconvert_exporter":"python","version":"3.6.4","file_extension":".py","codemirror_mode":{"name":"ipython","version":3},"name":"python","mimetype":"text/x-python"}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"<img src = \"https://www.yourmoney.com/wp-content/uploads/sites/3/2022/02/house-prices-scaled.jpg\" >\n\n# <center> House Prices Prediction with XGBoost </center>","metadata":{}},{"cell_type":"markdown","source":"# Introduction\n\nThis notebook is my **first submission** for a machine learning competition on Kaggle. <br><br>\nThe task is to create a machine learning model to **predict sales prices** for residential homes in Aimes, Iowa, and practice skills such as **feature engineering** and **exploratory data analysis**. <br><br>\nThe metric used to evaluate the model's performance is the **RMSE, Root Mean Squared Error**, as stated in the overview for this competition on Kaggle.<br><br>\nLet's go then!","metadata":{}},{"cell_type":"code","source":"# Importing libraries\nimport pandas as pd, numpy as np, seaborn as sns, matplotlib.pyplot as plt, plotly.express as px, xgboost as XGB\nfrom scipy.stats import norm, skew\nfrom sklearn.model_selection import train_test_split, cross_val_score,KFold\nfrom sklearn.metrics import mean_squared_error\nimport warnings\nwarnings.filterwarnings('ignore')","metadata":{"execution":{"iopub.status.busy":"2022-07-06T01:45:36.496393Z","iopub.execute_input":"2022-07-06T01:45:36.496754Z","iopub.status.idle":"2022-07-06T01:45:36.504127Z","shell.execute_reply.started":"2022-07-06T01:45:36.496724Z","shell.execute_reply":"2022-07-06T01:45:36.502726Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Lodading datasets\ntrain = pd.read_csv(\"../input/house-prices-advanced-regression-techniques/train.csv\")\ntest = pd.read_csv(\"../input/house-prices-advanced-regression-techniques/test.csv\")","metadata":{"execution":{"iopub.status.busy":"2022-07-06T01:20:29.626274Z","iopub.execute_input":"2022-07-06T01:20:29.626832Z","iopub.status.idle":"2022-07-06T01:20:29.677114Z","shell.execute_reply.started":"2022-07-06T01:20:29.626783Z","shell.execute_reply":"2022-07-06T01:20:29.675688Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"print(f\"The training set has {train.shape[0]} rows and {train.shape[1]} columns\\n\")\nprint(f\"The testing set has {test.shape[0]} rows and {test.shape[1]} columns\")","metadata":{"execution":{"iopub.status.busy":"2022-07-06T01:20:29.739429Z","iopub.execute_input":"2022-07-06T01:20:29.739837Z","iopub.status.idle":"2022-07-06T01:20:29.746195Z","shell.execute_reply.started":"2022-07-06T01:20:29.739801Z","shell.execute_reply":"2022-07-06T01:20:29.744925Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"train.head()","metadata":{"execution":{"iopub.status.busy":"2022-07-06T01:20:29.814177Z","iopub.execute_input":"2022-07-06T01:20:29.814567Z","iopub.status.idle":"2022-07-06T01:20:29.843424Z","shell.execute_reply.started":"2022-07-06T01:20:29.814520Z","shell.execute_reply":"2022-07-06T01:20:29.842132Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"test.head()","metadata":{"execution":{"iopub.status.busy":"2022-07-06T01:20:29.879315Z","iopub.execute_input":"2022-07-06T01:20:29.879708Z","iopub.status.idle":"2022-07-06T01:20:29.909973Z","shell.execute_reply.started":"2022-07-06T01:20:29.879669Z","shell.execute_reply":"2022-07-06T01:20:29.908689Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Right away we can see a lot of null values. We'll deal with them later on.","metadata":{}},{"cell_type":"markdown","source":"## Analysing Target Variable: SalePrice","metadata":{}},{"cell_type":"code","source":"print(train.SalePrice.describe())\nplt.figure(figsize=(20,10))\nsns.distplot(train.SalePrice)","metadata":{"execution":{"iopub.status.busy":"2022-07-06T01:20:29.966559Z","iopub.execute_input":"2022-07-06T01:20:29.966939Z","iopub.status.idle":"2022-07-06T01:20:30.307774Z","shell.execute_reply.started":"2022-07-06T01:20:29.966905Z","shell.execute_reply":"2022-07-06T01:20:30.306464Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Just by looking at the distribution plot, we can notice that our target variable doesn't follow a normal distribution of values. <br><br>\n\nIn statistics, the degree of asymmetry observed in a distribution that deviates from the normal distribution (bell curve) in a certain dataset is called **skewness**.<br><br>\n\nIn a normal distribution, the left-hand side and right-hand side contain the same number of observations. In our case, we have a positive skewed distribution, where data are mostly concentrated on the lef-hand side of the chart and the mean is higher than the median.<br><br>\n\nTo avoid having issues with our model's predicting abilities, we must transform this data and bring it closer to a normal distribution.<br><br>","metadata":{}},{"cell_type":"code","source":"print(\"Skeweness: \", train.SalePrice.skew())","metadata":{"execution":{"iopub.status.busy":"2022-07-06T01:20:30.309723Z","iopub.execute_input":"2022-07-06T01:20:30.310160Z","iopub.status.idle":"2022-07-06T01:20:30.316354Z","shell.execute_reply.started":"2022-07-06T01:20:30.310125Z","shell.execute_reply":"2022-07-06T01:20:30.315488Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Transforming target variable distribution\ntrain.SalePrice = np.log1p(train.SalePrice)","metadata":{"execution":{"iopub.status.busy":"2022-07-06T01:20:30.317684Z","iopub.execute_input":"2022-07-06T01:20:30.318684Z","iopub.status.idle":"2022-07-06T01:20:30.330451Z","shell.execute_reply.started":"2022-07-06T01:20:30.318647Z","shell.execute_reply":"2022-07-06T01:20:30.329438Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"print(\"Skeweness: \", train.SalePrice.skew())\nplt.figure(figsize=(20,10))\nsns.distplot(train.SalePrice, fit = norm)","metadata":{"execution":{"iopub.status.busy":"2022-07-06T01:20:30.332451Z","iopub.execute_input":"2022-07-06T01:20:30.333813Z","iopub.status.idle":"2022-07-06T01:20:30.651755Z","shell.execute_reply.started":"2022-07-06T01:20:30.333771Z","shell.execute_reply":"2022-07-06T01:20:30.650392Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Now we have a skewness closer to 0, bringing the values closer to a normal distribution.","metadata":{}},{"cell_type":"code","source":"# Printing more relevant correlations\ncorr = train.corr()\nmost_correlated_features = corr.index[abs(corr.SalePrice) > 0.5]\nplt.figure(figsize = (15,10))\nmask = np.triu(np.ones_like(train[most_correlated_features].corr(), dtype = np.bool))\ng = sns.heatmap(train[most_correlated_features].corr(), annot = True, mask = mask)","metadata":{"execution":{"iopub.status.busy":"2022-07-06T01:20:30.653555Z","iopub.execute_input":"2022-07-06T01:20:30.653971Z","iopub.status.idle":"2022-07-06T01:20:31.298757Z","shell.execute_reply.started":"2022-07-06T01:20:30.653937Z","shell.execute_reply":"2022-07-06T01:20:31.297277Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"- The highest correlation we have for Sale Price is 'OverallQual' with 0.82. <br><br>\n- '1stFlrSF' and 'TotalBsmtSF' have a correlation of 0.82.<br><br>\n- 'TotRmsAbvGrd' and 'GrLivArea'  have a correlation of 0.83.<br><br>\n- 'GarageYrBlt' and 'YearBuilt' have a correlation of 0.83.<br><br>\n- 'GarageArea' and 'GarageCars' have a correlation of 0.88.<br><br>","metadata":{}},{"cell_type":"markdown","source":"# Dealing with Missing Data\n\nTo deal with missing data, let's concatenate the train set and test set into a unified dataset.","metadata":{}},{"cell_type":"code","source":"y_train = train.SalePrice\nid_test = test.Id\ndf = pd.concat([train, test], axis = 0, sort = False)\ndf = df.drop(['SalePrice','Id'], axis = 1)\ndf.head()","metadata":{"execution":{"iopub.status.busy":"2022-07-06T01:20:31.300536Z","iopub.execute_input":"2022-07-06T01:20:31.301296Z","iopub.status.idle":"2022-07-06T01:20:31.345901Z","shell.execute_reply.started":"2022-07-06T01:20:31.301261Z","shell.execute_reply":"2022-07-06T01:20:31.345096Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Showing missing values \nMissing_Values = df.isnull().sum().sort_values(ascending=False)\nPercentage = (df.isnull().sum() / df.isnull().count()).sort_values(ascending=False)\nmissing_data = pd.concat([Missing_Values, Percentage], axis=1, keys=['Missing_Values', 'Percentage'])\nmissing_data.head(30)","metadata":{"execution":{"iopub.status.busy":"2022-07-06T01:20:31.347204Z","iopub.execute_input":"2022-07-06T01:20:31.347491Z","iopub.status.idle":"2022-07-06T01:20:31.409911Z","shell.execute_reply.started":"2022-07-06T01:20:31.347462Z","shell.execute_reply":"2022-07-06T01:20:31.408789Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Some features are mostly composed of missing values. 99.65% of 'PoolQC' is missing data. <br><br>\nThere's no point in maintaining these features in the dataset, so we can remove them.","metadata":{}},{"cell_type":"code","source":"# Dropping features with more than 5 missing values\ndf.drop((missing_data[missing_data['Missing_Values'] > 5]).index, axis = 1, inplace = True)","metadata":{"execution":{"iopub.status.busy":"2022-07-06T01:20:31.412201Z","iopub.execute_input":"2022-07-06T01:20:31.412536Z","iopub.status.idle":"2022-07-06T01:20:31.420867Z","shell.execute_reply.started":"2022-07-06T01:20:31.412507Z","shell.execute_reply":"2022-07-06T01:20:31.419915Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"df.isnull().sum().sort_values(ascending=False).head(15)","metadata":{"execution":{"iopub.status.busy":"2022-07-06T01:20:51.436220Z","iopub.execute_input":"2022-07-06T01:20:51.436608Z","iopub.status.idle":"2022-07-06T01:20:51.456939Z","shell.execute_reply.started":"2022-07-06T01:20:51.436571Z","shell.execute_reply":"2022-07-06T01:20:51.455676Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Filling in categorical data\ncat_col = ['Exterior1st',\n          'Exterior2nd',\n          'Functional',\n          'MSZoning',\n          'KitchenQual',\n          'Electrical',\n          'SaleType',\n          'Utilities']\nfor feature in cat_col:\n    df[feature] = df[feature].fillna(df[feature].mode()[0])","metadata":{"execution":{"iopub.status.busy":"2022-07-06T01:27:08.401110Z","iopub.execute_input":"2022-07-06T01:27:08.401499Z","iopub.status.idle":"2022-07-06T01:27:08.420476Z","shell.execute_reply.started":"2022-07-06T01:27:08.401464Z","shell.execute_reply":"2022-07-06T01:27:08.419485Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Filling in numeric data\nnum_col = ['BsmtFinSF1',\n          'BsmtFinSF2',\n          'BsmtUnfSF',\n          'BsmtFullBath',\n          'BsmtHalfBath',\n          'GarageArea',\n          'GarageCars',\n          'TotalBsmtSF']\n\nfor feature in num_col:\n    df[feature] = df[feature].fillna(0)","metadata":{"execution":{"iopub.status.busy":"2022-07-06T01:27:12.288660Z","iopub.execute_input":"2022-07-06T01:27:12.289098Z","iopub.status.idle":"2022-07-06T01:27:12.299041Z","shell.execute_reply.started":"2022-07-06T01:27:12.289056Z","shell.execute_reply":"2022-07-06T01:27:12.298078Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Fixing skewness in other numeric features\nnumeric_inputs = df.dtypes[df.dtypes != 'object'].index\nskewed = df[numeric_inputs].apply(lambda x: skew(x)).sort_values(ascending = False)\nprint(skewed[abs(skewed > 0.5)])","metadata":{"execution":{"iopub.status.busy":"2022-07-06T01:30:16.696134Z","iopub.execute_input":"2022-07-06T01:30:16.696602Z","iopub.status.idle":"2022-07-06T01:30:16.720839Z","shell.execute_reply.started":"2022-07-06T01:30:16.696566Z","shell.execute_reply":"2022-07-06T01:30:16.719267Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Normalizing \nhighskewed = skewed[abs(skewed) > 0.5]\n\nfor feature in highskewed.index:\n    df[feature] = np.log1p(df[feature])","metadata":{"execution":{"iopub.status.busy":"2022-07-06T01:32:36.081095Z","iopub.execute_input":"2022-07-06T01:32:36.081505Z","iopub.status.idle":"2022-07-06T01:32:36.102077Z","shell.execute_reply.started":"2022-07-06T01:32:36.081469Z","shell.execute_reply":"2022-07-06T01:32:36.100695Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Encoding Categoric Variables","metadata":{}},{"cell_type":"code","source":"df = pd.get_dummies(df)\ndf","metadata":{"execution":{"iopub.status.busy":"2022-07-06T01:33:21.900948Z","iopub.execute_input":"2022-07-06T01:33:21.901346Z","iopub.status.idle":"2022-07-06T01:33:21.970253Z","shell.execute_reply.started":"2022-07-06T01:33:21.901313Z","shell.execute_reply":"2022-07-06T01:33:21.969099Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Splitting data again\nX_train = df[:len(y_train)]\nX_test = df[len(y_train):]","metadata":{"execution":{"iopub.status.busy":"2022-07-06T01:35:03.011626Z","iopub.execute_input":"2022-07-06T01:35:03.012056Z","iopub.status.idle":"2022-07-06T01:35:03.018544Z","shell.execute_reply.started":"2022-07-06T01:35:03.012004Z","shell.execute_reply":"2022-07-06T01:35:03.017079Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Applying XGBRegressor for Predictions","metadata":{}},{"cell_type":"code","source":"XGBRegressor = XGB.XGBRegressor(colsample_bytree=0.4603, gamma=0.0468, \n                             learning_rate=0.05, max_depth=3, \n                             min_child_weight=1.7817, n_estimators=2200,\n                             reg_alpha=0.4640, reg_lambda=0.8571,\n                             subsample=0.5213, random_state =7, nthread = -1)\nXGBRegressor.fit(X_train, y_train)","metadata":{"execution":{"iopub.status.busy":"2022-07-06T01:45:46.001553Z","iopub.execute_input":"2022-07-06T01:45:46.001955Z","iopub.status.idle":"2022-07-06T01:45:58.032008Z","shell.execute_reply.started":"2022-07-06T01:45:46.001921Z","shell.execute_reply":"2022-07-06T01:45:58.030691Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"y_pred = np.floor(np.expm1(XGBRegressor.predict(X_test)))\ny_pred","metadata":{"execution":{"iopub.status.busy":"2022-07-06T01:46:30.583956Z","iopub.execute_input":"2022-07-06T01:46:30.584385Z","iopub.status.idle":"2022-07-06T01:46:30.619933Z","shell.execute_reply.started":"2022-07-06T01:46:30.584347Z","shell.execute_reply":"2022-07-06T01:46:30.618967Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"submission = pd.DataFrame()\nsubmission['Id'] = id_test\nsubmission['SalePrice'] = y_pred\nsubmission.to_csv('submission.csv', index = False)","metadata":{"execution":{"iopub.status.busy":"2022-07-06T01:48:25.899067Z","iopub.execute_input":"2022-07-06T01:48:25.899419Z","iopub.status.idle":"2022-07-06T01:48:25.914690Z","shell.execute_reply.started":"2022-07-06T01:48:25.899389Z","shell.execute_reply":"2022-07-06T01:48:25.913631Z"},"trusted":true},"execution_count":null,"outputs":[]}]}