{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"pygments_lexer":"ipython3","nbconvert_exporter":"python","version":"3.6.4","file_extension":".py","codemirror_mode":{"name":"ipython","version":3},"name":"python","mimetype":"text/x-python"}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"# **Housing Prices - Advanced Regression Techniques**\n\nGoal: Create a model that is able to predict the sales price of a given house, given other variables. In the model it must be able to predict the \"SalePrice\" variable value\n\n# **Overview**\n1. Understanding the shape of the data\n2. Data cleaning and exploration\n3. Feature engineering\n4. Data preprocessing for model\n5. Basic model building\n6. Model Tuning\n7. Ensemble model building\n8. Final Results","metadata":{}},{"cell_type":"code","source":"# !pip install -q pycaret","metadata":{"execution":{"iopub.status.busy":"2022-07-21T07:50:38.297930Z","iopub.execute_input":"2022-07-21T07:50:38.298236Z","iopub.status.idle":"2022-07-21T07:50:38.302269Z","shell.execute_reply.started":"2022-07-21T07:50:38.298203Z","shell.execute_reply":"2022-07-21T07:50:38.301407Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Importing libraries to start\n\nimport pandas as pd\npd.set_option(\"max_columns\",None)\npd.set_option(\"max_rows\",90)\nimport numpy as np\n\nimport matplotlib.pyplot as plt\nimport seaborn as sns\n\n# --------------------------------------------------------Collection libraries used later (for easy reference)---------------------------------------------------------\n\n# For data cleaning\n# from sklearn.neighbors import KNeighborsRegressor\n\n# For feature transformation\n# import scipy.stats\n\n# For Scaling\n# from sklearn.preprocessing import StandardScaler\n\n# For model building\n# from pycaret.regression import setup, compare_models\n\n# For model selection\n# from catboost import CatBoostRegressor\n# from sklearn.linear_model import BayesianRidge, HuberRegressor, Ridge, OrthogonalMatchingPursuit\n# from lightgbm import LGBMRegressor\n# from sklearn.ensemble import GradientBoostingRegressor\n# from xgboost import XGBRegressor\n# from sklearn.model_selection import KFold, cross_val_score\n\n# For Model Tuning\n# import optuna","metadata":{"execution":{"iopub.status.busy":"2022-07-21T07:50:38.303761Z","iopub.execute_input":"2022-07-21T07:50:38.303978Z","iopub.status.idle":"2022-07-21T07:50:38.315497Z","shell.execute_reply.started":"2022-07-21T07:50:38.303954Z","shell.execute_reply":"2022-07-21T07:50:38.314631Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"train0= pd.read_csv(\"../input/house-prices-advanced-regression-techniques/train.csv\")\ntest0= pd.read_csv(\"../input/house-prices-advanced-regression-techniques/test.csv\")\nsample_submission= pd.read_csv(\"../input/house-prices-advanced-regression-techniques/sample_submission.csv\")","metadata":{"execution":{"iopub.status.busy":"2022-07-21T07:50:38.316729Z","iopub.execute_input":"2022-07-21T07:50:38.316981Z","iopub.status.idle":"2022-07-21T07:50:38.366472Z","shell.execute_reply.started":"2022-07-21T07:50:38.316950Z","shell.execute_reply":"2022-07-21T07:50:38.365608Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Display training data\ntrain0","metadata":{"_kg_hide-output":false,"_kg_hide-input":false,"execution":{"iopub.status.busy":"2022-07-21T07:50:38.368073Z","iopub.execute_input":"2022-07-21T07:50:38.368703Z","iopub.status.idle":"2022-07-21T07:50:38.435106Z","shell.execute_reply.started":"2022-07-21T07:50:38.368653Z","shell.execute_reply":"2022-07-21T07:50:38.434398Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Display test data\ntest0","metadata":{"_kg_hide-output":true,"execution":{"iopub.status.busy":"2022-07-21T07:50:38.436260Z","iopub.execute_input":"2022-07-21T07:50:38.436489Z","iopub.status.idle":"2022-07-21T07:50:38.510604Z","shell.execute_reply.started":"2022-07-21T07:50:38.436460Z","shell.execute_reply":"2022-07-21T07:50:38.509509Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"#What the final result should look like\nsample_submission","metadata":{"execution":{"iopub.status.busy":"2022-07-21T07:50:38.512364Z","iopub.execute_input":"2022-07-21T07:50:38.512595Z","iopub.status.idle":"2022-07-21T07:50:38.524015Z","shell.execute_reply.started":"2022-07-21T07:50:38.512566Z","shell.execute_reply":"2022-07-21T07:50:38.523347Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# **1. Understanding the shape of the data**\n* .info to understand data types and null values\n* .describe to understand the numeric data, the central tendencies of the data","metadata":{}},{"cell_type":"code","source":"train0.info()\ntrain0.describe()","metadata":{"_kg_hide-input":false,"_kg_hide-output":true,"execution":{"iopub.status.busy":"2022-07-21T07:50:38.525144Z","iopub.execute_input":"2022-07-21T07:50:38.525502Z","iopub.status.idle":"2022-07-21T07:50:38.642347Z","shell.execute_reply.started":"2022-07-21T07:50:38.525463Z","shell.execute_reply":"2022-07-21T07:50:38.641596Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"test0.info()","metadata":{"_kg_hide-input":false,"_kg_hide-output":true,"execution":{"iopub.status.busy":"2022-07-21T07:50:38.643520Z","iopub.execute_input":"2022-07-21T07:50:38.643753Z","iopub.status.idle":"2022-07-21T07:50:38.662626Z","shell.execute_reply.started":"2022-07-21T07:50:38.643718Z","shell.execute_reply":"2022-07-21T07:50:38.661818Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# **2. Data Cleaning & Exploration**\n   1. Removing not-useful features\n   2. Locate missing values\n   2. Checking Data Types\n   3. Imputing\n        * Mode imputing (categoricals)\n        * Constant value imputing (categoricals)\n        * K-NN imputing (Numericals)\n   4. Feature Engineering (Later stage explored)","metadata":{}},{"cell_type":"code","source":"# Missing value in train set\ndisplay(train0.isna().sum())\n\n# Missing value in test set\ndisplay(test0.isna().sum())","metadata":{"_kg_hide-input":false,"_kg_hide-output":true,"execution":{"iopub.status.busy":"2022-07-21T07:50:38.664282Z","iopub.execute_input":"2022-07-21T07:50:38.664569Z","iopub.status.idle":"2022-07-21T07:50:38.684398Z","shell.execute_reply.started":"2022-07-21T07:50:38.664532Z","shell.execute_reply":"2022-07-21T07:50:38.683784Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Cleaning up the data (preliminary observations). Do not need Sales Price and Id for missing values imputing\ntarget = train0['SalePrice']  #Assigned to variables to add back later, aka y_train\ntest_ids = test0['Id']\n\ntrain1 = train0.drop(['Id','SalePrice'],axis=1)  # Drop Id column and Saleprice column\ntest1 = test0.drop(['Id'], axis=1)           # Only drop Id column because test data does not have Saleprice column (what we are suppose to figure out)\n\n\n\n# Observing the train and test data shows both sets have missing values, to impute missing values with the mean of said varaible would be more accurate to use the full data set\n# Therefore will be combining the train and test dataset, by concat\n# After preprocessing will split them up again, but not shuffling them, so no train test split.\n\ndata1 = pd.concat([train1,test1],axis=0).reset_index(drop=True)\ndisplay(data1)","metadata":{"execution":{"iopub.status.busy":"2022-07-21T07:50:38.685268Z","iopub.execute_input":"2022-07-21T07:50:38.685949Z","iopub.status.idle":"2022-07-21T07:50:38.782724Z","shell.execute_reply.started":"2022-07-21T07:50:38.685893Z","shell.execute_reply":"2022-07-21T07:50:38.780475Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## Ensure proper data types\n* Categorical or Numerical","metadata":{"execution":{"iopub.status.busy":"2022-02-21T13:37:00.241405Z","iopub.execute_input":"2022-02-21T13:37:00.241801Z","iopub.status.idle":"2022-02-21T13:37:00.248807Z","shell.execute_reply.started":"2022-02-21T13:37:00.241769Z","shell.execute_reply":"2022-02-21T13:37:00.247951Z"}}},{"cell_type":"code","source":"data1.select_dtypes(np.number)  #show all numerical data.\n# we want to look out for numbers that are actually categorical and set them to strings\n","metadata":{"execution":{"iopub.status.busy":"2022-07-21T07:50:38.785093Z","iopub.execute_input":"2022-07-21T07:50:38.785304Z","iopub.status.idle":"2022-07-21T07:50:38.822263Z","shell.execute_reply.started":"2022-07-21T07:50:38.785279Z","shell.execute_reply":"2022-07-21T07:50:38.821572Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"data2=data1.copy()\ndata2['MSSubClass'] = data2['MSSubClass'].astype(str)  #Converting to string as it is categorical","metadata":{"execution":{"iopub.status.busy":"2022-07-21T07:50:38.823283Z","iopub.execute_input":"2022-07-21T07:50:38.823596Z","iopub.status.idle":"2022-07-21T07:50:38.830568Z","shell.execute_reply.started":"2022-07-21T07:50:38.823571Z","shell.execute_reply":"2022-07-21T07:50:38.829939Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## Now we can impute missing data\n\n**Fill categorical missing values**\n- Go through each feature (variable) and decide how to impute (mean? mode? etc.)\nhttps://towardsdatascience.com/6-different-ways-to-compensate-for-missing-values-data-imputation-with-examples-6022d9ca0779\n- This is important because categorical variables, some NA have a specific meaning, whereas some do not (numericals). **Always have to do both**\n- In this example, **categoricals are replaced by constants and modes, numericals are replaced using KNN-imputer****","metadata":{}},{"cell_type":"code","source":"# Identify which columns have NA\n# simple can just manually use\n# data2.isna().sum()\n\n# more advance:\ndata2.select_dtypes('object').loc[:,data2.isna().sum()>0].columns","metadata":{"execution":{"iopub.status.busy":"2022-07-21T07:50:38.831533Z","iopub.execute_input":"2022-07-21T07:50:38.831844Z","iopub.status.idle":"2022-07-21T07:50:38.853773Z","shell.execute_reply.started":"2022-07-21T07:50:38.831819Z","shell.execute_reply":"2022-07-21T07:50:38.852777Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Impute using a constant value   (Categoricals where NA means None )\nfor column in ['Alley','BsmtQual','BsmtCond','BsmtExposure','BsmtFinType1','BsmtFinType2','FireplaceQu','GarageType','GarageFinish','GarageQual', \n               'GarageCond','PoolQC','Fence','MiscFeature']:\n    data2[column] = data2[column].fillna('None')\n\n# Impute using the column mode   \nfor column in ['MSZoning','Utilities','Exterior1st','Exterior2nd','MasVnrType','Electrical','KitchenQual','Functional','SaleType']:\n    data2[column] = data2[column].mode()[0]  #0 because it can return more than 1 value","metadata":{"execution":{"iopub.status.busy":"2022-07-21T07:50:38.855270Z","iopub.execute_input":"2022-07-21T07:50:38.855951Z","iopub.status.idle":"2022-07-21T07:50:38.877713Z","shell.execute_reply.started":"2022-07-21T07:50:38.855889Z","shell.execute_reply":"2022-07-21T07:50:38.877030Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"data2.isna().sum()  \ndata2.select_dtypes('object').isna().sum()  # it is all 0\n\n\ndisplay(data2.select_dtypes(np.number).isna().sum())","metadata":{"execution":{"iopub.status.busy":"2022-07-21T07:50:38.878831Z","iopub.execute_input":"2022-07-21T07:50:38.879419Z","iopub.status.idle":"2022-07-21T07:50:38.908790Z","shell.execute_reply.started":"2022-07-21T07:50:38.879377Z","shell.execute_reply":"2022-07-21T07:50:38.907946Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## KNN Imputing","metadata":{}},{"cell_type":"code","source":"data3 = data2.copy()\n\nfrom sklearn.neighbors import KNeighborsRegressor\n\n\ndef knn_impute(df, na_target):   \n    df = df.copy()\n    \n    numeric_df = df.select_dtypes(np.number)\n    non_na_columns = numeric_df.loc[:,numeric_df.isna().sum() == 0].columns\n    \n    y_train = numeric_df.loc[numeric_df[na_target].isna() == False,na_target]      \n    x_train = numeric_df.loc[numeric_df[na_target].isna() == False,non_na_columns] \n    x_test = numeric_df.loc[numeric_df[na_target].isna() == True,non_na_columns]  \n    \n    knn = KNeighborsRegressor()\n    knn.fit(x_train, y_train)          \n    \n    y_pred = knn.predict(x_test)       \n    \n    df.loc[numeric_df[na_target].isna() == True, na_target] = y_pred   \n    \n    return df","metadata":{"execution":{"iopub.status.busy":"2022-07-21T07:50:38.909784Z","iopub.execute_input":"2022-07-21T07:50:38.910100Z","iopub.status.idle":"2022-07-21T07:50:38.918101Z","shell.execute_reply.started":"2022-07-21T07:50:38.910070Z","shell.execute_reply":"2022-07-21T07:50:38.917282Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# To understand whats happening in y_train basically\ndata3.loc[data3['LotFrontage'].isna() == False,'LotFrontage']\n# locating in all the NA = False (basically mean where there is no NA) in \"LotFrontage\" column in data3","metadata":{"execution":{"iopub.status.busy":"2022-07-21T07:50:38.919241Z","iopub.execute_input":"2022-07-21T07:50:38.919834Z","iopub.status.idle":"2022-07-21T07:50:38.932884Z","shell.execute_reply.started":"2022-07-21T07:50:38.919791Z","shell.execute_reply":"2022-07-21T07:50:38.932235Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"for column in ['LotFrontage', 'MasVnrArea', 'BsmtFinSF1', 'BsmtFinSF2', 'BsmtUnfSF',\n       'TotalBsmtSF', 'BsmtFullBath', 'BsmtHalfBath', 'GarageYrBlt',\n       'GarageCars', 'GarageArea']:\n    \n    data3 = knn_impute(data3,column)","metadata":{"execution":{"iopub.status.busy":"2022-07-21T07:50:38.934024Z","iopub.execute_input":"2022-07-21T07:50:38.934571Z","iopub.status.idle":"2022-07-21T07:50:39.175519Z","shell.execute_reply.started":"2022-07-21T07:50:38.934538Z","shell.execute_reply":"2022-07-21T07:50:39.174651Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"data3.isna().sum()  # Checked, no more missing values.","metadata":{"execution":{"iopub.status.busy":"2022-07-21T07:50:39.176689Z","iopub.execute_input":"2022-07-21T07:50:39.177515Z","iopub.status.idle":"2022-07-21T07:50:39.192513Z","shell.execute_reply.started":"2022-07-21T07:50:39.177478Z","shell.execute_reply":"2022-07-21T07:50:39.191973Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"data4 = data3.copy()\ndata4\n# Fresh set to work on, data cleaned","metadata":{"execution":{"iopub.status.busy":"2022-07-21T07:50:39.193565Z","iopub.execute_input":"2022-07-21T07:50:39.194132Z","iopub.status.idle":"2022-07-21T07:50:39.269814Z","shell.execute_reply.started":"2022-07-21T07:50:39.194093Z","shell.execute_reply":"2022-07-21T07:50:39.268973Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## **Feature Engineering / Selection**\n\nFeature engineering is done based on insights obtained from data exploration, create new variables to increase model predictions. \n\nFeature selection is the removal of features that does not help with predictions, ie. closely correlated features (Correlation matrix). However for this example, since columns are <1000, selection of features is not neccesary","metadata":{}},{"cell_type":"code","source":"data4[\"SqFtPerRoom\"] = data4[\"GrLivArea\"] / (data4[\"TotRmsAbvGrd\"] +\n                                                       data4[\"FullBath\"] +\n                                                       data4[\"HalfBath\"] +\n                                                       data4[\"KitchenAbvGr\"])\n\ndata4[\"Total_Home_Quality\"] = data4[\"OverallQual\"] + data4[\"OverallCond\"]\n\ndata4[\"Total_Bathrooms\"] = (data4[\"FullBath\"] + (0.5 * data4[\"HalfBath\"])+ data4[\"BsmtFullBath\"]+ (0.5 * data4[\"BsmtHalfBath\"]))\n\ndata4[\"HighQualSF\"] = data4[\"1stFlrSF\"] + data4[\"2ndFlrSF\"] ","metadata":{"execution":{"iopub.status.busy":"2022-07-21T07:50:39.271040Z","iopub.execute_input":"2022-07-21T07:50:39.271267Z","iopub.status.idle":"2022-07-21T07:50:39.282386Z","shell.execute_reply.started":"2022-07-21T07:50:39.271239Z","shell.execute_reply":"2022-07-21T07:50:39.281695Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# **3. Feature Transformations**\n\nCertain models will perform better, when data is normal distributed. You cannot guarantee a normal distribution for a feature column. One thing that can be done is to look at a skew of a column, where data leans to one side (Data exploration step).\n\n1. Identifying features to be transformed\n2. Transforming features\n      * Log Transform\n      * Cosine Transform","metadata":{}},{"cell_type":"markdown","source":"## Identify features to be transformed","metadata":{}},{"cell_type":"code","source":"import scipy.stats  # library to check a skew of a numeric feature","metadata":{"execution":{"iopub.status.busy":"2022-07-21T07:50:39.283406Z","iopub.execute_input":"2022-07-21T07:50:39.283693Z","iopub.status.idle":"2022-07-21T07:50:39.293601Z","shell.execute_reply.started":"2022-07-21T07:50:39.283669Z","shell.execute_reply":"2022-07-21T07:50:39.293049Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"scipy.stats.skew(data4['LotFrontage'])\n\n# 0 means its not skewed. Positive value means it is right skewed. Negative value means it is left skewed.\n# Typically when the data is over 0.5, you can assume that the data is skewed, and may need of a transformation.","metadata":{"execution":{"iopub.status.busy":"2022-07-21T07:50:39.294595Z","iopub.execute_input":"2022-07-21T07:50:39.295059Z","iopub.status.idle":"2022-07-21T07:50:39.310472Z","shell.execute_reply.started":"2022-07-21T07:50:39.295018Z","shell.execute_reply":"2022-07-21T07:50:39.309588Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"\nskew_df = pd.DataFrame(data4.select_dtypes(np.number).columns, columns = ['Feature'])\nskew_df\n#-------------------------------------------------------------------------------------------------------------------------------------\n\nskew_df['Skew'] = skew_df['Feature'].apply(lambda feature: scipy.stats.skew(data4[feature])) \n\n# -------------------------------------------------------------------------------------------------------------------------------------------------------------------\n\nskew_df['Absolute Skew'] = skew_df['Skew'].apply(abs)  # to obtain magnitude of skew irrelevant of direction\nskew_df['Skewed']= skew_df['Absolute Skew'].apply(lambda x: True if x>= 0.5 else False)\nskew_df","metadata":{"execution":{"iopub.status.busy":"2022-07-21T07:50:39.312232Z","iopub.execute_input":"2022-07-21T07:50:39.312530Z","iopub.status.idle":"2022-07-21T07:50:39.345792Z","shell.execute_reply.started":"2022-07-21T07:50:39.312490Z","shell.execute_reply":"2022-07-21T07:50:39.344771Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"skew_df.query(\"Skewed == True\")[\"Feature\"]   # careful when using '' vs \"\", one is for strings only \"\"","metadata":{"execution":{"iopub.status.busy":"2022-07-21T07:50:39.346991Z","iopub.execute_input":"2022-07-21T07:50:39.347226Z","iopub.status.idle":"2022-07-21T07:50:39.357547Z","shell.execute_reply.started":"2022-07-21T07:50:39.347198Z","shell.execute_reply":"2022-07-21T07:50:39.356687Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## Transforming the identified features\n\nApply log tranform, help spread out very small value to a very nice range.\\\nBeware of 0 values as, log(0) = Undefined. In this case use log(x+1), in python function -> log1p aka log 1 plus.\n\n1. Log transform for skewed features\n2. Cosine transform for cyclical features","metadata":{}},{"cell_type":"code","source":"for column in skew_df.query(\"Skewed == True\")['Feature'].values:\n    data4[column] = np.log1p(data4[column])","metadata":{"execution":{"iopub.status.busy":"2022-07-21T07:50:39.361208Z","iopub.execute_input":"2022-07-21T07:50:39.361445Z","iopub.status.idle":"2022-07-21T07:50:39.384514Z","shell.execute_reply.started":"2022-07-21T07:50:39.361417Z","shell.execute_reply":"2022-07-21T07:50:39.383670Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Run again to check    \nskew_df['Skew'] = skew_df['Feature'].apply(lambda feature: scipy.stats.skew(data4[feature]))\nskew_df['Absolute Skew'] = skew_df['Skew'].apply(abs)  # to obtain magnitude of skew irrelevant of direction\nskew_df['Skewed']= skew_df['Absolute Skew'].apply(lambda x: True if x>= 0.5 else False)\nskew_df\n\n# Some still skewed but it is ok. Most of the time 1 log transform is ok","metadata":{"execution":{"iopub.status.busy":"2022-07-21T07:50:39.385964Z","iopub.execute_input":"2022-07-21T07:50:39.386262Z","iopub.status.idle":"2022-07-21T07:50:39.414316Z","shell.execute_reply.started":"2022-07-21T07:50:39.386223Z","shell.execute_reply":"2022-07-21T07:50:39.413633Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"The MoSold column will have to do its own feature engineering.\\\nSpecific feature engineering will require more indepth understanding of the dataset.\n\nFor the MoSold column, Month is a cylical feature, where 12 and 1 are closely link. However, to the model it will see it as 12 > 1.\\\nAs an example, in terms of temperature at different time of year, temperature during December and January are same, whereas middle of the cycle is the most different.\\\nThis complex relationship is not easy for the model to figure out. Would be good to implement a function for the model to understand that.\n\nFor this example we can use Sin & Cos, as this function forms waves. Where we can set point 0 and 12 on the x axis on the cos wave to obtain the same y value.\nSin will give 0 = 12, but 6 would be also on the same y value, which is not what we want.\nCos will give 0 = 12, where 6 would be on the opposite peak of the cos wave.","metadata":{}},{"cell_type":"code","source":"display(data4['MoSold'].unique())\n\nnp.cos(data4[\"MoSold\"]) #this gives the value without adjusted frequency, however we want to adjust so the stand and end of each wave ends 0 & 12 respectively\ndisplay(-np.cos(0.5236 * data4[\"MoSold\"]))\ndisplay(np.min(-np.cos(0.5236 * data4[\"MoSold\"])))  # should be very close to -1   (y value)\ndisplay(np.max(-np.cos(0.5236 * data4[\"MoSold\"])))  # should be very close to +1   (y value)\n\ndata4[\"MoSold\"]=(-np.cos(0.5236 * data4[\"MoSold\"])) # Applying to dataframe\ndata4[\"MoSold\"]   # Check that column has been applied","metadata":{"execution":{"iopub.status.busy":"2022-07-21T07:50:39.415255Z","iopub.execute_input":"2022-07-21T07:50:39.415610Z","iopub.status.idle":"2022-07-21T07:50:39.437608Z","shell.execute_reply.started":"2022-07-21T07:50:39.415577Z","shell.execute_reply":"2022-07-21T07:50:39.436852Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"data5 = data4.copy()\n# int_list = [\"LotArea\",\"OverallCond\",\"YearBuilt\",\"1stFlrSF\",\"2ndFlrSF\",\"LowQualFinSF\",\"GrLivArea\",\"HalfBath\",\n#           \"KitchenAbvGr\",\"TotRmsAbvGrd\",\"Fireplaces\",\"OpenPorchSF\",\"EnclosedPorch\",\"3SsnPorch\",\"ScreenPorch\",\"PoolArea\",\"MiscVal\"]\n\ndata5[\"LotArea\"] = pd.to_numeric(data5[\"LotArea\"], downcast=\"float\")\ndata5[\"OverallCond\"] = pd.to_numeric(data5[\"OverallCond\"], downcast=\"float\")\ndata5[\"YearBuilt\"] = pd.to_numeric(data5[\"YearBuilt\"], downcast=\"float\")\ndata5[\"1stFlrSF\"] = pd.to_numeric(data5[\"1stFlrSF\"], downcast=\"float\")\ndata5[\"2ndFlrSF\"] = pd.to_numeric(data5[\"2ndFlrSF\"], downcast=\"float\")\ndata5[\"LowQualFinSF\"] = pd.to_numeric(data5[\"LowQualFinSF\"], downcast=\"float\")\ndata5[\"GrLivArea\"] = pd.to_numeric(data5[\"GrLivArea\"], downcast=\"float\")\ndata5[\"HalfBath\"] = pd.to_numeric(data5[\"HalfBath\"], downcast=\"float\")\ndata5[\"KitchenAbvGr\"] = pd.to_numeric(data5[\"KitchenAbvGr\"], downcast=\"float\")\ndata5[\"TotRmsAbvGrd\"] = pd.to_numeric(data5[\"TotRmsAbvGrd\"], downcast=\"float\")\ndata5[\"Fireplaces\"] = pd.to_numeric(data5[\"Fireplaces\"], downcast=\"float\")\ndata5[\"WoodDeckSF\"] = pd.to_numeric(data5[\"WoodDeckSF\"], downcast=\"float\")\ndata5[\"OpenPorchSF\"] = pd.to_numeric(data5[\"OpenPorchSF\"], downcast=\"float\")\ndata5[\"EnclosedPorch\"] = pd.to_numeric(data5[\"EnclosedPorch\"], downcast=\"float\")\ndata5[\"3SsnPorch\"] = pd.to_numeric(data5[\"3SsnPorch\"], downcast=\"float\")\ndata5[\"ScreenPorch\"] = pd.to_numeric(data5[\"ScreenPorch\"], downcast=\"float\")\ndata5[\"PoolArea\"] = pd.to_numeric(data5[\"PoolArea\"], downcast=\"float\")\ndata5[\"MiscVal\"] = pd.to_numeric(data5[\"MiscVal\"], downcast=\"float\")\ndata5\n# New dataframe to work with that has been feature engineered","metadata":{"execution":{"iopub.status.busy":"2022-07-21T07:50:39.438669Z","iopub.execute_input":"2022-07-21T07:50:39.438868Z","iopub.status.idle":"2022-07-21T07:50:39.540384Z","shell.execute_reply.started":"2022-07-21T07:50:39.438844Z","shell.execute_reply":"2022-07-21T07:50:39.539400Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# **4. Data Preprocessing**\n\n1. Encoding\n2. Scaling\n3. Target Transformation","metadata":{}},{"cell_type":"markdown","source":"##  Encoding\n\nCategoricals","metadata":{}},{"cell_type":"code","source":"data5 = pd.get_dummies(data5)\n# Automatically encodes all categoricals\ndata5 ","metadata":{"execution":{"iopub.status.busy":"2022-07-21T07:50:39.542097Z","iopub.execute_input":"2022-07-21T07:50:39.542406Z","iopub.status.idle":"2022-07-21T07:50:39.715810Z","shell.execute_reply.started":"2022-07-21T07:50:39.542365Z","shell.execute_reply":"2022-07-21T07:50:39.714753Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"data6 = data5.copy()\n# data6 is encoded version","metadata":{"execution":{"iopub.status.busy":"2022-07-21T07:50:39.717183Z","iopub.execute_input":"2022-07-21T07:50:39.717456Z","iopub.status.idle":"2022-07-21T07:50:39.725689Z","shell.execute_reply.started":"2022-07-21T07:50:39.717422Z","shell.execute_reply":"2022-07-21T07:50:39.724695Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## Scaling","metadata":{}},{"cell_type":"code","source":"from sklearn.preprocessing import StandardScaler\n# MinMax scaler vs standard scaler depends on the dataset. Can try out both to see which gives a better result","metadata":{"execution":{"iopub.status.busy":"2022-07-21T07:50:39.727290Z","iopub.execute_input":"2022-07-21T07:50:39.727506Z","iopub.status.idle":"2022-07-21T07:50:39.734391Z","shell.execute_reply.started":"2022-07-21T07:50:39.727480Z","shell.execute_reply":"2022-07-21T07:50:39.733691Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"scaler = StandardScaler()\nscaler.fit(data6)\n\ndata6 = pd.DataFrame(scaler.transform(data6), index=data6.index, columns=data6.columns)\n# transform returns a numpy array, therefore need to turn it back into Dataframe\ndata6","metadata":{"execution":{"iopub.status.busy":"2022-07-21T07:50:39.735635Z","iopub.execute_input":"2022-07-21T07:50:39.736533Z","iopub.status.idle":"2022-07-21T07:50:39.983564Z","shell.execute_reply.started":"2022-07-21T07:50:39.736489Z","shell.execute_reply":"2022-07-21T07:50:39.982689Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"data7 = data6.copy()    # Scaled Data","metadata":{"execution":{"iopub.status.busy":"2022-07-21T07:50:39.984669Z","iopub.execute_input":"2022-07-21T07:50:39.984981Z","iopub.status.idle":"2022-07-21T07:50:39.990149Z","shell.execute_reply.started":"2022-07-21T07:50:39.984950Z","shell.execute_reply":"2022-07-21T07:50:39.989263Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## Target Transformation\n\nThis is done separately from feature transfromation. Have to be very careful during a target transformation, as you are changing the unit where the model is making predictions. Any changes done to it must be converted back to the original","metadata":{}},{"cell_type":"code","source":"target  # REMINDER, \"TARGET\" IS AKA y_train","metadata":{"execution":{"iopub.status.busy":"2022-07-21T07:50:39.991282Z","iopub.execute_input":"2022-07-21T07:50:39.991501Z","iopub.status.idle":"2022-07-21T07:50:40.004618Z","shell.execute_reply.started":"2022-07-21T07:50:39.991474Z","shell.execute_reply":"2022-07-21T07:50:40.004002Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"plt.figure(figsize=(20,10))  # figsize (width,height)    - its in inches\n\nplt.subplot(1,2,1)          # 1 row, 2 column, position 1\nsns.distplot(target, kde=True, fit=scipy.stats.norm)\nplt.title(\"Without Log Transform\")\n\nplt.subplot(1,2,2)\nsns.distplot(np.log(target), kde=True, fit=scipy.stats.norm)\nplt.title(\"With Log Tranform\")\nplt.xlabel(\"Log SalePrice\")\n\nplt.show()","metadata":{"execution":{"iopub.status.busy":"2022-07-21T07:50:40.005737Z","iopub.execute_input":"2022-07-21T07:50:40.006282Z","iopub.status.idle":"2022-07-21T07:50:40.511035Z","shell.execute_reply.started":"2022-07-21T07:50:40.006231Z","shell.execute_reply":"2022-07-21T07:50:40.510183Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## Split the Train and test set\n\nNow that we have done the Pre-processing, time to split the data for model building","metadata":{}},{"cell_type":"code","source":"train_final = data7.loc[:train0.index.max(),:].copy()  # Inside data7(concated), loc from 0(start:) to train0-df max value (1460), all columns. Assign it to train_final\ntest_final = data7.loc[train0.index.max()+1:,:].reset_index(drop=True).copy() # Inside data7(concated), loc from train0-df max value +1 (1460+1) to the end(:), \n                                                                              #all column. Assign it to test final","metadata":{"execution":{"iopub.status.busy":"2022-07-21T07:50:40.512536Z","iopub.execute_input":"2022-07-21T07:50:40.512979Z","iopub.status.idle":"2022-07-21T07:50:40.519611Z","shell.execute_reply.started":"2022-07-21T07:50:40.512940Z","shell.execute_reply":"2022-07-21T07:50:40.519001Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"display(train_final)   # Rows matches train0-df\ndisplay(test_final)    # Row matches test0-df","metadata":{"execution":{"iopub.status.busy":"2022-07-21T07:50:40.520619Z","iopub.execute_input":"2022-07-21T07:50:40.521006Z","iopub.status.idle":"2022-07-21T07:50:40.954431Z","shell.execute_reply.started":"2022-07-21T07:50:40.520965Z","shell.execute_reply":"2022-07-21T07:50:40.953496Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# **5. Model Building**\n\n1. Compare Models\n2. Select Models\n3. Run chosen model\n4. Bagging ensemble","metadata":{}},{"cell_type":"markdown","source":"## Model Selection","metadata":{}},{"cell_type":"code","source":"# from pycaret.regression import setup, compare_models\n# Installation issues prevents notebook from saving properly (Still resolving)","metadata":{"execution":{"iopub.status.busy":"2022-07-21T07:50:40.955902Z","iopub.execute_input":"2022-07-21T07:50:40.956154Z","iopub.status.idle":"2022-07-21T07:50:40.959693Z","shell.execute_reply.started":"2022-07-21T07:50:40.956126Z","shell.execute_reply":"2022-07-21T07:50:40.958823Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"log_target=np.log(target)","metadata":{"execution":{"iopub.status.busy":"2022-07-21T07:50:40.960794Z","iopub.execute_input":"2022-07-21T07:50:40.961017Z","iopub.status.idle":"2022-07-21T07:50:40.973010Z","shell.execute_reply.started":"2022-07-21T07:50:40.960989Z","shell.execute_reply":"2022-07-21T07:50:40.972279Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# _ = setup(data=pd.concat([train_final, log_target], axis=1), target='SalePrice')     # i have to log the target because my features were normalised using log1p function","metadata":{"execution":{"iopub.status.busy":"2022-07-21T07:50:40.974067Z","iopub.execute_input":"2022-07-21T07:50:40.974482Z","iopub.status.idle":"2022-07-21T07:50:40.984278Z","shell.execute_reply.started":"2022-07-21T07:50:40.974448Z","shell.execute_reply":"2022-07-21T07:50:40.983373Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# compare_models()","metadata":{"execution":{"iopub.status.busy":"2022-07-21T07:50:40.985428Z","iopub.execute_input":"2022-07-21T07:50:40.985661Z","iopub.status.idle":"2022-07-21T07:50:40.996164Z","shell.execute_reply.started":"2022-07-21T07:50:40.985634Z","shell.execute_reply":"2022-07-21T07:50:40.995236Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## Select Models to use\n\nFrom pycaret model evaluation, top performing models were\n* catboost\n* bayesian ridge\n* huber\n* ridge\n* orthoganal matching pursuit\n* light gbm\n* gbr\n* xgboost","metadata":{}},{"cell_type":"markdown","source":"## Run chosen models\n\nRun models highlighted in cell above","metadata":{}},{"cell_type":"code","source":"# Import all libraries for models\nfrom catboost import CatBoostRegressor\nfrom sklearn.linear_model import BayesianRidge, HuberRegressor, Ridge, OrthogonalMatchingPursuit\nfrom lightgbm import LGBMRegressor\nfrom sklearn.ensemble import GradientBoostingRegressor\nfrom xgboost import XGBRegressor\n\n# To validate models, but because train_test split will randomise the split, causing different results each time\n# Will use cross validation\nfrom sklearn.model_selection import KFold, cross_val_score","metadata":{"execution":{"iopub.status.busy":"2022-07-21T07:50:40.997338Z","iopub.execute_input":"2022-07-21T07:50:40.997527Z","iopub.status.idle":"2022-07-21T07:50:41.008341Z","shell.execute_reply.started":"2022-07-21T07:50:40.997504Z","shell.execute_reply":"2022-07-21T07:50:41.007375Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## *Catboost Model (Baseline Model)*\n\n1. fit\n2. evaluate","metadata":{}},{"cell_type":"code","source":"baseline_model = CatBoostRegressor(verbose=0)","metadata":{"execution":{"iopub.status.busy":"2022-07-21T07:50:41.009696Z","iopub.execute_input":"2022-07-21T07:50:41.010446Z","iopub.status.idle":"2022-07-21T07:50:41.027889Z","shell.execute_reply.started":"2022-07-21T07:50:41.010409Z","shell.execute_reply":"2022-07-21T07:50:41.027281Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"baseline_model.fit(train_final, log_target)","metadata":{"execution":{"iopub.status.busy":"2022-07-21T07:50:41.029006Z","iopub.execute_input":"2022-07-21T07:50:41.029228Z","iopub.status.idle":"2022-07-21T07:50:44.919672Z","shell.execute_reply.started":"2022-07-21T07:50:41.029195Z","shell.execute_reply":"2022-07-21T07:50:44.918855Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"**Evaluate**","metadata":{}},{"cell_type":"code","source":"kf = KFold(n_splits=10)\n# KFold is use in cross validation to split data into sections to train and test, always be the same everytime you run. So score generated never changes unless K changes\n\nresults = cross_val_score(baseline_model, train_final, log_target, scoring='neg_mean_squared_error', cv=kf)\n# cross_val_score(estimator, training, target, scoring method, crossvalidation)","metadata":{"execution":{"iopub.status.busy":"2022-07-21T07:50:44.921031Z","iopub.execute_input":"2022-07-21T07:50:44.921304Z","iopub.status.idle":"2022-07-21T07:51:23.970515Z","shell.execute_reply.started":"2022-07-21T07:50:44.921274Z","shell.execute_reply":"2022-07-21T07:51:23.969570Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"plt.figure(figsize=(16,10))\n\nsns.displot(-results, bins=10 ,kde=True)","metadata":{"execution":{"iopub.status.busy":"2022-07-21T07:51:23.971615Z","iopub.execute_input":"2022-07-21T07:51:23.971841Z","iopub.status.idle":"2022-07-21T07:51:24.290713Z","shell.execute_reply.started":"2022-07-21T07:51:23.971813Z","shell.execute_reply":"2022-07-21T07:51:24.289657Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"print(np.mean(-results)) #mean of the mean squared error from across all folds\n\n# We have to convert this to a metric that has meaning. Convert this (Root it to get the error, then apply exponential since its log to get the actual error value)\nnp.exp(np.sqrt(np.mean(-results)))","metadata":{"execution":{"iopub.status.busy":"2022-07-21T07:51:24.292067Z","iopub.execute_input":"2022-07-21T07:51:24.292289Z","iopub.status.idle":"2022-07-21T07:51:24.299425Z","shell.execute_reply.started":"2022-07-21T07:51:24.292258Z","shell.execute_reply":"2022-07-21T07:51:24.298527Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"print(target.min())\nprint(target.max())","metadata":{"execution":{"iopub.status.busy":"2022-07-21T07:51:24.300765Z","iopub.execute_input":"2022-07-21T07:51:24.301218Z","iopub.status.idle":"2022-07-21T07:51:24.312999Z","shell.execute_reply.started":"2022-07-21T07:51:24.301189Z","shell.execute_reply":"2022-07-21T07:51:24.311936Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Error in Salesprice unit is 1.13 of 34900 or 755000, which is extremely accurate","metadata":{}},{"cell_type":"code","source":"# Predict SalesPrice in test set\nfinal_predictions = np.exp(baseline_model.predict(test_final))  # test_final is in log unit, requires exp","metadata":{"execution":{"iopub.status.busy":"2022-07-21T07:51:24.314721Z","iopub.execute_input":"2022-07-21T07:51:24.314961Z","iopub.status.idle":"2022-07-21T07:51:24.342192Z","shell.execute_reply.started":"2022-07-21T07:51:24.314935Z","shell.execute_reply":"2022-07-21T07:51:24.341228Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## **Make subsmission**\n\nconvert to relevant format","metadata":{}},{"cell_type":"code","source":"submission = pd.concat([test_ids, pd.Series(final_predictions, name=\"SalePrice\")], axis=1)\nsubmission","metadata":{"execution":{"iopub.status.busy":"2022-07-21T07:51:24.343949Z","iopub.execute_input":"2022-07-21T07:51:24.344467Z","iopub.status.idle":"2022-07-21T07:51:24.359456Z","shell.execute_reply.started":"2022-07-21T07:51:24.344420Z","shell.execute_reply":"2022-07-21T07:51:24.358566Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"submission.to_csv('./submission.csv', index=False, header= True)","metadata":{"execution":{"iopub.status.busy":"2022-07-21T07:51:24.360614Z","iopub.execute_input":"2022-07-21T07:51:24.360843Z","iopub.status.idle":"2022-07-21T07:51:24.372654Z","shell.execute_reply.started":"2022-07-21T07:51:24.360818Z","shell.execute_reply":"2022-07-21T07:51:24.371470Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# **Model Tuning**\n\n* Hyperparameter Optimisation w. Optuna (hyperparameters are parameters to control the learning process of the models ie alpha, lambda)","metadata":{}},{"cell_type":"code","source":"import optuna","metadata":{"execution":{"iopub.status.busy":"2022-07-21T07:51:24.374324Z","iopub.execute_input":"2022-07-21T07:51:24.374619Z","iopub.status.idle":"2022-07-21T07:51:24.865655Z","shell.execute_reply.started":"2022-07-21T07:51:24.374580Z","shell.execute_reply":"2022-07-21T07:51:24.864797Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# objective function is measure of how good the model's predictions are\n# Using Bayesian Ridge Regressor as an example\ndef br_objective(trial):\n    n_iter = trial.suggest_int('n_iter',50,600)\n    tol = trial.suggest_loguniform('tol',1e-8,10.0)\n    alpha_1 = trial.suggest_loguniform('alpha_1',1e-8,10.0)\n    alpha_2 = trial.suggest_loguniform('alpha_2',1e-8,10.0)\n    lambda_1 = trial.suggest_loguniform('lambda_1',1e-8,10.0)\n    lambda_2 = trial.suggest_loguniform('lamba_2',1e-8,10.0)\n    \n    model = BayesianRidge(\n        n_iter=n_iter,\n        tol=tol,\n        alpha_1=alpha_1,\n        alpha_2=alpha_2,\n        lambda_1=lambda_1,\n        lambda_2=lambda_2\n    )\n    \n    model.fit(train_final, log_target)\n    \n    cv_score = np.exp(np.sqrt(-cross_val_score(model, train_final, log_target, scoring='neg_mean_squared_error', cv=kf)))\n\n    return np.mean(cv_score)","metadata":{"execution":{"iopub.status.busy":"2022-07-21T07:51:24.868109Z","iopub.execute_input":"2022-07-21T07:51:24.868429Z","iopub.status.idle":"2022-07-21T07:51:24.876181Z","shell.execute_reply.started":"2022-07-21T07:51:24.868388Z","shell.execute_reply":"2022-07-21T07:51:24.874824Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"study = optuna.create_study(direction='minimize')\nstudy.optimize(br_objective, n_trials=100)","metadata":{"execution":{"iopub.status.busy":"2022-07-21T07:51:24.877682Z","iopub.execute_input":"2022-07-21T07:51:24.878148Z","iopub.status.idle":"2022-07-21T07:54:10.580823Z","shell.execute_reply.started":"2022-07-21T07:51:24.878102Z","shell.execute_reply":"2022-07-21T07:54:10.579831Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Based on the hyperparameter optimisation, out of the 100 trials, trial 38 had the hyperparameters which are the most accurate.\\\nTherefore will use trial 38 hyperparameters\n\n*Bayesian Ridge Regressor Hyperparameters*\\\nn_iter: 513\\\ntol: 1.6164198543424577e-05\\\nalpha_1: 1.1313430506577451e-07\\\nalpha_2: 9.89435956159145\\\nlambda_1: 0.023813710900796637\\\nlambda_2: 2.0965696650559703e-07","metadata":{}},{"cell_type":"markdown","source":"**For implementation not done here**\n\nApply optuna to all models used in ensembling to get best hyperparameters settings. Apply hyperparameters when doing ensembling w weights, for optimisation.","metadata":{}},{"cell_type":"markdown","source":"# **7. Ensembling**\nTechniques that create multiple models and then combine them to produce improved results","metadata":{}},{"cell_type":"markdown","source":"## **Bagging Ensemble**\n\nBagging ensemble is running all models and averaging all the predicted outputs, to obtain average predicted output.","metadata":{}},{"cell_type":"code","source":"models = {\n    \"catboost\": CatBoostRegressor(verbose=0),\n    \"br\": BayesianRidge(),\n#     \"huber\": HuberRegressor(),\n    \"lightgbm\":LGBMRegressor(),\n    \"ridge\": Ridge(),\n    \"omp\": OrthogonalMatchingPursuit()\n}","metadata":{"execution":{"iopub.status.busy":"2022-07-21T07:54:10.586523Z","iopub.execute_input":"2022-07-21T07:54:10.590354Z","iopub.status.idle":"2022-07-21T07:54:10.601570Z","shell.execute_reply.started":"2022-07-21T07:54:10.590288Z","shell.execute_reply":"2022-07-21T07:54:10.600716Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"for name, model in models.items():\n    model.fit(train_final, log_target)   # Reminder that target is y_train\n    print(name + \" trained.\")","metadata":{"execution":{"iopub.status.busy":"2022-07-21T07:54:10.606876Z","iopub.execute_input":"2022-07-21T07:54:10.610362Z","iopub.status.idle":"2022-07-21T07:54:15.508287Z","shell.execute_reply.started":"2022-07-21T07:54:10.610297Z","shell.execute_reply":"2022-07-21T07:54:15.505824Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"**Evaluate** bagging ensemble","metadata":{}},{"cell_type":"code","source":"results = {}\n\nkf = KFold(n_splits=10)\n\nfor name,model in models.items():\n    result = np.exp(np.sqrt(-cross_val_score(model, train_final, log_target, scoring=\"neg_mean_squared_error\", cv=kf)))\n    results[name] = result","metadata":{"execution":{"iopub.status.busy":"2022-07-21T07:54:15.510107Z","iopub.execute_input":"2022-07-21T07:54:15.510484Z","iopub.status.idle":"2022-07-21T07:55:00.297810Z","shell.execute_reply.started":"2022-07-21T07:54:15.510438Z","shell.execute_reply":"2022-07-21T07:55:00.296882Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"for name, result in results.items():\n    print(\"---------\\n\" + name + \"\\n-----------\")\n    print(np.mean(result))\n    print(np.std(result))","metadata":{"execution":{"iopub.status.busy":"2022-07-21T07:55:00.299866Z","iopub.execute_input":"2022-07-21T07:55:00.300538Z","iopub.status.idle":"2022-07-21T07:55:00.311239Z","shell.execute_reply.started":"2022-07-21T07:55:00.300489Z","shell.execute_reply":"2022-07-21T07:55:00.310339Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"from the output, huber seems to do very poorly. Replace it with LGBM.","metadata":{}},{"cell_type":"markdown","source":"## **Combine Ensemble**","metadata":{}},{"cell_type":"code","source":"# Obtaining all the predicted values and averaging the results among all 3  (Weighted equally)\n\n# final_predictions = (\n#     0.2 * np.exp(models['catboost'].predict(test_final)) +\n#     0.2 * np.exp(models['br'].predict(test_final)) +\n#     0.2 * np.exp(models['lightgbm'].predict(test_final)) +\n#     0.2 * np.exp(models['ridge'].predict(test_final)) +\n#     0.2 * np.exp(models['omp'].predict(test_final))\n# )\n#-------------------------------------------------------------------------------------------------------------------------------------------------------------------\n\n# Obtaining all the predicted values and averaging the results among all 3 (Weighted based on accuracy)\n\nfinal_predictions = (\n    0.4 * np.exp(models['catboost'].predict(test_final)) +\n    0.2 * np.exp(models['br'].predict(test_final)) +\n    0.2 * np.exp(models['lightgbm'].predict(test_final)) +\n    0.1 * np.exp(models['ridge'].predict(test_final)) +\n    0.1 * np.exp(models['omp'].predict(test_final))\n)","metadata":{"execution":{"iopub.status.busy":"2022-07-21T07:55:00.313220Z","iopub.execute_input":"2022-07-21T07:55:00.314055Z","iopub.status.idle":"2022-07-21T07:55:00.415004Z","shell.execute_reply.started":"2022-07-21T07:55:00.314005Z","shell.execute_reply":"2022-07-21T07:55:00.414007Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"final_predictions","metadata":{"execution":{"iopub.status.busy":"2022-07-21T07:55:00.417016Z","iopub.execute_input":"2022-07-21T07:55:00.417678Z","iopub.status.idle":"2022-07-21T07:55:00.425747Z","shell.execute_reply.started":"2022-07-21T07:55:00.417632Z","shell.execute_reply":"2022-07-21T07:55:00.424563Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## **8. Final Submission**","metadata":{}},{"cell_type":"code","source":"submission.to_csv('./submission.csv', index=False, header= True)\n\n#Previous scores:\n#000 : 0.12562\n#003 : 0.12377","metadata":{"execution":{"iopub.status.busy":"2022-07-21T07:55:00.427395Z","iopub.execute_input":"2022-07-21T07:55:00.427734Z","iopub.status.idle":"2022-07-21T07:55:00.445908Z","shell.execute_reply.started":"2022-07-21T07:55:00.427693Z","shell.execute_reply":"2022-07-21T07:55:00.444969Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## **Review/Thoughts**","metadata":{}},{"cell_type":"markdown","source":"There is some sort of a structure creating ML models, however it is not extremely fixed.\n\nTypically start with some data cleaning and exploration, followed by feature and target transformation.\\\nWith that create a baseline model.\n\nThink of ways to improve the models by adding components at different points of the code be it:\n* going back to add in feature engineering or selections\n* continue to do ensembling on baseline model to improve accuracy\n* adding weights to the ensembled model\n* optimising hyperparameters\n\nIdeally to improve models, do each implementation a step at a time, review results to decide if it benefits model or not. If it does, keep it, if it doesnt remove it. Record the results as one goes.\n\n(Check out optuna and how to hyperparameterise other model types)","metadata":{}},{"cell_type":"markdown","source":"## **References**\n\nGabriel Atkin - https://www.kaggle.com/code/gcdatkin/top-10-house-price-regression-competition-nb/notebook\n","metadata":{}},{"cell_type":"code","source":"","metadata":{},"execution_count":null,"outputs":[]}]}