{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"pygments_lexer":"ipython3","nbconvert_exporter":"python","version":"3.6.4","file_extension":".py","codemirror_mode":{"name":"ipython","version":3},"name":"python","mimetype":"text/x-python"}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"code","source":"# This Python 3 environment comes with many helpful analytics libraries installed\n# It is defined by the kaggle/python Docker image: https://github.com/kaggle/docker-python\n# For example, here's several helpful packages to load\n\nimport numpy as np # linear algebra\nimport pandas as pd # data processing, CSV file I/O (e.g. pd.read_csv)\n\n# Input data files are available in the read-only \"../input/\" directory\n# For example, running this (by clicking run or pressing Shift+Enter) will list all files under the input directory\n\nimport os\nfor dirname, _, filenames in os.walk('/kaggle/input'):\n    for filename in filenames:\n        print(os.path.join(dirname, filename))\n\n# You can write up to 20GB to the current directory (/kaggle/working/) that gets preserved as output when you create a version using \"Save & Run All\" \n# You can also write temporary files to /kaggle/temp/, but they won't be saved outside of the current session","metadata":{"_uuid":"8f2839f25d086af736a60e9eeb907d3b93b6e0e5","_cell_guid":"b1076dfc-b9ad-4769-8c92-a6c4dae69d19","execution":{"iopub.status.busy":"2022-07-14T11:20:18.459143Z","iopub.execute_input":"2022-07-14T11:20:18.459707Z","iopub.status.idle":"2022-07-14T11:20:18.468281Z","shell.execute_reply.started":"2022-07-14T11:20:18.459671Z","shell.execute_reply":"2022-07-14T11:20:18.467564Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Let's load the data: \n\ntrain_file_path = \"/kaggle/input/house-prices-advanced-regression-techniques/train.csv\"\ntrain = pd.read_csv(train_file_path)","metadata":{"execution":{"iopub.status.busy":"2022-07-14T11:20:18.482539Z","iopub.execute_input":"2022-07-14T11:20:18.482924Z","iopub.status.idle":"2022-07-14T11:20:18.508504Z","shell.execute_reply.started":"2022-07-14T11:20:18.482888Z","shell.execute_reply":"2022-07-14T11:20:18.507630Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# First look at the data\n\nFirst we'll have a brief look at the data and some basic properties...","metadata":{}},{"cell_type":"code","source":"train.head()","metadata":{"execution":{"iopub.status.busy":"2022-07-14T11:20:18.510004Z","iopub.execute_input":"2022-07-14T11:20:18.510721Z","iopub.status.idle":"2022-07-14T11:20:18.537464Z","shell.execute_reply.started":"2022-07-14T11:20:18.510684Z","shell.execute_reply":"2022-07-14T11:20:18.536585Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"train.describe()","metadata":{"execution":{"iopub.status.busy":"2022-07-14T11:20:18.539295Z","iopub.execute_input":"2022-07-14T11:20:18.540152Z","iopub.status.idle":"2022-07-14T11:20:18.633488Z","shell.execute_reply.started":"2022-07-14T11:20:18.540109Z","shell.execute_reply":"2022-07-14T11:20:18.632576Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"train.shape","metadata":{"execution":{"iopub.status.busy":"2022-07-14T11:20:18.635362Z","iopub.execute_input":"2022-07-14T11:20:18.636149Z","iopub.status.idle":"2022-07-14T11:20:18.643496Z","shell.execute_reply.started":"2022-07-14T11:20:18.636103Z","shell.execute_reply":"2022-07-14T11:20:18.642554Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"train.columns","metadata":{"execution":{"iopub.status.busy":"2022-07-14T11:20:18.645351Z","iopub.execute_input":"2022-07-14T11:20:18.646147Z","iopub.status.idle":"2022-07-14T11:20:18.656661Z","shell.execute_reply.started":"2022-07-14T11:20:18.646054Z","shell.execute_reply":"2022-07-14T11:20:18.655445Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"As we can see, there are lots of features and there is probably some dimensionality reduction that can be done. We will look at this when we come to the feature engineering. ","metadata":{}},{"cell_type":"code","source":"%matplotlib inline\nimport matplotlib.pyplot as plt\ntrain.hist(bins=50, figsize=(20,15))\nplt.show()","metadata":{"execution":{"iopub.status.busy":"2022-07-14T13:58:33.625285Z","iopub.execute_input":"2022-07-14T13:58:33.625744Z","iopub.status.idle":"2022-07-14T13:58:41.302036Z","shell.execute_reply.started":"2022-07-14T13:58:33.625708Z","shell.execute_reply":"2022-07-14T13:58:41.301202Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Dealing with missing data\n\nWe start by looking at the missing values from the data and how we can deal with this. As we will see, we will deal with it mostly by imputing the missing values in some way, such as with the mean of the relevant column. ","metadata":{}},{"cell_type":"code","source":"missing_values_count = train.isnull().sum()\n\nfor i in range(0, len(missing_values_count)):\n    if missing_values_count[i]!=0:\n        print(missing_values_count.index[i], missing_values_count[i])","metadata":{"execution":{"iopub.status.busy":"2022-07-14T11:20:18.658123Z","iopub.execute_input":"2022-07-14T11:20:18.659919Z","iopub.status.idle":"2022-07-14T11:20:18.680235Z","shell.execute_reply.started":"2022-07-14T11:20:18.659881Z","shell.execute_reply":"2022-07-14T11:20:18.679452Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# What proportion of cells are missing? \n\ntotal_cells = np.product(train.shape)\ntotal_missing = missing_values_count.sum()\npercentage_missing = 100*total_missing/total_cells\nprint(percentage_missing)","metadata":{"execution":{"iopub.status.busy":"2022-07-14T11:20:18.681877Z","iopub.execute_input":"2022-07-14T11:20:18.682112Z","iopub.status.idle":"2022-07-14T11:20:18.687166Z","shell.execute_reply.started":"2022-07-14T11:20:18.682080Z","shell.execute_reply":"2022-07-14T11:20:18.686510Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# First we'll remove the columns with almost all data missing as any imputation strategy is likely to be inaccurate anyway...\n\ntrain2 = train.drop([\"Alley\", \"PoolQC\", \"Fence\", \"MiscFeature\"], axis=1)","metadata":{"execution":{"iopub.status.busy":"2022-07-14T11:20:18.688547Z","iopub.execute_input":"2022-07-14T11:20:18.688941Z","iopub.status.idle":"2022-07-14T11:20:18.697242Z","shell.execute_reply.started":"2022-07-14T11:20:18.688905Z","shell.execute_reply":"2022-07-14T11:20:18.696558Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Check successful update:\n\ntrain2.shape","metadata":{"execution":{"iopub.status.busy":"2022-07-14T11:20:18.698656Z","iopub.execute_input":"2022-07-14T11:20:18.699219Z","iopub.status.idle":"2022-07-14T11:20:18.706996Z","shell.execute_reply.started":"2022-07-14T11:20:18.699183Z","shell.execute_reply":"2022-07-14T11:20:18.706204Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Remaining missing data:\n\nmissing_values_count2 = train2.isnull().sum()\n\nfor i in range(0, len(missing_values_count2)):\n    if missing_values_count2[i]!=0:\n        print(missing_values_count2.index[i], missing_values_count2[i])","metadata":{"execution":{"iopub.status.busy":"2022-07-14T11:20:18.708574Z","iopub.execute_input":"2022-07-14T11:20:18.708978Z","iopub.status.idle":"2022-07-14T11:20:18.728883Z","shell.execute_reply.started":"2022-07-14T11:20:18.708942Z","shell.execute_reply":"2022-07-14T11:20:18.728039Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Let's have a look at some of these columns with missing values to see how we should deal with them. \n\nprint(train2.LotFrontage)","metadata":{"execution":{"iopub.status.busy":"2022-07-14T11:20:18.731125Z","iopub.execute_input":"2022-07-14T11:20:18.731537Z","iopub.status.idle":"2022-07-14T11:20:18.738597Z","shell.execute_reply.started":"2022-07-14T11:20:18.731501Z","shell.execute_reply":"2022-07-14T11:20:18.737893Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# From the documentation: LotFrontage: Linear feet of street connected to property. \n# We'll just replace the missing values with the mean.\n\ntrain2.LotFrontage = train2.LotFrontage.fillna(train2.LotFrontage.mean())","metadata":{"execution":{"iopub.status.busy":"2022-07-14T11:20:18.739865Z","iopub.execute_input":"2022-07-14T11:20:18.740279Z","iopub.status.idle":"2022-07-14T11:20:18.751930Z","shell.execute_reply.started":"2022-07-14T11:20:18.740230Z","shell.execute_reply":"2022-07-14T11:20:18.750879Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"#Check it worked\n\ntrain2.LotFrontage.isnull().sum()","metadata":{"execution":{"iopub.status.busy":"2022-07-14T11:20:18.753201Z","iopub.execute_input":"2022-07-14T11:20:18.753587Z","iopub.status.idle":"2022-07-14T11:20:18.763053Z","shell.execute_reply.started":"2022-07-14T11:20:18.753552Z","shell.execute_reply":"2022-07-14T11:20:18.762276Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"\"\"\"Now let's look at the FireplaceQu column. From the documentation, we have: \n    FireplaceQu: Fireplace quality\n\n       Ex\tExcellent - Exceptional Masonry Fireplace\n       Gd\tGood - Masonry Fireplace in main level\n       TA\tAverage - Prefabricated Fireplace in main living area or Masonry Fireplace in basement\n       Fa\tFair - Prefabricated Fireplace in basement\n       Po\tPoor - Ben Franklin Stove\n       NA\tNo Fireplace\n\nTherefore, the missing values are probably because the house doesn't have a fireplace. However, we see that there is already\na value, NA, for missing fireplace, so it may be that the data for these cases was not put in correctly. Let's look closer\nat these cases...\n\"\"\"","metadata":{"execution":{"iopub.status.busy":"2022-07-14T11:20:18.764479Z","iopub.execute_input":"2022-07-14T11:20:18.764992Z","iopub.status.idle":"2022-07-14T11:20:18.775468Z","shell.execute_reply.started":"2022-07-14T11:20:18.764956Z","shell.execute_reply":"2022-07-14T11:20:18.774568Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"(train2.FireplaceQu == \"NA\").sum()","metadata":{"execution":{"iopub.status.busy":"2022-07-14T11:20:18.777094Z","iopub.execute_input":"2022-07-14T11:20:18.777653Z","iopub.status.idle":"2022-07-14T11:20:18.785633Z","shell.execute_reply.started":"2022-07-14T11:20:18.777616Z","shell.execute_reply":"2022-07-14T11:20:18.784877Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"train2.FireplaceQu.value_counts()","metadata":{"execution":{"iopub.status.busy":"2022-07-14T11:20:18.788801Z","iopub.execute_input":"2022-07-14T11:20:18.789183Z","iopub.status.idle":"2022-07-14T11:20:18.797353Z","shell.execute_reply.started":"2022-07-14T11:20:18.789148Z","shell.execute_reply":"2022-07-14T11:20:18.796568Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Let's replace all missing values with \"NA\":\n\ntrain2.FireplaceQu = train2.FireplaceQu.fillna(\"NA\")","metadata":{"execution":{"iopub.status.busy":"2022-07-14T11:20:18.798551Z","iopub.execute_input":"2022-07-14T11:20:18.799675Z","iopub.status.idle":"2022-07-14T11:20:18.804198Z","shell.execute_reply.started":"2022-07-14T11:20:18.799633Z","shell.execute_reply":"2022-07-14T11:20:18.803530Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"train2.FireplaceQu.value_counts()","metadata":{"execution":{"iopub.status.busy":"2022-07-14T11:20:18.805346Z","iopub.execute_input":"2022-07-14T11:20:18.805817Z","iopub.status.idle":"2022-07-14T11:20:18.816743Z","shell.execute_reply.started":"2022-07-14T11:20:18.805780Z","shell.execute_reply":"2022-07-14T11:20:18.815934Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Once again check what values are still missing: \n\nmissing_values_count2 = train2.isnull().sum()\n\nfor i in range(0, len(missing_values_count2)):\n    if missing_values_count2[i]!=0:\n        print(missing_values_count2.index[i], missing_values_count2[i])","metadata":{"execution":{"iopub.status.busy":"2022-07-14T11:20:18.817938Z","iopub.execute_input":"2022-07-14T11:20:18.818741Z","iopub.status.idle":"2022-07-14T11:20:18.836698Z","shell.execute_reply.started":"2022-07-14T11:20:18.818702Z","shell.execute_reply":"2022-07-14T11:20:18.836038Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"train2.head()","metadata":{"execution":{"iopub.status.busy":"2022-07-14T11:20:18.838489Z","iopub.execute_input":"2022-07-14T11:20:18.839217Z","iopub.status.idle":"2022-07-14T11:20:18.862753Z","shell.execute_reply.started":"2022-07-14T11:20:18.839178Z","shell.execute_reply":"2022-07-14T11:20:18.861949Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Let's deal with the rest of the missing values...\n\ntrain2.MasVnrType.value_counts()\n\n# Seems like the -1 values are recorded as missing for some reason, but don't actually seem to be any missing values here.","metadata":{"execution":{"iopub.status.busy":"2022-07-14T11:20:18.864420Z","iopub.execute_input":"2022-07-14T11:20:18.864702Z","iopub.status.idle":"2022-07-14T11:20:18.874074Z","shell.execute_reply.started":"2022-07-14T11:20:18.864665Z","shell.execute_reply":"2022-07-14T11:20:18.873165Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"train2.MasVnrArea = train2.MasVnrArea.fillna(train2.MasVnrArea.mean()) # Replace with mean","metadata":{"execution":{"iopub.status.busy":"2022-07-14T11:20:18.876070Z","iopub.execute_input":"2022-07-14T11:20:18.876821Z","iopub.status.idle":"2022-07-14T11:20:18.883562Z","shell.execute_reply.started":"2022-07-14T11:20:18.876785Z","shell.execute_reply":"2022-07-14T11:20:18.882894Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# For BsmtQual and BsmtCond, note that they both have the same number of missing values, so these must be houses which don't have a basement, so we will fill these with NA. Do the same for all other basement features...\n\ntrain2.BsmtQual = train2.BsmtQual.fillna(\"NA\")\ntrain2.BsmtCond = train2.BsmtCond.fillna(\"NA\")\ntrain2.BsmtExposure = train2.BsmtExposure.fillna(\"NA\")\ntrain2.BsmtFinType1 = train2.BsmtFinType1.fillna(\"NA\")\ntrain2.BsmtFinType2 = train2.BsmtFinType2.fillna(\"NA\")","metadata":{"execution":{"iopub.status.busy":"2022-07-14T11:20:18.884896Z","iopub.execute_input":"2022-07-14T11:20:18.885528Z","iopub.status.idle":"2022-07-14T11:20:18.896808Z","shell.execute_reply.started":"2022-07-14T11:20:18.885492Z","shell.execute_reply":"2022-07-14T11:20:18.896078Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# For electrical, there is only one missing value, and again it is in fact a -1, so we will leave it as it is for now.","metadata":{"execution":{"iopub.status.busy":"2022-07-14T11:20:18.901457Z","iopub.execute_input":"2022-07-14T11:20:18.902352Z","iopub.status.idle":"2022-07-14T11:20:18.909901Z","shell.execute_reply.started":"2022-07-14T11:20:18.902315Z","shell.execute_reply":"2022-07-14T11:20:18.908851Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# For all garage related features, just replace with NA since these should correspond to houses which don't have a garage.\ntrain2.GarageType = train2.GarageType.fillna(\"NA\")\ntrain2.GarageYrBlt = train2.GarageYrBlt.fillna(\"NA\")\ntrain2.GarageFinish = train2.GarageFinish.fillna(\"NA\")\ntrain2.GarageQual = train2.GarageQual.fillna(\"NA\")\ntrain2.GarageCond = train2.GarageCond.fillna(\"NA\")","metadata":{"execution":{"iopub.status.busy":"2022-07-14T11:20:18.922553Z","iopub.execute_input":"2022-07-14T11:20:18.922753Z","iopub.status.idle":"2022-07-14T11:20:18.934973Z","shell.execute_reply.started":"2022-07-14T11:20:18.922729Z","shell.execute_reply":"2022-07-14T11:20:18.934274Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"missing_values_count2 = train2.isnull().sum()\n\nfor i in range(0, len(missing_values_count2)):\n    if missing_values_count2[i]!=0:\n        print(missing_values_count2.index[i], missing_values_count2[i])\n        \n# We see that there are no longer any missing values (the -1 values seem to no longer be treated as missing, not sure what's going on there...)","metadata":{"execution":{"iopub.status.busy":"2022-07-14T11:20:18.939873Z","iopub.execute_input":"2022-07-14T11:20:18.940059Z","iopub.status.idle":"2022-07-14T11:20:18.955357Z","shell.execute_reply.started":"2022-07-14T11:20:18.940037Z","shell.execute_reply":"2022-07-14T11:20:18.954575Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Feature selection\n\nHere we briefly look at the features and try to gauge how important each of them is in determining the house price. We only look at the mutual information, but a more detailed analysis could involve looking at the correlation between different features and the target and perhaps doing a principal component analysis.","metadata":{}},{"cell_type":"code","source":"train2.columns.shape","metadata":{"execution":{"iopub.status.busy":"2022-07-14T11:20:18.959492Z","iopub.execute_input":"2022-07-14T11:20:18.959679Z","iopub.status.idle":"2022-07-14T11:20:18.967003Z","shell.execute_reply.started":"2022-07-14T11:20:18.959656Z","shell.execute_reply":"2022-07-14T11:20:18.964684Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"train2copy = train2.copy()\ny = train2copy.pop(\"SalePrice\")","metadata":{"execution":{"iopub.status.busy":"2022-07-14T11:20:18.984457Z","iopub.execute_input":"2022-07-14T11:20:18.984950Z","iopub.status.idle":"2022-07-14T11:20:18.995099Z","shell.execute_reply.started":"2022-07-14T11:20:18.984910Z","shell.execute_reply":"2022-07-14T11:20:18.993985Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"train2copy.shape","metadata":{"execution":{"iopub.status.busy":"2022-07-14T11:20:19.002031Z","iopub.execute_input":"2022-07-14T11:20:19.002760Z","iopub.status.idle":"2022-07-14T11:20:19.013035Z","shell.execute_reply.started":"2022-07-14T11:20:19.002716Z","shell.execute_reply":"2022-07-14T11:20:19.012027Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Let's have a look at the mutual information to identify the most useful features:\n\n\n# Label encoding for categoricals (turns them all into integer columns)\n\nfor colname in train2copy.select_dtypes(\"object\"):\n    train2copy[colname], _ = train2copy[colname].factorize()\n\n# All discrete features should now have integer dtypes (double-check this before using MI!)\n\ndiscrete_features = train2copy.dtypes == int","metadata":{"execution":{"iopub.status.busy":"2022-07-14T11:20:19.020474Z","iopub.execute_input":"2022-07-14T11:20:19.020883Z","iopub.status.idle":"2022-07-14T11:20:19.084736Z","shell.execute_reply.started":"2022-07-14T11:20:19.020846Z","shell.execute_reply":"2022-07-14T11:20:19.083672Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"discrete_features.head()","metadata":{"execution":{"iopub.status.busy":"2022-07-14T11:20:19.086801Z","iopub.execute_input":"2022-07-14T11:20:19.087212Z","iopub.status.idle":"2022-07-14T11:20:19.097427Z","shell.execute_reply.started":"2022-07-14T11:20:19.087175Z","shell.execute_reply":"2022-07-14T11:20:19.096679Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"y.shape","metadata":{"execution":{"iopub.status.busy":"2022-07-14T11:20:19.099327Z","iopub.execute_input":"2022-07-14T11:20:19.100500Z","iopub.status.idle":"2022-07-14T11:20:19.110705Z","shell.execute_reply.started":"2022-07-14T11:20:19.100461Z","shell.execute_reply":"2022-07-14T11:20:19.109780Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"y.dtypes","metadata":{"execution":{"iopub.status.busy":"2022-07-14T11:20:19.112864Z","iopub.execute_input":"2022-07-14T11:20:19.113383Z","iopub.status.idle":"2022-07-14T11:20:19.125763Z","shell.execute_reply.started":"2022-07-14T11:20:19.113308Z","shell.execute_reply":"2022-07-14T11:20:19.124978Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"from sklearn.feature_selection import mutual_info_regression\nfrom sklearn.feature_selection import mutual_info_classif\n\n\n# Define a function that will present the MI scores clearly:\n\ndef make_mi_scores(X, y, discrete_features):\n    mi_scores = mutual_info_classif(X, y, discrete_features=discrete_features, random_state=0)\n    mi_scores = pd.Series(mi_scores, name=\"MI Scores\", index=X.columns)\n    mi_scores = mi_scores.sort_values(ascending=False)\n    return mi_scores\n\nmi_scores = make_mi_scores(train2copy, y, discrete_features)\nprint(mi_scores)","metadata":{"execution":{"iopub.status.busy":"2022-07-14T11:20:19.127176Z","iopub.execute_input":"2022-07-14T11:20:19.127737Z","iopub.status.idle":"2022-07-14T11:20:20.717386Z","shell.execute_reply.started":"2022-07-14T11:20:19.127699Z","shell.execute_reply":"2022-07-14T11:20:20.716627Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Recall that higher MI scores mean that those features tell us more about the target. Given some feature x and target y, the MI is essentially the relative entropy between the joint distribution f(x,y) and the product of distributions f(x)f(y). If these are the same, i.e. if x and y are independent, then the MI is zero, i.e knowledge of x gives us no knowledge of y. ","metadata":{}},{"cell_type":"code","source":"mi_scores.head(21)","metadata":{"execution":{"iopub.status.busy":"2022-07-14T11:20:20.718679Z","iopub.execute_input":"2022-07-14T11:20:20.719069Z","iopub.status.idle":"2022-07-14T11:20:20.727612Z","shell.execute_reply.started":"2022-07-14T11:20:20.719033Z","shell.execute_reply":"2022-07-14T11:20:20.726785Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"import matplotlib.pyplot as plt\n\ndef plot_mi_scores(scores):\n    scores = scores.sort_values(ascending=True)\n    width = np.arange(len(scores))\n    ticks = list(scores.index)\n    plt.barh(width, scores)\n    plt.yticks(width, ticks)\n    plt.title(\"Mutual Information Scores\")","metadata":{"execution":{"iopub.status.busy":"2022-07-14T11:20:20.730335Z","iopub.execute_input":"2022-07-14T11:20:20.730747Z","iopub.status.idle":"2022-07-14T11:20:20.736047Z","shell.execute_reply.started":"2022-07-14T11:20:20.730712Z","shell.execute_reply":"2022-07-14T11:20:20.735322Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"plt.figure(dpi=100, figsize=(8, 5))\n\nplot_mi_scores(mi_scores.head(21))","metadata":{"execution":{"iopub.status.busy":"2022-07-14T11:20:20.737577Z","iopub.execute_input":"2022-07-14T11:20:20.738197Z","iopub.status.idle":"2022-07-14T11:20:21.089393Z","shell.execute_reply.started":"2022-07-14T11:20:20.738150Z","shell.execute_reply":"2022-07-14T11:20:21.088698Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"\"\"\"At this point, we could either take a closer look at these categories and try and find more precise relations between them.\nIn particular, we should see if any other features with low MIs have some significant interactions with some of these high\nMI features. We could also try and create new features that would better fit what we're trying to predict. However, for now, \nwe will simply take the first 20 features with the highest MI scores (not inc. the ID) and see what this does for us.\n\"\"\"\n\ntrain2.pop(\"Id\")\ny = train2.pop(\"SalePrice\")","metadata":{"execution":{"iopub.status.busy":"2022-07-14T11:20:21.090631Z","iopub.execute_input":"2022-07-14T11:20:21.090958Z","iopub.status.idle":"2022-07-14T11:20:21.100045Z","shell.execute_reply.started":"2022-07-14T11:20:21.090920Z","shell.execute_reply":"2022-07-14T11:20:21.099271Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"train2.shape","metadata":{"execution":{"iopub.status.busy":"2022-07-14T11:20:21.102213Z","iopub.execute_input":"2022-07-14T11:20:21.102962Z","iopub.status.idle":"2022-07-14T11:20:21.108688Z","shell.execute_reply.started":"2022-07-14T11:20:21.102920Z","shell.execute_reply":"2022-07-14T11:20:21.107949Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"y.shape","metadata":{"execution":{"iopub.status.busy":"2022-07-14T11:20:21.109940Z","iopub.execute_input":"2022-07-14T11:20:21.111017Z","iopub.status.idle":"2022-07-14T11:20:21.118800Z","shell.execute_reply.started":"2022-07-14T11:20:21.110981Z","shell.execute_reply":"2022-07-14T11:20:21.118116Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"train2.columns","metadata":{"execution":{"iopub.status.busy":"2022-07-14T11:20:21.120182Z","iopub.execute_input":"2022-07-14T11:20:21.120493Z","iopub.status.idle":"2022-07-14T11:20:21.129551Z","shell.execute_reply.started":"2022-07-14T11:20:21.120458Z","shell.execute_reply":"2022-07-14T11:20:21.128684Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"features = [\"LotArea\", \"GrLivArea\", \"1stFlrSF\", \"BsmtUnfSF\", \"TotalBsmtSF\", \"GarageArea\", \"BsmtFinSF1\", \"YearBuilt\", \"GarageYrBlt\", \"YearRemodAdd\", \"OpenPorchSF\", \"2ndFlrSF\", \"WoodDeckSF\", \"Neighborhood\", \"MoSold\", \"MSSubClass\", \"OverallQual\", \"Exterior2nd\", \"TotRmsAbvGrd\", \"Exterior1st\"]\n\nX = train2[features]","metadata":{"execution":{"iopub.status.busy":"2022-07-14T11:20:21.131043Z","iopub.execute_input":"2022-07-14T11:20:21.131617Z","iopub.status.idle":"2022-07-14T11:20:21.139833Z","shell.execute_reply.started":"2022-07-14T11:20:21.131567Z","shell.execute_reply":"2022-07-14T11:20:21.139126Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"X.shape","metadata":{"execution":{"iopub.status.busy":"2022-07-14T11:20:21.141393Z","iopub.execute_input":"2022-07-14T11:20:21.141839Z","iopub.status.idle":"2022-07-14T11:20:21.147758Z","shell.execute_reply.started":"2022-07-14T11:20:21.141802Z","shell.execute_reply":"2022-07-14T11:20:21.147056Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Let's take a look at the new dataframe with the 20 \"best\" features:\n\nX.head()","metadata":{"execution":{"iopub.status.busy":"2022-07-14T11:20:21.152274Z","iopub.execute_input":"2022-07-14T11:20:21.152913Z","iopub.status.idle":"2022-07-14T11:20:21.169480Z","shell.execute_reply.started":"2022-07-14T11:20:21.152883Z","shell.execute_reply":"2022-07-14T11:20:21.168773Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"X.describe()","metadata":{"execution":{"iopub.status.busy":"2022-07-14T11:20:21.170642Z","iopub.execute_input":"2022-07-14T11:20:21.171474Z","iopub.status.idle":"2022-07-14T11:20:21.223237Z","shell.execute_reply.started":"2022-07-14T11:20:21.171435Z","shell.execute_reply":"2022-07-14T11:20:21.222528Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Turn all categorical data into integers\n\nfor colname in X.select_dtypes(\"object\"):\n    X[colname], _ = X[colname].factorize()","metadata":{"execution":{"iopub.status.busy":"2022-07-14T11:20:21.224733Z","iopub.execute_input":"2022-07-14T11:20:21.225254Z","iopub.status.idle":"2022-07-14T11:20:21.235972Z","shell.execute_reply.started":"2022-07-14T11:20:21.225213Z","shell.execute_reply":"2022-07-14T11:20:21.235121Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Model selection and training\n\nHere we train the model for generating predictions. We use a range of models, mostly those constructed using random trees such as random forests and boosted tree models. We use all of the available data for training at first. ","metadata":{}},{"cell_type":"code","source":"# Now we need to decide which models to use. To start with we'll just use a random forest model\n\nfrom sklearn.ensemble import RandomForestRegressor\nfrom sklearn.metrics import mean_absolute_error\n\nforest_model = RandomForestRegressor(random_state=1)\nforest_model.fit(X, y)","metadata":{"execution":{"iopub.status.busy":"2022-07-14T11:20:21.237700Z","iopub.execute_input":"2022-07-14T11:20:21.238276Z","iopub.status.idle":"2022-07-14T11:20:22.463398Z","shell.execute_reply.started":"2022-07-14T11:20:21.238219Z","shell.execute_reply":"2022-07-14T11:20:22.462744Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Now get the test data for predictions. Need to format the test data first so that it is in the right form. \n\ntest_file_path = \"/kaggle/input/house-prices-advanced-regression-techniques/test.csv\"\ntest = pd.read_csv(test_file_path)","metadata":{"execution":{"iopub.status.busy":"2022-07-14T11:20:22.464739Z","iopub.execute_input":"2022-07-14T11:20:22.464964Z","iopub.status.idle":"2022-07-14T11:20:22.495031Z","shell.execute_reply.started":"2022-07-14T11:20:22.464932Z","shell.execute_reply":"2022-07-14T11:20:22.494372Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"test.shape","metadata":{"execution":{"iopub.status.busy":"2022-07-14T11:20:22.495939Z","iopub.execute_input":"2022-07-14T11:20:22.496121Z","iopub.status.idle":"2022-07-14T11:20:22.505640Z","shell.execute_reply.started":"2022-07-14T11:20:22.496097Z","shell.execute_reply":"2022-07-14T11:20:22.504931Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"test.head()","metadata":{"execution":{"iopub.status.busy":"2022-07-14T11:20:22.506900Z","iopub.execute_input":"2022-07-14T11:20:22.507549Z","iopub.status.idle":"2022-07-14T11:20:22.531931Z","shell.execute_reply.started":"2022-07-14T11:20:22.507513Z","shell.execute_reply":"2022-07-14T11:20:22.531312Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"sample_path = \"/kaggle/input/house-prices-advanced-regression-techniques/sample_submission.csv\"\nsample = pd.read_csv(sample_path)","metadata":{"execution":{"iopub.status.busy":"2022-07-14T11:20:22.533121Z","iopub.execute_input":"2022-07-14T11:20:22.533892Z","iopub.status.idle":"2022-07-14T11:20:22.544603Z","shell.execute_reply.started":"2022-07-14T11:20:22.533853Z","shell.execute_reply":"2022-07-14T11:20:22.543520Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"sample.head()","metadata":{"execution":{"iopub.status.busy":"2022-07-14T11:20:22.546023Z","iopub.execute_input":"2022-07-14T11:20:22.546951Z","iopub.status.idle":"2022-07-14T11:20:22.555909Z","shell.execute_reply.started":"2022-07-14T11:20:22.546906Z","shell.execute_reply":"2022-07-14T11:20:22.555104Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"\"\"\"We need to check the data to see if there are missing values. These may be in different columns to what we had for \nthe training data. We also need to remove the columns that we removed for the training data since our model is trained \nonly for specific input variables. \n\"\"\"\n\nmissing_count_test = test.isnull().sum()\n\nfor i in range(0, len(missing_count_test)):\n    if missing_count_test[i]!=0:\n        print(missing_count_test.index[i], missing_count_test[i])","metadata":{"execution":{"iopub.status.busy":"2022-07-14T11:20:22.557220Z","iopub.execute_input":"2022-07-14T11:20:22.557979Z","iopub.status.idle":"2022-07-14T11:20:22.582035Z","shell.execute_reply.started":"2022-07-14T11:20:22.557945Z","shell.execute_reply":"2022-07-14T11:20:22.581381Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"test2 = test.drop([\"Alley\", \"PoolQC\", \"Fence\", \"MiscFeature\"], axis=1)\n\nprint(features) \n\n# Since we only need to focus on these, let's just remove all other columns:\n\nX_test = test2[features]","metadata":{"execution":{"iopub.status.busy":"2022-07-14T11:20:22.583047Z","iopub.execute_input":"2022-07-14T11:20:22.583282Z","iopub.status.idle":"2022-07-14T11:20:22.594622Z","shell.execute_reply.started":"2022-07-14T11:20:22.583235Z","shell.execute_reply":"2022-07-14T11:20:22.593669Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"missing_count_X_test = X_test.isnull().sum()\n\nfor i in range(0, len(missing_count_X_test)):\n    if missing_count_X_test[i]!=0:\n        print(missing_count_X_test.index[i], missing_count_X_test[i])","metadata":{"execution":{"iopub.status.busy":"2022-07-14T11:20:22.596616Z","iopub.execute_input":"2022-07-14T11:20:22.596848Z","iopub.status.idle":"2022-07-14T11:20:22.606827Z","shell.execute_reply.started":"2022-07-14T11:20:22.596817Z","shell.execute_reply":"2022-07-14T11:20:22.605948Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Impute missing values:\n\nX_test.BsmtUnfSF.value_counts()","metadata":{"execution":{"iopub.status.busy":"2022-07-14T11:20:22.608430Z","iopub.execute_input":"2022-07-14T11:20:22.608695Z","iopub.status.idle":"2022-07-14T11:20:22.621806Z","shell.execute_reply.started":"2022-07-14T11:20:22.608663Z","shell.execute_reply":"2022-07-14T11:20:22.621142Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"X_test.BsmtUnfSF = X_test.BsmtUnfSF.fillna(X_test.BsmtUnfSF.mean())","metadata":{"execution":{"iopub.status.busy":"2022-07-14T11:20:22.624266Z","iopub.execute_input":"2022-07-14T11:20:22.624523Z","iopub.status.idle":"2022-07-14T11:20:22.635097Z","shell.execute_reply.started":"2022-07-14T11:20:22.624493Z","shell.execute_reply":"2022-07-14T11:20:22.634114Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"X_test.BsmtUnfSF.isnull().sum()","metadata":{"execution":{"iopub.status.busy":"2022-07-14T11:20:22.636791Z","iopub.execute_input":"2022-07-14T11:20:22.637347Z","iopub.status.idle":"2022-07-14T11:20:22.643821Z","shell.execute_reply.started":"2022-07-14T11:20:22.637312Z","shell.execute_reply":"2022-07-14T11:20:22.642930Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"X_test.TotalBsmtSF.value_counts()\n\nX_test.TotalBsmtSF = X_test.TotalBsmtSF.fillna(X_test.TotalBsmtSF.mean())\nX_test.GarageArea = X_test.GarageArea.fillna(X_test.GarageArea.mean())\nX_test.BsmtFinSF1 = X_test.BsmtFinSF1.fillna(X_test.BsmtFinSF1.mean())\n\n# Replace the missing value for the next two with the mode, since these are categorical. The [0] is needed because mode returns a series.\n\nX_test.Exterior1st = X_test.Exterior1st.fillna(X_test.Exterior1st.mode()[0])\nX_test.Exterior2nd = X_test.Exterior2nd.fillna(X_test.Exterior2nd.mode()[0])","metadata":{"execution":{"iopub.status.busy":"2022-07-14T11:20:22.645320Z","iopub.execute_input":"2022-07-14T11:20:22.645753Z","iopub.status.idle":"2022-07-14T11:20:22.657167Z","shell.execute_reply.started":"2022-07-14T11:20:22.645570Z","shell.execute_reply":"2022-07-14T11:20:22.656473Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"missing_count_X_test = X_test.isnull().sum()\n\nfor i in range(0, len(missing_count_X_test)):\n    if missing_count_X_test[i]!=0:\n        print(missing_count_X_test.index[i], missing_count_X_test[i])","metadata":{"execution":{"iopub.status.busy":"2022-07-14T11:20:22.658663Z","iopub.execute_input":"2022-07-14T11:20:22.658969Z","iopub.status.idle":"2022-07-14T11:20:22.668835Z","shell.execute_reply.started":"2022-07-14T11:20:22.658934Z","shell.execute_reply":"2022-07-14T11:20:22.667847Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"X_test.GarageYrBlt = X_test.GarageYrBlt.fillna(\"NA\")\n\nmissing_count_X_test = X_test.isnull().sum()\n\nfor i in range(0, len(missing_count_X_test)):\n    if missing_count_X_test[i]!=0:\n        print(missing_count_X_test.index[i], missing_count_X_test[i])\n\n# Now we should be ready to make predictions...","metadata":{"execution":{"iopub.status.busy":"2022-07-14T11:20:22.670194Z","iopub.execute_input":"2022-07-14T11:20:22.670690Z","iopub.status.idle":"2022-07-14T11:20:22.679941Z","shell.execute_reply.started":"2022-07-14T11:20:22.670655Z","shell.execute_reply":"2022-07-14T11:20:22.679133Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Encode categorical variables:\n\nfor colname in X_test.select_dtypes(\"object\"):\n    X_test[colname], _ = X_test[colname].factorize()","metadata":{"execution":{"iopub.status.busy":"2022-07-14T11:20:22.681146Z","iopub.execute_input":"2022-07-14T11:20:22.681533Z","iopub.status.idle":"2022-07-14T11:20:22.694117Z","shell.execute_reply.started":"2022-07-14T11:20:22.681499Z","shell.execute_reply":"2022-07-14T11:20:22.693164Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Let's make predictions using the random forest model:\n\npreds = forest_model.predict(X_test)","metadata":{"execution":{"iopub.status.busy":"2022-07-14T11:20:22.695532Z","iopub.execute_input":"2022-07-14T11:20:22.696061Z","iopub.status.idle":"2022-07-14T11:20:22.730508Z","shell.execute_reply.started":"2022-07-14T11:20:22.696024Z","shell.execute_reply":"2022-07-14T11:20:22.729867Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"preds","metadata":{"execution":{"iopub.status.busy":"2022-07-14T11:20:22.731603Z","iopub.execute_input":"2022-07-14T11:20:22.731835Z","iopub.status.idle":"2022-07-14T11:20:22.737609Z","shell.execute_reply.started":"2022-07-14T11:20:22.731801Z","shell.execute_reply":"2022-07-14T11:20:22.736828Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"test_id = np.array(test.Id)","metadata":{"execution":{"iopub.status.busy":"2022-07-14T11:20:22.739024Z","iopub.execute_input":"2022-07-14T11:20:22.739972Z","iopub.status.idle":"2022-07-14T11:20:22.744248Z","shell.execute_reply.started":"2022-07-14T11:20:22.739935Z","shell.execute_reply":"2022-07-14T11:20:22.743410Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Prepare for submission:\n\nframe = { 'Id': test_id, 'SalePrice': preds }\n  \nfirst_submission = pd.DataFrame(frame)","metadata":{"execution":{"iopub.status.busy":"2022-07-14T11:20:22.745552Z","iopub.execute_input":"2022-07-14T11:20:22.746274Z","iopub.status.idle":"2022-07-14T11:20:22.753400Z","shell.execute_reply.started":"2022-07-14T11:20:22.746181Z","shell.execute_reply":"2022-07-14T11:20:22.752693Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"first_submission.head()","metadata":{"execution":{"iopub.status.busy":"2022-07-14T11:20:22.755802Z","iopub.execute_input":"2022-07-14T11:20:22.756607Z","iopub.status.idle":"2022-07-14T11:20:22.767845Z","shell.execute_reply.started":"2022-07-14T11:20:22.756572Z","shell.execute_reply":"2022-07-14T11:20:22.767047Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"sample.head()","metadata":{"execution":{"iopub.status.busy":"2022-07-14T11:20:22.774618Z","iopub.execute_input":"2022-07-14T11:20:22.775248Z","iopub.status.idle":"2022-07-14T11:20:22.783443Z","shell.execute_reply.started":"2022-07-14T11:20:22.775214Z","shell.execute_reply":"2022-07-14T11:20:22.782641Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"first_submission.to_csv('sub1.csv', index=False)\n\n# Rank at time of submission: 3525/5382, error = 0.15524\n# To improve, should spend more time on feature engineering, maybe think more carefully about how missing values are filled in.\n# Try creating new features, look at more metrics to help with this, use more graphs to visualise which features are useful and which\n# features interact with each other. \n# Easy thing to try would be to use different models such as XGBoost or a Neural Network. ","metadata":{"execution":{"iopub.status.busy":"2022-07-14T11:20:22.784558Z","iopub.execute_input":"2022-07-14T11:20:22.785357Z","iopub.status.idle":"2022-07-14T11:20:22.800225Z","shell.execute_reply.started":"2022-07-14T11:20:22.785320Z","shell.execute_reply":"2022-07-14T11:20:22.799425Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"#ATTEMPT 2: XGBoost (gradient boosted tree regression model)\n\nfrom xgboost import XGBRegressor\n\nmy_model_2 = XGBRegressor()\n\n# Add silent=True to avoid printing out updates with each cycle.\n\nmy_model_2.fit(X, y, verbose=False)","metadata":{"execution":{"iopub.status.busy":"2022-07-14T11:20:22.801297Z","iopub.execute_input":"2022-07-14T11:20:22.802027Z","iopub.status.idle":"2022-07-14T11:20:26.302174Z","shell.execute_reply.started":"2022-07-14T11:20:22.801986Z","shell.execute_reply":"2022-07-14T11:20:26.301488Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"preds2 = my_model_2.predict(X_test)","metadata":{"execution":{"iopub.status.busy":"2022-07-14T11:20:26.303359Z","iopub.execute_input":"2022-07-14T11:20:26.304171Z","iopub.status.idle":"2022-07-14T11:20:26.315033Z","shell.execute_reply.started":"2022-07-14T11:20:26.304133Z","shell.execute_reply":"2022-07-14T11:20:26.314430Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"frame2 = { 'Id': test_id, 'SalePrice': preds2 }\n  \nsecond_submission = pd.DataFrame(frame2)\n\nsecond_submission.to_csv('sub2.csv', index=False)\n\n\"\"\"\nerror = 0.16480, worse than the random forest model. \n\"\"\"","metadata":{"execution":{"iopub.status.busy":"2022-07-14T11:20:26.316411Z","iopub.execute_input":"2022-07-14T11:20:26.316649Z","iopub.status.idle":"2022-07-14T11:20:26.331221Z","shell.execute_reply.started":"2022-07-14T11:20:26.316622Z","shell.execute_reply":"2022-07-14T11:20:26.330311Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"#ATTEMPT 3: XGBoost with n_estimators=1000, learning_rate=0.05, early_stopping_rounds=5\n\nmy_model_3 = XGBRegressor(n_estimators = 1000, learning_rate=0.05)\n\nmy_model_3.fit(X, y, early_stopping_rounds=5, eval_set=[(X, y)], verbose=False)\n\npreds3 = my_model_3.predict(X_test)\n\nframe3 = { 'Id': test_id, 'SalePrice': preds3 }\n  \nsubmission3 = pd.DataFrame(frame3)\n\nsubmission3.to_csv('sub3.csv', index=False)\n\n\"\"\"error = 0.15476, best so far, but only slightly better than random forest model. Seems like it would be worth playing\naround with the hyperparameters. Can do so using GridSearchCV for example.\n\"\"\"","metadata":{"execution":{"iopub.status.busy":"2022-07-14T11:20:26.333514Z","iopub.execute_input":"2022-07-14T11:20:26.333906Z","iopub.status.idle":"2022-07-14T11:20:29.691922Z","shell.execute_reply.started":"2022-07-14T11:20:26.333879Z","shell.execute_reply":"2022-07-14T11:20:29.691391Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"#ATTEMPT 4: Ridge regression (i.e. L2-regularised linear regression)\n\nfrom sklearn import linear_model\nmy_model_4 = linear_model.Ridge(alpha=.5) # alpha is the regularisation/penalisation parameter.\nmy_model_4.fit(X,y)\n\npreds4 = my_model_4.predict(X_test)\n\nframe4 = { 'Id': test_id, 'SalePrice': preds4 }\n  \nsubmission4 = pd.DataFrame(frame4)\n\nsubmission4.to_csv('sub4.csv', index=False)\n\n\"\"\"error = 0.59955, much worse than tree-based models above. Seems like a linear fit is probably too simple. \n\"\"\"","metadata":{"execution":{"iopub.status.busy":"2022-07-14T11:20:29.695171Z","iopub.execute_input":"2022-07-14T11:20:29.695898Z","iopub.status.idle":"2022-07-14T11:20:29.731317Z","shell.execute_reply.started":"2022-07-14T11:20:29.695862Z","shell.execute_reply":"2022-07-14T11:20:29.730628Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Back to feature selection\n\nLet's try to improve on the above by using scikit-learn's in-built SelectFromModel feature on ExtraTreesRegressor to automatically identify useful features. ","metadata":{}},{"cell_type":"code","source":"# Let's try to preprocess the input data better before we train the model. We'll use a built-in feature selection tool. \n\nX_full = train2.copy() # This is the full dataset we started with (with some columns removed and missing values replaced).\n\nfor colname in X_full.select_dtypes(\"object\"):\n    X_full[colname], _ = X_full[colname].factorize()","metadata":{"execution":{"iopub.status.busy":"2022-07-14T11:20:29.732352Z","iopub.execute_input":"2022-07-14T11:20:29.732706Z","iopub.status.idle":"2022-07-14T11:20:29.786627Z","shell.execute_reply.started":"2022-07-14T11:20:29.732672Z","shell.execute_reply":"2022-07-14T11:20:29.785848Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"X_full.shape","metadata":{"execution":{"iopub.status.busy":"2022-07-14T11:20:29.787696Z","iopub.execute_input":"2022-07-14T11:20:29.788097Z","iopub.status.idle":"2022-07-14T11:20:29.794188Z","shell.execute_reply.started":"2022-07-14T11:20:29.788059Z","shell.execute_reply":"2022-07-14T11:20:29.793316Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"X.shape","metadata":{"execution":{"iopub.status.busy":"2022-07-14T11:20:29.795572Z","iopub.execute_input":"2022-07-14T11:20:29.796301Z","iopub.status.idle":"2022-07-14T11:20:29.805158Z","shell.execute_reply.started":"2022-07-14T11:20:29.796250Z","shell.execute_reply":"2022-07-14T11:20:29.804310Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"from sklearn.ensemble import ExtraTreesRegressor\nfrom sklearn.feature_selection import SelectFromModel\n\nclf = ExtraTreesRegressor(n_estimators=50)\nclf = clf.fit(X_full, y)\n\nmodel = SelectFromModel(clf, prefit=True)\nX_new = model.transform(X_full)","metadata":{"execution":{"iopub.status.busy":"2022-07-14T12:15:42.083400Z","iopub.execute_input":"2022-07-14T12:15:42.084194Z","iopub.status.idle":"2022-07-14T12:15:42.824944Z","shell.execute_reply.started":"2022-07-14T12:15:42.084153Z","shell.execute_reply":"2022-07-14T12:15:42.824236Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"X_new.shape","metadata":{"execution":{"iopub.status.busy":"2022-07-14T12:15:47.538829Z","iopub.execute_input":"2022-07-14T12:15:47.539633Z","iopub.status.idle":"2022-07-14T12:15:47.545188Z","shell.execute_reply.started":"2022-07-14T12:15:47.539585Z","shell.execute_reply":"2022-07-14T12:15:47.544453Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"model.get_support()","metadata":{"execution":{"iopub.status.busy":"2022-07-14T12:16:25.077839Z","iopub.execute_input":"2022-07-14T12:16:25.078245Z","iopub.status.idle":"2022-07-14T12:16:25.103199Z","shell.execute_reply.started":"2022-07-14T12:16:25.078209Z","shell.execute_reply":"2022-07-14T12:16:25.102415Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"a=(X_full.columns * model.get_support()).tolist()\nclf_features = []\nfor col in a:\n    if len(col)!=0:\n        clf_features.append(col)\nprint(clf_features)","metadata":{"execution":{"iopub.status.busy":"2022-07-14T12:16:29.408129Z","iopub.execute_input":"2022-07-14T12:16:29.408585Z","iopub.status.idle":"2022-07-14T12:16:29.423063Z","shell.execute_reply.started":"2022-07-14T12:16:29.408550Z","shell.execute_reply":"2022-07-14T12:16:29.422327Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"X2 = X_full[clf_features]","metadata":{"execution":{"iopub.status.busy":"2022-07-14T12:17:36.443045Z","iopub.execute_input":"2022-07-14T12:17:36.443327Z","iopub.status.idle":"2022-07-14T12:17:36.448219Z","shell.execute_reply.started":"2022-07-14T12:17:36.443297Z","shell.execute_reply":"2022-07-14T12:17:36.447549Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Test data treatment and model training\n\nAs above, we get the test data into the appropriate form and train a model. Here, we use the same XGBoost model as above.","metadata":{}},{"cell_type":"code","source":"# Now we have the new preprocessed input data, but we need to do the same to the test data. First, we will replace\n# all missing values from the full test data frame. We start with test2, which is the test data with the columns removed\n# that had most values missing. \n\ntest2copy = test2.copy()\n\nfor colname in test2copy.select_dtypes(\"object\"):\n    test2copy[colname], _ = test2copy[colname].factorize()\n    \nmissing_test2copy = test2copy.isnull().sum()\n\nfor i in range(0, len(missing_test2copy)):\n    if missing_test2copy[i]!=0:\n        print(missing_test2copy.index[i], missing_test2copy[i])","metadata":{"execution":{"iopub.status.busy":"2022-07-14T12:17:40.587658Z","iopub.execute_input":"2022-07-14T12:17:40.588483Z","iopub.status.idle":"2022-07-14T12:17:40.632500Z","shell.execute_reply.started":"2022-07-14T12:17:40.588440Z","shell.execute_reply":"2022-07-14T12:17:40.631654Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"test2copy.BsmtUnfSF = test2copy.BsmtUnfSF.fillna(test2copy.BsmtUnfSF.mean())\ntest2copy.TotalBsmtSF =test2copy.TotalBsmtSF.fillna(test2copy.TotalBsmtSF.mean())\ntest2copy.GarageArea = test2copy.GarageArea.fillna(test2copy.GarageArea.mean())\ntest2copy.BsmtFinSF1 = test2copy.BsmtFinSF1.fillna(test2copy.BsmtFinSF1.mean())\ntest2copy.GarageYrBlt = test2copy.GarageYrBlt.fillna(\"NA\")","metadata":{"execution":{"iopub.status.busy":"2022-07-14T12:17:49.112169Z","iopub.execute_input":"2022-07-14T12:17:49.112631Z","iopub.status.idle":"2022-07-14T12:17:49.123700Z","shell.execute_reply.started":"2022-07-14T12:17:49.112595Z","shell.execute_reply":"2022-07-14T12:17:49.122944Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"missing_test2copy = test2copy.isnull().sum()\n\nfor i in range(0, len(missing_test2copy)):\n    if missing_test2copy[i]!=0:\n        print(missing_test2copy.index[i], missing_test2copy[i])","metadata":{"execution":{"iopub.status.busy":"2022-07-14T12:17:51.866854Z","iopub.execute_input":"2022-07-14T12:17:51.867117Z","iopub.status.idle":"2022-07-14T12:17:51.884397Z","shell.execute_reply.started":"2022-07-14T12:17:51.867087Z","shell.execute_reply":"2022-07-14T12:17:51.883312Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"test2copy.LotFrontage = test2copy.LotFrontage.fillna(test2copy.LotFrontage.mean())\ntest2copy.MasVnrArea = test2copy.MasVnrArea.fillna(test2copy.MasVnrArea.mean())\ntest2copy.BsmtFinSF2 = test2copy.BsmtFinSF2.fillna(test2copy.BsmtFinSF2.mean())\ntest2copy.BsmtFullBath = test2copy.BsmtFullBath.fillna(test2copy.BsmtFullBath.mode()[0])\ntest2copy.BsmtHalfBath = test2copy.BsmtHalfBath.fillna(test2copy.BsmtHalfBath.mode()[0])\ntest2copy.GarageCars = test2copy.GarageCars.fillna(test2copy.GarageCars.mode()[0])","metadata":{"execution":{"iopub.status.busy":"2022-07-14T12:17:53.959241Z","iopub.execute_input":"2022-07-14T12:17:53.959832Z","iopub.status.idle":"2022-07-14T12:17:53.970994Z","shell.execute_reply.started":"2022-07-14T12:17:53.959790Z","shell.execute_reply":"2022-07-14T12:17:53.970205Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"for colname in test2copy.select_dtypes(\"object\"):\n    test2copy[colname], _ = test2copy[colname].factorize()\n\ntest_full = test2copy \n\n# This is now the full data set with missing values replaces and appropriate columns removed. This is \n# what we will start with from now on when we create any new preprocessing structure. ","metadata":{"execution":{"iopub.status.busy":"2022-07-14T12:17:57.073764Z","iopub.execute_input":"2022-07-14T12:17:57.074642Z","iopub.status.idle":"2022-07-14T12:17:57.083574Z","shell.execute_reply.started":"2022-07-14T12:17:57.074587Z","shell.execute_reply":"2022-07-14T12:17:57.082680Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"test_full.head()","metadata":{"execution":{"iopub.status.busy":"2022-07-14T12:18:00.521210Z","iopub.execute_input":"2022-07-14T12:18:00.522037Z","iopub.status.idle":"2022-07-14T12:18:00.541898Z","shell.execute_reply.started":"2022-07-14T12:18:00.521991Z","shell.execute_reply":"2022-07-14T12:18:00.541239Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"test_full.shape","metadata":{"execution":{"iopub.status.busy":"2022-07-14T12:18:15.440981Z","iopub.execute_input":"2022-07-14T12:18:15.446443Z","iopub.status.idle":"2022-07-14T12:18:15.455006Z","shell.execute_reply.started":"2022-07-14T12:18:15.446386Z","shell.execute_reply":"2022-07-14T12:18:15.454391Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Now we can fit the model as usual. We will use the XGBoost that gave the best performance before. \n\nX_test_2 = test_full[clf_features]","metadata":{"execution":{"iopub.status.busy":"2022-07-14T12:18:39.734726Z","iopub.execute_input":"2022-07-14T12:18:39.734979Z","iopub.status.idle":"2022-07-14T12:18:39.740949Z","shell.execute_reply.started":"2022-07-14T12:18:39.734950Z","shell.execute_reply":"2022-07-14T12:18:39.740298Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# ATTEMPT 5: XGBoost with SelectFromModel feature selector (tree regressor)\n\n\nmy_model_5 = XGBRegressor(n_estimators = 1000, learning_rate=0.05)\n\nmy_model_5.fit(X2, y, early_stopping_rounds=5, eval_set=[(X2, y)], verbose=False)\n\npreds5 = my_model_5.predict(X_test_2)\n\nframe5 = { 'Id': test_id, 'SalePrice': preds5 }\n  \nsubmission5 = pd.DataFrame(frame5)\n\nsubmission5.to_csv('sub5.csv', index=False)\n\n\"\"\"\nerror = 0.16250, worse than our previous attempt with the features selected with MI, so seems that was better. \n\"\"\"","metadata":{"execution":{"iopub.status.busy":"2022-07-14T12:35:09.612726Z","iopub.execute_input":"2022-07-14T12:35:09.612988Z","iopub.status.idle":"2022-07-14T12:35:12.179297Z","shell.execute_reply.started":"2022-07-14T12:35:09.612960Z","shell.execute_reply":"2022-07-14T12:35:12.178763Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# ATTEMPT 6: Same as 5, but now we use separate parts of the data for training and validation to see if there is a problem with overfitting.\n\nmy_model_6 = XGBRegressor(n_estimators = 1000, learning_rate=0.05)\n\nmy_model_6.fit(X2[0:1001], y[0:1001], early_stopping_rounds=5, eval_set=[(X2[1001:1460], y[1001:1460])], verbose=False)\n\npreds6 = my_model_6.predict(X_test_2)\n\nframe6 = { 'Id': test_id, 'SalePrice': preds6 }\n  \nsubmission6 = pd.DataFrame(frame6)\n\n\n\n# error = 0.15014, slightly worse than 5, so there doesn't seem to be any issue with overfitting. \n\n\nsubmission6.to_csv('sub6.csv', index=False)","metadata":{"execution":{"iopub.status.busy":"2022-07-14T13:00:14.585312Z","iopub.execute_input":"2022-07-14T13:00:14.586254Z","iopub.status.idle":"2022-07-14T13:00:14.933960Z","shell.execute_reply.started":"2022-07-14T13:00:14.586182Z","shell.execute_reply":"2022-07-14T13:00:14.933227Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"\"\"\"\ndef make_mi_scores(X, y, discrete_features):\n    mi_scores = mutual_info_classif(X, y, discrete_features=discrete_features, random_state=0)\n    mi_scores = pd.Series(mi_scores, name=\"MI Scores\", index=X.columns)\n    mi_scores = mi_scores.sort_values(ascending=False)\n    return mi_scores\n\"\"\"\n\ndiscrete_features_2 = X2.dtypes == int\n\nmi_scores_2 = make_mi_scores(X2, y, discrete_features_2)\nprint(mi_scores_2)","metadata":{"execution":{"iopub.status.busy":"2022-07-14T12:28:22.244420Z","iopub.execute_input":"2022-07-14T12:28:22.245164Z","iopub.status.idle":"2022-07-14T12:28:22.277330Z","shell.execute_reply.started":"2022-07-14T12:28:22.245127Z","shell.execute_reply":"2022-07-14T12:28:22.276472Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"import seaborn as sns","metadata":{"execution":{"iopub.status.busy":"2022-07-14T12:30:02.056246Z","iopub.execute_input":"2022-07-14T12:30:02.056523Z","iopub.status.idle":"2022-07-14T12:30:02.062457Z","shell.execute_reply.started":"2022-07-14T12:30:02.056492Z","shell.execute_reply":"2022-07-14T12:30:02.061701Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"\"\"\"IGNORE!!\n\nfrom sklearn.cluster import KMeans\n\nplt.style.use(\"seaborn-whitegrid\")\nplt.rc(\"figure\", autolayout=True)\nplt.rc(\n    \"axes\",\n    labelweight=\"bold\",\n    labelsize=\"large\",\n    titleweight=\"bold\",\n    titlesize=14,\n    titlepad=10,\n)\n\n\nsns.relplot(\n    x=\"LotArea\", y=\"GrLivArea\", hue=\"OverallQual\", data=X2, height=6,\n)\n\nSeem to be some outliers here, and most of the houses seem to have similar LotArea, so maybe this isn't a very useful feature \nto have? There does seem to be a fairly clear relation between GrLivArea and OverallQual though.\n\"\"\"","metadata":{"execution":{"iopub.status.busy":"2022-07-14T12:30:34.618046Z","iopub.execute_input":"2022-07-14T12:30:34.618863Z","iopub.status.idle":"2022-07-14T12:30:34.624956Z","shell.execute_reply.started":"2022-07-14T12:30:34.618815Z","shell.execute_reply":"2022-07-14T12:30:34.624316Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"sns.catplot(x=\"GrLivArea\", y=\"OverallQual\", data=X2, kind=\"boxen\", height=10)\n\n# Weird looking, but there does at least seem to be a general upwards diagonal trend, as hoped. ","metadata":{"execution":{"iopub.status.busy":"2022-07-14T12:29:11.510600Z","iopub.execute_input":"2022-07-14T12:29:11.510858Z","iopub.status.idle":"2022-07-14T12:29:27.350535Z","shell.execute_reply.started":"2022-07-14T12:29:11.510831Z","shell.execute_reply":"2022-07-14T12:29:27.349834Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Given that model 3 is the best one so far, let's return to the features we were using there, but let's try and search for the best hyperparameters.","metadata":{}},{"cell_type":"code","source":"submission5.to_csv('sub5.csv', index=False)","metadata":{"execution":{"iopub.status.busy":"2022-07-14T12:29:32.539625Z","iopub.execute_input":"2022-07-14T12:29:32.540108Z","iopub.status.idle":"2022-07-14T12:29:32.550318Z","shell.execute_reply.started":"2022-07-14T12:29:32.540071Z","shell.execute_reply":"2022-07-14T12:29:32.549535Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Hyperparameter search\n\nIn the above, we have made little effort to choose good hyperparameters. Here, we try to rectify this by using the GridSearchCV method. We do this for both the random forest model and the XGBoost model.","metadata":{}},{"cell_type":"code","source":"from sklearn.model_selection import GridSearchCV\n\nparam_grid = [\n {'n_estimators': [10, 50, 100, 200], 'max_features': [2, 4, 6, 8]},\n {'bootstrap': [False], 'n_estimators': [10, 100], 'max_features': [2, 4, 6]},\n ]\n\nforest = RandomForestRegressor()\n\ngrid_search = GridSearchCV(forest, param_grid, cv=5,\n scoring='neg_mean_squared_error')\n\ngrid_search.fit(X, y)","metadata":{"execution":{"iopub.status.busy":"2022-07-14T13:07:17.437732Z","iopub.execute_input":"2022-07-14T13:07:17.437986Z","iopub.status.idle":"2022-07-14T13:07:55.377778Z","shell.execute_reply.started":"2022-07-14T13:07:17.437958Z","shell.execute_reply":"2022-07-14T13:07:55.377097Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Now let's take the best parameters: \n\ngrid_search.best_params_","metadata":{"execution":{"iopub.status.busy":"2022-07-14T13:09:25.673577Z","iopub.execute_input":"2022-07-14T13:09:25.673857Z","iopub.status.idle":"2022-07-14T13:09:25.679035Z","shell.execute_reply.started":"2022-07-14T13:09:25.673828Z","shell.execute_reply":"2022-07-14T13:09:25.678328Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"grid_search.best_estimator_","metadata":{"execution":{"iopub.status.busy":"2022-07-14T13:09:32.537668Z","iopub.execute_input":"2022-07-14T13:09:32.538435Z","iopub.status.idle":"2022-07-14T13:09:32.545949Z","shell.execute_reply.started":"2022-07-14T13:09:32.538387Z","shell.execute_reply":"2022-07-14T13:09:32.544184Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"#ATTEMPT 7: Random forest with grid search to determine hyperparameters. \n\nmy_model_7 = grid_search.best_estimator_\n\nmy_model_7.fit(X, y)\n\npreds7 = my_model_7.predict(X_test)\n\nframe7 = { 'Id': test_id, 'SalePrice': preds7 }\n  \nsubmission7 = pd.DataFrame(frame7)\n\nsubmission7.to_csv('sub7.csv', index=False)\n\n\n# error=0.15276, so better than the base XGBoost model we used above.","metadata":{"execution":{"iopub.status.busy":"2022-07-14T13:12:12.791592Z","iopub.execute_input":"2022-07-14T13:12:12.791862Z","iopub.status.idle":"2022-07-14T13:12:13.361924Z","shell.execute_reply.started":"2022-07-14T13:12:12.791833Z","shell.execute_reply":"2022-07-14T13:12:13.361206Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"The random forest with grid search performed quite well. Let's see if the grid search improves the performance of the XGBoost model.","metadata":{}},{"cell_type":"code","source":"# Now let's do the same for XGBoost:\n\nparam_grid_xgb = [\n {'n_estimators': [10, 100, 500, 1000], 'max_depth': [2, 4, 6, 8], 'learning_rate': [0.01, 0.05, 0.2]},\n ]\n\nxgb = XGBRegressor()\n\ngrid_search_xgb = GridSearchCV(xgb, param_grid_xgb, cv=5,\n scoring='neg_mean_squared_error')\n\ngrid_search_xgb.fit(X, y)","metadata":{"execution":{"iopub.status.busy":"2022-07-14T13:32:51.113022Z","iopub.execute_input":"2022-07-14T13:32:51.113629Z","iopub.status.idle":"2022-07-14T13:36:42.918918Z","shell.execute_reply.started":"2022-07-14T13:32:51.113595Z","shell.execute_reply":"2022-07-14T13:36:42.918124Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"grid_search_xgb.best_estimator_","metadata":{"execution":{"iopub.status.busy":"2022-07-14T13:37:52.606023Z","iopub.execute_input":"2022-07-14T13:37:52.606304Z","iopub.status.idle":"2022-07-14T13:37:52.615993Z","shell.execute_reply.started":"2022-07-14T13:37:52.606254Z","shell.execute_reply":"2022-07-14T13:37:52.615250Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"#ATTEMPT 8: XGBoost with grid search to determine hyperparameters. \n\nmy_model_8 = grid_search_xgb.best_estimator_\n\nmy_model_8.fit(X[0:1001], y[0:1001], early_stopping_rounds=5, eval_set=[(X[1001:1460], y[1001:1460])], verbose=False)\n\npreds8 = my_model_8.predict(X_test)\n\nframe8 = { 'Id': test_id, 'SalePrice': preds8 }\n  \nsubmission8 = pd.DataFrame(frame8)\n\nsubmission8.to_csv('sub8.csv', index=False)\n\n\n# error=0.15848. This is actually higher than our base estimator without the grid search.\n# This could indicate that there is some overfitting, so it might be an idea to do something like cross-validation over\n# a series of different parameter choices, though this would be quite computationally intensive. ","metadata":{"execution":{"iopub.status.busy":"2022-07-14T13:40:29.277325Z","iopub.execute_input":"2022-07-14T13:40:29.277871Z","iopub.status.idle":"2022-07-14T13:40:29.392809Z","shell.execute_reply.started":"2022-07-14T13:40:29.277835Z","shell.execute_reply":"2022-07-14T13:40:29.392271Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"So far, the best model seems to be model 7, i.e. the random forest with hyperparameters determined by grid search. Some further steps could be to use RandomizedSearchCV instead of grid search, since this automates the parameter search better. It would also be good to have a method for validation to test overfitting. Additionally, we could try using different models like a neural network. ","metadata":{}},{"cell_type":"code","source":"submission7.to_csv('sub7.csv', index=False)","metadata":{"execution":{"iopub.status.busy":"2022-07-14T13:56:08.669445Z","iopub.execute_input":"2022-07-14T13:56:08.669804Z","iopub.status.idle":"2022-07-14T13:56:08.680760Z","shell.execute_reply.started":"2022-07-14T13:56:08.669771Z","shell.execute_reply":"2022-07-14T13:56:08.679913Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"","metadata":{},"execution_count":null,"outputs":[]}]}