{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"pygments_lexer":"ipython3","nbconvert_exporter":"python","version":"3.6.4","file_extension":".py","codemirror_mode":{"name":"ipython","version":3},"name":"python","mimetype":"text/x-python"}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"# **Background**\n\nHouse Prices - Advanced Regression Techniques is a Kaggle competition. The goal is to determine, thanks to the data provided, the price of Ames in Iowa. To achieve this goal, I used feature extraction and advanced regression techniques. For more details about the background or the data, please refer to the [overview](https://www.kaggle.com/c/house-prices-advanced-regression-techniques/overview) of the project","metadata":{}},{"cell_type":"markdown","source":"# **Determine the problematic**","metadata":{}},{"cell_type":"markdown","source":"* **Task :** Determine thanks to the data what is the price of a house\n* **Performance metric :** RMSE of the log of the value predicted vs the real value\n* **Solution proposal :** A regression algorithm written with Scikit-learn","metadata":{}},{"cell_type":"markdown","source":"# **Inport**","metadata":{}},{"cell_type":"markdown","source":"## Working environment","metadata":{}},{"cell_type":"code","source":"import os\nfor dirname, _, filenames in os.walk('/kaggle/input'):\n    for filename in filenames:\n        print(os.path.join(dirname, filename))","metadata":{"execution":{"iopub.status.busy":"2022-07-08T13:18:16.567863Z","iopub.execute_input":"2022-07-08T13:18:16.568242Z","iopub.status.idle":"2022-07-08T13:18:16.598939Z","shell.execute_reply.started":"2022-07-08T13:18:16.568122Z","shell.execute_reply":"2022-07-08T13:18:16.598054Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"import numpy as np\nimport pandas as pd\nfrom pandas.plotting import scatter_matrix\nfrom scipy import stats\nfrom scipy.stats import norm, skew","metadata":{"execution":{"iopub.status.busy":"2022-07-08T13:18:16.699424Z","iopub.execute_input":"2022-07-08T13:18:16.699700Z","iopub.status.idle":"2022-07-08T13:18:17.521988Z","shell.execute_reply.started":"2022-07-08T13:18:16.699663Z","shell.execute_reply":"2022-07-08T13:18:17.521146Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"from sklearn.base import BaseEstimator, TransformerMixin\n\nfrom sklearn.preprocessing import FunctionTransformer\nfrom sklearn.preprocessing import OneHotEncoder, OrdinalEncoder\nfrom sklearn.preprocessing import LabelEncoder\nfrom sklearn.preprocessing import StandardScaler\nfrom sklearn.preprocessing import PolynomialFeatures\n\nfrom sklearn.impute import SimpleImputer\nfrom sklearn.pipeline import Pipeline\nfrom sklearn.compose import ColumnTransformer\n\nfrom sklearn.model_selection import GridSearchCV\nfrom sklearn.model_selection import train_test_split\nfrom sklearn.model_selection import cross_val_score\nfrom sklearn.metrics import mean_squared_error\n\nfrom sklearn.linear_model import SGDRegressor\nfrom sklearn.linear_model import LinearRegression\nfrom sklearn.linear_model import Ridge\nfrom sklearn.linear_model import ElasticNet\nfrom sklearn.tree import DecisionTreeRegressor\nfrom sklearn.ensemble import RandomForestRegressor\nfrom sklearn.svm import SVR\nfrom sklearn.ensemble import GradientBoostingRegressor\n\nfrom sklearn.metrics import mean_squared_error","metadata":{"execution":{"iopub.status.busy":"2022-07-08T13:18:17.523496Z","iopub.execute_input":"2022-07-08T13:18:17.523745Z","iopub.status.idle":"2022-07-08T13:18:17.939673Z","shell.execute_reply.started":"2022-07-08T13:18:17.523719Z","shell.execute_reply":"2022-07-08T13:18:17.939034Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"import sys\nsys.version_info","metadata":{"execution":{"iopub.status.busy":"2022-07-08T13:18:17.940719Z","iopub.execute_input":"2022-07-08T13:18:17.941205Z","iopub.status.idle":"2022-07-08T13:18:17.948494Z","shell.execute_reply.started":"2022-07-08T13:18:17.941138Z","shell.execute_reply":"2022-07-08T13:18:17.947842Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"import sklearn\nsklearn.__version__","metadata":{"execution":{"iopub.status.busy":"2022-07-08T13:18:17.950041Z","iopub.execute_input":"2022-07-08T13:18:17.950443Z","iopub.status.idle":"2022-07-08T13:18:17.966048Z","shell.execute_reply.started":"2022-07-08T13:18:17.950411Z","shell.execute_reply":"2022-07-08T13:18:17.965279Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"%matplotlib inline\nimport matplotlib as mpl\nimport matplotlib.pyplot as plt\nmpl.rc('axes', labelsize=14)\nmpl.rc('xtick', labelsize=12)\nmpl.rc('ytick', labelsize=12)\nimport seaborn as sns","metadata":{"execution":{"iopub.status.busy":"2022-07-08T13:18:17.967652Z","iopub.execute_input":"2022-07-08T13:18:17.968102Z","iopub.status.idle":"2022-07-08T13:18:18.130029Z","shell.execute_reply.started":"2022-07-08T13:18:17.968074Z","shell.execute_reply":"2022-07-08T13:18:18.129001Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## Data","metadata":{}},{"cell_type":"code","source":"train_df = pd.read_csv(\"/kaggle/input/house-prices-advanced-regression-techniques/train.csv\")\ntest_df = pd.read_csv(\"/kaggle/input/house-prices-advanced-regression-techniques/test.csv\")\ndf = train_df.append(test_df)","metadata":{"execution":{"iopub.status.busy":"2022-07-08T13:18:18.131308Z","iopub.execute_input":"2022-07-08T13:18:18.131886Z","iopub.status.idle":"2022-07-08T13:18:18.223049Z","shell.execute_reply.started":"2022-07-08T13:18:18.131846Z","shell.execute_reply":"2022-07-08T13:18:18.222178Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# **Explore**","metadata":{}},{"cell_type":"markdown","source":"## Overview","metadata":{}},{"cell_type":"code","source":"train_df.shape, test_df.shape, df.shape","metadata":{"execution":{"iopub.status.busy":"2022-07-08T13:18:18.224263Z","iopub.execute_input":"2022-07-08T13:18:18.224563Z","iopub.status.idle":"2022-07-08T13:18:18.230658Z","shell.execute_reply.started":"2022-07-08T13:18:18.224534Z","shell.execute_reply":"2022-07-08T13:18:18.229952Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"train_df.head()","metadata":{"execution":{"iopub.status.busy":"2022-07-08T13:18:18.231989Z","iopub.execute_input":"2022-07-08T13:18:18.232313Z","iopub.status.idle":"2022-07-08T13:18:18.270139Z","shell.execute_reply.started":"2022-07-08T13:18:18.232283Z","shell.execute_reply":"2022-07-08T13:18:18.269263Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"test_df.head()","metadata":{"execution":{"iopub.status.busy":"2022-07-08T13:18:18.271655Z","iopub.execute_input":"2022-07-08T13:18:18.272101Z","iopub.status.idle":"2022-07-08T13:18:18.295238Z","shell.execute_reply.started":"2022-07-08T13:18:18.272065Z","shell.execute_reply":"2022-07-08T13:18:18.294268Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"SalePrice is our column labels","metadata":{}},{"cell_type":"code","source":"df.info()","metadata":{"execution":{"iopub.status.busy":"2022-07-08T13:18:18.298260Z","iopub.execute_input":"2022-07-08T13:18:18.298686Z","iopub.status.idle":"2022-07-08T13:18:18.341547Z","shell.execute_reply.started":"2022-07-08T13:18:18.298652Z","shell.execute_reply":"2022-07-08T13:18:18.340446Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"**Missing values** will be a big part of preprocessing data","metadata":{}},{"cell_type":"code","source":"# Manually plot every feature\ncolumn_id = 78\ndf.iloc[:, column_id].hist(bins=50, figsize=(8,5))\nplt.show()\ndf.iloc[:, column_id].name","metadata":{"execution":{"iopub.status.busy":"2022-07-08T13:18:18.342885Z","iopub.execute_input":"2022-07-08T13:18:18.343101Z","iopub.status.idle":"2022-07-08T13:18:18.767826Z","shell.execute_reply.started":"2022-07-08T13:18:18.343077Z","shell.execute_reply":"2022-07-08T13:18:18.766784Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"**General notes :**\n - A lot of irrelevant features -> Feature extraction to do\n - Numerical features which are categorical features\n - Numerical features with a lot of 0 -> add extra relevant features\n - Skewed features -> use log\n - Some outliers are presents -> remove them after feature extraction","metadata":{}},{"cell_type":"markdown","source":"## Numerical features","metadata":{}},{"cell_type":"code","source":"df.describe()","metadata":{"execution":{"iopub.status.busy":"2022-07-08T13:18:18.770920Z","iopub.execute_input":"2022-07-08T13:18:18.771173Z","iopub.status.idle":"2022-07-08T13:18:18.867429Z","shell.execute_reply.started":"2022-07-08T13:18:18.771135Z","shell.execute_reply":"2022-07-08T13:18:18.866590Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"df.hist(bins=50, figsize=(20,15))\nplt.show()","metadata":{"execution":{"iopub.status.busy":"2022-07-08T13:18:18.868512Z","iopub.execute_input":"2022-07-08T13:18:18.868803Z","iopub.status.idle":"2022-07-08T13:18:25.637430Z","shell.execute_reply.started":"2022-07-08T13:18:18.868768Z","shell.execute_reply":"2022-07-08T13:18:25.636291Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## Categorical features","metadata":{}},{"cell_type":"code","source":"df.describe(include=['O'])","metadata":{"execution":{"iopub.status.busy":"2022-07-08T13:18:25.639144Z","iopub.execute_input":"2022-07-08T13:18:25.639498Z","iopub.status.idle":"2022-07-08T13:18:25.729289Z","shell.execute_reply.started":"2022-07-08T13:18:25.639456Z","shell.execute_reply":"2022-07-08T13:18:25.728189Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## Skewed features","metadata":{}},{"cell_type":"code","source":"# List real numerical features\nreal_num_features = []\nfor column in df.columns:\n    if (df[column].drop_duplicates()).shape[0]>max(df.describe(include=['O']).loc[\"unique\"]):\n        # Assuming than we don't have more than 25 categories in categorical feature (and more than 25 for num feature)\n        real_num_features.append(df[column].name)\n# Adding PoolArea manually because two many missing values\nreal_num_features.append(\"PoolArea\")\nreal_num_features","metadata":{"execution":{"iopub.status.busy":"2022-07-08T13:18:25.730423Z","iopub.execute_input":"2022-07-08T13:18:25.730656Z","iopub.status.idle":"2022-07-08T13:18:30.842603Z","shell.execute_reply.started":"2022-07-08T13:18:25.730632Z","shell.execute_reply":"2022-07-08T13:18:30.841524Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Check the skew of all numerical features\nfeatures_skewness = df[real_num_features].skew().sort_values(ascending=False)\nfeatures_skewness","metadata":{"execution":{"iopub.status.busy":"2022-07-08T13:18:30.845503Z","iopub.execute_input":"2022-07-08T13:18:30.845984Z","iopub.status.idle":"2022-07-08T13:18:30.860176Z","shell.execute_reply.started":"2022-07-08T13:18:30.845949Z","shell.execute_reply":"2022-07-08T13:18:30.858929Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Transform all features with abs(skewness)>0.7 in 1+log to normalize them\nskewed_features = list(features_skewness[abs(features_skewness)>0.7].index)\nskewed_features","metadata":{"execution":{"iopub.status.busy":"2022-07-08T13:18:30.861620Z","iopub.execute_input":"2022-07-08T13:18:30.862056Z","iopub.status.idle":"2022-07-08T13:18:30.881189Z","shell.execute_reply.started":"2022-07-08T13:18:30.862000Z","shell.execute_reply":"2022-07-08T13:18:30.880428Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## Missing values","metadata":{}},{"cell_type":"code","source":"def missing_values(df):\n    return (df.loc[:, df.isnull().any()].isna().sum()).sort_values()\nmissing_values(train_df),missing_values(df)","metadata":{"execution":{"iopub.status.busy":"2022-07-08T13:18:30.882217Z","iopub.execute_input":"2022-07-08T13:18:30.883038Z","iopub.status.idle":"2022-07-08T13:18:30.919043Z","shell.execute_reply.started":"2022-07-08T13:18:30.882997Z","shell.execute_reply":"2022-07-08T13:18:30.918149Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"**For categorical features**\n - Some features have a NA category, and NaN values must be placed in NA category\n - Some features have a NA category, but NaN values must not be placed in NA category\n - Some features have not NA category","metadata":{}},{"cell_type":"markdown","source":"**For numerical features**\n - Some features have to be filled by 0\n - In other case, the information is just missing","metadata":{}},{"cell_type":"markdown","source":"**Strategy adopted :** We have two choices\n - Doing a feature reduction : easier and faster but we'll have more errors\n - Not doing this : must fill all missing values not present in the train set (option selected)","metadata":{}},{"cell_type":"code","source":"# Plot each feature with missing values to see how to fill missing values\ndf[\"PoolQC\"].hist(bins=50, figsize=(8,5))\nplt.show()","metadata":{"execution":{"iopub.status.busy":"2022-07-08T13:18:30.920428Z","iopub.execute_input":"2022-07-08T13:18:30.920956Z","iopub.status.idle":"2022-07-08T13:18:31.433080Z","shell.execute_reply.started":"2022-07-08T13:18:30.920915Z","shell.execute_reply":"2022-07-08T13:18:31.432119Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":" - **Fill with most common value or with NA :** (see the documentation of the data for more details)[\"Electrical\",\"GarageCars\",\"Exterior1st\",\"Exterior2nd\",\"KitchenQual\",\"SaleType\",\"BsmtUnfSF\",\"Utilities\",\"Functional\",\"BsmtHalfBath\",\"BsmtFullBath\",\"MSZoning\",\"MasVnrType\",\"BsmtFinType1\",\"BsmtFinType2\",\"BsmtQual\",\"BsmtExposure\",\"BsmtCond\",\"GarageType\",\"GarageCond\",\"GarageQual\",\"GarageFinish\",\"FireplaceQu\"]\n - **Fill with the median :** [\"GarageArea\",\"TotalBsmtSF\",\"GarageYrBlt\",\"LotFrontage\",\"Fence\",\"Alley\",\"MiscFeature\",\"PoolQC\"]\n - **Fill with 0 :**[\"BsmtFinSF1\",\"BsmtFinSF2\",\"MasVnrArea\"]","metadata":{}},{"cell_type":"markdown","source":"**PoolQC**","metadata":{}},{"cell_type":"code","source":"condition = train_df[\"PoolArea\"]==0\nlen(train_df[condition])","metadata":{"execution":{"iopub.status.busy":"2022-07-08T13:18:31.435370Z","iopub.execute_input":"2022-07-08T13:18:31.435894Z","iopub.status.idle":"2022-07-08T13:18:31.443777Z","shell.execute_reply.started":"2022-07-08T13:18:31.435841Z","shell.execute_reply":"2022-07-08T13:18:31.443112Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Compared PoolQC to PoolArea : no real missing value in the training set. Fill  NaN by NA","metadata":{}},{"cell_type":"markdown","source":"**FireplaceQu**","metadata":{}},{"cell_type":"code","source":"condition = df[\"Fireplaces\"]==0\nlen(df[\"Fireplaces\"][condition])","metadata":{"execution":{"iopub.status.busy":"2022-07-08T13:18:31.444844Z","iopub.execute_input":"2022-07-08T13:18:31.445251Z","iopub.status.idle":"2022-07-08T13:18:31.459289Z","shell.execute_reply.started":"2022-07-08T13:18:31.445220Z","shell.execute_reply":"2022-07-08T13:18:31.458292Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Compared Fireplaces and FireplaceQu : no value really missing. Fill NaN by NA","metadata":{}},{"cell_type":"markdown","source":"**LotFrontage**","metadata":{}},{"cell_type":"code","source":"corr_matrix = df.corr()","metadata":{"execution":{"iopub.status.busy":"2022-07-08T13:18:31.460735Z","iopub.execute_input":"2022-07-08T13:18:31.461279Z","iopub.status.idle":"2022-07-08T13:18:31.489417Z","shell.execute_reply.started":"2022-07-08T13:18:31.461222Z","shell.execute_reply":"2022-07-08T13:18:31.488682Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"corr_series_lot_frontage = corr_matrix[\"LotFrontage\"].sort_values(ascending=False)\ncorr_series_lot_frontage","metadata":{"execution":{"iopub.status.busy":"2022-07-08T13:18:31.490997Z","iopub.execute_input":"2022-07-08T13:18:31.491313Z","iopub.status.idle":"2022-07-08T13:18:31.500487Z","shell.execute_reply.started":"2022-07-08T13:18:31.491277Z","shell.execute_reply":"2022-07-08T13:18:31.499569Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Can be predicted with a regression model. Easier to fill with median (seems to have street access for all houses because of street feature). Need also a 1+log transformation and remove outliers","metadata":{}},{"cell_type":"markdown","source":"**Garages**","metadata":{}},{"cell_type":"markdown","source":"Seems to have two real missing values in the test set. Fill every missing values by NA, and GarageYrBuilt by his median value (note there is one outlier in the test set)","metadata":{}},{"cell_type":"markdown","source":"**Bsmt**","metadata":{}},{"cell_type":"code","source":"isna_bsmt_qual = train_df[\"BsmtQual\"].isna()\nisna_bsmt_exposure = train_df[\"BsmtExposure\"].isna()\nreal_missing_bsmt_exposure = list(set(train_df[\"BsmtExposure\"][isna_bsmt_exposure].index) - set(train_df[\"BsmtQual\"][isna_bsmt_qual].index))\nreal_missing_bsmt_exposure","metadata":{"execution":{"iopub.status.busy":"2022-07-08T13:18:31.502210Z","iopub.execute_input":"2022-07-08T13:18:31.502774Z","iopub.status.idle":"2022-07-08T13:18:31.520009Z","shell.execute_reply.started":"2022-07-08T13:18:31.502736Z","shell.execute_reply":"2022-07-08T13:18:31.519042Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Seems to have only one real missing value in the training_set. We'll fill this row and fill the rest by NA","metadata":{}},{"cell_type":"markdown","source":"**MasVnr**","metadata":{}},{"cell_type":"code","source":"isna_masvnr_type = train_df[\"MasVnrType\"].isna()\nisna_masvnr_area = train_df[\"MasVnrArea\"].isna()\nprint(train_df[\"MasVnrArea\"][isna_masvnr_area].index,train_df[\"MasVnrType\"][isna_masvnr_type].index)\nreal_missing_masvnr = list(train_df[\"MasVnrArea\"][isna_masvnr_area].index)","metadata":{"execution":{"iopub.status.busy":"2022-07-08T13:18:31.521540Z","iopub.execute_input":"2022-07-08T13:18:31.521883Z","iopub.status.idle":"2022-07-08T13:18:31.533404Z","shell.execute_reply.started":"2022-07-08T13:18:31.521846Z","shell.execute_reply":"2022-07-08T13:18:31.532398Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Seems to be real missing values. For testing set, we'll fill by median (i.e. 0) and by most common (i.e. None) idem for Electrical.","metadata":{}},{"cell_type":"markdown","source":"**Electrical**","metadata":{}},{"cell_type":"code","source":"condition = train_df[\"Electrical\"].isna()\nreal_missing_electrical = list(train_df[condition].index)","metadata":{"execution":{"iopub.status.busy":"2022-07-08T13:18:31.537948Z","iopub.execute_input":"2022-07-08T13:18:31.538883Z","iopub.status.idle":"2022-07-08T13:18:31.551999Z","shell.execute_reply.started":"2022-07-08T13:18:31.538846Z","shell.execute_reply":"2022-07-08T13:18:31.551087Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# **Fill missing values**","metadata":{}},{"cell_type":"code","source":"class DataFrameSelector(BaseEstimator, TransformerMixin):\n    def __init__(self, attribute_names):\n        self.attribute_names = attribute_names\n    def fit(self, X, y=None):\n        return self\n    def transform(self, X):\n        return X[self.attribute_names]","metadata":{"execution":{"iopub.status.busy":"2022-07-08T13:18:31.553540Z","iopub.execute_input":"2022-07-08T13:18:31.554074Z","iopub.status.idle":"2022-07-08T13:18:31.564961Z","shell.execute_reply.started":"2022-07-08T13:18:31.554034Z","shell.execute_reply":"2022-07-08T13:18:31.564269Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"def fill_missing_values(features_name,strategy,fill_value):\n    return Pipeline([\n        (\"select_cat\", DataFrameSelector(features_name)),\n        (\"replace_nan\", SimpleImputer(strategy=strategy,fill_value=fill_value)),\n    ])","metadata":{"execution":{"iopub.status.busy":"2022-07-08T13:18:31.566209Z","iopub.execute_input":"2022-07-08T13:18:31.566910Z","iopub.status.idle":"2022-07-08T13:18:31.580450Z","shell.execute_reply.started":"2022-07-08T13:18:31.566856Z","shell.execute_reply":"2022-07-08T13:18:31.579576Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"**NaN replaces by NA**","metadata":{}},{"cell_type":"code","source":"fill_nan_by_na_features = [\"BsmtFinType1\",\"BsmtFinType2\",\"BsmtQual\",\"BsmtExposure\",\"BsmtCond\",\"GarageType\",\"GarageCond\",\"GarageQual\",\"GarageFinish\",\"FireplaceQu\",\n                           \"Fence\",\"Alley\",\"MiscFeature\",\"PoolQC\"]\nlen(fill_nan_by_na_features)","metadata":{"execution":{"iopub.status.busy":"2022-07-08T13:18:31.581748Z","iopub.execute_input":"2022-07-08T13:18:31.582067Z","iopub.status.idle":"2022-07-08T13:18:31.596925Z","shell.execute_reply.started":"2022-07-08T13:18:31.582035Z","shell.execute_reply":"2022-07-08T13:18:31.595760Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"fill_nan_by_na_pipeline = fill_missing_values(fill_nan_by_na_features,\"constant\",\"NA\")","metadata":{"execution":{"iopub.status.busy":"2022-07-08T13:18:31.598012Z","iopub.execute_input":"2022-07-08T13:18:31.598747Z","iopub.status.idle":"2022-07-08T13:18:31.611657Z","shell.execute_reply.started":"2022-07-08T13:18:31.598712Z","shell.execute_reply":"2022-07-08T13:18:31.610696Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"**NaN replaces by most_common**","metadata":{}},{"cell_type":"code","source":"fill_nan_by_common_features = [\"Electrical\",\"GarageCars\",\"Exterior1st\",\"Exterior2nd\",\"KitchenQual\",\"SaleType\",\"BsmtUnfSF\",\"Utilities\",\"Functional\",\n                               \"BsmtHalfBath\",\"BsmtFullBath\",\"MSZoning\",\"MasVnrType\"]\nlen(fill_nan_by_common_features)","metadata":{"execution":{"iopub.status.busy":"2022-07-08T13:18:31.612723Z","iopub.execute_input":"2022-07-08T13:18:31.613286Z","iopub.status.idle":"2022-07-08T13:18:31.629492Z","shell.execute_reply.started":"2022-07-08T13:18:31.613252Z","shell.execute_reply":"2022-07-08T13:18:31.628655Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"fill_nan_by_common_pipeline = fill_missing_values(fill_nan_by_common_features,\"most_frequent\",\"most_frequent\")","metadata":{"execution":{"iopub.status.busy":"2022-07-08T13:18:31.630553Z","iopub.execute_input":"2022-07-08T13:18:31.631269Z","iopub.status.idle":"2022-07-08T13:18:31.642958Z","shell.execute_reply.started":"2022-07-08T13:18:31.631237Z","shell.execute_reply":"2022-07-08T13:18:31.642127Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"**NaN replaces by median**","metadata":{}},{"cell_type":"code","source":"fill_nan_by_median_features = [\"GarageArea\",\"TotalBsmtSF\",\"GarageYrBlt\",\"LotFrontage\",]\nlen(fill_nan_by_median_features)","metadata":{"execution":{"iopub.status.busy":"2022-07-08T13:18:31.644422Z","iopub.execute_input":"2022-07-08T13:18:31.644671Z","iopub.status.idle":"2022-07-08T13:18:31.660150Z","shell.execute_reply.started":"2022-07-08T13:18:31.644638Z","shell.execute_reply":"2022-07-08T13:18:31.659308Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"fill_nan_by_median_pipeline = fill_missing_values(fill_nan_by_median_features,\"median\",\"median\")","metadata":{"execution":{"iopub.status.busy":"2022-07-08T13:18:31.661849Z","iopub.execute_input":"2022-07-08T13:18:31.662290Z","iopub.status.idle":"2022-07-08T13:18:31.675065Z","shell.execute_reply.started":"2022-07-08T13:18:31.662260Z","shell.execute_reply":"2022-07-08T13:18:31.673997Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"**NaN replaces by 0**","metadata":{}},{"cell_type":"code","source":"fill_nan_by_0_features = [\"BsmtFinSF1\",\"BsmtFinSF2\",\"MasVnrArea\"]\nlen(fill_nan_by_0_features)","metadata":{"execution":{"iopub.status.busy":"2022-07-08T13:18:31.676628Z","iopub.execute_input":"2022-07-08T13:18:31.677047Z","iopub.status.idle":"2022-07-08T13:18:31.692367Z","shell.execute_reply.started":"2022-07-08T13:18:31.677009Z","shell.execute_reply":"2022-07-08T13:18:31.691112Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"fill_nan_by_0_pipeline = fill_missing_values(fill_nan_by_0_features,\"constant\",0)","metadata":{"execution":{"iopub.status.busy":"2022-07-08T13:18:31.693731Z","iopub.execute_input":"2022-07-08T13:18:31.696423Z","iopub.status.idle":"2022-07-08T13:18:31.704408Z","shell.execute_reply.started":"2022-07-08T13:18:31.696374Z","shell.execute_reply":"2022-07-08T13:18:31.703566Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"**Not NaN features (drop the label)**","metadata":{}},{"cell_type":"code","source":"class DataFrameDrop(BaseEstimator, TransformerMixin):\n    def __init__(self, attribute_names):\n        self.attribute_names = attribute_names\n    def fit(self, X, y=None):\n        return self\n    def transform(self, X):\n        X.drop(self.attribute_names, axis=1, inplace=True)\n        return X","metadata":{"execution":{"iopub.status.busy":"2022-07-08T13:18:31.706108Z","iopub.execute_input":"2022-07-08T13:18:31.706640Z","iopub.status.idle":"2022-07-08T13:18:31.725623Z","shell.execute_reply.started":"2022-07-08T13:18:31.706607Z","shell.execute_reply":"2022-07-08T13:18:31.724471Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"not_nan_list = list(set(df.columns) - set(fill_nan_by_na_features) - set(fill_nan_by_common_features) - set(fill_nan_by_median_features) - set(fill_nan_by_0_features))\nlen(not_nan_list)","metadata":{"execution":{"iopub.status.busy":"2022-07-08T13:18:31.726978Z","iopub.execute_input":"2022-07-08T13:18:31.727360Z","iopub.status.idle":"2022-07-08T13:18:31.742671Z","shell.execute_reply.started":"2022-07-08T13:18:31.727240Z","shell.execute_reply":"2022-07-08T13:18:31.741577Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"not_nan_pipeline = Pipeline([(\"select_cat\", DataFrameSelector(not_nan_list)),\n                            (\"drop\", DataFrameDrop([\"SalePrice\"])), ])","metadata":{"execution":{"iopub.status.busy":"2022-07-08T13:18:31.744018Z","iopub.execute_input":"2022-07-08T13:18:31.744559Z","iopub.status.idle":"2022-07-08T13:18:31.755443Z","shell.execute_reply.started":"2022-07-08T13:18:31.744518Z","shell.execute_reply":"2022-07-08T13:18:31.754726Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"**Gather pipelines**","metadata":{}},{"cell_type":"code","source":"nan_pipeline = ColumnTransformer([\n        (\"num_median_nan\", fill_nan_by_median_pipeline, fill_nan_by_median_features),\n        (\"num_0_nan\", fill_nan_by_0_pipeline, fill_nan_by_0_features),\n        (\"cat_na_nan\", fill_nan_by_na_pipeline, fill_nan_by_na_features),\n        (\"cat_common_nan\", fill_nan_by_common_pipeline, fill_nan_by_common_features),\n        (\"not_nan\", not_nan_pipeline, not_nan_list)\n    ])","metadata":{"execution":{"iopub.status.busy":"2022-07-08T13:18:31.756601Z","iopub.execute_input":"2022-07-08T13:18:31.757198Z","iopub.status.idle":"2022-07-08T13:18:31.772504Z","shell.execute_reply.started":"2022-07-08T13:18:31.757138Z","shell.execute_reply":"2022-07-08T13:18:31.771605Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"nan_pipeline.fit_transform(df).shape","metadata":{"execution":{"iopub.status.busy":"2022-07-08T13:18:31.773930Z","iopub.execute_input":"2022-07-08T13:18:31.774197Z","iopub.status.idle":"2022-07-08T13:18:31.865356Z","shell.execute_reply.started":"2022-07-08T13:18:31.774145Z","shell.execute_reply":"2022-07-08T13:18:31.864496Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"fill_nan_by_median_features","metadata":{"execution":{"iopub.status.busy":"2022-07-08T13:18:31.866705Z","iopub.execute_input":"2022-07-08T13:18:31.867641Z","iopub.status.idle":"2022-07-08T13:18:31.873492Z","shell.execute_reply.started":"2022-07-08T13:18:31.867607Z","shell.execute_reply":"2022-07-08T13:18:31.872433Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"ordered_features = fill_nan_by_median_features + fill_nan_by_0_features + fill_nan_by_na_features + fill_nan_by_common_features + not_nan_list\nordered_features.remove(\"SalePrice\")\nprint(len(ordered_features))\nall_data_filled = pd.DataFrame(nan_pipeline.fit_transform(df),columns=ordered_features)\nall_data_filled.shape","metadata":{"execution":{"iopub.status.busy":"2022-07-08T13:18:31.874737Z","iopub.execute_input":"2022-07-08T13:18:31.875265Z","iopub.status.idle":"2022-07-08T13:18:31.965642Z","shell.execute_reply.started":"2022-07-08T13:18:31.875222Z","shell.execute_reply":"2022-07-08T13:18:31.964663Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Check if all missing values have been filled\nmissing_values = (all_data_filled.loc[:, all_data_filled.isnull().any()].isna().sum())\nmissing_values","metadata":{"execution":{"iopub.status.busy":"2022-07-08T13:18:31.967377Z","iopub.execute_input":"2022-07-08T13:18:31.967698Z","iopub.status.idle":"2022-07-08T13:18:31.989140Z","shell.execute_reply.started":"2022-07-08T13:18:31.967658Z","shell.execute_reply":"2022-07-08T13:18:31.988307Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"columns_except_label = list(df.columns)\ncolumns_except_label.remove(\"SalePrice\")","metadata":{"execution":{"iopub.status.busy":"2022-07-08T13:18:31.990477Z","iopub.execute_input":"2022-07-08T13:18:31.991316Z","iopub.status.idle":"2022-07-08T13:18:32.002764Z","shell.execute_reply.started":"2022-07-08T13:18:31.991270Z","shell.execute_reply":"2022-07-08T13:18:32.001299Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"train_df_filled = pd.DataFrame(all_data_filled[:1460], columns = columns_except_label)\ntest_df_filled = pd.DataFrame(all_data_filled[1460:], columns = columns_except_label)\ndf_filled = train_df_filled.append(test_df_filled)","metadata":{"execution":{"iopub.status.busy":"2022-07-08T13:18:32.004498Z","iopub.execute_input":"2022-07-08T13:18:32.005152Z","iopub.status.idle":"2022-07-08T13:18:32.023891Z","shell.execute_reply.started":"2022-07-08T13:18:32.005069Z","shell.execute_reply":"2022-07-08T13:18:32.022841Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"df_filled.head()","metadata":{"execution":{"iopub.status.busy":"2022-07-08T13:18:32.026103Z","iopub.execute_input":"2022-07-08T13:18:32.026454Z","iopub.status.idle":"2022-07-08T13:18:32.050697Z","shell.execute_reply.started":"2022-07-08T13:18:32.026413Z","shell.execute_reply":"2022-07-08T13:18:32.049552Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Transform","metadata":{}},{"cell_type":"markdown","source":"**Normalize skewed features** before continuing the vizualisation of the data (from the train_df)","metadata":{}},{"cell_type":"code","source":"class DataFrameLog(BaseEstimator, TransformerMixin):\n    def __init__(self, attribute_names):\n        self.attribute_names = attribute_names\n    def fit(self, X, y=None):\n        return self\n    def transform(self, X):\n        return np.log(X[self.attribute_names].astype('float32') + 1)","metadata":{"execution":{"iopub.status.busy":"2022-07-08T13:18:32.052121Z","iopub.execute_input":"2022-07-08T13:18:32.052614Z","iopub.status.idle":"2022-07-08T13:18:32.058050Z","shell.execute_reply.started":"2022-07-08T13:18:32.052570Z","shell.execute_reply":"2022-07-08T13:18:32.057185Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"skewed_features.remove(\"SalePrice\")","metadata":{"execution":{"iopub.status.busy":"2022-07-08T13:18:32.059372Z","iopub.execute_input":"2022-07-08T13:18:32.059611Z","iopub.status.idle":"2022-07-08T13:18:32.077284Z","shell.execute_reply.started":"2022-07-08T13:18:32.059587Z","shell.execute_reply":"2022-07-08T13:18:32.076615Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"skewed_pipeline = Pipeline([\n                            (\"log\", DataFrameLog(skewed_features)),\n                            ])","metadata":{"execution":{"iopub.status.busy":"2022-07-08T13:18:32.080744Z","iopub.execute_input":"2022-07-08T13:18:32.081002Z","iopub.status.idle":"2022-07-08T13:18:32.090367Z","shell.execute_reply.started":"2022-07-08T13:18:32.080969Z","shell.execute_reply":"2022-07-08T13:18:32.089266Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"test_df_filled[skewed_features].describe()","metadata":{"execution":{"iopub.status.busy":"2022-07-08T13:18:32.091419Z","iopub.execute_input":"2022-07-08T13:18:32.091766Z","iopub.status.idle":"2022-07-08T13:18:32.151124Z","shell.execute_reply.started":"2022-07-08T13:18:32.091719Z","shell.execute_reply":"2022-07-08T13:18:32.150211Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"skew_df = skewed_pipeline.fit_transform(df_filled)","metadata":{"execution":{"iopub.status.busy":"2022-07-08T13:18:32.152405Z","iopub.execute_input":"2022-07-08T13:18:32.152635Z","iopub.status.idle":"2022-07-08T13:18:32.161794Z","shell.execute_reply.started":"2022-07-08T13:18:32.152610Z","shell.execute_reply":"2022-07-08T13:18:32.160849Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"skew_df","metadata":{"execution":{"iopub.status.busy":"2022-07-08T13:18:32.163181Z","iopub.execute_input":"2022-07-08T13:18:32.163726Z","iopub.status.idle":"2022-07-08T13:18:32.203015Z","shell.execute_reply.started":"2022-07-08T13:18:32.163684Z","shell.execute_reply":"2022-07-08T13:18:32.202140Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"skew_df.hist(bins=50, figsize=(20,15))\nplt.show()","metadata":{"execution":{"iopub.status.busy":"2022-07-08T13:18:32.204601Z","iopub.execute_input":"2022-07-08T13:18:32.204921Z","iopub.status.idle":"2022-07-08T13:18:35.755627Z","shell.execute_reply.started":"2022-07-08T13:18:32.204882Z","shell.execute_reply":"2022-07-08T13:18:35.754610Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Plot the result\nfor feature in skew_df:\n    sns.distplot(skew_df[feature], fit=norm) \n\n    # Get the fitted parameters used by the function\n    (mu, sigma) = norm.fit(skew_df[feature])\n    print( '\\n mu = {:.2f} and sigma = {:.2f}\\n'.format(mu, sigma))\n\n    #Now plot the distribution\n    plt.legend(['Normal dist. ($\\mu=$ {:.2f} and $\\sigma=$ {:.2f} )'.format(mu, sigma)],\n                loc='best')\n    plt.ylabel('Frequency')\n    plt.title('Column distribution')\n\n    #Get also the QQ-plot\n    fig = plt.figure()\n    res = stats.probplot(skew_df[feature], plot=plt)\n    plt.show()","metadata":{"execution":{"iopub.status.busy":"2022-07-08T13:18:35.757371Z","iopub.execute_input":"2022-07-08T13:18:35.757750Z","iopub.status.idle":"2022-07-08T13:18:45.649696Z","shell.execute_reply.started":"2022-07-08T13:18:35.757712Z","shell.execute_reply":"2022-07-08T13:18:45.649152Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"df_norm = df_filled\ndf_norm[skewed_features] = skew_df\ndf_norm","metadata":{"execution":{"iopub.status.busy":"2022-07-08T13:18:45.655335Z","iopub.execute_input":"2022-07-08T13:18:45.655745Z","iopub.status.idle":"2022-07-08T13:18:45.715386Z","shell.execute_reply.started":"2022-07-08T13:18:45.655713Z","shell.execute_reply":"2022-07-08T13:18:45.714264Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"train_df_norm = df_norm[:1460]\ntest_df_norm = df_norm[1460:]","metadata":{"execution":{"iopub.status.busy":"2022-07-08T13:18:45.716824Z","iopub.execute_input":"2022-07-08T13:18:45.717099Z","iopub.status.idle":"2022-07-08T13:18:45.721886Z","shell.execute_reply.started":"2022-07-08T13:18:45.717062Z","shell.execute_reply":"2022-07-08T13:18:45.721313Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## Outliers","metadata":{}},{"cell_type":"code","source":"scatter_matrix(df_norm, figsize=(50, 50))","metadata":{"execution":{"iopub.status.busy":"2022-07-08T13:18:45.722791Z","iopub.execute_input":"2022-07-08T13:18:45.723049Z","iopub.status.idle":"2022-07-08T13:19:10.961980Z","shell.execute_reply.started":"2022-07-08T13:18:45.723024Z","shell.execute_reply":"2022-07-08T13:19:10.961301Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## Correlations","metadata":{}},{"cell_type":"code","source":"corr_matrix = df_norm.corr()\nplt.subplots(figsize=(12,9))\nsns.heatmap(corr_matrix, vmax=0.9, square=True)","metadata":{"execution":{"iopub.status.busy":"2022-07-08T13:19:10.963131Z","iopub.execute_input":"2022-07-08T13:19:10.963678Z","iopub.status.idle":"2022-07-08T13:19:11.838932Z","shell.execute_reply.started":"2022-07-08T13:19:10.963645Z","shell.execute_reply":"2022-07-08T13:19:11.838016Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"- 1stFlrSF and 2ndFlrSF are strongly correlated\n- Idem for GarageCars and GarageAreas\n- Idem for GarageYrBlt and YearBuilt\n- Idem for TotRmsAbvGrd and GrLivArea\n- OverallQual and GrLivArea seems to be the best features to describe SalePrice (with linear correlation)\n- Some features are not present because they are categorical","metadata":{}},{"cell_type":"code","source":"corr_matrix = df.corr()\ncorr_series = corr_matrix[\"SalePrice\"].sort_values(ascending=False)\ncorr_series","metadata":{"execution":{"iopub.status.busy":"2022-07-08T13:19:11.840352Z","iopub.execute_input":"2022-07-08T13:19:11.840720Z","iopub.status.idle":"2022-07-08T13:19:11.862421Z","shell.execute_reply.started":"2022-07-08T13:19:11.840682Z","shell.execute_reply":"2022-07-08T13:19:11.861857Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# **Categorical features**","metadata":{}},{"cell_type":"code","source":"# List categorical features\nfalse_num_features = set(corr_series.index) - set(real_num_features)\nprint(false_num_features)\n\ncat_features = list(set(df_filled.columns) - set(real_num_features))\nlen(cat_features)","metadata":{"execution":{"iopub.status.busy":"2022-07-08T13:19:11.863515Z","iopub.execute_input":"2022-07-08T13:19:11.863885Z","iopub.status.idle":"2022-07-08T13:19:11.871783Z","shell.execute_reply.started":"2022-07-08T13:19:11.863860Z","shell.execute_reply":"2022-07-08T13:19:11.870864Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"len(real_num_features)","metadata":{"execution":{"iopub.status.busy":"2022-07-08T13:19:11.872924Z","iopub.execute_input":"2022-07-08T13:19:11.873134Z","iopub.status.idle":"2022-07-08T13:19:11.885017Z","shell.execute_reply.started":"2022-07-08T13:19:11.873113Z","shell.execute_reply":"2022-07-08T13:19:11.884369Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Change \"numerical\" features into categorical one\ndef num_to_cat(df):\n    return df.astype(str)\n\nnum_to_cat_pipeline = Pipeline([(\"select\", DataFrameSelector(false_num_features)),\n                                (\"num_to_cat\", FunctionTransformer(num_to_cat)), ])","metadata":{"execution":{"iopub.status.busy":"2022-07-08T13:19:11.886246Z","iopub.execute_input":"2022-07-08T13:19:11.886562Z","iopub.status.idle":"2022-07-08T13:19:11.898487Z","shell.execute_reply.started":"2022-07-08T13:19:11.886536Z","shell.execute_reply":"2022-07-08T13:19:11.897740Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"false_num_features","metadata":{"execution":{"iopub.status.busy":"2022-07-08T13:19:11.900347Z","iopub.execute_input":"2022-07-08T13:19:11.900995Z","iopub.status.idle":"2022-07-08T13:19:11.916614Z","shell.execute_reply.started":"2022-07-08T13:19:11.900965Z","shell.execute_reply":"2022-07-08T13:19:11.915741Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"num_to_cat_df = pd.DataFrame(num_to_cat_pipeline.fit_transform(df_filled[false_num_features]), columns = false_num_features)\nnum_to_cat_df.describe()","metadata":{"execution":{"iopub.status.busy":"2022-07-08T13:19:11.917822Z","iopub.execute_input":"2022-07-08T13:19:11.918269Z","iopub.status.idle":"2022-07-08T13:19:11.979730Z","shell.execute_reply.started":"2022-07-08T13:19:11.918242Z","shell.execute_reply":"2022-07-08T13:19:11.978831Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Gathered categorical features which have an order in their categories (manually)\ncat_order_features = [\"LotShape\",\"LandContour\",\"Utilities\",\"LotConfig\",\"LandSlope\",\"BldgType\",\"HouseStyle\",\"OverallQual\",\"OverallCond\",\"ExterQual\",\"ExterCond\",\n                      \"BsmtQual\",\"BsmtCond\",\"BsmtExposure\",\"BsmtFinType1\",\"BsmtFinType2\",\"HeatingQC\",\"Electrical\",\"KitchenQual\",\"Functional\",\"FireplaceQu\",\n                      \"GarageFinish\",\"GarageQual\",\"GarageCond\",\"PavedDrive\",\"PoolQC\",\"Fence\"]\n\ncat_order_pipeline = Pipeline([(\"select\", DataFrameSelector(cat_order_features)),\n                               (\"encoding\", OrdinalEncoder()), ])","metadata":{"execution":{"iopub.status.busy":"2022-07-08T13:19:11.980933Z","iopub.execute_input":"2022-07-08T13:19:11.981189Z","iopub.status.idle":"2022-07-08T13:19:11.987266Z","shell.execute_reply.started":"2022-07-08T13:19:11.981138Z","shell.execute_reply":"2022-07-08T13:19:11.986478Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"cat_order_df = pd.DataFrame(cat_order_pipeline.fit_transform(df_filled[cat_order_features]), columns = cat_order_features)\ncat_order_df.describe()","metadata":{"execution":{"iopub.status.busy":"2022-07-08T13:19:11.988427Z","iopub.execute_input":"2022-07-08T13:19:11.988976Z","iopub.status.idle":"2022-07-08T13:19:12.101110Z","shell.execute_reply.started":"2022-07-08T13:19:11.988950Z","shell.execute_reply":"2022-07-08T13:19:12.100213Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"cat_not_order_features = list(set(cat_features) - set(cat_order_features))\n\ncat_not_order_pipeline = Pipeline([(\"select\", DataFrameSelector(cat_not_order_features)),\n                                   (\"num_to_cat\", FunctionTransformer(num_to_cat)),\n                                   (\"encoding\", OneHotEncoder()), ])","metadata":{"execution":{"iopub.status.busy":"2022-07-08T13:19:12.102314Z","iopub.execute_input":"2022-07-08T13:19:12.102544Z","iopub.status.idle":"2022-07-08T13:19:12.106909Z","shell.execute_reply.started":"2022-07-08T13:19:12.102519Z","shell.execute_reply":"2022-07-08T13:19:12.106381Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"cat_not_order_pipeline.fit_transform(df_filled[cat_not_order_features])","metadata":{"execution":{"iopub.status.busy":"2022-07-08T13:19:12.108061Z","iopub.execute_input":"2022-07-08T13:19:12.108311Z","iopub.status.idle":"2022-07-08T13:19:12.174309Z","shell.execute_reply.started":"2022-07-08T13:19:12.108278Z","shell.execute_reply":"2022-07-08T13:19:12.173499Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Implement StandardScaler here if needed\nnum_pipeline = Pipeline([(\"select\", DataFrameSelector(real_num_features)),\n                        ])","metadata":{"execution":{"iopub.status.busy":"2022-07-08T13:19:12.175663Z","iopub.execute_input":"2022-07-08T13:19:12.176628Z","iopub.status.idle":"2022-07-08T13:19:12.181686Z","shell.execute_reply.started":"2022-07-08T13:19:12.176587Z","shell.execute_reply":"2022-07-08T13:19:12.180813Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"real_num_features.remove(\"SalePrice\")\nreal_num_features.remove(\"Id\")","metadata":{"execution":{"iopub.status.busy":"2022-07-08T13:19:12.182941Z","iopub.execute_input":"2022-07-08T13:19:12.183147Z","iopub.status.idle":"2022-07-08T13:19:12.194893Z","shell.execute_reply.started":"2022-07-08T13:19:12.183124Z","shell.execute_reply":"2022-07-08T13:19:12.194167Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"num_df = pd.DataFrame(num_pipeline.fit_transform(df_filled[real_num_features]), columns = real_num_features)\nnum_df.head()","metadata":{"execution":{"iopub.status.busy":"2022-07-08T13:19:12.196231Z","iopub.execute_input":"2022-07-08T13:19:12.196938Z","iopub.status.idle":"2022-07-08T13:19:12.235957Z","shell.execute_reply.started":"2022-07-08T13:19:12.196900Z","shell.execute_reply":"2022-07-08T13:19:12.235334Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# **Tranform the dataset**","metadata":{}},{"cell_type":"code","source":"len(real_num_features) + len(cat_order_features) + len(cat_not_order_features) ","metadata":{"execution":{"iopub.status.busy":"2022-07-08T13:19:12.237016Z","iopub.execute_input":"2022-07-08T13:19:12.237392Z","iopub.status.idle":"2022-07-08T13:19:12.242822Z","shell.execute_reply.started":"2022-07-08T13:19:12.237363Z","shell.execute_reply":"2022-07-08T13:19:12.241960Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"full_pipeline = ColumnTransformer([\n    (\"real_num\", num_pipeline, real_num_features),\n    (\"num_to_cat\", num_to_cat_pipeline, list(false_num_features)),\n    (\"cat_order\", cat_order_pipeline, cat_order_features),\n    (\"cat_not_order\", cat_not_order_pipeline, cat_not_order_features)\n    ])","metadata":{"execution":{"iopub.status.busy":"2022-07-08T13:19:12.244040Z","iopub.execute_input":"2022-07-08T13:19:12.244936Z","iopub.status.idle":"2022-07-08T13:19:12.261515Z","shell.execute_reply.started":"2022-07-08T13:19:12.244781Z","shell.execute_reply":"2022-07-08T13:19:12.260721Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"all_data = full_pipeline.fit_transform(df_filled)\nall_data","metadata":{"execution":{"iopub.status.busy":"2022-07-08T13:19:12.263041Z","iopub.execute_input":"2022-07-08T13:19:12.265348Z","iopub.status.idle":"2022-07-08T13:19:12.405564Z","shell.execute_reply.started":"2022-07-08T13:19:12.265298Z","shell.execute_reply":"2022-07-08T13:19:12.404737Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"train_df_prepared = all_data[:1460]\ntest_df_prepared = all_data[1460:]","metadata":{"execution":{"iopub.status.busy":"2022-07-08T13:19:12.406670Z","iopub.execute_input":"2022-07-08T13:19:12.406882Z","iopub.status.idle":"2022-07-08T13:19:12.411132Z","shell.execute_reply.started":"2022-07-08T13:19:12.406859Z","shell.execute_reply":"2022-07-08T13:19:12.410321Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Test models","metadata":{}},{"cell_type":"markdown","source":"**First step :** Determine relevant machine learning models. Have a few datas, and few features, but the problem is relatively complex. We can begin with a\n**linear regression**, then **polynomial regression**","metadata":{}},{"cell_type":"code","source":"X_train = train_df_prepared\ny_train = np.log(train_df[\"SalePrice\"] + 1)\ny_train.shape","metadata":{"execution":{"iopub.status.busy":"2022-07-08T13:19:12.412261Z","iopub.execute_input":"2022-07-08T13:19:12.412471Z","iopub.status.idle":"2022-07-08T13:19:12.430217Z","shell.execute_reply.started":"2022-07-08T13:19:12.412449Z","shell.execute_reply":"2022-07-08T13:19:12.429384Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"**1 - Linear Regression (ridge)**","metadata":{}},{"cell_type":"code","source":"param_grid = [{'alpha': [0.05, 0.1, 0.3, 1, 3, 5, 10, 15, 30, 50, 75]},]\ngrid_search = GridSearchCV(Ridge(random_state=42), param_grid, cv=5, verbose=0, scoring=\"neg_root_mean_squared_error\")\ngrid_search.fit(X_train, y_train)","metadata":{"execution":{"iopub.status.busy":"2022-07-08T13:19:12.431460Z","iopub.execute_input":"2022-07-08T13:19:12.432046Z","iopub.status.idle":"2022-07-08T13:19:14.872060Z","shell.execute_reply.started":"2022-07-08T13:19:12.432008Z","shell.execute_reply":"2022-07-08T13:19:14.871081Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"grid_search.best_params_","metadata":{"execution":{"iopub.status.busy":"2022-07-08T13:19:14.873567Z","iopub.execute_input":"2022-07-08T13:19:14.874114Z","iopub.status.idle":"2022-07-08T13:19:14.882016Z","shell.execute_reply.started":"2022-07-08T13:19:14.874070Z","shell.execute_reply":"2022-07-08T13:19:14.881094Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"grid_search.best_score_","metadata":{"execution":{"iopub.status.busy":"2022-07-08T13:19:14.883665Z","iopub.execute_input":"2022-07-08T13:19:14.884240Z","iopub.status.idle":"2022-07-08T13:19:14.897423Z","shell.execute_reply.started":"2022-07-08T13:19:14.884198Z","shell.execute_reply":"2022-07-08T13:19:14.896430Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"The value of alpha seems to have not influence on the scoring. Reference score of **0.13**","metadata":{}},{"cell_type":"markdown","source":"**2 - Linear regression (elastic net)**","metadata":{}},{"cell_type":"code","source":"param_grid = [{'alpha': [0.0001,0.0003, 0.001, 0.003],\"l1_ratio\": [0.3,0.4,0.5,0.6,0.7,0.8,0.9]},]\ngrid_search = GridSearchCV(ElasticNet(random_state=42, max_iter=100000), param_grid, cv=5, verbose=0, scoring=\"neg_root_mean_squared_error\")\ngrid_search.fit(X_train, y_train)","metadata":{"execution":{"iopub.status.busy":"2022-07-08T13:19:14.898842Z","iopub.execute_input":"2022-07-08T13:19:14.899919Z","iopub.status.idle":"2022-07-08T13:20:58.612133Z","shell.execute_reply.started":"2022-07-08T13:19:14.899865Z","shell.execute_reply":"2022-07-08T13:20:58.611152Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"grid_search.best_params_","metadata":{"execution":{"iopub.status.busy":"2022-07-08T13:20:58.613836Z","iopub.execute_input":"2022-07-08T13:20:58.614428Z","iopub.status.idle":"2022-07-08T13:20:58.621832Z","shell.execute_reply.started":"2022-07-08T13:20:58.614380Z","shell.execute_reply":"2022-07-08T13:20:58.620955Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"grid_search.best_score_","metadata":{"execution":{"iopub.status.busy":"2022-07-08T13:20:58.623902Z","iopub.execute_input":"2022-07-08T13:20:58.624568Z","iopub.status.idle":"2022-07-08T13:20:58.640996Z","shell.execute_reply.started":"2022-07-08T13:20:58.624523Z","shell.execute_reply":"2022-07-08T13:20:58.640079Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"**New best score of 0.125**","metadata":{}},{"cell_type":"markdown","source":"**3 - Polynomial regression**","metadata":{}},{"cell_type":"code","source":"poly_features = PolynomialFeatures(degree=2, include_bias=False)\nX_poly = poly_features.fit_transform(X_train)\nX_poly[0]","metadata":{"execution":{"iopub.status.busy":"2022-07-08T13:20:58.642646Z","iopub.execute_input":"2022-07-08T13:20:58.643223Z","iopub.status.idle":"2022-07-08T13:20:59.078759Z","shell.execute_reply.started":"2022-07-08T13:20:58.643183Z","shell.execute_reply":"2022-07-08T13:20:59.077800Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"param_grid = [{},]\ngrid_search = GridSearchCV(LinearRegression(), param_grid, cv=5, verbose=0, scoring=\"neg_root_mean_squared_error\")\ngrid_search.fit(X_poly, y_train)","metadata":{"execution":{"iopub.status.busy":"2022-07-08T13:20:59.080149Z","iopub.execute_input":"2022-07-08T13:20:59.080863Z","iopub.status.idle":"2022-07-08T13:21:55.678495Z","shell.execute_reply.started":"2022-07-08T13:20:59.080819Z","shell.execute_reply":"2022-07-08T13:21:55.677682Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"grid_search.best_params_","metadata":{"execution":{"iopub.status.busy":"2022-07-08T13:21:55.679924Z","iopub.execute_input":"2022-07-08T13:21:55.680400Z","iopub.status.idle":"2022-07-08T13:21:55.687256Z","shell.execute_reply.started":"2022-07-08T13:21:55.680362Z","shell.execute_reply":"2022-07-08T13:21:55.686312Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"grid_search.best_score_","metadata":{"execution":{"iopub.status.busy":"2022-07-08T13:21:55.688962Z","iopub.execute_input":"2022-07-08T13:21:55.689935Z","iopub.status.idle":"2022-07-08T13:21:55.703977Z","shell.execute_reply.started":"2022-07-08T13:21:55.689886Z","shell.execute_reply":"2022-07-08T13:21:55.703083Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Not better and time consuming","metadata":{}},{"cell_type":"markdown","source":"**4 - DecisionTreeRegressor**","metadata":{}},{"cell_type":"code","source":"param_grid = [{'max_leaf_nodes' : [32,64,128], 'max_depth' : [4,8,16]}]\ngrid_search = GridSearchCV(DecisionTreeRegressor(random_state=42), param_grid, cv=5, verbose=0, scoring=\"neg_root_mean_squared_error\")\ngrid_search.fit(X_train, y_train)","metadata":{"execution":{"iopub.status.busy":"2022-07-08T13:21:55.705848Z","iopub.execute_input":"2022-07-08T13:21:55.706744Z","iopub.status.idle":"2022-07-08T13:21:57.572077Z","shell.execute_reply.started":"2022-07-08T13:21:55.706693Z","shell.execute_reply":"2022-07-08T13:21:57.571450Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"grid_search.best_params_","metadata":{"execution":{"iopub.status.busy":"2022-07-08T13:21:57.573253Z","iopub.execute_input":"2022-07-08T13:21:57.573660Z","iopub.status.idle":"2022-07-08T13:21:57.578421Z","shell.execute_reply.started":"2022-07-08T13:21:57.573631Z","shell.execute_reply":"2022-07-08T13:21:57.577687Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"grid_search.best_score_","metadata":{"execution":{"iopub.status.busy":"2022-07-08T13:21:57.579806Z","iopub.execute_input":"2022-07-08T13:21:57.580040Z","iopub.status.idle":"2022-07-08T13:21:57.593339Z","shell.execute_reply.started":"2022-07-08T13:21:57.580014Z","shell.execute_reply":"2022-07-08T13:21:57.592475Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Not good enough","metadata":{}},{"cell_type":"markdown","source":"**5 - RandomForestRegressor**","metadata":{}},{"cell_type":"code","source":"param_grid = [{'max_leaf_nodes' : [32,64,128], 'max_depth' : [4,8,16]}]\ngrid_search = GridSearchCV(RandomForestRegressor(n_estimators=100, n_jobs=-1, random_state=42), param_grid, cv=5, verbose=0, scoring=\"neg_root_mean_squared_error\")\ngrid_search.fit(X_train, y_train)","metadata":{"execution":{"iopub.status.busy":"2022-07-08T13:21:57.594513Z","iopub.execute_input":"2022-07-08T13:21:57.594935Z","iopub.status.idle":"2022-07-08T13:22:44.522521Z","shell.execute_reply.started":"2022-07-08T13:21:57.594897Z","shell.execute_reply":"2022-07-08T13:22:44.521232Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"grid_search.best_params_","metadata":{"execution":{"iopub.status.busy":"2022-07-08T13:22:44.524540Z","iopub.execute_input":"2022-07-08T13:22:44.524777Z","iopub.status.idle":"2022-07-08T13:22:44.532437Z","shell.execute_reply.started":"2022-07-08T13:22:44.524738Z","shell.execute_reply":"2022-07-08T13:22:44.531527Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"grid_search.best_score_","metadata":{"execution":{"iopub.status.busy":"2022-07-08T13:22:44.533668Z","iopub.execute_input":"2022-07-08T13:22:44.534114Z","iopub.status.idle":"2022-07-08T13:22:44.550297Z","shell.execute_reply.started":"2022-07-08T13:22:44.534081Z","shell.execute_reply":"2022-07-08T13:22:44.549361Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Not good enough","metadata":{}},{"cell_type":"markdown","source":"**6 - Linear SVM**","metadata":{}},{"cell_type":"code","source":"param_grid = [{'C': [0.1]}]\ngrid_search = GridSearchCV(SVR(kernel=\"linear\"), param_grid, cv=5, verbose=0, scoring=\"neg_root_mean_squared_error\")\ngrid_search.fit(X_train, y_train)","metadata":{"execution":{"iopub.status.busy":"2022-07-08T13:22:44.551796Z","iopub.execute_input":"2022-07-08T13:22:44.552549Z","iopub.status.idle":"2022-07-08T13:31:54.369329Z","shell.execute_reply.started":"2022-07-08T13:22:44.552512Z","shell.execute_reply":"2022-07-08T13:31:54.368404Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"grid_search.best_params_","metadata":{"execution":{"iopub.status.busy":"2022-07-08T13:31:54.370729Z","iopub.execute_input":"2022-07-08T13:31:54.371047Z","iopub.status.idle":"2022-07-08T13:31:54.378019Z","shell.execute_reply.started":"2022-07-08T13:31:54.371015Z","shell.execute_reply":"2022-07-08T13:31:54.376845Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"grid_search.best_score_","metadata":{"execution":{"iopub.status.busy":"2022-07-08T13:31:54.379516Z","iopub.execute_input":"2022-07-08T13:31:54.379992Z","iopub.status.idle":"2022-07-08T13:31:54.391894Z","shell.execute_reply.started":"2022-07-08T13:31:54.379965Z","shell.execute_reply":"2022-07-08T13:31:54.390910Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Might be better with a StandardScaler","metadata":{}},{"cell_type":"markdown","source":"**7 - Polynomial SVM**","metadata":{}},{"cell_type":"code","source":"param_grid = [{'C': [1], 'degree': [5], \"coef0\": [10]},]\ngrid_search = GridSearchCV(SVR(kernel=\"poly\"), param_grid, cv=5, verbose=0, scoring=\"neg_root_mean_squared_error\")\ngrid_search.fit(X_train, y_train)","metadata":{"execution":{"iopub.status.busy":"2022-07-08T13:31:54.393009Z","iopub.execute_input":"2022-07-08T13:31:54.393391Z","iopub.status.idle":"2022-07-08T13:32:34.487272Z","shell.execute_reply.started":"2022-07-08T13:31:54.393362Z","shell.execute_reply":"2022-07-08T13:32:34.486421Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"grid_search.best_params_","metadata":{"execution":{"iopub.status.busy":"2022-07-08T13:32:34.488426Z","iopub.execute_input":"2022-07-08T13:32:34.488640Z","iopub.status.idle":"2022-07-08T13:32:34.495658Z","shell.execute_reply.started":"2022-07-08T13:32:34.488616Z","shell.execute_reply":"2022-07-08T13:32:34.494730Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"grid_search.best_score_","metadata":{"execution":{"iopub.status.busy":"2022-07-08T13:32:34.497487Z","iopub.execute_input":"2022-07-08T13:32:34.498103Z","iopub.status.idle":"2022-07-08T13:32:34.511985Z","shell.execute_reply.started":"2022-07-08T13:32:34.498060Z","shell.execute_reply":"2022-07-08T13:32:34.511191Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"**8 - RBF SVM**","metadata":{}},{"cell_type":"code","source":"param_grid = [{'C': [10,100], \"gamma\": [0.00001,0.0001]},]\ngrid_search = GridSearchCV(SVR(kernel=\"rbf\"), param_grid, cv=5, verbose=0, scoring=\"neg_root_mean_squared_error\")\ngrid_search.fit(X_train, y_train)","metadata":{"execution":{"iopub.status.busy":"2022-07-08T13:32:34.513694Z","iopub.execute_input":"2022-07-08T13:32:34.513954Z","iopub.status.idle":"2022-07-08T13:32:48.002142Z","shell.execute_reply.started":"2022-07-08T13:32:34.513927Z","shell.execute_reply":"2022-07-08T13:32:48.001250Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"grid_search.best_params_","metadata":{"execution":{"iopub.status.busy":"2022-07-08T13:32:48.003358Z","iopub.execute_input":"2022-07-08T13:32:48.003745Z","iopub.status.idle":"2022-07-08T13:32:48.010250Z","shell.execute_reply.started":"2022-07-08T13:32:48.003715Z","shell.execute_reply":"2022-07-08T13:32:48.009372Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"grid_search.best_score_","metadata":{"execution":{"iopub.status.busy":"2022-07-08T13:32:48.011489Z","iopub.execute_input":"2022-07-08T13:32:48.011817Z","iopub.status.idle":"2022-07-08T13:32:48.025479Z","shell.execute_reply.started":"2022-07-08T13:32:48.011789Z","shell.execute_reply":"2022-07-08T13:32:48.024635Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Not better","metadata":{}},{"cell_type":"markdown","source":"**9 - Gradient Boosting**","metadata":{}},{"cell_type":"code","source":"param_grid = [{'max_leaf_nodes': [4,8,16],'max_depth': [3,4,5,6]},]\ngrid_search = GridSearchCV(GradientBoostingRegressor(random_state=42), param_grid, cv=5, verbose=0, scoring=\"neg_root_mean_squared_error\")\ngrid_search.fit(X_train, y_train)","metadata":{"execution":{"iopub.status.busy":"2022-07-08T13:32:48.027056Z","iopub.execute_input":"2022-07-08T13:32:48.027395Z","iopub.status.idle":"2022-07-08T13:34:13.617126Z","shell.execute_reply.started":"2022-07-08T13:32:48.027367Z","shell.execute_reply":"2022-07-08T13:34:13.616087Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"grid_search.best_params_","metadata":{"execution":{"iopub.status.busy":"2022-07-08T13:34:13.618426Z","iopub.execute_input":"2022-07-08T13:34:13.618748Z","iopub.status.idle":"2022-07-08T13:34:13.624653Z","shell.execute_reply.started":"2022-07-08T13:34:13.618718Z","shell.execute_reply":"2022-07-08T13:34:13.623904Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"grid_search.best_score_","metadata":{"execution":{"iopub.status.busy":"2022-07-08T13:34:13.625978Z","iopub.execute_input":"2022-07-08T13:34:13.626206Z","iopub.status.idle":"2022-07-08T13:34:13.643782Z","shell.execute_reply.started":"2022-07-08T13:34:13.626181Z","shell.execute_reply":"2022-07-08T13:34:13.643145Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"**Quite the same score than Elastic Net**","metadata":{}},{"cell_type":"markdown","source":"## Bias and Variance","metadata":{}},{"cell_type":"code","source":"def plot_learning_curves(model, X, y):\n    X_train, X_val, y_train, y_val = train_test_split(X, y, test_size=0.2, random_state=10)\n    train_errors, val_errors = [], []\n    for m in range(1, len(X_train)):\n        model.fit(X_train[:m], y_train[:m])\n        y_train_predict = model.predict(X_train[:m])\n        y_val_predict = model.predict(X_val)\n        train_errors.append(mean_squared_error(y_train[:m], y_train_predict))\n        val_errors.append(mean_squared_error(y_val, y_val_predict))\n\n    plt.plot(np.sqrt(train_errors), \"r-+\", linewidth=2, label=\"train\")\n    plt.plot(np.sqrt(val_errors), \"b-\", linewidth=3, label=\"val\")\n    plt.legend(loc=\"upper right\", fontsize=14)\n    plt.xlabel(\"Training set size\", fontsize=14)\n    plt.ylabel(\"RMSE\", fontsize=14)","metadata":{"execution":{"iopub.status.busy":"2022-07-08T13:34:13.645082Z","iopub.execute_input":"2022-07-08T13:34:13.645631Z","iopub.status.idle":"2022-07-08T13:34:13.660504Z","shell.execute_reply.started":"2022-07-08T13:34:13.645599Z","shell.execute_reply":"2022-07-08T13:34:13.659610Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"plot_learning_curves(ElasticNet(random_state=42, max_iter=100000, alpha=0.001, l1_ratio=0.7), X_train, y_train)\nplt.axis([0, 1200, 0, 1])\nplt.show()","metadata":{"execution":{"iopub.status.busy":"2022-07-08T13:34:13.661633Z","iopub.execute_input":"2022-07-08T13:34:13.661862Z","iopub.status.idle":"2022-07-08T13:36:11.830444Z","shell.execute_reply.started":"2022-07-08T13:34:13.661839Z","shell.execute_reply":"2022-07-08T13:36:11.829552Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"plot_learning_curves(GradientBoostingRegressor(random_state=42,max_depth=5, max_leaf_nodes=8), X_train, y_train)\nplt.axis([0, 1200, 0, 1])\nplt.show()","metadata":{"execution":{"iopub.status.busy":"2022-07-08T13:36:11.831780Z","iopub.execute_input":"2022-07-08T13:36:11.832661Z","iopub.status.idle":"2022-07-08T13:52:15.422963Z","shell.execute_reply.started":"2022-07-08T13:36:11.832621Z","shell.execute_reply":"2022-07-08T13:52:15.422121Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# **Propose a solution**","metadata":{}},{"cell_type":"markdown","source":"**Implement the Random Forest**","metadata":{}},{"cell_type":"code","source":"reg_model = ElasticNet(random_state=42, max_iter=100000, alpha=0.001, l1_ratio=0.7)\nreg_model.fit(X_train, y_train)","metadata":{"execution":{"iopub.status.busy":"2022-07-08T13:52:15.424043Z","iopub.execute_input":"2022-07-08T13:52:15.424396Z","iopub.status.idle":"2022-07-08T13:52:16.451989Z","shell.execute_reply.started":"2022-07-08T13:52:15.424363Z","shell.execute_reply":"2022-07-08T13:52:16.450879Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"y_pred_test = np.exp(reg_model.predict(test_df_prepared))\ny_pred_test.T.shape","metadata":{"execution":{"iopub.status.busy":"2022-07-08T13:52:16.457476Z","iopub.execute_input":"2022-07-08T13:52:16.458124Z","iopub.status.idle":"2022-07-08T13:52:16.607641Z","shell.execute_reply.started":"2022-07-08T13:52:16.458078Z","shell.execute_reply":"2022-07-08T13:52:16.606529Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"example_data = pd.read_csv(\"/kaggle/input/house-prices-advanced-regression-techniques/sample_submission.csv\")\nnp.array(example_data[\"Id\"]).shape","metadata":{"execution":{"iopub.status.busy":"2022-07-08T13:52:16.614358Z","iopub.execute_input":"2022-07-08T13:52:16.618078Z","iopub.status.idle":"2022-07-08T13:52:16.655669Z","shell.execute_reply.started":"2022-07-08T13:52:16.618002Z","shell.execute_reply":"2022-07-08T13:52:16.654583Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"example_data.head(10)","metadata":{"execution":{"iopub.status.busy":"2022-07-08T13:52:16.663621Z","iopub.execute_input":"2022-07-08T13:52:16.667754Z","iopub.status.idle":"2022-07-08T13:52:16.690314Z","shell.execute_reply.started":"2022-07-08T13:52:16.667684Z","shell.execute_reply":"2022-07-08T13:52:16.689370Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"np.array([np.array(example_data[\"Id\"]),y_pred_test.T]).T.shape","metadata":{"execution":{"iopub.status.busy":"2022-07-08T13:52:16.696258Z","iopub.execute_input":"2022-07-08T13:52:16.699058Z","iopub.status.idle":"2022-07-08T13:52:16.715619Z","shell.execute_reply.started":"2022-07-08T13:52:16.698999Z","shell.execute_reply":"2022-07-08T13:52:16.714682Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"id_sample = np.array(example_data[\"Id\"],dtype=np.int32)","metadata":{"execution":{"iopub.status.busy":"2022-07-08T13:52:16.717539Z","iopub.execute_input":"2022-07-08T13:52:16.719012Z","iopub.status.idle":"2022-07-08T13:52:16.724262Z","shell.execute_reply.started":"2022-07-08T13:52:16.718978Z","shell.execute_reply":"2022-07-08T13:52:16.723309Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"df = pd.DataFrame(np.array([id_sample,y_pred_test.T]).T,columns=[\"Id\",\"SalePrice\"])\ndf = df.astype({'Id': 'int32'})","metadata":{"execution":{"iopub.status.busy":"2022-07-08T13:52:16.725884Z","iopub.execute_input":"2022-07-08T13:52:16.727016Z","iopub.status.idle":"2022-07-08T13:52:16.742096Z","shell.execute_reply.started":"2022-07-08T13:52:16.726980Z","shell.execute_reply":"2022-07-08T13:52:16.741092Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"df.dtypes","metadata":{"execution":{"iopub.status.busy":"2022-07-08T13:52:16.743571Z","iopub.execute_input":"2022-07-08T13:52:16.744458Z","iopub.status.idle":"2022-07-08T13:52:16.760870Z","shell.execute_reply.started":"2022-07-08T13:52:16.744418Z","shell.execute_reply":"2022-07-08T13:52:16.759979Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"df.head(10)","metadata":{"execution":{"iopub.status.busy":"2022-07-08T13:52:16.762246Z","iopub.execute_input":"2022-07-08T13:52:16.762549Z","iopub.status.idle":"2022-07-08T13:52:16.776314Z","shell.execute_reply.started":"2022-07-08T13:52:16.762511Z","shell.execute_reply":"2022-07-08T13:52:16.775296Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"df.to_csv('./submission.csv', index=False)","metadata":{"execution":{"iopub.status.busy":"2022-07-08T13:52:16.777742Z","iopub.execute_input":"2022-07-08T13:52:16.778409Z","iopub.status.idle":"2022-07-08T13:52:16.801626Z","shell.execute_reply.started":"2022-07-08T13:52:16.778288Z","shell.execute_reply":"2022-07-08T13:52:16.800818Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# **References**\n - [Regularized Linear Models](https://www.kaggle.com/code/apapiu/regularized-linear-models) by Alexandru Papiu\n - Hands-On Machine Learning with Scikit-Learn and TensorFlow: Concepts, Tools, and Techniques for Building Intelligent Systems by Aurélien Géron","metadata":{}}]}