{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"pygments_lexer":"ipython3","nbconvert_exporter":"python","version":"3.6.4","file_extension":".py","codemirror_mode":{"name":"ipython","version":3},"name":"python","mimetype":"text/x-python"}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"# Introduction #\n\nWelcome to the feature engineering project for the [House Prices - Advanced Regression Techniques](https://www.kaggle.com/c/house-prices-advanced-regression-techniques) competition! This competition uses nearly the same data you used in the exercises of the [Feature Engineering](https://www.kaggle.com/learn/feature-engineering) course. We'll collect together the work you did into a complete project which you can build off of with ideas of your own.\n\n<blockquote style=\"margin-right:auto; margin-left:auto; background-color: #ebf9ff; padding: 1em; margin:24px;\">\n    <strong>Fork This Notebook!</strong><br>\nCreate your own editable copy of this notebook by clicking on the <strong>Copy and Edit</strong> button in the top right corner.\n</blockquote>\n\n# Step 1 - Preliminaries #\n## Imports and Configuration ##\n\nWe'll start by importing the packages we used in the exercises and setting some notebook defaults. Unhide this cell if you'd like to see the libraries we'll use:","metadata":{}},{"cell_type":"code","source":"\nimport os\nimport warnings\nfrom pathlib import Path\n\nimport matplotlib.pyplot as plt\nimport numpy as np\nimport pandas as pd\nimport seaborn as sns\nfrom IPython.display import display\nfrom pandas.api.types import CategoricalDtype\n\nfrom category_encoders import MEstimateEncoder\nfrom sklearn.cluster import KMeans\nfrom sklearn.decomposition import PCA\nfrom sklearn.feature_selection import mutual_info_regression\nfrom sklearn.model_selection import KFold, cross_val_score\nfrom xgboost import XGBRegressor\n\n\n# Set Matplotlib defaults\nplt.style.use(\"seaborn-whitegrid\")\nplt.rc(\"figure\", autolayout=True)\nplt.rc(\n    \"axes\",\n    labelweight=\"bold\",\n    labelsize=\"large\",\n    titleweight=\"bold\",\n    titlesize=14,\n    titlepad=10,\n)\n\n# Mute warnings\nwarnings.filterwarnings('ignore')\n","metadata":{"_kg_hide-input":true,"execution":{"iopub.status.busy":"2022-07-12T13:04:55.602885Z","iopub.execute_input":"2022-07-12T13:04:55.603580Z","iopub.status.idle":"2022-07-12T13:04:57.514714Z","shell.execute_reply.started":"2022-07-12T13:04:55.603463Z","shell.execute_reply":"2022-07-12T13:04:57.513506Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## Data Preprocessing ##\n\nBefore we can do any feature engineering, we need to *preprocess* the data to get it in a form suitable for analysis. The data we used in the course was a bit simpler than the competition data. For the *Ames* competition dataset, we'll need to:\n- **Load** the data from CSV files\n- **Clean** the data to fix any errors or inconsistencies\n- **Encode** the statistical data type (numeric, categorical)\n- **Impute** any missing values\n\nWe'll wrap all these steps up in a function, which will make easy for you to get a fresh dataframe whenever you need. After reading the CSV file, we'll apply three preprocessing steps, `clean`, `encode`, and `impute`, and then create the data splits: one (`df_train`) for training the model, and one (`df_test`) for making the predictions that you'll submit to the competition for scoring on the leaderboard.","metadata":{}},{"cell_type":"code","source":"def load_data():\n    # Read data\n    data_dir = Path(\"../input/house-prices-advanced-regression-techniques/\")\n    df_train = pd.read_csv(data_dir / \"train.csv\", index_col=\"Id\")\n    df_test = pd.read_csv(data_dir / \"test.csv\", index_col=\"Id\")\n    # Merge the splits so we can process them together\n    df = pd.concat([df_train, df_test])\n    # Preprocessing\n    df = clean(df)\n    df = encode(df)\n    df = impute(df)\n    # Reform splits\n    df_train = df.loc[df_train.index, :]\n    df_test = df.loc[df_test.index, :]\n    return df_train, df_test\n","metadata":{"execution":{"iopub.status.busy":"2022-07-12T13:05:04.039739Z","iopub.execute_input":"2022-07-12T13:05:04.040154Z","iopub.status.idle":"2022-07-12T13:05:04.047258Z","shell.execute_reply.started":"2022-07-12T13:05:04.040120Z","shell.execute_reply":"2022-07-12T13:05:04.046112Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### Clean Data ###\n\nSome of the categorical features in this dataset have what are apparently typos in their categories:","metadata":{}},{"cell_type":"code","source":"data_dir = Path(\"../input/house-prices-advanced-regression-techniques/\")\ndf = pd.read_csv(data_dir / \"train.csv\", index_col=\"Id\")\n\ndf.Exterior2nd.unique()","metadata":{"execution":{"iopub.status.busy":"2022-07-12T13:05:09.407964Z","iopub.execute_input":"2022-07-12T13:05:09.408331Z","iopub.status.idle":"2022-07-12T13:05:09.469237Z","shell.execute_reply.started":"2022-07-12T13:05:09.408300Z","shell.execute_reply":"2022-07-12T13:05:09.468046Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Comparing these to `data_description.txt` shows us what needs cleaning. We'll take care of a couple of issues here, but you might want to evaluate this data further.","metadata":{}},{"cell_type":"code","source":"def clean(df):\n    df[\"Exterior2nd\"] = df[\"Exterior2nd\"].replace({\"Brk Cmn\": \"BrkComm\"})\n    # Some values of GarageYrBlt are corrupt, so we'll replace them\n    # with the year the house was built\n    df[\"GarageYrBlt\"] = df[\"GarageYrBlt\"].where(df.GarageYrBlt <= 2010, df.YearBuilt)\n    # Names beginning with numbers are awkward to work with\n    df.rename(columns={\n        \"1stFlrSF\": \"FirstFlrSF\",\n        \"2ndFlrSF\": \"SecondFlrSF\",\n        \"3SsnPorch\": \"Threeseasonporch\",\n    }, inplace=True,\n    )\n    return df\n","metadata":{"execution":{"iopub.status.busy":"2022-07-12T13:08:31.487475Z","iopub.execute_input":"2022-07-12T13:08:31.487964Z","iopub.status.idle":"2022-07-12T13:08:31.494702Z","shell.execute_reply.started":"2022-07-12T13:08:31.487929Z","shell.execute_reply":"2022-07-12T13:08:31.493463Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### Encode the Statistical Data Type ###\n\nPandas has Python types corresponding to the standard statistical types (numeric, categorical, etc.). Encoding each feature with its correct type helps ensure each feature is treated appropriately by whatever functions we use, and makes it easier for us to apply transformations consistently. This hidden cell defines the `encode` function:","metadata":{}},{"cell_type":"code","source":"# The numeric features are already encoded correctly (`float` for\n# continuous, `int` for discrete), but the categoricals we'll need to\n# do ourselves. Note in particular, that the `MSSubClass` feature is\n# read as an `int` type, but is actually a (nominative) categorical.\n\n# The nominative (unordered) categorical features\nfeatures_nom = [\"MSSubClass\", \"MSZoning\", \"Street\", \"Alley\", \"LandContour\", \"LotConfig\", \"Neighborhood\", \"Condition1\", \"Condition2\", \"BldgType\", \"HouseStyle\", \"RoofStyle\", \"RoofMatl\", \"Exterior1st\", \"Exterior2nd\", \"MasVnrType\", \"Foundation\", \"Heating\", \"CentralAir\", \"GarageType\", \"MiscFeature\", \"SaleType\", \"SaleCondition\"]\n\n\n# The ordinal (ordered) categorical features \n\n# Pandas calls the categories \"levels\"\nfive_levels = [\"Po\", \"Fa\", \"TA\", \"Gd\", \"Ex\"]\nten_levels = list(range(10))\n\nordered_levels = {\n    \"OverallQual\": ten_levels,\n    \"OverallCond\": ten_levels,\n    \"ExterQual\": five_levels,\n    \"ExterCond\": five_levels,\n    \"BsmtQual\": five_levels,\n    \"BsmtCond\": five_levels,\n    \"HeatingQC\": five_levels,\n    \"KitchenQual\": five_levels,\n    \"FireplaceQu\": five_levels,\n    \"GarageQual\": five_levels,\n    \"GarageCond\": five_levels,\n    \"PoolQC\": five_levels,\n    \"LotShape\": [\"Reg\", \"IR1\", \"IR2\", \"IR3\"],\n    \"LandSlope\": [\"Sev\", \"Mod\", \"Gtl\"],\n    \"BsmtExposure\": [\"No\", \"Mn\", \"Av\", \"Gd\"],\n    \"BsmtFinType1\": [\"Unf\", \"LwQ\", \"Rec\", \"BLQ\", \"ALQ\", \"GLQ\"],\n    \"BsmtFinType2\": [\"Unf\", \"LwQ\", \"Rec\", \"BLQ\", \"ALQ\", \"GLQ\"],\n    \"Functional\": [\"Sal\", \"Sev\", \"Maj1\", \"Maj2\", \"Mod\", \"Min2\", \"Min1\", \"Typ\"],\n    \"GarageFinish\": [\"Unf\", \"RFn\", \"Fin\"],\n    \"PavedDrive\": [\"N\", \"P\", \"Y\"],\n    \"Utilities\": [\"NoSeWa\", \"NoSewr\", \"AllPub\"],\n    \"CentralAir\": [\"N\", \"Y\"],\n    \"Electrical\": [\"Mix\", \"FuseP\", \"FuseF\", \"FuseA\", \"SBrkr\"],\n    \"Fence\": [\"MnWw\", \"GdWo\", \"MnPrv\", \"GdPrv\"],\n}\n\n# Add a None level for missing values\nordered_levels = {key: [\"None\"] + value for key, value in\n                  ordered_levels.items()}\n\n\ndef encode(df):\n    # Nominal categories\n    for name in features_nom:\n        df[name] = df[name].astype(\"category\")\n        # Add a None category for missing values\n        if \"None\" not in df[name].cat.categories:\n            df[name].cat.add_categories(\"None\", inplace=True)\n    # Ordinal categories\n    for name, levels in ordered_levels.items():\n        df[name] = df[name].astype(CategoricalDtype(levels,\n                                                    ordered=True))\n    return df\n","metadata":{"_kg_hide-input":true,"execution":{"iopub.status.busy":"2022-07-12T13:11:08.291092Z","iopub.execute_input":"2022-07-12T13:11:08.291816Z","iopub.status.idle":"2022-07-12T13:11:08.304077Z","shell.execute_reply.started":"2022-07-12T13:11:08.291776Z","shell.execute_reply":"2022-07-12T13:11:08.303227Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### Handle Missing Values ###\n\nHandling missing values now will make the feature engineering go more smoothly. We'll impute `0` for missing numeric values and `\"None\"` for missing categorical values. You might like to experiment with other imputation strategies. In particular, you could try creating \"missing value\" indicators: `1` whenever a value was imputed and `0` otherwise.","metadata":{}},{"cell_type":"code","source":"def impute(df):\n    for name in df.select_dtypes(\"number\"):\n        df[name] = df[name].fillna(0)\n    for name in df.select_dtypes(\"category\"):\n        df[name] = df[name].fillna(\"None\")\n    return df\n","metadata":{"execution":{"iopub.status.busy":"2022-07-12T13:23:21.798624Z","iopub.execute_input":"2022-07-12T13:23:21.799074Z","iopub.status.idle":"2022-07-12T13:23:21.805613Z","shell.execute_reply.started":"2022-07-12T13:23:21.799040Z","shell.execute_reply":"2022-07-12T13:23:21.804239Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## Load Data ##\n\nAnd now we can call the data loader and get the processed data splits:","metadata":{}},{"cell_type":"code","source":"df_train, df_test = load_data()","metadata":{"execution":{"iopub.status.busy":"2022-07-12T13:23:29.530257Z","iopub.execute_input":"2022-07-12T13:23:29.530625Z","iopub.status.idle":"2022-07-12T13:23:29.695942Z","shell.execute_reply.started":"2022-07-12T13:23:29.530595Z","shell.execute_reply":"2022-07-12T13:23:29.694791Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Uncomment and run this cell if you'd like to see what they contain. Notice that `df_test` is\nmissing values for `SalePrice`. (`NA`s were willed with 0's in the imputation step.)","metadata":{}},{"cell_type":"code","source":"# Peek at the values\n#display(df_train)\n#display(df_test)\n\n# Display information about dtypes and missing values\n#display(df_train.info())\n#display(df_test.info())","metadata":{},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## Establish Baseline ##\n\nFinally, let's establish a baseline score to judge our feature engineering against.\n\nHere is the function we created in Lesson 1 that will compute the cross-validated RMSLE score for a feature set. We've used XGBoost for our model, but you might want to experiment with other models.\n","metadata":{}},{"cell_type":"code","source":"\ndef score_dataset(X, y, model=XGBRegressor()):\n    # Label encoding for categoricals\n    #\n    # Label encoding is good for XGBoost and RandomForest, but one-hot\n    # would be better for models like Lasso or Ridge. The `cat.codes`\n    # attribute holds the category levels.\n    for colname in X.select_dtypes([\"category\"]):\n        X[colname] = X[colname].cat.codes\n    # Metric for Housing competition is RMSLE (Root Mean Squared Log Error)\n    log_y = np.log(y)\n    score = cross_val_score(\n        model, X, log_y, cv=5, scoring=\"neg_mean_squared_error\",\n    )\n    score = -1 * score.mean()\n    score = np.sqrt(score)\n    return score\n","metadata":{"_kg_hide-input":true,"execution":{"iopub.status.busy":"2022-07-12T13:30:33.846162Z","iopub.execute_input":"2022-07-12T13:30:33.846590Z","iopub.status.idle":"2022-07-12T13:30:33.853857Z","shell.execute_reply.started":"2022-07-12T13:30:33.846540Z","shell.execute_reply":"2022-07-12T13:30:33.852483Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"We can reuse this scoring function anytime we want to try out a new feature set. We'll run it now on the processed data with no additional features and get a baseline score:","metadata":{}},{"cell_type":"code","source":"X = df_train.copy()\ny = X.pop(\"SalePrice\")\n\nbaseline_score = score_dataset(X, y)\nprint(f\"Baseline score: {baseline_score:.5f} RMSLE\")","metadata":{"execution":{"iopub.status.busy":"2022-07-12T13:31:07.740534Z","iopub.execute_input":"2022-07-12T13:31:07.740929Z","iopub.status.idle":"2022-07-12T13:31:10.740914Z","shell.execute_reply.started":"2022-07-12T13:31:07.740897Z","shell.execute_reply":"2022-07-12T13:31:10.739979Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"This baseline score helps us to know whether some set of features we've assembled has actually led to any improvement or not.\n\n# Step 2 - Feature Utility Scores #\n\nIn Lesson 2 we saw how to use mutual information to compute a *utility score* for a feature, giving you an indication of how much potential the feature has. This hidden cell defines the two utility functions we used, `make_mi_scores` and `plot_mi_scores`: ","metadata":{}},{"cell_type":"code","source":"\ndef make_mi_scores(X, y):\n    X = X.copy()\n    for colname in X.select_dtypes([\"object\", \"category\"]):\n        X[colname], _ = X[colname].factorize()\n    # All discrete features should now have integer dtypes\n    discrete_features = [pd.api.types.is_integer_dtype(t) for t in X.dtypes]\n    mi_scores = mutual_info_regression(X, y, discrete_features=discrete_features, random_state=0)\n    mi_scores = pd.Series(mi_scores, name=\"MI Scores\", index=X.columns)\n    mi_scores = mi_scores.sort_values(ascending=False)\n    return mi_scores\n\n\ndef plot_mi_scores(scores):\n    scores = scores.sort_values(ascending=True)\n    width = np.arange(len(scores))\n    ticks = list(scores.index)\n    plt.barh(width, scores)\n    plt.yticks(width, ticks)\n    plt.title(\"Mutual Information Scores\")","metadata":{"_kg_hide-input":true,"execution":{"iopub.status.busy":"2022-07-12T13:38:24.922271Z","iopub.execute_input":"2022-07-12T13:38:24.923343Z","iopub.status.idle":"2022-07-12T13:38:24.932094Z","shell.execute_reply.started":"2022-07-12T13:38:24.923287Z","shell.execute_reply":"2022-07-12T13:38:24.930923Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Let's look at our feature scores again:","metadata":{}},{"cell_type":"code","source":"X = df_train.copy()\ny = X.pop(\"SalePrice\")\n\nmi_scores = make_mi_scores(X, y)\nmi_scores","metadata":{"execution":{"iopub.status.busy":"2022-07-12T13:38:27.577827Z","iopub.execute_input":"2022-07-12T13:38:27.578793Z","iopub.status.idle":"2022-07-12T13:38:29.451135Z","shell.execute_reply.started":"2022-07-12T13:38:27.578749Z","shell.execute_reply":"2022-07-12T13:38:29.449882Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"You can see that we have a number of features that are highly informative and also some that don't seem to be informative at all (at least by themselves). As we talked about in Tutorial 2, the top scoring features will usually pay-off the most during feature development, so it could be a good idea to focus your efforts on those. On the other hand, training on uninformative features can lead to overfitting. So, the features with 0.0 scores we'll drop entirely:","metadata":{}},{"cell_type":"code","source":"def drop_uninformative(df, mi_scores):\n    return df.loc[:, mi_scores > 0.0]\n","metadata":{"execution":{"iopub.status.busy":"2022-07-12T13:40:14.701358Z","iopub.execute_input":"2022-07-12T13:40:14.701765Z","iopub.status.idle":"2022-07-12T13:40:14.707254Z","shell.execute_reply.started":"2022-07-12T13:40:14.701733Z","shell.execute_reply":"2022-07-12T13:40:14.706138Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Removing them does lead to a modest performance gain:","metadata":{}},{"cell_type":"code","source":"X = df_train.copy()\ny = X.pop(\"SalePrice\")\nX = drop_uninformative(X, mi_scores)\n\nscore_dataset(X, y)","metadata":{"execution":{"iopub.status.busy":"2022-07-12T13:40:28.371777Z","iopub.execute_input":"2022-07-12T13:40:28.372161Z","iopub.status.idle":"2022-07-12T13:40:32.107593Z","shell.execute_reply.started":"2022-07-12T13:40:28.372132Z","shell.execute_reply":"2022-07-12T13:40:32.106400Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Later, we'll add the `drop_uninformative` function to our feature-creation pipeline.\n\n# Step 3 - Create Features #\n\nNow we'll start developing our feature set.\n\nTo make our feature engineering workflow more modular, we'll define a function that will take a prepared dataframe and pass it through a pipeline of transformations to get the final feature set. It will look something like this:\n\n```\ndef create_features(df):\n    X = df.copy()\n    y = X.pop(\"SalePrice\")\n    X = X.join(create_features_1(X))\n    X = X.join(create_features_2(X))\n    X = X.join(create_features_3(X))\n    # ...\n    return X\n```\n\nLet's go ahead and define one transformation now, a [label encoding](https://www.kaggle.com/alexisbcook/categorical-variables) for the categorical features:","metadata":{}},{"cell_type":"code","source":"def label_encode(df):\n    X = df.copy()\n    for colname in X.select_dtypes([\"category\"]):\n        X[colname] = X[colname].cat.codes\n    return X\n","metadata":{"execution":{"iopub.status.busy":"2022-07-12T13:43:12.860089Z","iopub.execute_input":"2022-07-12T13:43:12.860460Z","iopub.status.idle":"2022-07-12T13:43:12.866226Z","shell.execute_reply.started":"2022-07-12T13:43:12.860428Z","shell.execute_reply":"2022-07-12T13:43:12.864843Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"A label encoding is okay for any kind of categorical feature when you're using a tree-ensemble like XGBoost, even for unordered categories. If you wanted to try a linear regression model (also popular in this competition), you would instead want to use a one-hot encoding, especially for the features with unordered categories.\n\n## Create Features with Pandas ##\n\nThis cell reproduces the work you did in Exercise 3, where you applied strategies for creating features in Pandas. Modify or add to these functions to try out other feature combinations.","metadata":{}},{"cell_type":"code","source":"\ndef mathematical_transforms(df):\n    X = pd.DataFrame()  # dataframe to hold new features\n    X[\"LivLotRatio\"] = df.GrLivArea / df.LotArea\n    X[\"Spaciousness\"] = (df.FirstFlrSF + df.SecondFlrSF) / df.TotRmsAbvGrd\n    # This feature ended up not helping performance\n    # X[\"TotalOutsideSF\"] = \\\n    #     df.WoodDeckSF + df.OpenPorchSF + df.EnclosedPorch + \\\n    #     df.Threeseasonporch + df.ScreenPorch\n    return X\n\n\ndef interactions(df):\n    X = pd.get_dummies(df.BldgType, prefix=\"Bldg\")\n    X = X.mul(df.GrLivArea, axis=0)\n    return X\n\n\ndef counts(df):\n    X = pd.DataFrame()\n    X[\"PorchTypes\"] = df[[\n        \"WoodDeckSF\",\n        \"OpenPorchSF\",\n        \"EnclosedPorch\",\n        \"Threeseasonporch\",\n        \"ScreenPorch\",\n    ]].gt(0.0).sum(axis=1)\n    return X\n\n\ndef break_down(df):\n    X = pd.DataFrame()\n    X[\"MSClass\"] = df.MSSubClass.str.split(\"_\", n=1, expand=True)[0]\n    return X\n\n\ndef group_transforms(df):\n    X = pd.DataFrame()\n    X[\"MedNhbdArea\"] = df.groupby(\"Neighborhood\")[\"GrLivArea\"].transform(\"median\")\n    return X\n","metadata":{"_kg_hide-input":true,"execution":{"iopub.status.busy":"2022-07-12T13:43:18.211038Z","iopub.execute_input":"2022-07-12T13:43:18.211433Z","iopub.status.idle":"2022-07-12T13:43:18.220339Z","shell.execute_reply.started":"2022-07-12T13:43:18.211388Z","shell.execute_reply":"2022-07-12T13:43:18.219514Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Here are some ideas for other transforms you could explore:\n- Interactions between the quality `Qual` and condition `Cond` features. `OverallQual`, for instance, was a high-scoring feature. You could try combining it with `OverallCond` by converting both to integer type and taking a product.\n- Square roots of area features. This would convert units of square feet to just feet.\n- Logarithms of numeric features. If a feature has a skewed distribution, applying a logarithm can help normalize it.\n- Interactions between numeric and categorical features that describe the same thing. You could look at interactions between `BsmtQual` and `TotalBsmtSF`, for instance.\n- Other group statistics in `Neighboorhood`. We did the median of `GrLivArea`. Looking at `mean`, `std`, or `count` could be interesting. You could also try combining the group statistics with other features. Maybe the *difference* of `GrLivArea` and the median is important?\n\n## k-Means Clustering ##\n\nThe first unsupervised algorithm we used to create features was k-means clustering. We saw that you could either use the cluster labels as a feature (a column with `0, 1, 2, ...`) or you could use the *distance* of the observations to each cluster. We saw how these features can sometimes be effective at untangling complicated spatial relationships.","metadata":{}},{"cell_type":"code","source":"\ncluster_features = [\n    \"LotArea\",\n    \"TotalBsmtSF\",\n    \"FirstFlrSF\",\n    \"SecondFlrSF\",\n    \"GrLivArea\",\n]\n\n\ndef cluster_labels(df, features, n_clusters=20):\n    X = df.copy()\n    X_scaled = X.loc[:, features]\n    X_scaled = (X_scaled - X_scaled.mean(axis=0)) / X_scaled.std(axis=0)\n    kmeans = KMeans(n_clusters=n_clusters, n_init=50, random_state=0)\n    X_new = pd.DataFrame()\n    X_new[\"Cluster\"] = kmeans.fit_predict(X_scaled)\n    return X_new\n\n\ndef cluster_distance(df, features, n_clusters=20):\n    X = df.copy()\n    X_scaled = X.loc[:, features]\n    X_scaled = (X_scaled - X_scaled.mean(axis=0)) / X_scaled.std(axis=0)\n    kmeans = KMeans(n_clusters=20, n_init=50, random_state=0)\n    X_cd = kmeans.fit_transform(X_scaled)\n    # Label features and join to dataset\n    X_cd = pd.DataFrame(\n        X_cd, columns=[f\"Centroid_{i}\" for i in range(X_cd.shape[1])]\n    )\n    return X_cd\n","metadata":{"_kg_hide-input":true,"lines_to_next_cell":2,"execution":{"iopub.status.busy":"2022-07-12T13:55:25.613134Z","iopub.execute_input":"2022-07-12T13:55:25.614039Z","iopub.status.idle":"2022-07-12T13:55:25.624305Z","shell.execute_reply.started":"2022-07-12T13:55:25.613998Z","shell.execute_reply":"2022-07-12T13:55:25.622944Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## Principal Component Analysis ##\n\nPCA was the second unsupervised model we used for feature creation. We saw how it could be used to decompose the variational structure in the data. The PCA algorithm gave us *loadings* which described each component of variation, and also the *components* which were the transformed datapoints. The loadings can suggest features to create and the components we can use as features directly.\n\nHere are the utility functions from the PCA lesson:","metadata":{}},{"cell_type":"code","source":"\ndef apply_pca(X, standardize=True):\n    # Standardize\n    if standardize:\n        X = (X - X.mean(axis=0)) / X.std(axis=0)\n    # Create principal components\n    pca = PCA()\n    X_pca = pca.fit_transform(X)\n    # Convert to dataframe\n    component_names = [f\"PC{i+1}\" for i in range(X_pca.shape[1])]\n    X_pca = pd.DataFrame(X_pca, columns=component_names)\n    # Create loadings\n    loadings = pd.DataFrame(\n        pca.components_.T,  # transpose the matrix of loadings\n        columns=component_names,  # so the columns are the principal components\n        index=X.columns,  # and the rows are the original features\n    )\n    return pca, X_pca, loadings\n\n\ndef plot_variance(pca, width=8, dpi=100):\n    # Create figure\n    fig, axs = plt.subplots(1, 2)\n    n = pca.n_components_\n    grid = np.arange(1, n + 1)\n    # Explained variance\n    evr = pca.explained_variance_ratio_\n    axs[0].bar(grid, evr)\n    axs[0].set(\n        xlabel=\"Component\", title=\"% Explained Variance\", ylim=(0.0, 1.0)\n    )\n    # Cumulative Variance\n    cv = np.cumsum(evr)\n    axs[1].plot(np.r_[0, grid], np.r_[0, cv], \"o-\")\n    axs[1].set(\n        xlabel=\"Component\", title=\"% Cumulative Variance\", ylim=(0.0, 1.0)\n    )\n    # Set up figure\n    fig.set(figwidth=8, dpi=100)\n    return axs\n","metadata":{"_kg_hide-input":true,"execution":{"iopub.status.busy":"2022-07-12T13:59:54.278809Z","iopub.execute_input":"2022-07-12T13:59:54.279192Z","iopub.status.idle":"2022-07-12T13:59:54.290406Z","shell.execute_reply.started":"2022-07-12T13:59:54.279162Z","shell.execute_reply":"2022-07-12T13:59:54.289226Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"And here are transforms that produce the features from the Exercise 5. You might want to change these if you came up with a different answer.\n","metadata":{}},{"cell_type":"code","source":"\ndef pca_inspired(df):\n    X = pd.DataFrame()\n    X[\"Feature1\"] = df.GrLivArea + df.TotalBsmtSF\n    X[\"Feature2\"] = df.YearRemodAdd * df.TotalBsmtSF\n    return X\n\n\ndef pca_components(df, features):\n    X = df.loc[:, features]\n    _, X_pca, _ = apply_pca(X)\n    return X_pca\n\n\npca_features = [\n    \"GarageArea\",\n    \"YearRemodAdd\",\n    \"TotalBsmtSF\",\n    \"GrLivArea\",\n]","metadata":{"_kg_hide-input":true,"execution":{"iopub.status.busy":"2022-07-12T14:00:24.627895Z","iopub.execute_input":"2022-07-12T14:00:24.628237Z","iopub.status.idle":"2022-07-12T14:00:24.635241Z","shell.execute_reply.started":"2022-07-12T14:00:24.628210Z","shell.execute_reply":"2022-07-12T14:00:24.633940Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"These are only a couple ways you could use the principal components. You could also try clustering using one or more components. One thing to note is that PCA doesn't change the distance between points -- it's just like a rotation. So clustering with the full set of components is the same as clustering with the original features. Instead, pick some subset of components, maybe those with the most variance or the highest MI scores.\n\nFor further analysis, you might want to look at a correlation matrix for the dataset:","metadata":{}},{"cell_type":"code","source":"def corrplot(df, method=\"pearson\", annot=True, **kwargs):\n    sns.clustermap(\n        df.corr(method),\n        vmin=-1.0,\n        vmax=1.0,\n        cmap=\"icefire\",\n        method=\"complete\",\n        annot=annot,\n        **kwargs,\n    )\n\n\ncorrplot(df_train, annot=None)","metadata":{"execution":{"iopub.status.busy":"2022-07-12T14:01:50.754279Z","iopub.execute_input":"2022-07-12T14:01:50.754706Z","iopub.status.idle":"2022-07-12T14:01:51.954611Z","shell.execute_reply.started":"2022-07-12T14:01:50.754670Z","shell.execute_reply":"2022-07-12T14:01:51.953402Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Groups of highly correlated features often yield interesting loadings.\n\n### PCA Application - Indicate Outliers ###\n\nIn Exercise 5, you applied PCA to determine houses that were **outliers**, that is, houses having values not well represented in the rest of the data. You saw that there was a group of houses in the `Edwards` neighborhood having a `SaleCondition` of `Partial` whose values were especially extreme.\n\nSome models can benefit from having these outliers indicated, which is what this next transform will do.","metadata":{}},{"cell_type":"code","source":"def indicate_outliers(df):\n    X_new = pd.DataFrame()\n    X_new[\"Outlier\"] = (df.Neighborhood == \"Edwards\") & (df.SaleCondition == \"Partial\")\n    return X_new\n","metadata":{"execution":{"iopub.status.busy":"2022-07-12T14:06:12.612169Z","iopub.execute_input":"2022-07-12T14:06:12.612605Z","iopub.status.idle":"2022-07-12T14:06:12.618047Z","shell.execute_reply.started":"2022-07-12T14:06:12.612569Z","shell.execute_reply":"2022-07-12T14:06:12.617268Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"You could also consider applying some sort of robust scaler from scikit-learn's `sklearn.preprocessing` module to the outlying values, especially those in `GrLivArea`. [Here](https://scikit-learn.org/stable/auto_examples/preprocessing/plot_all_scaling.html) is a tutorial illustrating some of them. Another option could be to create a feature of \"outlier scores\" using one of scikit-learn's [outlier detectors](https://scikit-learn.org/stable/modules/outlier_detection.html).","metadata":{}},{"cell_type":"markdown","source":"## Target Encoding ##\n\nNeeding a separate holdout set to create a target encoding is rather wasteful of data. In *Tutorial 6* we used 25% of our dataset just to encode a single feature, `Zipcode`. The data from the other features in that 25% we didn't get to use at all.\n\nThere is, however, a way you can use target encoding without having to use held-out encoding data. It's basically the same trick used in cross-validation:\n1. Split the data into folds, each fold having two splits of the dataset.\n2. Train the encoder on one split but transform the values of the other.\n3. Repeat for all the splits.\n\nThis way, training and transformation always take place on independent sets of data, just like when you use a holdout set but without any data going to waste.\n\nIn the next hidden cell is a wrapper you can use with any target encoder:","metadata":{}},{"cell_type":"code","source":"\nclass CrossFoldEncoder:\n    def __init__(self, encoder, **kwargs):\n        self.encoder_ = encoder\n        self.kwargs_ = kwargs  # keyword arguments for the encoder\n        self.cv_ = KFold(n_splits=5)\n\n    # Fit an encoder on one split and transform the feature on the\n    # other. Iterating over the splits in all folds gives a complete\n    # transformation. We also now have one trained encoder on each\n    # fold.\n    def fit_transform(self, X, y, cols):\n        self.fitted_encoders_ = []\n        self.cols_ = cols\n        X_encoded = []\n        for idx_encode, idx_train in self.cv_.split(X):\n            fitted_encoder = self.encoder_(cols=cols, **self.kwargs_)\n            fitted_encoder.fit(\n                X.iloc[idx_encode, :], y.iloc[idx_encode],\n            )\n            X_encoded.append(fitted_encoder.transform(X.iloc[idx_train, :])[cols])\n            self.fitted_encoders_.append(fitted_encoder)\n        X_encoded = pd.concat(X_encoded)\n        X_encoded.columns = [name + \"_encoded\" for name in X_encoded.columns]\n        return X_encoded\n\n    # To transform the test data, average the encodings learned from\n    # each fold.\n    def transform(self, X):\n        from functools import reduce\n\n        X_encoded_list = []\n        for fitted_encoder in self.fitted_encoders_:\n            X_encoded = fitted_encoder.transform(X)\n            X_encoded_list.append(X_encoded[self.cols_])\n        X_encoded = reduce(\n            lambda x, y: x.add(y, fill_value=0), X_encoded_list\n        ) / len(X_encoded_list)\n        X_encoded.columns = [name + \"_encoded\" for name in X_encoded.columns]\n        return X_encoded\n","metadata":{"_kg_hide-input":true,"execution":{"iopub.status.busy":"2022-07-12T14:20:28.241968Z","iopub.execute_input":"2022-07-12T14:20:28.242331Z","iopub.status.idle":"2022-07-12T14:20:28.253920Z","shell.execute_reply.started":"2022-07-12T14:20:28.242302Z","shell.execute_reply":"2022-07-12T14:20:28.252229Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Use it like:\n\n```\nencoder = CrossFoldEncoder(MEstimateEncoder, m=1)\nX_encoded = encoder.fit_transform(X, y, cols=[\"MSSubClass\"]))\n```\n\nYou can turn any of the encoders from the [`category_encoders`](http://contrib.scikit-learn.org/category_encoders/) library into a cross-fold encoder. The [`CatBoostEncoder`](http://contrib.scikit-learn.org/category_encoders/catboost.html) would be worth trying. It's similar to `MEstimateEncoder` but uses some tricks to better prevent overfitting. Its smoothing parameter is called `a` instead of `m`.\n\n## Create Final Feature Set ##\n\nNow let's combine everything together. Putting the transformations into separate functions makes it easier to experiment with various combinations. The ones I left uncommented I found gave the best results. You should experiment with you own ideas though! Modify any of these transformations or come up with some of your own to add to the pipeline.","metadata":{}},{"cell_type":"code","source":"def create_features(df, df_test=None):\n    X = df.copy()\n    y = X.pop(\"SalePrice\")\n    mi_scores = make_mi_scores(X, y)\n\n    # Combine splits if test data is given\n    #\n    # If we're creating features for test set predictions, we should\n    # use all the data we have available. After creating our features,\n    # we'll recreate the splits.\n    if df_test is not None:\n        X_test = df_test.copy()\n        X_test.pop(\"SalePrice\")\n        X = pd.concat([X, X_test])\n\n    # Lesson 2 - Mutual Information\n    X = drop_uninformative(X, mi_scores)\n\n    # Lesson 3 - Transformations\n    X = X.join(mathematical_transforms(X))\n    X = X.join(interactions(X))\n    X = X.join(counts(X))\n    # X = X.join(break_down(X))\n    X = X.join(group_transforms(X))\n\n    # Lesson 4 - Clustering\n    # X = X.join(cluster_labels(X, cluster_features, n_clusters=20))\n    # X = X.join(cluster_distance(X, cluster_features, n_clusters=20))\n\n    # Lesson 5 - PCA\n    X = X.join(pca_inspired(X))\n    # X = X.join(pca_components(X, pca_features))\n    # X = X.join(indicate_outliers(X))\n\n    X = label_encode(X)\n\n    # Reform splits\n    if df_test is not None:\n        X_test = X.loc[df_test.index, :]\n        X.drop(df_test.index, inplace=True)\n\n    # Lesson 6 - Target Encoder\n    encoder = CrossFoldEncoder(MEstimateEncoder, m=1)\n    X = X.join(encoder.fit_transform(X, y, cols=[\"MSSubClass\"]))\n    if df_test is not None:\n        X_test = X_test.join(encoder.transform(X_test))\n\n    if df_test is not None:\n        return X, X_test\n    else:\n        return X\n\n\ndf_train, df_test = load_data()\nX_train = create_features(df_train)\ny_train = df_train.loc[:, \"SalePrice\"]\n\nscore_dataset(X_train, y_train)","metadata":{"execution":{"iopub.status.busy":"2022-07-12T14:20:32.115077Z","iopub.execute_input":"2022-07-12T14:20:32.115467Z","iopub.status.idle":"2022-07-12T14:20:37.341047Z","shell.execute_reply.started":"2022-07-12T14:20:32.115434Z","shell.execute_reply":"2022-07-12T14:20:37.340223Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Step 4 - Hyperparameter Tuning #\n\nAt this stage, you might like to do some hyperparameter tuning with XGBoost before creating your final submission.","metadata":{}},{"cell_type":"code","source":"X_train = create_features(df_train)\ny_train = df_train.loc[:, \"SalePrice\"]\n\nxgb_params = dict(\n    max_depth=6,           # maximum depth of each tree - try 2 to 10\n    learning_rate=0.01,    # effect of each tree - try 0.0001 to 0.1\n    n_estimators=1000,     # number of trees (that is, boosting rounds) - try 1000 to 8000\n    min_child_weight=1,    # minimum number of houses in a leaf - try 1 to 10\n    colsample_bytree=0.7,  # fraction of features (columns) per tree - try 0.2 to 1.0\n    subsample=0.7,         # fraction of instances (rows) per tree - try 0.2 to 1.0\n    reg_alpha=0.5,         # L1 regularization (like LASSO) - try 0.0 to 10.0\n    reg_lambda=1.0,        # L2 regularization (like Ridge) - try 0.0 to 10.0\n    num_parallel_tree=1,   # set > 1 for boosted random forests\n)\n\nxgb = XGBRegressor(**xgb_params)\nscore_dataset(X_train, y_train, xgb)","metadata":{"execution":{"iopub.status.busy":"2022-07-12T14:20:49.328980Z","iopub.execute_input":"2022-07-12T14:20:49.329348Z","iopub.status.idle":"2022-07-12T14:21:16.893164Z","shell.execute_reply.started":"2022-07-12T14:20:49.329309Z","shell.execute_reply":"2022-07-12T14:21:16.892025Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Just tuning these by hand can give you great results. However, you might like to try using one of scikit-learn's automatic [hyperparameter tuners](https://scikit-learn.org/stable/modules/grid_search.html). Or you could explore more advanced tuning libraries like [Optuna](https://optuna.readthedocs.io/en/stable/index.html) or [scikit-optimize](https://scikit-optimize.github.io/stable/).\n\nHere is how you can use Optuna with XGBoost:\n\n```\nimport optuna\n\ndef objective(trial):\n    xgb_params = dict(\n        max_depth=trial.suggest_int(\"max_depth\", 2, 10),\n        learning_rate=trial.suggest_float(\"learning_rate\", 1e-4, 1e-1, log=True),\n        n_estimators=trial.suggest_int(\"n_estimators\", 1000, 8000),\n        min_child_weight=trial.suggest_int(\"min_child_weight\", 1, 10),\n        colsample_bytree=trial.suggest_float(\"colsample_bytree\", 0.2, 1.0),\n        subsample=trial.suggest_float(\"subsample\", 0.2, 1.0),\n        reg_alpha=trial.suggest_float(\"reg_alpha\", 1e-4, 1e2, log=True),\n        reg_lambda=trial.suggest_float(\"reg_lambda\", 1e-4, 1e2, log=True),\n    )\n    xgb = XGBRegressor(**xgb_params)\n    return score_dataset(X_train, y_train, xgb)\n\nstudy = optuna.create_study(direction=\"minimize\")\nstudy.optimize(objective, n_trials=20)\nxgb_params = study.best_params\n```\n\nCopy this into a code cell if you'd like to use it, but be aware that it will take quite a while to run. After it's done, you might enjoy using some of [Optuna's visualizations](https://optuna.readthedocs.io/en/stable/tutorial/10_key_features/005_visualization.html).\n\n# Step 5 - Train Model and Create Submissions #\n\nOnce you're satisfied with everything, it's time to create your final predictions! This cell will:\n- create your feature set from the original data\n- train XGBoost on the training data\n- use the trained model to make predictions from the test set\n- save the predictions to a CSV file","metadata":{}},{"cell_type":"code","source":"X_train, X_test = create_features(df_train, df_test)\ny_train = df_train.loc[:, \"SalePrice\"]\n\nxgb = XGBRegressor(**xgb_params)\n# XGB minimizes MSE, but competition loss is RMSLE\n# So, we need to log-transform y to train and exp-transform the predictions\nxgb.fit(X_train, np.log(y))\npredictions = np.exp(xgb.predict(X_test))\n\noutput = pd.DataFrame({'Id': X_test.index, 'SalePrice': predictions})\noutput.to_csv('my_submission.csv', index=False)\nprint(\"Your submission was successfully saved!\")","metadata":{"execution":{"iopub.status.busy":"2022-07-12T14:22:10.155395Z","iopub.execute_input":"2022-07-12T14:22:10.155788Z","iopub.status.idle":"2022-07-12T14:22:17.563400Z","shell.execute_reply.started":"2022-07-12T14:22:10.155757Z","shell.execute_reply":"2022-07-12T14:22:17.562272Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"To submit these predictions to the competition, follow these steps:\n\n1. Begin by clicking on the blue **Save Version** button in the top right corner of the window.  This will generate a pop-up window.\n2. Ensure that the **Save and Run All** option is selected, and then click on the blue **Save** button.\n3. This generates a window in the bottom left corner of the notebook.  After it has finished running, click on the number to the right of the **Save Version** button.  This pulls up a list of versions on the right of the screen.  Click on the ellipsis **(...)** to the right of the most recent version, and select **Open in Viewer**.  This brings you into view mode of the same page. You will need to scroll down to get back to these instructions.\n4. Click on the **Output** tab on the right of the screen.  Then, click on the file you would like to submit, and click on the blue **Submit** button to submit your results to the leaderboard.\n\nYou have now successfully submitted to the competition!\n\n# Next Steps #\n\nIf you want to keep working to improve your performance, select the blue **Edit** button in the top right of the screen. Then you can change your code and repeat the process. There's a lot of room to improve, and you will climb up the leaderboard as you work.\n\nBe sure to check out [other users' notebooks](https://www.kaggle.com/c/house-prices-advanced-regression-techniques/notebooks) in this competition. You'll find lots of great ideas for new features and as well as other ways to discover more things about the dataset or make better predictions. There's also the [discussion forum](https://www.kaggle.com/c/house-prices-advanced-regression-techniques/discussion), where you can share ideas with other Kagglers.\n\nHave fun Kaggling!","metadata":{}},{"cell_type":"markdown","source":"---\n\n\n\n\n*Have questions or comments? Visit the [course discussion forum](https://www.kaggle.com/learn/feature-engineering/discussion) to chat with other learners.*","metadata":{}}]}