{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"pygments_lexer":"ipython3","nbconvert_exporter":"python","version":"3.6.4","file_extension":".py","codemirror_mode":{"name":"ipython","version":3},"name":"python","mimetype":"text/x-python"}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"<center><img src=https://www.mihaileric.com/static/model-selection-meme-bd4a6a86f615583d1a1bbc497ca4640e-67414.jpeg></center>\n<center>img source: towards data science</center>\n\n# <center><b>(nested) Cross-Validation⚙️</b></center>\n\n**What you can expect from this notebook:** This notebook covers cross-validation in theory and practice and emphasises why nested cross-validation can and often should be used instead of simple cross-validation. \n\n**libraries utilised:**","metadata":{"_uuid":"8f2839f25d086af736a60e9eeb907d3b93b6e0e5","_cell_guid":"b1076dfc-b9ad-4769-8c92-a6c4dae69d19"}},{"cell_type":"code","source":"#  for implementation\nimport numpy as np  # linear algebra\nfrom copy import copy  # deep copies of objects\n\n#  for evaluation\nfrom sklearn.linear_model import Ridge\nfrom sklearn.tree import DecisionTreeRegressor\nfrom sklearn.ensemble import RandomForestRegressor, GradientBoostingRegressor\n\n#  for data handeling and scaling\nimport pandas as pd  # data handeling\nfrom sklearn.preprocessing import StandardScaler  # Standartise data","metadata":{"execution":{"iopub.status.busy":"2022-07-02T09:40:30.210119Z","iopub.execute_input":"2022-07-02T09:40:30.210496Z","iopub.status.idle":"2022-07-02T09:40:30.217111Z","shell.execute_reply.started":"2022-07-02T09:40:30.210464Z","shell.execute_reply":"2022-07-02T09:40:30.215901Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"****\n# <b>1 <span style=\"color:#ebd1a4\">|</span> Basic cross-validation intuition</b>\n\n**Main idea:** utilising ***all*** data avaliable for training ***and*** testing.\n\n**Use cases:**\n* Model evaluation (how does my model perform taking into account *all* my data?)\n* Model selection (which model performs best taking into account *all* my data?)\n* Hyperparameter tuning (Gridsearch -> like model selection but between the same model, just with different hyperparameter settings)\n\n<center><img src=https://www.dummies.com/wp-content/uploads/9781119245513-fg1104.jpg></center>\n<center>img source: ww.dumies.com</center>\n\n**How it works:**\n* **Partition** data into groups\n* In a **rotating** manner:\n    * **Set aside** a group for testing\n    * Train models on the **rest**\n    * Test these models with the testing group and keep track of the scores\n    * Repeat but with **another group** being set aside as testing group\n* After each group was used for testing:\n    * **Average scores** to obtain an estimate **based on the whole data**\n    \n**Advantages:**\n* The obtained scores are a better estimate of how the model will perform on real-world data than the ones one could obtain by just testing on a fixed subset of the data\n\n**Disadvantages:**\n* Each model has to be trained multiple times \n\n**Some common cross-validation variants:**\n* ***K-Fold Cross-Validation:*** The most widely used one and also the one focussed on in this notebook:\n    * Partition data into k folds(groups) of equal size\n    * Test on one, train on the rest\n    * Then rotate\n    * **visualisation above**\n* ***Leave one out Cross-Validation:***\n    * Similar to K-Fold Cross-Validation\n    * But partitions into n folds equal to the number of observations in the dataset\n    * resulting in single item testing sets\n* ***Stratfield Cross-Validation:***\n    * Mainly used for classification problems\n    * Similar to K-Fold Cross-Validation, it tests on a set percentage of the data\n    * But partitions the data in a way that the mean response value is roughly equal across folds\n    \n<br>\n\n#### **Python implementation of K-Fold Cross-Validation:**\nThis implementation will be used in an example after the next section.","metadata":{}},{"cell_type":"code","source":"#  sample implementation\ndef k_fold_crossval(models, X, y, k):\n    #  initialise empty lists to store scores and trained models\n    error_scores = []\n    trained_models = []\n    #  split dataset into folds\n    folds_X = np.array_split(X, k)\n    folds_y = np.array_split(y, k)\n    #  iterate over folds\n    for test_fold_i in range(k):\n        #  log\n        print(f\"fold {test_fold_i+1}: ---------------------------------------------\")\n        #  specify train/test data\n        X_train = folds_X[:test_fold_i] + folds_X[test_fold_i+1:]\n        y_train = folds_y[:test_fold_i] + folds_y[test_fold_i+1:]\n        X_test = folds_X[test_fold_i]\n        y_test = folds_y[test_fold_i]\n        #  train on everything but the test_fold\n        fold_scores = []\n        fold_models = []\n        for model in self.models:\n            #  make deep copy of model\n            model = copy(model)\n            #  fit and test\n            model.fit(np.concatenate(X_train), np.concatenate(y_train))\n            fold_models.append(model)\n            score = model.score(X_test, y_test)\n            fold_scores.append(score)\n            #  log\n            print(f\"{model} trained, score = {score}\")\n        #  save fold scores and models\n        error_scores.append(np.array(fold_scores))\n        trained_models.append(np.array(fold_models, dtype=object))\n    return error_scores, trained_models","metadata":{"execution":{"iopub.status.busy":"2022-07-02T08:55:10.551819Z","iopub.execute_input":"2022-07-02T08:55:10.552260Z","iopub.status.idle":"2022-07-02T08:55:10.562708Z","shell.execute_reply.started":"2022-07-02T08:55:10.552226Z","shell.execute_reply":"2022-07-02T08:55:10.561761Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"****\n\n# <b>2 <span style=\"color:#ebd1a4\">|</span> The often accuring problem</b>\n\n**Cross-validation can be used for both model evaluation and selection/hyperparameter tuning, but not for both at the same time without running into potential problems.**<br>\n\n**Example Situation:**\n* Someone performs 5-fold cross-validation to select between a catboost, xgboost and lightgbm algorithm based on which one obtains the lowest loss/highest score\n* Since this already resulted in a loss value he reports that loss as an estimate of how the model is going to perform in production\n* Unfortunately the model ends up performing slightly worse than expected\n\n**What happened?**\n* When selecting the best model, **bias** was introduced, since with that he did also choose the lowest loss as an estimate\n* The obtained loss values are just estimates and have a sampling distribution\n* When choosing the lowest value, it increases the chance of it coming from a lower end of that distribution \n* -> **After selecting a model based on loss, that loss estimate is biased**","metadata":{}},{"cell_type":"markdown","source":"****\n\n# <b>3 <span style=\"color:#ebd1a4\">|</span> How it can be fixed with nested cross-validation</b>\n\n**The solution is stunningly simple:** The evaluation process is separated from the selection process, splitting into train, validation and test set, where the validation set is used for the model selection/hyperparameter optimization and the test set for the final evaluation.\n\n**This can be done using all of the data for training, validation and testing, by utilising 2 cross-validation loops, nested into each other:**\n\n<center><img src=https://vitalflux.com/wp-content/uploads/2020/08/Screenshot-2020-08-30-at-6.33.47-PM.png></center>\n<center>img source: vitalflux.com</center>\n\n**Step for step:**\n* partition training data into folds (outer loop), **rotating testing fold**:\n    * set aside testing set\n    * partition the training folds in each iteration into folds (inner loop), **rotating the validation fold**:\n        * set aside validation set\n        * fit models on rest of training set\n        * **evaluate on validation set**\n        * repeat till all the training data from the outer loop was used for validation\n    * **choose best model** based on scores of inner loop\n    * **train** that model **on the whole training set** from the outer loop\n    * evaluate on testing sett\n    * repeat till all the data was used for testing\n    \n* This results in one(or more) best model(s) with an unbiased loss estimate\n    \n<br>\n\n#### **Python implementation of K-Fold Cross-Validation:**\nTaking the single loop created in the last section, it is extended in a way that it can be nested into each other. This is done by making it possible to be treated similar to a sklearn model(with .fit and .score):","metadata":{}},{"cell_type":"code","source":"#  turn into class to work with sklearn model pipeline and allow for easy nesting of cross-validation loops\nclass KfoldCrossval:\n    def __init__(self, k, models, name=\"k_fold_crossval\"):\n        self.k = k\n        self.models = models\n        self.name = name\n        \n    def evaluate(self, X, y):\n        self.fit(X, y)\n        return self.error_scores, self.trained_models\n        \n    def fit(self, X, y, verbose=True):\n        #  initialise empty lists to store scores and trained models\n        self.error_scores = []\n        self.trained_models = []\n        #self.fold_scores = []\n        #  split dataset into folds\n        folds_X = np.array_split(X, self.k)\n        folds_y = np.array_split(y, self.k)\n        #  iterate over folds\n        for test_fold_i in range(self.k):\n            #  log\n            if verbose:\n                print(f\"{self.name} fold {test_fold_i+1}: ---------------------------------------------\")\n            #  specify train/test data\n            X_train = folds_X[:test_fold_i] + folds_X[test_fold_i+1:]\n            y_train = folds_y[:test_fold_i] + folds_y[test_fold_i+1:]\n            X_test = folds_X[test_fold_i]\n            y_test = folds_y[test_fold_i]\n            #  train on everything but the test_fold\n            fold_scores = []\n            fold_models = []\n            for model in self.models:\n                #  make deep copy of model\n                model = copy(model)\n                #  fit and test\n                model.fit(np.concatenate(X_train), np.concatenate(y_train))\n                fold_models.append(model)\n                score = model.score(X_test, y_test)\n                fold_scores.append(score)\n                #  log\n                if verbose:\n                    print(f\"{model} trained, score = {score}\")\n            #  save fold scores and models\n            self.error_scores.append(np.array(fold_scores))\n            self.trained_models.append(np.array(fold_models, dtype=object))\n        \n        #  just neccessary if inner loop:\n        #  find best model\n        scores = np.stack(self.error_scores).mean(axis=0)\n        self.best_model = copy(np.array(self.models)[scores == scores.max()][0])\n        #  fit best model on whole training data\n        self.best_model.fit(X, y)\n    \n    def score(self, X_test, y_test):\n        return self.best_model.score(X_test, y_test)","metadata":{"execution":{"iopub.status.busy":"2022-07-02T09:04:05.754888Z","iopub.execute_input":"2022-07-02T09:04:05.755425Z","iopub.status.idle":"2022-07-02T09:04:05.775263Z","shell.execute_reply.started":"2022-07-02T09:04:05.755379Z","shell.execute_reply":"2022-07-02T09:04:05.773273Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"```\nmodels = [model1, model2, ..., modeln]\n\n#  example usage\ninner = KfoldCrossval(10, models)\nouter = KfoldCrossval(10, [inner])\n\nouter.evaluate(X, y)\n```","metadata":{}},{"cell_type":"markdown","source":"****\n\n# <b>4 <span style=\"color:#ebd1a4\">|</span> Sample usage of nested cross-validation</b>","metadata":{}},{"cell_type":"code","source":"house_data = pd.read_csv(\"../input/house-prices-advanced-regression-techniques/train.csv\")\nscaler = StandardScaler()\n\n# encode categorical features, fill nan values\nfor feature in house_data.columns:\n    if house_data[feature].dtype == \"object\":\n        #  categorical encoding: turns [a, b, b, c] into [1, 2, 2, 3]\n        house_data[feature] = house_data[feature].astype(\"category\").cat.codes \n        if house_data[feature].isna().sum() != 0:\n            #  impute missing values with the mode of the corresponding variable\n            house_data[feature].fillna(house_data[feature].mode(), inplace=True)\n    else:\n        if house_data[feature].isna().sum() != 0:\n            #  impute missing values with the mean of the corresponding variable\n            house_data[feature].fillna(house_data[feature].mean(), inplace=True)\n            \nfeatures = house_data.loc[:, house_data.columns!=\"SalePrice\"].to_numpy()\nlabels = house_data.loc[:, \"SalePrice\"].to_numpy()\n\nscaler.fit(features)\nfeatures = scaler.transform(features)","metadata":{"execution":{"iopub.status.busy":"2022-07-02T09:39:02.269713Z","iopub.execute_input":"2022-07-02T09:39:02.270124Z","iopub.status.idle":"2022-07-02T09:39:02.378272Z","shell.execute_reply.started":"2022-07-02T09:39:02.270082Z","shell.execute_reply":"2022-07-02T09:39:02.377186Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"house_data.head()","metadata":{"execution":{"iopub.status.busy":"2022-07-02T09:39:03.510507Z","iopub.execute_input":"2022-07-02T09:39:03.510884Z","iopub.status.idle":"2022-07-02T09:39:03.537675Z","shell.execute_reply.started":"2022-07-02T09:39:03.510854Z","shell.execute_reply":"2022-07-02T09:39:03.536520Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"#  extreamly basic configuration, just for demonstation purposes\nmodels = [\n    Ridge(1),\n    DecisionTreeRegressor(max_depth=50, min_samples_split=10),\n    RandomForestRegressor(n_estimators=50),\n    #GradientBoostingRegressor(loss=\"squared_error\", min_samples_split=10)\n]","metadata":{"execution":{"iopub.status.busy":"2022-07-02T09:39:18.714375Z","iopub.execute_input":"2022-07-02T09:39:18.714836Z","iopub.status.idle":"2022-07-02T09:39:18.721680Z","shell.execute_reply.started":"2022-07-02T09:39:18.714795Z","shell.execute_reply":"2022-07-02T09:39:18.720280Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"**6-fold inner, 6-fold outer cross-validation:**","metadata":{}},{"cell_type":"code","source":"inner = KfoldCrossval(6, models, \"inner\")\nouter = KfoldCrossval(6, [inner], \"outer\")\n\nbest_scores, best_models = outer.evaluate(features, labels)","metadata":{"execution":{"iopub.status.busy":"2022-07-02T09:39:19.233911Z","iopub.execute_input":"2022-07-02T09:39:19.234823Z","iopub.status.idle":"2022-07-02T09:40:04.406669Z","shell.execute_reply.started":"2022-07-02T09:39:19.234758Z","shell.execute_reply":"2022-07-02T09:40:04.405407Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"best_models = [crossval[0].best_model for crossval in best_models]\nprint(best_models, best_scores)","metadata":{"execution":{"iopub.status.busy":"2022-07-02T09:40:05.784972Z","iopub.execute_input":"2022-07-02T09:40:05.785678Z","iopub.status.idle":"2022-07-02T09:40:05.794594Z","shell.execute_reply.started":"2022-07-02T09:40:05.785639Z","shell.execute_reply":"2022-07-02T09:40:05.793207Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"As expectable the random forest regressor showed to be the best model, the mean of the outer crossvalidation score gives an estimate on how it will perform on real world data:","metadata":{}},{"cell_type":"code","source":"#  get average score -> better estimate\nnp.stack(best_scores).mean()","metadata":{"execution":{"iopub.status.busy":"2022-07-02T09:40:11.045082Z","iopub.execute_input":"2022-07-02T09:40:11.045534Z","iopub.status.idle":"2022-07-02T09:40:11.054250Z","shell.execute_reply.started":"2022-07-02T09:40:11.045498Z","shell.execute_reply":"2022-07-02T09:40:11.052836Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"**That's all for this notebook, have a great day and happy learning!👋**","metadata":{}}]}