{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"pygments_lexer":"ipython3","nbconvert_exporter":"python","version":"3.6.4","file_extension":".py","codemirror_mode":{"name":"ipython","version":3},"name":"python","mimetype":"text/x-python"}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"**This notebook is an exercise in the [Intermediate Machine Learning](https://www.kaggle.com/learn/intermediate-machine-learning) course.  You can reference the tutorial at [this link](https://www.kaggle.com/alexisbcook/cross-validation).**\n\n---\n","metadata":{}},{"cell_type":"markdown","source":"In this exercise, you will leverage what you've learned to tune a machine learning model with **cross-validation**.\n\n# Setup\n\nThe questions below will give you feedback on your work. Run the following cell to set up the feedback system.","metadata":{}},{"cell_type":"code","source":"# Set up code checking\nimport os\nif not os.path.exists(\"../input/train.csv\"):\n    os.symlink(\"../input/home-data-for-ml-course/train.csv\", \"../input/train.csv\")  \n    os.symlink(\"../input/home-data-for-ml-course/test.csv\", \"../input/test.csv\") \nfrom learntools.core import binder\nbinder.bind(globals())\nfrom learntools.ml_intermediate.ex5 import *\nprint(\"Setup Complete\")","metadata":{"execution":{"iopub.status.busy":"2022-08-05T05:13:49.145593Z","iopub.execute_input":"2022-08-05T05:13:49.146010Z","iopub.status.idle":"2022-08-05T05:13:49.156270Z","shell.execute_reply.started":"2022-08-05T05:13:49.145978Z","shell.execute_reply":"2022-08-05T05:13:49.155466Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"You will work with the [Housing Prices Competition for Kaggle Learn Users](https://www.kaggle.com/c/home-data-for-ml-course) from the previous exercise. \n\n![Ames Housing dataset image](https://i.imgur.com/lTJVG4e.png)\n\nRun the next code cell without changes to load the training and test data in `X` and `X_test`.  For simplicity, we drop categorical variables.","metadata":{}},{"cell_type":"code","source":"import pandas as pd\nfrom sklearn.model_selection import train_test_split\n\n# Read the data\ntrain_data = pd.read_csv('../input/train.csv', index_col='Id')\ntest_data = pd.read_csv('../input/test.csv', index_col='Id')\n\n# Remove rows with missing target, separate target from predictors\ntrain_data.dropna(axis=0, subset=['SalePrice'], inplace=True)\ny = train_data.SalePrice              \ntrain_data.drop(['SalePrice'], axis=1, inplace=True)\n\n# Select numeric columns only\nnumeric_cols = [cname for cname in train_data.columns if train_data[cname].dtype in ['int64', 'float64']]\nX = train_data[numeric_cols].copy()\nX_test = test_data[numeric_cols].copy()","metadata":{"execution":{"iopub.status.busy":"2022-08-05T05:13:51.688035Z","iopub.execute_input":"2022-08-05T05:13:51.689067Z","iopub.status.idle":"2022-08-05T05:13:51.748800Z","shell.execute_reply.started":"2022-08-05T05:13:51.689001Z","shell.execute_reply":"2022-08-05T05:13:51.747897Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Use the next code cell to print the first several rows of the data.","metadata":{}},{"cell_type":"code","source":"X.head()","metadata":{"execution":{"iopub.status.busy":"2022-08-05T05:13:53.789794Z","iopub.execute_input":"2022-08-05T05:13:53.790758Z","iopub.status.idle":"2022-08-05T05:13:53.819331Z","shell.execute_reply.started":"2022-08-05T05:13:53.790709Z","shell.execute_reply":"2022-08-05T05:13:53.817944Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"So far, you've learned how to build pipelines with scikit-learn.  For instance, the pipeline below will use [`SimpleImputer()`](https://scikit-learn.org/stable/modules/generated/sklearn.impute.SimpleImputer.html) to replace missing values in the data, before using [`RandomForestRegressor()`](https://scikit-learn.org/stable/modules/generated/sklearn.ensemble.RandomForestRegressor.html) to train a random forest model to make predictions.  We set the number of trees in the random forest model with the `n_estimators` parameter, and setting `random_state` ensures reproducibility.","metadata":{}},{"cell_type":"code","source":"from sklearn.ensemble import RandomForestRegressor\nfrom sklearn.pipeline import Pipeline\nfrom sklearn.impute import SimpleImputer\n\nmy_pipeline = Pipeline(steps=[\n    ('preprocessor', SimpleImputer()),\n    ('model', RandomForestRegressor(n_estimators=50, random_state=0))\n])","metadata":{"execution":{"iopub.status.busy":"2022-08-05T05:13:57.609666Z","iopub.execute_input":"2022-08-05T05:13:57.610312Z","iopub.status.idle":"2022-08-05T05:13:57.615874Z","shell.execute_reply.started":"2022-08-05T05:13:57.610276Z","shell.execute_reply":"2022-08-05T05:13:57.614822Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"You have also learned how to use pipelines in cross-validation.  The code below uses the [`cross_val_score()`](https://scikit-learn.org/stable/modules/generated/sklearn.model_selection.cross_val_score.html) function to obtain the mean absolute error (MAE), averaged across five different folds.  Recall we set the number of folds with the `cv` parameter.","metadata":{}},{"cell_type":"code","source":"from sklearn.model_selection import cross_val_score\n\n# Multiply by -1 since sklearn calculates *negative* MAE\nscores = -1 * cross_val_score(my_pipeline, X, y,\n                              cv=5,\n                              scoring='neg_mean_absolute_error')\n\nprint(\"Average MAE score:\", scores.mean())","metadata":{"execution":{"iopub.status.busy":"2022-08-05T05:14:01.739399Z","iopub.execute_input":"2022-08-05T05:14:01.739860Z","iopub.status.idle":"2022-08-05T05:14:04.786349Z","shell.execute_reply.started":"2022-08-05T05:14:01.739826Z","shell.execute_reply":"2022-08-05T05:14:04.785262Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Step 1: Write a useful function\n\nIn this exercise, you'll use cross-validation to select parameters for a machine learning model.\n\nBegin by writing a function `get_score()` that reports the average (over three cross-validation folds) MAE of a machine learning pipeline that uses:\n- the data in `X` and `y` to create folds,\n- `SimpleImputer()` (with all parameters left as default) to replace missing values, and\n- `RandomForestRegressor()` (with `random_state=0`) to fit a random forest model.\n\nThe `n_estimators` parameter supplied to `get_score()` is used when setting the number of trees in the random forest model.  ","metadata":{}},{"cell_type":"code","source":"def get_score(n_estimators):\n    \"\"\"Return the average MAE over 3 CV folds of random forest model.\n    \n    Keyword argument:\n    n_estimators -- the number of trees in the forest\n    \"\"\"\n    my_pipeline = Pipeline(steps=[ ('preprocessor', SimpleImputer()),\n                                  ('model', RandomForestRegressor(n_estimators, random_state=0))])\n    scores = cross_val_score(my_pipeline, X, y,\n                              cv=3,\n                              scoring='neg_mean_absolute_error')\n    mean_score = -1* scores.mean()\n    # Replace this body with your own code\n    return mean_score \n\n# Check your answer\nstep_1.check()","metadata":{"execution":{"iopub.status.busy":"2022-08-05T05:29:16.433575Z","iopub.execute_input":"2022-08-05T05:29:16.433998Z","iopub.status.idle":"2022-08-05T05:29:17.428761Z","shell.execute_reply.started":"2022-08-05T05:29:16.433962Z","shell.execute_reply":"2022-08-05T05:29:17.427781Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"import sklearn\nsorted(sklearn.metrics.SCORERS.keys())","metadata":{"execution":{"iopub.status.busy":"2022-08-05T05:27:09.417279Z","iopub.execute_input":"2022-08-05T05:27:09.418210Z","iopub.status.idle":"2022-08-05T05:27:09.426959Z","shell.execute_reply.started":"2022-08-05T05:27:09.418171Z","shell.execute_reply":"2022-08-05T05:27:09.425466Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Lines below will give you a hint or solution code\nstep_1.hint()\nstep_1.solution()","metadata":{"execution":{"iopub.status.busy":"2022-08-05T05:13:04.488483Z","iopub.execute_input":"2022-08-05T05:13:04.489391Z","iopub.status.idle":"2022-08-05T05:13:04.502216Z","shell.execute_reply.started":"2022-08-05T05:13:04.489351Z","shell.execute_reply":"2022-08-05T05:13:04.500718Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Step 2: Test different parameter values\n\nNow, you will use the function that you defined in Step 1 to evaluate the model performance corresponding to eight different values for the number of trees in the random forest: 50, 100, 150, ..., 300, 350, 400.\n\nStore your results in a Python dictionary `results`, where `results[i]` is the average MAE returned by `get_score(i)`.","metadata":{}},{"cell_type":"code","source":"results = {i:get_score(n_estimators = i) for i in range(50,401,50)} # Your code here\n\n# Check your answer\nstep_2.check()","metadata":{"execution":{"iopub.status.busy":"2022-08-05T05:31:52.898380Z","iopub.execute_input":"2022-08-05T05:31:52.898805Z","iopub.status.idle":"2022-08-05T05:32:47.608840Z","shell.execute_reply.started":"2022-08-05T05:31:52.898772Z","shell.execute_reply":"2022-08-05T05:32:47.608109Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Lines below will give you a hint or solution code\nstep_2.hint()\nstep_2.solution()","metadata":{"execution":{"iopub.status.busy":"2022-08-05T05:40:10.643097Z","iopub.execute_input":"2022-08-05T05:40:10.643491Z","iopub.status.idle":"2022-08-05T05:40:10.656318Z","shell.execute_reply.started":"2022-08-05T05:40:10.643462Z","shell.execute_reply":"2022-08-05T05:40:10.655491Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Use the next cell to visualize your results from Step 2.  Run the code without changes.","metadata":{}},{"cell_type":"code","source":"import matplotlib.pyplot as plt\n%matplotlib inline\n\nplt.plot(list(results.keys()), list(results.values()))\nplt.show()","metadata":{"execution":{"iopub.status.busy":"2022-08-05T05:40:25.843003Z","iopub.execute_input":"2022-08-05T05:40:25.843424Z","iopub.status.idle":"2022-08-05T05:40:26.236863Z","shell.execute_reply.started":"2022-08-05T05:40:25.843376Z","shell.execute_reply":"2022-08-05T05:40:26.236052Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Step 3: Find the best parameter value\n\nGiven the results, which value for `n_estimators` seems best for the random forest model?  Use your answer to set the value of `n_estimators_best`.","metadata":{}},{"cell_type":"code","source":"n_estimators_best = 200\n\n# Check your answer\nstep_3.check()","metadata":{"execution":{"iopub.status.busy":"2022-08-05T05:40:44.344965Z","iopub.execute_input":"2022-08-05T05:40:44.345353Z","iopub.status.idle":"2022-08-05T05:40:44.354277Z","shell.execute_reply.started":"2022-08-05T05:40:44.345320Z","shell.execute_reply":"2022-08-05T05:40:44.353174Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Lines below will give you a hint or solution code\n#step_3.hint()\n#step_3.solution()","metadata":{},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"In this exercise, you have explored one method for choosing appropriate parameters in a machine learning model.  \n\nIf you'd like to learn more about [hyperparameter optimization](https://en.wikipedia.org/wiki/Hyperparameter_optimization), you're encouraged to start with **grid search**, which is a straightforward method for determining the best _combination_ of parameters for a machine learning model.  Thankfully, scikit-learn also contains a built-in function [`GridSearchCV()`](https://scikit-learn.org/stable/modules/generated/sklearn.model_selection.GridSearchCV.html) that can make your grid search code very efficient!\n\n# Keep going\n\nContinue to learn about **[gradient boosting](https://www.kaggle.com/alexisbcook/xgboost)**, a powerful technique that achieves state-of-the-art results on a variety of datasets.","metadata":{}},{"cell_type":"markdown","source":"---\n\n\n\n\n*Have questions or comments? Visit the [course discussion forum](https://www.kaggle.com/learn/intermediate-machine-learning/discussion) to chat with other learners.*","metadata":{}}]}