{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"pygments_lexer":"ipython3","nbconvert_exporter":"python","version":"3.6.4","file_extension":".py","codemirror_mode":{"name":"ipython","version":3},"name":"python","mimetype":"text/x-python"}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"**This notebook is an exercise in the [Intermediate Machine Learning](https://www.kaggle.com/learn/intermediate-machine-learning) course.  You can reference the tutorial at [this link](https://www.kaggle.com/alexisbcook/cross-validation).**\n\n---\n","metadata":{}},{"cell_type":"markdown","source":"In this exercise, you will leverage what you've learned to tune a machine learning model with **cross-validation**.\n\n# Setup\n\nThe questions below will give you feedback on your work. Run the following cell to set up the feedback system.","metadata":{}},{"cell_type":"code","source":"# Set up code checking\nimport os\nif not os.path.exists(\"../input/train.csv\"):\n    os.symlink(\"../input/home-data-for-ml-course/train.csv\", \"../input/train.csv\")  \n    os.symlink(\"../input/home-data-for-ml-course/test.csv\", \"../input/test.csv\") \nfrom learntools.core import binder\nbinder.bind(globals())\nfrom learntools.ml_intermediate.ex5 import *\nprint(\"Setup Complete\")","metadata":{"execution":{"iopub.status.busy":"2022-07-22T07:44:05.646181Z","iopub.execute_input":"2022-07-22T07:44:05.647617Z","iopub.status.idle":"2022-07-22T07:44:05.731545Z","shell.execute_reply.started":"2022-07-22T07:44:05.647484Z","shell.execute_reply":"2022-07-22T07:44:05.730306Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"You will work with the [Housing Prices Competition for Kaggle Learn Users](https://www.kaggle.com/c/home-data-for-ml-course) from the previous exercise. \n\n![Ames Housing dataset image](https://i.imgur.com/lTJVG4e.png)\n\nRun the next code cell without changes to load the training and test data in `X` and `X_test`.  For simplicity, we drop categorical variables.","metadata":{}},{"cell_type":"code","source":"import pandas as pd\nfrom sklearn.model_selection import train_test_split\n\n# Read the data\ntrain_data = pd.read_csv('../input/train.csv', index_col='Id')\ntest_data = pd.read_csv('../input/test.csv', index_col='Id')\n\n# Remove rows with missing target, separate target from predictors\ntrain_data.dropna(axis=0, subset=['SalePrice'], inplace=True)\ny = train_data.SalePrice              \ntrain_data.drop(['SalePrice'], axis=1, inplace=True)\n\n# Select numeric columns only\nnumeric_cols = [cname for cname in train_data.columns if train_data[cname].dtype in ['int64', 'float64']]\nX = train_data[numeric_cols].copy()\nX_test = test_data[numeric_cols].copy()","metadata":{"execution":{"iopub.status.busy":"2022-07-22T07:44:26.975095Z","iopub.execute_input":"2022-07-22T07:44:26.975613Z","iopub.status.idle":"2022-07-22T07:44:27.683241Z","shell.execute_reply.started":"2022-07-22T07:44:26.975571Z","shell.execute_reply":"2022-07-22T07:44:27.681871Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Use the next code cell to print the first several rows of the data.","metadata":{}},{"cell_type":"code","source":"X.head()","metadata":{"execution":{"iopub.status.busy":"2022-07-22T07:44:50.466937Z","iopub.execute_input":"2022-07-22T07:44:50.467404Z","iopub.status.idle":"2022-07-22T07:44:50.500021Z","shell.execute_reply.started":"2022-07-22T07:44:50.467366Z","shell.execute_reply":"2022-07-22T07:44:50.498754Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"So far, you've learned how to build pipelines with scikit-learn.  For instance, the pipeline below will use [`SimpleImputer()`](https://scikit-learn.org/stable/modules/generated/sklearn.impute.SimpleImputer.html) to replace missing values in the data, before using [`RandomForestRegressor()`](https://scikit-learn.org/stable/modules/generated/sklearn.ensemble.RandomForestRegressor.html) to train a random forest model to make predictions.  We set the number of trees in the random forest model with the `n_estimators` parameter, and setting `random_state` ensures reproducibility.","metadata":{}},{"cell_type":"code","source":"from sklearn.ensemble import RandomForestRegressor\nfrom sklearn.pipeline import Pipeline\nfrom sklearn.impute import SimpleImputer\n\nmy_pipeline = Pipeline(steps=[\n    ('preprocessor', SimpleImputer()),\n    ('model', RandomForestRegressor(n_estimators=50, random_state=0))\n])","metadata":{"execution":{"iopub.status.busy":"2022-07-22T07:45:02.202730Z","iopub.execute_input":"2022-07-22T07:45:02.203163Z","iopub.status.idle":"2022-07-22T07:45:02.490254Z","shell.execute_reply.started":"2022-07-22T07:45:02.203128Z","shell.execute_reply":"2022-07-22T07:45:02.488613Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"You have also learned how to use pipelines in cross-validation.  The code below uses the [`cross_val_score()`](https://scikit-learn.org/stable/modules/generated/sklearn.model_selection.cross_val_score.html) function to obtain the mean absolute error (MAE), averaged across five different folds.  Recall we set the number of folds with the `cv` parameter.","metadata":{}},{"cell_type":"code","source":"from sklearn.model_selection import cross_val_score\n\n# Multiply by -1 since sklearn calculates *negative* MAE\nscores = -1 * cross_val_score(my_pipeline, X, y,\n                              cv=5,\n                              scoring='neg_mean_absolute_error')\n\nprint(\"Average MAE score:\", scores.mean())","metadata":{"execution":{"iopub.status.busy":"2022-07-22T07:51:13.099609Z","iopub.execute_input":"2022-07-22T07:51:13.100146Z","iopub.status.idle":"2022-07-22T07:51:16.306696Z","shell.execute_reply.started":"2022-07-22T07:51:13.100106Z","shell.execute_reply":"2022-07-22T07:51:16.304998Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Step 1: Write a useful function\n\nIn this exercise, you'll use cross-validation to select parameters for a machine learning model.\n\nBegin by writing a function `get_score()` that reports the average (over three cross-validation folds) MAE of a machine learning pipeline that uses:\n- the data in `X` and `y` to create folds,\n- `SimpleImputer()` (with all parameters left as default) to replace missing values, and\n- `RandomForestRegressor()` (with `random_state=0`) to fit a random forest model.\n\nThe `n_estimators` parameter supplied to `get_score()` is used when setting the number of trees in the random forest model.  ","metadata":{}},{"cell_type":"code","source":"def get_score(n_estimators):\n    \"\"\"Return the average MAE over 3 CV folds of random forest model.\n    \n    Keyword argument:\n    n_estimators -- the number of trees in the forest\n    \"\"\"\n    # Replace this body with your own code\n    my_pipeline = Pipeline(steps=[('preprocessor',SimpleImputer()),\n                                 ('model', RandomForestRegressor(n_estimators, random_state=0))])\n    scores = -1 * cross_val_score(my_pipeline, X, y,\n                              cv=3,\n                              scoring='neg_mean_absolute_error')\n    return scores.mean()\n\n\n# Check your answer\nstep_1.check()","metadata":{"execution":{"iopub.status.busy":"2022-07-22T07:51:16.408750Z","iopub.execute_input":"2022-07-22T07:51:16.409199Z","iopub.status.idle":"2022-07-22T07:51:17.502621Z","shell.execute_reply.started":"2022-07-22T07:51:16.409164Z","shell.execute_reply":"2022-07-22T07:51:17.501190Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Lines below will give you a hint or solution code\n#step_1.hint()\nstep_1.solution()","metadata":{"execution":{"iopub.status.busy":"2022-07-22T07:46:58.438394Z","iopub.execute_input":"2022-07-22T07:46:58.438872Z","iopub.status.idle":"2022-07-22T07:46:58.449538Z","shell.execute_reply.started":"2022-07-22T07:46:58.438837Z","shell.execute_reply":"2022-07-22T07:46:58.447969Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Step 2: Test different parameter values\n\nNow, you will use the function that you defined in Step 1 to evaluate the model performance corresponding to eight different values for the number of trees in the random forest: 50, 100, 150, ..., 300, 350, 400.\n\nStore your results in a Python dictionary `results`, where `results[i]` is the average MAE returned by `get_score(i)`.","metadata":{}},{"cell_type":"code","source":"results = {}\nfor i in range(1,9):\n    results[50*i] = get_score(50*i) \n\n# Check your answer\nstep_2.check()","metadata":{"execution":{"iopub.status.busy":"2022-07-22T07:56:38.983555Z","iopub.execute_input":"2022-07-22T07:56:38.983995Z","iopub.status.idle":"2022-07-22T07:57:36.726488Z","shell.execute_reply.started":"2022-07-22T07:56:38.983962Z","shell.execute_reply":"2022-07-22T07:57:36.725244Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Lines below will give you a hint or solution code\n#step_2.hint()\nstep_2.solution()","metadata":{"execution":{"iopub.status.busy":"2022-07-22T07:53:19.142434Z","iopub.execute_input":"2022-07-22T07:53:19.142826Z","iopub.status.idle":"2022-07-22T07:53:19.154565Z","shell.execute_reply.started":"2022-07-22T07:53:19.142794Z","shell.execute_reply":"2022-07-22T07:53:19.153107Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Use the next cell to visualize your results from Step 2.  Run the code without changes.","metadata":{}},{"cell_type":"code","source":"import matplotlib.pyplot as plt\n%matplotlib inline\n\nplt.plot(list(results.keys()), list(results.values()))\nplt.show()","metadata":{"execution":{"iopub.status.busy":"2022-07-22T07:57:36.729395Z","iopub.execute_input":"2022-07-22T07:57:36.729942Z","iopub.status.idle":"2022-07-22T07:57:36.981647Z","shell.execute_reply.started":"2022-07-22T07:57:36.729873Z","shell.execute_reply":"2022-07-22T07:57:36.979979Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Step 3: Find the best parameter value\n\nGiven the results, which value for `n_estimators` seems best for the random forest model?  Use your answer to set the value of `n_estimators_best`.","metadata":{}},{"cell_type":"code","source":"n_estimators_best = min(results, key = results.get)\n\n# Check your answer\nstep_3.check()","metadata":{"execution":{"iopub.status.busy":"2022-07-22T08:05:05.134773Z","iopub.execute_input":"2022-07-22T08:05:05.135280Z","iopub.status.idle":"2022-07-22T08:05:05.148976Z","shell.execute_reply.started":"2022-07-22T08:05:05.135241Z","shell.execute_reply":"2022-07-22T08:05:05.147604Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Lines below will give you a hint or solution code\n#step_3.hint()\n#step_3.solution()","metadata":{"execution":{"iopub.status.busy":"2022-07-22T07:52:37.662845Z","iopub.execute_input":"2022-07-22T07:52:37.663227Z","iopub.status.idle":"2022-07-22T07:52:37.668864Z","shell.execute_reply.started":"2022-07-22T07:52:37.663197Z","shell.execute_reply":"2022-07-22T07:52:37.667387Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"In this exercise, you have explored one method for choosing appropriate parameters in a machine learning model.  \n\nIf you'd like to learn more about [hyperparameter optimization](https://en.wikipedia.org/wiki/Hyperparameter_optimization), you're encouraged to start with **grid search**, which is a straightforward method for determining the best _combination_ of parameters for a machine learning model.  Thankfully, scikit-learn also contains a built-in function [`GridSearchCV()`](https://scikit-learn.org/stable/modules/generated/sklearn.model_selection.GridSearchCV.html) that can make your grid search code very efficient!\n\n# Keep going\n\nContinue to learn about **[gradient boosting](https://www.kaggle.com/alexisbcook/xgboost)**, a powerful technique that achieves state-of-the-art results on a variety of datasets.","metadata":{}},{"cell_type":"markdown","source":"---\n\n\n\n\n*Have questions or comments? Visit the [course discussion forum](https://www.kaggle.com/learn/intermediate-machine-learning/discussion) to chat with other learners.*","metadata":{}}]}