{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"pygments_lexer":"ipython3","nbconvert_exporter":"python","version":"3.6.4","file_extension":".py","codemirror_mode":{"name":"ipython","version":3},"name":"python","mimetype":"text/x-python"}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"**[Machine Learning Micro-Course Home Page](https://www.kaggle.com/learn/intro-to-machine-learning)**\n\n---\n","metadata":{}},{"cell_type":"markdown","source":"## Recap\nYou've built a model. In this exercise you will test how good your model is.\n\nRun the cell below to set up your coding environment where the previous exercise left off.","metadata":{}},{"cell_type":"code","source":"# Code you have previously used to load data\nimport pandas as pd\nfrom sklearn.tree import DecisionTreeRegressor\n\n# Path of the file to read\niowa_file_path = '../input/home-data-for-ml-course/train.csv'\n\nhome_data = pd.read_csv(iowa_file_path)\ny = home_data.SalePrice\nfeature_columns = ['LotArea', 'YearBuilt', '1stFlrSF', '2ndFlrSF', 'FullBath', 'BedroomAbvGr', 'TotRmsAbvGrd']\nX = home_data[feature_columns]\n\n# Specify Model\niowa_model = DecisionTreeRegressor()\n# Fit Model\niowa_model.fit(X, y)\n\nprint(\"First in-sample predictions:\", iowa_model.predict(X.head()))\nprint(\"Actual target values for those homes:\", y.head().tolist())\n\n# Set up code checking\nfrom learntools.core import binder\nbinder.bind(globals())\nfrom learntools.machine_learning.ex4 import *\nprint(\"Setup Complete\")","metadata":{"execution":{"iopub.status.busy":"2022-07-24T12:07:12.351895Z","iopub.execute_input":"2022-07-24T12:07:12.352435Z","iopub.status.idle":"2022-07-24T12:07:13.745948Z","shell.execute_reply.started":"2022-07-24T12:07:12.352360Z","shell.execute_reply":"2022-07-24T12:07:13.744783Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Exercises\n\n## Step 1: Split Your Data\nUse the `train_test_split` function to split up your data.\n\nGive it the argument `random_state=1` so the `check` functions know what to expect when verifying your code.\n\nRecall, your features are loaded in the DataFrame **X** and your target is loaded in **y**.\n","metadata":{}},{"cell_type":"code","source":"# Import the train_test_split function and uncomment\nfrom sklearn.model_selection import train_test_split\n\n# fill in and uncomment\ntrain_X, val_X, train_y, val_y = train_test_split(X, y, random_state=1)\n\nstep_1.check()","metadata":{"execution":{"iopub.status.busy":"2022-07-24T12:07:24.332474Z","iopub.execute_input":"2022-07-24T12:07:24.332853Z","iopub.status.idle":"2022-07-24T12:07:24.353549Z","shell.execute_reply.started":"2022-07-24T12:07:24.332788Z","shell.execute_reply":"2022-07-24T12:07:24.351943Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# The lines below will show you a hint or the solution.\n# step_1.hint() \n# step_1.solution()\n","metadata":{"execution":{"iopub.status.busy":"2022-07-24T12:07:30.756701Z","iopub.execute_input":"2022-07-24T12:07:30.757035Z","iopub.status.idle":"2022-07-24T12:07:30.760819Z","shell.execute_reply.started":"2022-07-24T12:07:30.756988Z","shell.execute_reply":"2022-07-24T12:07:30.759859Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## Step 2: Specify and Fit the Model\n\nCreate a `DecisionTreeRegressor` model and fit it to the relevant data.\nSet `random_state` to 1 again when creating the model.","metadata":{}},{"cell_type":"code","source":"# You imported DecisionTreeRegressor in your last exercise\n# and that code has been copied to the setup code above. So, no need to\n# import it again\n\n# Specify the model\niowa_model = DecisionTreeRegressor(random_state=1)\n\n# Fit iowa_model with the training data.\niowa_model.fit(train_X, train_y)\nstep_2.check()","metadata":{"execution":{"iopub.status.busy":"2022-07-24T12:08:16.441717Z","iopub.execute_input":"2022-07-24T12:08:16.442043Z","iopub.status.idle":"2022-07-24T12:08:16.469097Z","shell.execute_reply.started":"2022-07-24T12:08:16.442003Z","shell.execute_reply":"2022-07-24T12:08:16.468042Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# step_2.hint()\n# step_2.solution()","metadata":{},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## Step 3: Make Predictions with Validation data\n","metadata":{}},{"cell_type":"code","source":"# Predict with all validation observations\nval_predictions = iowa_model.predict(val_X)\n\nstep_3.check()","metadata":{"execution":{"iopub.status.busy":"2022-07-24T12:08:37.154687Z","iopub.execute_input":"2022-07-24T12:08:37.155197Z","iopub.status.idle":"2022-07-24T12:08:37.167300Z","shell.execute_reply.started":"2022-07-24T12:08:37.155148Z","shell.execute_reply":"2022-07-24T12:08:37.166325Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# step_3.hint()\n# step_3.solution()","metadata":{},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Inspect your predictions and actual values from validation data.","metadata":{}},{"cell_type":"code","source":"# print the top few validation predictions\nprint(iowa_model.predict(val_X.head()))\n# print the top few actual prices from validation data\nprint(val_y.head().tolist())","metadata":{"execution":{"iopub.status.busy":"2022-07-24T12:08:51.215941Z","iopub.execute_input":"2022-07-24T12:08:51.216440Z","iopub.status.idle":"2022-07-24T12:08:51.224465Z","shell.execute_reply.started":"2022-07-24T12:08:51.216392Z","shell.execute_reply":"2022-07-24T12:08:51.222837Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"What do you notice that is different from what you saw with in-sample predictions (which are printed after the top code cell in this page).\n\nDo you remember why validation predictions differ from in-sample (or training) predictions? This is an important idea from the last lesson.\n\n## Step 4: Calculate the Mean Absolute Error in Validation Data\n","metadata":{}},{"cell_type":"code","source":"from sklearn.metrics import mean_absolute_error\nval_mae = mean_absolute_error(val_y, val_predictions)\n\n# uncomment following line to see the validation_mae\nprint(val_mae)\nstep_4.check()","metadata":{"execution":{"iopub.status.busy":"2022-07-24T12:09:03.714094Z","iopub.execute_input":"2022-07-24T12:09:03.714571Z","iopub.status.idle":"2022-07-24T12:09:03.723646Z","shell.execute_reply.started":"2022-07-24T12:09:03.714509Z","shell.execute_reply":"2022-07-24T12:09:03.722779Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# step_4.hint()\n# step_4.solution()","metadata":{"execution":{"iopub.status.busy":"2022-07-24T12:09:22.048100Z","iopub.execute_input":"2022-07-24T12:09:22.048765Z","iopub.status.idle":"2022-07-24T12:09:22.053277Z","shell.execute_reply.started":"2022-07-24T12:09:22.048694Z","shell.execute_reply":"2022-07-24T12:09:22.051961Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Is that MAE good?  There isn't a general rule for what values are good that applies across applications. But you'll see how to use (and improve) this number in the next step.\n\n# Keep Going\n\nYou are ready for **[Underfitting and Overfitting](https://www.kaggle.com/dansbecker/underfitting-and-overfitting).**\n","metadata":{}},{"cell_type":"markdown","source":"---\n**[Machine Learning Micro-Course Home Page](https://www.kaggle.com/learn/intro-to-machine-learning)**\n\n","metadata":{}}]}