{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"pygments_lexer":"ipython3","nbconvert_exporter":"python","version":"3.6.4","file_extension":".py","codemirror_mode":{"name":"ipython","version":3},"name":"python","mimetype":"text/x-python"}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"**This notebook is an exercise in the [Introduction to Machine Learning](https://www.kaggle.com/learn/intro-to-machine-learning) course.  You can reference the tutorial at [this link](https://www.kaggle.com/dansbecker/model-validation).**\n\n---\n","metadata":{}},{"cell_type":"markdown","source":"## Recap\nYou've built a model. In this exercise you will test how good your model is.\n\nRun the cell below to set up your coding environment where the previous exercise left off.","metadata":{}},{"cell_type":"code","source":"# Code you have previously used to load data\nimport pandas as pd\nfrom sklearn.tree import DecisionTreeRegressor\n\n# Path of the file to read\niowa_file_path = '../input/home-data-for-ml-course/train.csv'\n\nhome_data = pd.read_csv(iowa_file_path)\ny = home_data.SalePrice\nfeature_columns = ['LotArea', 'YearBuilt', '1stFlrSF', '2ndFlrSF', 'FullBath', 'BedroomAbvGr', 'TotRmsAbvGrd']\nX = home_data[feature_columns]\n\n# Specify Model\niowa_model = DecisionTreeRegressor()\n# Fit Model\niowa_model.fit(X, y)\n\nprint(\"First in-sample predictions:\", iowa_model.predict(X.head()))\nprint(\"Actual target values for those homes:\", y.head().tolist())\n\n# Set up code checking\nfrom learntools.core import binder\nbinder.bind(globals())\nfrom learntools.machine_learning.ex4 import *\nprint(\"Setup Complete\")","metadata":{"execution":{"iopub.status.busy":"2022-07-28T02:56:05.295036Z","iopub.execute_input":"2022-07-28T02:56:05.295486Z","iopub.status.idle":"2022-07-28T02:56:06.467966Z","shell.execute_reply.started":"2022-07-28T02:56:05.295393Z","shell.execute_reply":"2022-07-28T02:56:06.466666Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Exercises\n\n## Step 1: Split Your Data\nUse the `train_test_split` function to split up your data.\n\nGive it the argument `random_state=1` so the `check` functions know what to expect when verifying your code.\n\nRecall, your features are loaded in the DataFrame **X** and your target is loaded in **y**.\n","metadata":{}},{"cell_type":"code","source":"# Import the train_test_split function and uncomment\nfrom sklearn.model_selection import train_test_split\n\n# fill in and uncomment\ntrain_X, val_X, train_y, val_y = train_test_split(X,y, random_state=1)\n\n# Check your answer\nstep_1.check()","metadata":{"execution":{"iopub.status.busy":"2022-07-28T02:58:20.677350Z","iopub.execute_input":"2022-07-28T02:58:20.677801Z","iopub.status.idle":"2022-07-28T02:58:20.693970Z","shell.execute_reply.started":"2022-07-28T02:58:20.677764Z","shell.execute_reply":"2022-07-28T02:58:20.692945Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# The lines below will show you a hint or the solution.\n# step_1.hint() \n# step_1.solution()\n","metadata":{},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## Step 2: Specify and Fit the Model\n\nCreate a `DecisionTreeRegressor` model and fit it to the relevant data.\nSet `random_state` to 1 again when creating the model.","metadata":{}},{"cell_type":"code","source":"# You imported DecisionTreeRegressor in your last exercise\n# and that code has been copied to the setup code above. So, no need to\n# import it again\n\n# Specify the model\niowa_model = DecisionTreeRegressor(random_state=1)\n\n# Fit iowa_model with the training data.\niowa_model.fit(train_X,train_y)\n\n# Check your answer\nstep_2.check()","metadata":{"execution":{"iopub.status.busy":"2022-07-28T03:01:49.760064Z","iopub.execute_input":"2022-07-28T03:01:49.760475Z","iopub.status.idle":"2022-07-28T03:01:49.789339Z","shell.execute_reply.started":"2022-07-28T03:01:49.760431Z","shell.execute_reply":"2022-07-28T03:01:49.788548Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# step_2.hint()\n# step_2.solution()","metadata":{"execution":{"iopub.status.busy":"2022-07-28T03:01:20.196959Z","iopub.execute_input":"2022-07-28T03:01:20.197728Z","iopub.status.idle":"2022-07-28T03:01:20.202125Z","shell.execute_reply.started":"2022-07-28T03:01:20.197689Z","shell.execute_reply":"2022-07-28T03:01:20.201012Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## Step 3: Make Predictions with Validation data\n","metadata":{}},{"cell_type":"code","source":"from sklearn.metrics import mean_absolute_error\n\n# Predict with all validation observations\nval_predictions = iowa_model.predict(val_X)\nprint(mean_absolute_error(val_y,val_predictions))\n\n# Check your answer\nstep_3.check()","metadata":{"execution":{"iopub.status.busy":"2022-07-28T03:05:18.596313Z","iopub.execute_input":"2022-07-28T03:05:18.596738Z","iopub.status.idle":"2022-07-28T03:05:18.612550Z","shell.execute_reply.started":"2022-07-28T03:05:18.596704Z","shell.execute_reply":"2022-07-28T03:05:18.611346Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# step_3.hint()\n# step_3.solution()","metadata":{},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Inspect your predictions and actual values from validation data.","metadata":{}},{"cell_type":"code","source":"# print the top few validation predictions\nprint(val_X)\n# print the top few actual prices from validation data\nprint(\"======================\")\nprint(val_y)","metadata":{"execution":{"iopub.status.busy":"2022-07-28T03:06:43.616416Z","iopub.execute_input":"2022-07-28T03:06:43.617317Z","iopub.status.idle":"2022-07-28T03:06:43.628441Z","shell.execute_reply.started":"2022-07-28T03:06:43.617275Z","shell.execute_reply":"2022-07-28T03:06:43.626959Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"What do you notice that is different from what you saw with in-sample predictions (which are printed after the top code cell in this page).\n\nDo you remember why validation predictions differ from in-sample (or training) predictions? This is an important idea from the last lesson.\n\n## Step 4: Calculate the Mean Absolute Error in Validation Data\n","metadata":{}},{"cell_type":"code","source":"from sklearn.metrics import mean_absolute_error\nval_mae = mean_absolute_error(val_y,val_predictions)\n\n# uncomment following line to see the validation_mae\nprint(val_mae)\n\n# Check your answer\nstep_4.check()","metadata":{"execution":{"iopub.status.busy":"2022-07-28T03:09:33.113777Z","iopub.execute_input":"2022-07-28T03:09:33.114218Z","iopub.status.idle":"2022-07-28T03:09:33.126445Z","shell.execute_reply.started":"2022-07-28T03:09:33.114180Z","shell.execute_reply":"2022-07-28T03:09:33.125269Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# step_4.hint()\n# step_4.solution()","metadata":{"execution":{"iopub.status.busy":"2022-07-28T03:09:24.507946Z","iopub.execute_input":"2022-07-28T03:09:24.508315Z","iopub.status.idle":"2022-07-28T03:09:24.512816Z","shell.execute_reply.started":"2022-07-28T03:09:24.508284Z","shell.execute_reply":"2022-07-28T03:09:24.511772Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Is that MAE good?  There isn't a general rule for what values are good that applies across applications. But you'll see how to use (and improve) this number in the next step.\n\n# Keep Going\n\nYou are ready for **[Underfitting and Overfitting](https://www.kaggle.com/dansbecker/underfitting-and-overfitting).**\n","metadata":{}},{"cell_type":"markdown","source":"---\n\n\n\n\n*Have questions or comments? Visit the [course discussion forum](https://www.kaggle.com/learn/intro-to-machine-learning/discussion) to chat with other learners.*","metadata":{}}]}