{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"pygments_lexer":"ipython3","nbconvert_exporter":"python","version":"3.6.4","file_extension":".py","codemirror_mode":{"name":"ipython","version":3},"name":"python","mimetype":"text/x-python"}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"**This notebook is an exercise in the [Introduction to Machine Learning](https://www.kaggle.com/learn/intro-to-machine-learning) course.  You can reference the tutorial at [this link](https://www.kaggle.com/dansbecker/model-validation).**\n\n---\n","metadata":{}},{"cell_type":"markdown","source":"## Recap\nYou've built a model. In this exercise you will test how good your model is.\n\nRun the cell below to set up your coding environment where the previous exercise left off.","metadata":{}},{"cell_type":"code","source":"# Code you have previously used to load data\nimport pandas as pd\nfrom sklearn.tree import DecisionTreeRegressor\n\n# Path of the file to read\niowa_file_path = '../input/home-data-for-ml-course/train.csv'\n\nhome_data = pd.read_csv(iowa_file_path)\ny = home_data.SalePrice\nfeature_columns = ['LotArea', 'YearBuilt', '1stFlrSF', '2ndFlrSF', 'FullBath', 'BedroomAbvGr', 'TotRmsAbvGrd']\nX = home_data[feature_columns]\n\n# Specify Model\niowa_model = DecisionTreeRegressor()\n# Fit Model\niowa_model.fit(X, y)\n\nprint(\"First in-sample predictions:\", iowa_model.predict(X.head()))\nprint(\"Actual target values for those homes:\", y.head().tolist())\n\n# Set up code checking\nfrom learntools.core import binder\nbinder.bind(globals())\nfrom learntools.machine_learning.ex4 import *\nprint(\"Setup Complete\")","metadata":{"execution":{"iopub.status.busy":"2022-05-28T17:56:40.601343Z","iopub.execute_input":"2022-05-28T17:56:40.60202Z","iopub.status.idle":"2022-05-28T17:56:40.65391Z","shell.execute_reply.started":"2022-05-28T17:56:40.601984Z","shell.execute_reply":"2022-05-28T17:56:40.652902Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Exercises\n\n## Step 1: Split Your Data\nUse the `train_test_split` function to split up your data.\n\nGive it the argument `random_state=1` so the `check` functions know what to expect when verifying your code.\n\nRecall, your features are loaded in the DataFrame **X** and your target is loaded in **y**.\n","metadata":{}},{"cell_type":"code","source":"# Import the train_test_split function and uncomment\nfrom sklearn.model_selection import train_test_split as tts\n\n# fill in and uncomment\ntrain_X, val_X, train_y, val_y = tts(X,y,random_state=1)\n\n# Check your answer\nstep_1.check()","metadata":{"execution":{"iopub.status.busy":"2022-05-28T17:56:46.747802Z","iopub.execute_input":"2022-05-28T17:56:46.748271Z","iopub.status.idle":"2022-05-28T17:56:46.760931Z","shell.execute_reply.started":"2022-05-28T17:56:46.748235Z","shell.execute_reply":"2022-05-28T17:56:46.759879Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# The lines below will show you a hint or the solution.\n# step_1.hint() \n# step_1.solution()\n","metadata":{},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## Step 2: Specify and Fit the Model\n\nCreate a `DecisionTreeRegressor` model and fit it to the relevant data.\nSet `random_state` to 1 again when creating the model.","metadata":{}},{"cell_type":"code","source":"# You imported DecisionTreeRegressor in your last exercise\n# and that code has been copied to the setup code above. So, no need to\n# import it again\ntrain_X, val_X, train_y, val_y = tts(X,y,random_state=1)\n# Specify the model\niowa_model = DecisionTreeRegressor(random_state=1)\n\n# Fit iowa_model with the training data.\niowa_model.fit(train_X, train_y)\n\n\n# Check your answer\nstep_2.check()","metadata":{"execution":{"iopub.status.busy":"2022-05-28T18:11:11.378847Z","iopub.execute_input":"2022-05-28T18:11:11.379546Z","iopub.status.idle":"2022-05-28T18:11:11.411706Z","shell.execute_reply.started":"2022-05-28T18:11:11.379499Z","shell.execute_reply":"2022-05-28T18:11:11.410854Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# step_2.hint()\n# step_2.solution()","metadata":{},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## Step 3: Make Predictions with Validation data\n","metadata":{}},{"cell_type":"code","source":"# Predict with all validation observations\nval_predictions = iowa_model.predict(val_X)\n\n# Check your answer\nstep_3.check()","metadata":{"execution":{"iopub.status.busy":"2022-05-28T18:01:38.131229Z","iopub.execute_input":"2022-05-28T18:01:38.131632Z","iopub.status.idle":"2022-05-28T18:01:38.143444Z","shell.execute_reply.started":"2022-05-28T18:01:38.1316Z","shell.execute_reply":"2022-05-28T18:01:38.142664Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# step_3.hint()\n# step_3.solution()","metadata":{},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Inspect your predictions and actual values from validation data.","metadata":{}},{"cell_type":"code","source":"# print the top few validation predictions\nprint('the top few validation predictions\\n',iowa_model.predict(val_X))\n# print the top few actual prices from validation data\nprint('the top few actual prices\\n',val_y.head())","metadata":{"execution":{"iopub.status.busy":"2022-05-28T18:11:20.779871Z","iopub.execute_input":"2022-05-28T18:11:20.780315Z","iopub.status.idle":"2022-05-28T18:11:20.794308Z","shell.execute_reply.started":"2022-05-28T18:11:20.78028Z","shell.execute_reply":"2022-05-28T18:11:20.793275Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"What do you notice that is different from what you saw with in-sample predictions (which are printed after the top code cell in this page).\n\nDo you remember why validation predictions differ from in-sample (or training) predictions? This is an important idea from the last lesson.\n\n## Step 4: Calculate the Mean Absolute Error in Validation Data\n","metadata":{}},{"cell_type":"code","source":"from sklearn.metrics import mean_absolute_error\nval_mae = mean_absolute_error(val_y,iowa_model.predict(val_X))\n\n# uncomment following line to see the validation_mae\nprint(val_mae)\n\n# Check your answer\nstep_4.check()","metadata":{"execution":{"iopub.status.busy":"2022-05-28T18:22:47.534451Z","iopub.execute_input":"2022-05-28T18:22:47.535101Z","iopub.status.idle":"2022-05-28T18:22:47.546002Z","shell.execute_reply.started":"2022-05-28T18:22:47.535068Z","shell.execute_reply":"2022-05-28T18:22:47.545041Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# step_4.hint()\n# step_4.solution()","metadata":{},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Is that MAE good?  There isn't a general rule for what values are good that applies across applications. But you'll see how to use (and improve) this number in the next step.\n\n# Keep Going\n\nYou are ready for **[Underfitting and Overfitting](https://www.kaggle.com/dansbecker/underfitting-and-overfitting).**\n","metadata":{}},{"cell_type":"markdown","source":"---\n\n\n\n\n*Have questions or comments? Visit the [course discussion forum](https://www.kaggle.com/learn/intro-to-machine-learning/discussion) to chat with other learners.*","metadata":{}}]}