{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"pygments_lexer":"ipython3","nbconvert_exporter":"python","version":"3.6.4","file_extension":".py","codemirror_mode":{"name":"ipython","version":3},"name":"python","mimetype":"text/x-python"}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"**This notebook is an exercise in the [Introduction to Machine Learning](https://www.kaggle.com/learn/intro-to-machine-learning) course.  You can reference the tutorial at [this link](https://www.kaggle.com/dansbecker/model-validation).**\n\n---\n","metadata":{}},{"cell_type":"markdown","source":"## Recap\nYou've built a model. In this exercise you will test how good your model is.\n\nRun the cell below to set up your coding environment where the previous exercise left off.","metadata":{}},{"cell_type":"code","source":"# Code you have previously used to load data\nimport pandas as pd\nfrom sklearn.tree import DecisionTreeRegressor\n\n# Path of the file to read\niowa_file_path = '../input/home-data-for-ml-course/train.csv'\n\nhome_data = pd.read_csv(iowa_file_path)\ny = home_data.SalePrice\nfeature_columns = ['LotArea', 'YearBuilt', '1stFlrSF', '2ndFlrSF', 'FullBath', 'BedroomAbvGr', 'TotRmsAbvGrd']\nX = home_data[feature_columns]\n\n# Specify Model\niowa_model = DecisionTreeRegressor()\n# Fit Model\niowa_model.fit(X, y)\n\nprint(\"First in-sample predictions:\", iowa_model.predict(X.head()))\nprint(\"Actual target values for those homes:\", y.head().tolist())\n\n# Set up code checking\nfrom learntools.core import binder\nbinder.bind(globals())\nfrom learntools.machine_learning.ex4 import *\nprint(\"Setup Complete\")","metadata":{"execution":{"iopub.status.busy":"2022-07-21T05:38:48.651712Z","iopub.execute_input":"2022-07-21T05:38:48.652215Z","iopub.status.idle":"2022-07-21T05:38:50.218806Z","shell.execute_reply.started":"2022-07-21T05:38:48.652088Z","shell.execute_reply":"2022-07-21T05:38:50.217961Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Exercises\n\n## Step 1: Split Your Data\nUse the `train_test_split` function to split up your data.\n\nGive it the argument `random_state=1` so the `check` functions know what to expect when verifying your code.\n\nRecall, your features are loaded in the DataFrame **X** and your target is loaded in **y**.\n","metadata":{}},{"cell_type":"code","source":"# Import the train_test_split function and uncomment\nfrom sklearn.model_selection import train_test_split\n\n# fill in and uncomment\ntrain_X, val_X, train_y, val_y = train_test_split(X,y, random_state=1)\n\n# Check your answer\nstep_1.check()","metadata":{"execution":{"iopub.status.busy":"2022-07-21T05:41:42.632092Z","iopub.execute_input":"2022-07-21T05:41:42.632900Z","iopub.status.idle":"2022-07-21T05:41:42.644798Z","shell.execute_reply.started":"2022-07-21T05:41:42.632863Z","shell.execute_reply":"2022-07-21T05:41:42.643911Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# The lines below will show you a hint or the solution.\n# step_1.hint() \n#step_1.solution()\n","metadata":{"execution":{"iopub.status.busy":"2022-07-21T05:41:55.633073Z","iopub.execute_input":"2022-07-21T05:41:55.633468Z","iopub.status.idle":"2022-07-21T05:41:55.638082Z","shell.execute_reply.started":"2022-07-21T05:41:55.633437Z","shell.execute_reply":"2022-07-21T05:41:55.636915Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## Step 2: Specify and Fit the Model\n\nCreate a `DecisionTreeRegressor` model and fit it to the relevant data.\nSet `random_state` to 1 again when creating the model.","metadata":{}},{"cell_type":"code","source":"# You imported DecisionTreeRegressor in your last exercise\n# and that code has been copied to the setup code above. So, no need to\n# import it again\n\n# Specify the model\niowa_model = DecisionTreeRegressor(random_state=1)\n\n# Fit iowa_model with the training data.\niowa_model.fit(train_X,train_y)\n\n# Check your answer\nstep_2.check()","metadata":{"execution":{"iopub.status.busy":"2022-07-21T05:43:04.321004Z","iopub.execute_input":"2022-07-21T05:43:04.321450Z","iopub.status.idle":"2022-07-21T05:43:04.353419Z","shell.execute_reply.started":"2022-07-21T05:43:04.321401Z","shell.execute_reply":"2022-07-21T05:43:04.352067Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# step_2.hint()\n#step_2.solution()","metadata":{"execution":{"iopub.status.busy":"2022-07-21T05:43:17.537820Z","iopub.execute_input":"2022-07-21T05:43:17.538242Z","iopub.status.idle":"2022-07-21T05:43:17.542937Z","shell.execute_reply.started":"2022-07-21T05:43:17.538193Z","shell.execute_reply":"2022-07-21T05:43:17.541893Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## Step 3: Make Predictions with Validation data\n","metadata":{}},{"cell_type":"code","source":"# Predict with all validation observations\nval_predictions = iowa_model.predict(val_X)\n\n# Check your answer\nstep_3.check()","metadata":{"execution":{"iopub.status.busy":"2022-07-21T05:43:57.780820Z","iopub.execute_input":"2022-07-21T05:43:57.781272Z","iopub.status.idle":"2022-07-21T05:43:57.794127Z","shell.execute_reply.started":"2022-07-21T05:43:57.781235Z","shell.execute_reply":"2022-07-21T05:43:57.793225Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# step_3.hint()\n# step_3.solution()","metadata":{"execution":{"iopub.status.busy":"2022-07-21T05:44:12.666261Z","iopub.execute_input":"2022-07-21T05:44:12.666640Z","iopub.status.idle":"2022-07-21T05:44:12.672109Z","shell.execute_reply.started":"2022-07-21T05:44:12.666610Z","shell.execute_reply":"2022-07-21T05:44:12.670319Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Inspect your predictions and actual values from validation data.","metadata":{}},{"cell_type":"code","source":"# print the top few validation predictions\nprint(val_predictions)\n# print the top few actual prices from validation data\nprint(val_X)","metadata":{"execution":{"iopub.status.busy":"2022-07-21T05:45:45.913472Z","iopub.execute_input":"2022-07-21T05:45:45.913892Z","iopub.status.idle":"2022-07-21T05:45:45.931484Z","shell.execute_reply.started":"2022-07-21T05:45:45.913859Z","shell.execute_reply":"2022-07-21T05:45:45.930029Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"What do you notice that is different from what you saw with in-sample predictions (which are printed after the top code cell in this page).\n\nDo you remember why validation predictions differ from in-sample (or training) predictions? This is an important idea from the last lesson.\n\n## Step 4: Calculate the Mean Absolute Error in Validation Data\n","metadata":{}},{"cell_type":"code","source":"from sklearn.metrics import mean_absolute_error\nval_mae = mean_absolute_error(val_y,val_predictions)\n\n# uncomment following line to see the validation_mae\nprint(val_mae)\n\n# Check your answer\nstep_4.check()","metadata":{"execution":{"iopub.status.busy":"2022-07-21T06:05:49.962839Z","iopub.execute_input":"2022-07-21T06:05:49.963329Z","iopub.status.idle":"2022-07-21T06:05:49.975514Z","shell.execute_reply.started":"2022-07-21T06:05:49.963285Z","shell.execute_reply":"2022-07-21T06:05:49.974526Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# step_4.hint()\nstep_4.solution()","metadata":{"execution":{"iopub.status.busy":"2022-07-21T05:45:31.359013Z","iopub.execute_input":"2022-07-21T05:45:31.359444Z","iopub.status.idle":"2022-07-21T05:45:31.369198Z","shell.execute_reply.started":"2022-07-21T05:45:31.359414Z","shell.execute_reply":"2022-07-21T05:45:31.368059Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Is that MAE good?  There isn't a general rule for what values are good that applies across applications. But you'll see how to use (and improve) this number in the next step.\n\n# Keep Going\n\nYou are ready for **[Underfitting and Overfitting](https://www.kaggle.com/dansbecker/underfitting-and-overfitting).**\n","metadata":{}},{"cell_type":"markdown","source":"---\n\n\n\n\n*Have questions or comments? Visit the [course discussion forum](https://www.kaggle.com/learn/intro-to-machine-learning/discussion) to chat with other learners.*","metadata":{}}]}