{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"pygments_lexer":"ipython3","nbconvert_exporter":"python","version":"3.6.4","file_extension":".py","codemirror_mode":{"name":"ipython","version":3},"name":"python","mimetype":"text/x-python"}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"**This notebook is an exercise in the [Introduction to Machine Learning](https://www.kaggle.com/learn/intro-to-machine-learning) course.  You can reference the tutorial at [this link](https://www.kaggle.com/dansbecker/your-first-machine-learning-model).**\n\n---\n","metadata":{}},{"cell_type":"markdown","source":"## Recap\nSo far, you have loaded your data and reviewed it with the following code. Run this cell to set up your coding environment where the previous step left off.","metadata":{}},{"cell_type":"code","source":"# Code you have previously used to load data\nimport pandas as pd\n\n# Path of the file to read\niowa_file_path = '../input/home-data-for-ml-course/train.csv'\n\nhome_data = pd.read_csv(iowa_file_path)\n\n# Set up code checking\nfrom learntools.core import binder\nbinder.bind(globals())\nfrom learntools.machine_learning.ex3 import *\n\nprint(\"Setup Complete\")","metadata":{"execution":{"iopub.status.busy":"2022-07-26T09:23:22.202688Z","iopub.execute_input":"2022-07-26T09:23:22.203111Z","iopub.status.idle":"2022-07-26T09:23:23.557504Z","shell.execute_reply.started":"2022-07-26T09:23:22.203005Z","shell.execute_reply":"2022-07-26T09:23:23.554218Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Exercises\n\n## Step 1: Specify Prediction Target\nSelect the target variable, which corresponds to the sales price. Save this to a new variable called `y`. You'll need to print a list of the columns to find the name of the column you need.\n","metadata":{}},{"cell_type":"code","source":"# print the list of columns in the dataset to find the name of the prediction target from the dataset #\n\nhome_data.columns\n","metadata":{"execution":{"iopub.status.busy":"2022-07-26T09:24:23.731476Z","iopub.execute_input":"2022-07-26T09:24:23.731848Z","iopub.status.idle":"2022-07-26T09:24:23.741248Z","shell.execute_reply.started":"2022-07-26T09:24:23.731817Z","shell.execute_reply":"2022-07-26T09:24:23.740283Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"y = home_data.SalePrice\n\n# Check your answer\nstep_1.check()","metadata":{"execution":{"iopub.status.busy":"2022-07-26T09:25:52.394923Z","iopub.execute_input":"2022-07-26T09:25:52.395314Z","iopub.status.idle":"2022-07-26T09:25:52.408336Z","shell.execute_reply.started":"2022-07-26T09:25:52.395283Z","shell.execute_reply":"2022-07-26T09:25:52.407393Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# The lines below will show you a hint or the solution.\n# step_1.hint() \n# step_1.solution()","metadata":{"execution":{"iopub.status.busy":"2022-07-24T21:03:56.624023Z","iopub.execute_input":"2022-07-24T21:03:56.625068Z","iopub.status.idle":"2022-07-24T21:03:56.630342Z","shell.execute_reply.started":"2022-07-24T21:03:56.625014Z","shell.execute_reply":"2022-07-24T21:03:56.628646Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## Step 2: Create X\nNow you will create a DataFrame called `X` holding the predictive features.\n\nSince you want only some columns from the original data, you'll first create a list with the names of the columns you want in `X`.\n\nYou'll use just the following columns in the list (you can copy and paste the whole list to save some typing, though you'll still need to add quotes):\n  * LotArea\n  * YearBuilt\n  * 1stFlrSF\n  * 2ndFlrSF\n  * FullBath\n  * BedroomAbvGr\n  * TotRmsAbvGrd\n\nAfter you've created that list of features, use it to create the DataFrame that you'll use to fit the model.","metadata":{}},{"cell_type":"code","source":"# Create the list of features below\nfeature_names = [\"LotArea\", \"YearBuilt\", \"1stFlrSF\", \"2ndFlrSF\", \"FullBath\", \"BedroomAbvGr\", \"TotRmsAbvGrd\"]\n\n# Select data corresponding to features in feature_names\nX = home_data[feature_names]\n\n# Check your answer\nstep_2.check()","metadata":{"execution":{"iopub.status.busy":"2022-07-26T09:26:52.914883Z","iopub.execute_input":"2022-07-26T09:26:52.915275Z","iopub.status.idle":"2022-07-26T09:26:52.930390Z","shell.execute_reply.started":"2022-07-26T09:26:52.915242Z","shell.execute_reply":"2022-07-26T09:26:52.929065Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# step_2.hint()\n# step_2.solution()","metadata":{"execution":{"iopub.status.busy":"2022-07-24T21:03:56.654440Z","iopub.execute_input":"2022-07-24T21:03:56.655073Z","iopub.status.idle":"2022-07-24T21:03:56.659301Z","shell.execute_reply.started":"2022-07-24T21:03:56.655033Z","shell.execute_reply":"2022-07-24T21:03:56.658312Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## Review Data\nBefore building a model, take a quick look at **X** to verify it looks sensible","metadata":{}},{"cell_type":"code","source":"# Review data\n# print description or statistics from X\nprint(X.describe())\n\n# print the top few lines\nprint(X.head())","metadata":{"execution":{"iopub.status.busy":"2022-07-24T21:03:56.660922Z","iopub.execute_input":"2022-07-24T21:03:56.661581Z","iopub.status.idle":"2022-07-24T21:03:56.702774Z","shell.execute_reply.started":"2022-07-24T21:03:56.661521Z","shell.execute_reply":"2022-07-24T21:03:56.701686Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## Step 3: Specify and Fit Model\nCreate a `DecisionTreeRegressor` and save it iowa_model. Ensure you've done the relevant import from sklearn to run this command.\n\nThen fit the model you just created using the data in `X` and `y` that you saved above.","metadata":{}},{"cell_type":"code","source":"from sklearn.tree import DecisionTreeRegressor\n#specify the model. \n#For model reproducibility, set a numeric value for random_state when specifying the model\niowa_model = DecisionTreeRegressor (random_state = 1)\n\n# Fit the model\niowa_model.fit(X,y)\n# Check your answer\nstep_3.check()","metadata":{"execution":{"iopub.status.busy":"2022-07-26T09:31:08.209313Z","iopub.execute_input":"2022-07-26T09:31:08.209714Z","iopub.status.idle":"2022-07-26T09:31:08.237174Z","shell.execute_reply.started":"2022-07-26T09:31:08.209682Z","shell.execute_reply":"2022-07-26T09:31:08.236358Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# step_3.hint()\n# step_3.solution()","metadata":{"execution":{"iopub.status.busy":"2022-07-24T21:03:56.739772Z","iopub.execute_input":"2022-07-24T21:03:56.740450Z","iopub.status.idle":"2022-07-24T21:03:56.752674Z","shell.execute_reply.started":"2022-07-24T21:03:56.740406Z","shell.execute_reply":"2022-07-24T21:03:56.745191Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## Step 4: Make Predictions\nMake predictions with the model's `predict` command using `X` as the data. Save the results to a variable called `predictions`.","metadata":{}},{"cell_type":"code","source":"# predicting using the features #\n\npredictions = iowa_model.predict(X)\nprint(predictions)\n\n# Check your answer\nstep_4.check()","metadata":{"execution":{"iopub.status.busy":"2022-07-26T09:32:10.187619Z","iopub.execute_input":"2022-07-26T09:32:10.188001Z","iopub.status.idle":"2022-07-26T09:32:10.205060Z","shell.execute_reply.started":"2022-07-26T09:32:10.187971Z","shell.execute_reply":"2022-07-26T09:32:10.203504Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# step_4.hint()\n# step_4.solution()","metadata":{"execution":{"iopub.status.busy":"2022-07-24T21:03:56.786077Z","iopub.execute_input":"2022-07-24T21:03:56.786936Z","iopub.status.idle":"2022-07-24T21:03:56.791883Z","shell.execute_reply.started":"2022-07-24T21:03:56.786889Z","shell.execute_reply":"2022-07-24T21:03:56.790622Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## Think About Your Results\n\nUse the `head` method to compare the top few predictions to the actual home values (in `y`) for those same homes. Anything surprising?\n","metadata":{}},{"cell_type":"code","source":"predictions = pd.DataFrame({\"Column1\": predictions})\nprint(predictions.head(5))","metadata":{"execution":{"iopub.status.busy":"2022-07-24T21:03:56.793893Z","iopub.execute_input":"2022-07-24T21:03:56.794752Z","iopub.status.idle":"2022-07-24T21:03:56.811024Z","shell.execute_reply.started":"2022-07-24T21:03:56.794705Z","shell.execute_reply":"2022-07-24T21:03:56.809721Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"y.head()","metadata":{"execution":{"iopub.status.busy":"2022-07-26T09:33:36.665591Z","iopub.execute_input":"2022-07-26T09:33:36.667003Z","iopub.status.idle":"2022-07-26T09:33:36.674751Z","shell.execute_reply.started":"2022-07-26T09:33:36.666959Z","shell.execute_reply":"2022-07-26T09:33:36.673635Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"It's natural to ask how accurate the model's predictions will be and how you can improve that. That will be you're next step.\n\n# Keep Going\n\nYou are ready for **[Model Validation](https://www.kaggle.com/dansbecker/model-validation).**\n","metadata":{}},{"cell_type":"markdown","source":"---\n\n\n\n\n*Have questions or comments? Visit the [course discussion forum](https://www.kaggle.com/learn/intro-to-machine-learning/discussion) to chat with other learners.*","metadata":{}}]}