{"cells":[{"metadata":{"_uuid":"3ac329a5894510af21f1c998ba7007b230e9d7a8"},"cell_type":"markdown","source":"## Recap\nSo far, you have loaded your data and reviewed it with the following code. Run this cell to set up your coding environment where the previous step left off."},{"metadata":{"collapsed":true,"_uuid":"3ae8c5ecce19e2c3892679c760c5d842be358fab","trusted":false},"cell_type":"code","source":"# Code you have previously used to load data\nimport pandas as pd\n\n# Path of the file to read\niowa_file_path = '../input/home-data-for-ml-course/train.csv'\n\nhome_data = pd.read_csv(iowa_file_path)\n\n# Set up code checking\nfrom learntools.core import binder\nbinder.bind(globals())\nfrom learntools.machine_learning.ex3 import *\n\nprint(\"Setup Complete\")","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"7c2598bccde8b09c6f4b59b157225a2ad4614724"},"cell_type":"markdown","source":"# Exercises\n\n## Step 1: Specify Prediction Target\nSelect the target variable, which corresponds to the sales price. Save this to a new variable called `y`. You'll need to print a list of the columns to find the name of the column you need.\n"},{"metadata":{"collapsed":true,"_uuid":"3704efeef3fa39316b763474d354e22dc8fb49ee","trusted":false},"cell_type":"code","source":"# print the list of columns in the dataset to find the name of the prediction target\n","execution_count":null,"outputs":[]},{"metadata":{"collapsed":true,"_uuid":"d6fa300c11d6925600d9b26c2043fb331b639cf7","trusted":false},"cell_type":"code","source":"#y = _\n\nstep_1.check()","execution_count":null,"outputs":[]},{"metadata":{"collapsed":true,"_uuid":"1ddc23337edfa7bb4a80d06eea9332a75f3a2e70","trusted":false},"cell_type":"code","source":"# The lines below will show you a hint or the solution.\n# step_1.hint() \n# step_1.solution()\n","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"4532f862b848ec1cd3e40a66f672c72b0f74417c"},"cell_type":"markdown","source":"## Step 2: Create X\nNow you will create a DataFrame called `X` holding the predictive features.\n\nSince you want only some columns from the original data, you'll first create a list with the names of the columns you want in `X`.\n\nYou'll use just the following columns in the list (you can copy and paste the whole list to save some typing, though you'll still need to add quotes):\n    * LotArea\n    * YearBuilt\n    * 1stFlrSF\n    * 2ndFlrSF\n    * FullBath\n    * BedroomAbvGr\n    * TotRmsAbvGrd\n\nAfter you've created that list of features, use it to create the DataFrame that you'll use to fit the model."},{"metadata":{"collapsed":true,"_uuid":"ad44a0fac79c6412dde372c76bbc8e3cd15415d1","trusted":false},"cell_type":"code","source":"# Create the list of features below\n# feature_names = ___\n\n# select data corresponding to features in feature_names\n#X = _\n\nstep_2.check()","execution_count":null,"outputs":[]},{"metadata":{"collapsed":true,"_uuid":"63db266226c2123fec6f39ac03407ebf2aaa89bb","trusted":false},"cell_type":"code","source":"# step_2.hint()\n# step_2.solution()","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"884d24995b72826c9b9fb3c7ce9291ad161cb2bb"},"cell_type":"markdown","source":"## Review Data\nBefore building a model, take a quick look at **X** to verify it looks sensible"},{"metadata":{"collapsed":true,"_uuid":"91f1d54403556707dadc71c532ff1dc54cb4d91f","trusted":false},"cell_type":"code","source":"# Review data\n# print description or statistics from X\n#print(_)\n\n# print the top few lines\n#print(_)\n","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"1ffc09adebffb5f53699b844eb27d5c76e3ea83d"},"cell_type":"markdown","source":"## Step 3: Specify and Fit Model\nCreate a `DecisionTreeRegressor` and save it iowa_model. Ensure you've done the relevant import from sklearn to run this command.\n\nThen fit the model you just created using the data in `X` and `y` that you saved above."},{"metadata":{"collapsed":true,"_uuid":"f4140fd4cab84d353861867ac26e6eeee551f170","trusted":false},"cell_type":"code","source":"# from _ import _\n#specify the model. \n#For model reproducibility, set a numeric value for random_state when specifying the model\niowa_model = _\n\n# Fit the model\n_\n\nstep_3.check()","execution_count":null,"outputs":[]},{"metadata":{"collapsed":true,"_uuid":"89516f40fe6b8de9ce0567fc3121c68ef3fda3d6","trusted":false},"cell_type":"code","source":"# step_3.hint()\n# step_3.solution()","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"eae0b9788785b5aec763d244ca992875c18b0e78"},"cell_type":"markdown","source":"## Step 4: Make Predictions\nMake predictions with the model's `predict` command using `X` as the data. Save the results to a variable called `predictions`."},{"metadata":{"collapsed":true,"_uuid":"c0093106a41af3e8ca883f2200667acf6b1b8d9c","trusted":false},"cell_type":"code","source":"predictions = _\nprint(predictions)\nstep_4.check()","execution_count":null,"outputs":[]},{"metadata":{"collapsed":true,"_uuid":"0b3912cbfaf58b6de4c199082e9ff697a5f8bc51","trusted":false},"cell_type":"code","source":"# step_4.hint()\n# step_4.solution()","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"b74a799dda0159d93b8db981438e915d096577c2"},"cell_type":"markdown","source":"## Think About Your Results\n\nUse the `head` method to compare the top few predictions to the actual home values (in `y`) for those same homes. Anything surprising?\n\nYou'll understand why this happened if you keep going."},{"metadata":{"_uuid":"f687f1c0e4485509bb70ba12bc93ab2da458c6bd"},"cell_type":"markdown","source":"\n## Keep Going\nYou've built a decision tree model.  It's natural to ask how accurate the model's predictions will be and how you can improve that. Learn how to do that with **[Model Validation](https://www.kaggle.com/dansbecker/model-validation)**.\n\n---\n**[Course Home Page](https://www.kaggle.com/learn/machine-learning)**\n\n\n"},{"metadata":{"trusted":true,"_uuid":"c317fa76d43c4ff4c0541afb1c3300c58fb92840"},"cell_type":"code","source":"import pandas as pd\nfrom sklearn.ensemble import RandomForestRegressor\nfrom sklearn.metrics import mean_absolute_error\nfrom sklearn.model_selection import train_test_split\n#load data\niowa_train_path='../input/train.csv'\niowa_test_path='../input/test.csv'\niowa_train_data=pd.read_csv(iowa_train_path)\niowa_test_data=pd.read_csv(iowa_test_path)\n#set the values \ntrain_x = iowa_train_data.drop(['SalePrice'], axis=1)\ntrain_y=iowa_train_data.SalePrice\ntest_x=iowa_test_data\n#we need to drop the non numerical data also  for both test and train \ntrain_x = train_x.select_dtypes(exclude=['object'])\ntest_x=test_x.select_dtypes(exclude=['object'])\n\n#test_y is to be predicted \n\n#import the imputer \nfrom sklearn.impute import SimpleImputer\nmy_imputer = SimpleImputer()\n\nimputed_train_x_plus = train_x.copy()\nimputed_test_x_plus=test_x.copy()\ncols_with_missing = {col for col in train_x.columns \n                                 if train_x[col].isnull().any()}\nfor col in cols_with_missing:\n    imputed_train_x_plus[col + '_was_missing'] = imputed_train_x_plus[col].isnull()\n    imputed_test_x_plus[col + '_was_missing'] = imputed_test_x_plus[col].isnull()\n\n\n#start with the imputer operation\nimputed_train_x_plus = my_imputer.fit_transform(imputed_train_x_plus)\nimputed_test_x_plus=my_imputer.transform(imputed_test_x_plus)\n#now generate the model random forest \nmodel = RandomForestRegressor(random_state=1)\nmodel.fit(imputed_train_x_plus, train_y)\npredictions = model.predict(imputed_test_x_plus)\n\n#getting the data ready for submission \noutput = pd.DataFrame({'Id': test_x.Id,\n                       'SalePrice': predictions})\noutput.to_csv('submission.csv', index=False)","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"a56d4e7f0cb465b1a0be6ec1eca410d61c262d96"},"cell_type":"code","source":"","execution_count":null,"outputs":[]}],"metadata":{"kernelspec":{"display_name":"Python 3","language":"python","name":"python3"},"language_info":{"name":"python","version":"3.6.6","mimetype":"text/x-python","codemirror_mode":{"name":"ipython","version":3},"pygments_lexer":"ipython3","nbconvert_exporter":"python","file_extension":".py"}},"nbformat":4,"nbformat_minor":1}