{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"pygments_lexer":"ipython3","nbconvert_exporter":"python","version":"3.6.4","file_extension":".py","codemirror_mode":{"name":"ipython","version":3},"name":"python","mimetype":"text/x-python"}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"**This notebook is an exercise in the [Intermediate Machine Learning](https://www.kaggle.com/learn/intermediate-machine-learning) course.  You can reference the tutorial at [this link](https://www.kaggle.com/alexisbcook/introduction).**\n\n---\n","metadata":{}},{"cell_type":"markdown","source":"As a warm-up, you'll review some machine learning fundamentals and submit your initial results to a Kaggle competition.\n\n# Setup\n\nThe questions below will give you feedback on your work. Run the following cell to set up the feedback system.","metadata":{}},{"cell_type":"code","source":"# Set up code checking\nimport os\nif not os.path.exists(\"../input/train.csv\"):\n    os.symlink(\"../input/home-data-for-ml-course/train.csv\", \"../input/train.csv\")  \n    os.symlink(\"../input/home-data-for-ml-course/test.csv\", \"../input/test.csv\")  \nfrom learntools.core import binder\nbinder.bind(globals())\nfrom learntools.ml_intermediate.ex1 import *\nprint(\"Setup Complete\")","metadata":{"execution":{"iopub.status.busy":"2022-07-28T22:12:06.191842Z","iopub.execute_input":"2022-07-28T22:12:06.192301Z","iopub.status.idle":"2022-07-28T22:12:07.584373Z","shell.execute_reply.started":"2022-07-28T22:12:06.192187Z","shell.execute_reply":"2022-07-28T22:12:07.580942Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"You will work with data from the [Housing Prices Competition for Kaggle Learn Users](https://www.kaggle.com/c/home-data-for-ml-course) to predict home prices in Iowa using 79 explanatory variables describing (almost) every aspect of the homes.  \n\n![Ames Housing dataset image](https://i.imgur.com/lTJVG4e.png)\n\nRun the next code cell without changes to load the training and validation features in `X_train` and `X_valid`, along with the prediction targets in `y_train` and `y_valid`.  The test features are loaded in `X_test`.  (_If you need to review **features** and **prediction targets**, please check out [this short tutorial](https://www.kaggle.com/dansbecker/your-first-machine-learning-model).  To read about model **validation**, look [here](https://www.kaggle.com/dansbecker/model-validation).  Alternatively, if you'd prefer to look through a full course to review all of these topics, start [here](https://www.kaggle.com/learn/machine-learning).)_","metadata":{}},{"cell_type":"code","source":"import pandas as pd\nfrom sklearn.model_selection import train_test_split\n\n# Read the data\nX_full = pd.read_csv('../input/train.csv', index_col='Id')\nX_test_full = pd.read_csv('../input/test.csv', index_col='Id')\n\n# Obtain target and predictors\ny = X_full.SalePrice\nfeatures = ['LotArea', 'YearBuilt', '1stFlrSF', '2ndFlrSF', 'FullBath', 'BedroomAbvGr', 'TotRmsAbvGrd']\nX = X_full[features].copy()\nX_test = X_test_full[features].copy()\n\n# Break off validation set from training data\nX_train, X_valid, y_train, y_valid = train_test_split(X, y, train_size=0.8, test_size=0.2,\n                                                      random_state=0)","metadata":{"execution":{"iopub.status.busy":"2022-07-28T22:15:19.114251Z","iopub.execute_input":"2022-07-28T22:15:19.114862Z","iopub.status.idle":"2022-07-28T22:15:19.222760Z","shell.execute_reply.started":"2022-07-28T22:15:19.114828Z","shell.execute_reply":"2022-07-28T22:15:19.221511Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Use the next cell to print the first several rows of the data. It's a nice way to get an overview of the data you will use in your price prediction model.","metadata":{}},{"cell_type":"code","source":"X_train.head()","metadata":{"execution":{"iopub.status.busy":"2022-07-28T22:15:48.409900Z","iopub.execute_input":"2022-07-28T22:15:48.410246Z","iopub.status.idle":"2022-07-28T22:15:48.422377Z","shell.execute_reply.started":"2022-07-28T22:15:48.410218Z","shell.execute_reply":"2022-07-28T22:15:48.421544Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"The next code cell defines five different random forest models.  Run this code cell without changes.  (_To review **random forests**, look [here](https://www.kaggle.com/dansbecker/random-forests)._)","metadata":{}},{"cell_type":"code","source":"from sklearn.ensemble import RandomForestRegressor\n\n# Define the models\nmodel_1 = RandomForestRegressor(n_estimators=50, random_state=0)\nmodel_2 = RandomForestRegressor(n_estimators=100, random_state=0)\nmodel_3 = RandomForestRegressor(n_estimators=100, criterion='absolute_error', random_state=0)\nmodel_4 = RandomForestRegressor(n_estimators=200, min_samples_split=20, random_state=0)\nmodel_5 = RandomForestRegressor(n_estimators=100, max_depth=7, random_state=0)\n\nmodels = [model_1, model_2, model_3, model_4, model_5]","metadata":{"execution":{"iopub.status.busy":"2022-07-28T22:16:43.570465Z","iopub.execute_input":"2022-07-28T22:16:43.570762Z","iopub.status.idle":"2022-07-28T22:16:43.577216Z","shell.execute_reply.started":"2022-07-28T22:16:43.570740Z","shell.execute_reply":"2022-07-28T22:16:43.576348Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"To select the best model out of the five, we define a function `score_model()` below.  This function returns the mean absolute error (MAE) from the validation set.  Recall that the best model will obtain the lowest MAE.  (_To review **mean absolute error**, look [here](https://www.kaggle.com/dansbecker/model-validation).)_\n\nRun the code cell without changes.","metadata":{}},{"cell_type":"code","source":"from sklearn.metrics import mean_absolute_error\n\n# Function for comparing different models\ndef score_model(model, X_t=X_train, X_v=X_valid, y_t=y_train, y_v=y_valid):\n    model.fit(X_t, y_t)\n    preds = model.predict(X_v)\n    return mean_absolute_error(y_v, preds)\n\nfor i in range(0, len(models)):\n    mae = score_model(models[i])\n    print(\"Model %d MAE: %d\" % (i+1, mae))","metadata":{"execution":{"iopub.status.busy":"2022-07-28T22:17:52.558326Z","iopub.execute_input":"2022-07-28T22:17:52.558630Z","iopub.status.idle":"2022-07-28T22:17:59.796042Z","shell.execute_reply.started":"2022-07-28T22:17:52.558606Z","shell.execute_reply":"2022-07-28T22:17:59.794816Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Step 1: Evaluate several models\n\nUse the above results to fill in the line below.  Which model is the best model?  Your answer should be one of `model_1`, `model_2`, `model_3`, `model_4`, or `model_5`.","metadata":{}},{"cell_type":"code","source":"# Fill in the best model\nbest_model = model_3\n\n# Check your answer\nstep_1.check()","metadata":{"execution":{"iopub.status.busy":"2022-07-28T22:19:00.667038Z","iopub.execute_input":"2022-07-28T22:19:00.667416Z","iopub.status.idle":"2022-07-28T22:19:00.675559Z","shell.execute_reply.started":"2022-07-28T22:19:00.667387Z","shell.execute_reply":"2022-07-28T22:19:00.674664Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Step 2: Generate test predictions\n\nGreat. You know how to evaluate what makes an accurate model. Now it's time to go through the modeling process and make predictions. In the line below, create a Random Forest model with the variable name `my_model`.","metadata":{}},{"cell_type":"code","source":"# Define a model\nmy_model = RandomForestRegressor(n_estimators=100, criterion='absolute_error', random_state=0) # Your code here\n\n# Check your answer\nstep_2.check()","metadata":{"execution":{"iopub.status.busy":"2022-07-28T22:20:08.282908Z","iopub.execute_input":"2022-07-28T22:20:08.283311Z","iopub.status.idle":"2022-07-28T22:20:08.293438Z","shell.execute_reply.started":"2022-07-28T22:20:08.283282Z","shell.execute_reply":"2022-07-28T22:20:08.292314Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Run the next code cell without changes.  The code fits the model to the training and validation data, and then generates test predictions that are saved to a CSV file.  These test predictions can be submitted directly to the competition!","metadata":{}},{"cell_type":"code","source":"# Fit the model to the training data\nmy_model.fit(X, y)\n\n# Generate test predictions\npreds_test = my_model.predict(X_test)\n\n# Save predictions in format used for competition scoring\noutput = pd.DataFrame({'Id': X_test.index,\n                       'SalePrice': preds_test})\noutput.to_csv('submission.csv', index=False)","metadata":{"execution":{"iopub.status.busy":"2022-07-28T22:20:42.422812Z","iopub.execute_input":"2022-07-28T22:20:42.423177Z","iopub.status.idle":"2022-07-28T22:20:46.257148Z","shell.execute_reply.started":"2022-07-28T22:20:42.423147Z","shell.execute_reply":"2022-07-28T22:20:46.256534Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Submit your results\n\nOnce you have successfully completed Step 2, you're ready to submit your results to the leaderboard!  First, you'll need to join the competition if you haven't already.  So open a new window by clicking on [this link](https://www.kaggle.com/c/home-data-for-ml-course).  Then click on the **Join Competition** button.  _(If you see a \"Submit Predictions\" button instead of a \"Join Competition\" button, you have already joined the competition, and don't need to do so again.)_\n\nNext, follow the instructions below:\n1. Begin by clicking on the **Save Version** button in the top right corner of the window.  This will generate a pop-up window.  \n2. Ensure that the **Save and Run All** option is selected, and then click on the **Save** button.\n3. This generates a window in the bottom left corner of the notebook.  After it has finished running, click on the number to the right of the **Save Version** button.  This pulls up a list of versions on the right of the screen.  Click on the ellipsis **(...)** to the right of the most recent version, and select **Open in Viewer**.  This brings you into view mode of the same page. You will need to scroll down to get back to these instructions.\n4. Click on the **Output** tab on the right of the screen.  Then, click on the file you would like to submit, and click on the **Submit** button to submit your results to the leaderboard.\n\nYou have now successfully submitted to the competition!\n\nIf you want to keep working to improve your performance, select the **Edit** button in the top right of the screen. Then you can change your code and repeat the process. There's a lot of room to improve, and you will climb up the leaderboard as you work.\n","metadata":{}},{"cell_type":"markdown","source":"# Keep going\n\nYou've made your first model. But how can you quickly make it better?\n\nLearn how to improve your competition results by incorporating columns with **[missing values](https://www.kaggle.com/alexisbcook/missing-values)**.","metadata":{}},{"cell_type":"markdown","source":"---\n\n\n\n\n*Have questions or comments? Visit the [course discussion forum](https://www.kaggle.com/learn/intermediate-machine-learning/discussion) to chat with other learners.*","metadata":{}}]}