{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"pygments_lexer":"ipython3","nbconvert_exporter":"python","version":"3.6.4","file_extension":".py","codemirror_mode":{"name":"ipython","version":3},"name":"python","mimetype":"text/x-python"}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"**This notebook is an exercise in the [Intermediate Machine Learning](https://www.kaggle.com/learn/intermediate-machine-learning) course.  You can reference the tutorial at [this link](https://www.kaggle.com/alexisbcook/categorical-variables).**\n\n---\n","metadata":{}},{"cell_type":"markdown","source":"By encoding **categorical variables**, you'll obtain your best results thus far!\n\n# Setup\n\nThe questions below will give you feedback on your work. Run the following cell to set up the feedback system.","metadata":{}},{"cell_type":"code","source":"# Set up code checking\nimport os\nif not os.path.exists(\"../input/train.csv\"):\n    os.symlink(\"../input/home-data-for-ml-course/train.csv\", \"../input/train.csv\")  \n    os.symlink(\"../input/home-data-for-ml-course/test.csv\", \"../input/test.csv\") \nfrom learntools.core import binder\nbinder.bind(globals())\nfrom learntools.ml_intermediate.ex3 import *\nprint(\"Setup Complete\")","metadata":{"execution":{"iopub.status.busy":"2022-07-09T16:20:46.462734Z","iopub.execute_input":"2022-07-09T16:20:46.463536Z","iopub.status.idle":"2022-07-09T16:20:46.537361Z","shell.execute_reply.started":"2022-07-09T16:20:46.463444Z","shell.execute_reply":"2022-07-09T16:20:46.536496Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"In this exercise, you will work with data from the [Housing Prices Competition for Kaggle Learn Users](https://www.kaggle.com/c/home-data-for-ml-course). \n\n![Ames Housing dataset image](https://i.imgur.com/lTJVG4e.png)\n\nRun the next code cell without changes to load the training and validation sets in `X_train`, `X_valid`, `y_train`, and `y_valid`.  The test set is loaded in `X_test`.","metadata":{}},{"cell_type":"code","source":"import pandas as pd\nfrom sklearn.model_selection import train_test_split\n\n# Read the data\nX = pd.read_csv('../input/train.csv', index_col='Id') \nX_test = pd.read_csv('../input/test.csv', index_col='Id')\n\n# Remove rows with missing target, separate target from predictors\nX.dropna(axis=0, subset=['SalePrice'], inplace=True)\ny = X.SalePrice\nX.drop(['SalePrice'], axis=1, inplace=True)\n\n# To keep things simple, we'll drop columns with missing values\ncols_with_missing = [col for col in X.columns if X[col].isnull().any()]\n\n# Add test columns with missing values\n#test_cols_with_missing = \n\nX.drop(cols_with_missing, axis=1, inplace=True)\nX_test.drop(cols_with_missing, axis=1, inplace=True)\n\n# Create copies of X and X_test\nX2 = X.copy()\nX2_test = X_test.copy()\n\n# Break off validation set from training data\nX_train, X_valid, y_train, y_valid = train_test_split(X, y,\n                                                      train_size=0.8, test_size=0.2,\n                                                      random_state=0)","metadata":{"execution":{"iopub.status.busy":"2022-07-09T16:20:46.540992Z","iopub.execute_input":"2022-07-09T16:20:46.541536Z","iopub.status.idle":"2022-07-09T16:20:47.646292Z","shell.execute_reply.started":"2022-07-09T16:20:46.541488Z","shell.execute_reply":"2022-07-09T16:20:47.645281Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Categorical columns in test set that have missing values\ntest_cols_with_missing = [col for col in X2_test.columns if X2_test[col].isnull().any()]\n\n# Drop columns with missing test values from copies of training and test sets \nX2.drop(test_cols_with_missing, axis = 1, inplace = True)\nX2_test.drop(test_cols_with_missing, axis = 1, inplace = True)\n\nprint(X2.shape)","metadata":{"execution":{"iopub.status.busy":"2022-07-09T16:39:24.214778Z","iopub.execute_input":"2022-07-09T16:39:24.21624Z","iopub.status.idle":"2022-07-09T16:39:24.234619Z","shell.execute_reply.started":"2022-07-09T16:39:24.216167Z","shell.execute_reply":"2022-07-09T16:39:24.233387Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Use the next code cell to print the first five rows of the data.","metadata":{}},{"cell_type":"code","source":"X_train.head()","metadata":{"execution":{"iopub.status.busy":"2022-07-09T16:20:47.675757Z","iopub.execute_input":"2022-07-09T16:20:47.676824Z","iopub.status.idle":"2022-07-09T16:20:47.711201Z","shell.execute_reply.started":"2022-07-09T16:20:47.676789Z","shell.execute_reply":"2022-07-09T16:20:47.710469Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Notice that the dataset contains both numerical and categorical variables.  You'll need to encode the categorical data before training a model.\n\nTo compare different models, you'll use the same `score_dataset()` function from the tutorial.  This function reports the [mean absolute error](https://en.wikipedia.org/wiki/Mean_absolute_error) (MAE) from a random forest model.","metadata":{}},{"cell_type":"code","source":"from sklearn.ensemble import RandomForestRegressor\nfrom sklearn.metrics import mean_absolute_error\n\n# function for comparing different approaches\ndef score_dataset(X_train, X_valid, y_train, y_valid):\n    model = RandomForestRegressor(n_estimators=100, random_state=0)\n    model.fit(X_train, y_train)\n    preds = model.predict(X_valid)\n    return mean_absolute_error(y_valid, preds)","metadata":{"execution":{"iopub.status.busy":"2022-07-09T16:20:47.71362Z","iopub.execute_input":"2022-07-09T16:20:47.714709Z","iopub.status.idle":"2022-07-09T16:20:47.928526Z","shell.execute_reply.started":"2022-07-09T16:20:47.714665Z","shell.execute_reply":"2022-07-09T16:20:47.927722Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Step 1: Drop columns with categorical data\n\nYou'll get started with the most straightforward approach.  Use the code cell below to preprocess the data in `X_train` and `X_valid` to remove columns with categorical data.  Set the preprocessed DataFrames to `drop_X_train` and `drop_X_valid`, respectively.  ","metadata":{}},{"cell_type":"code","source":"# Drop categorical columns\n\ndrop_X_train = X_train.select_dtypes(exclude = ['object'])\ndrop_X_valid = X_valid.select_dtypes(exclude = ['object'])\n\n# Check answer\nstep_1.check()","metadata":{"execution":{"iopub.status.busy":"2022-07-09T16:20:47.929695Z","iopub.execute_input":"2022-07-09T16:20:47.930121Z","iopub.status.idle":"2022-07-09T16:20:47.943297Z","shell.execute_reply.started":"2022-07-09T16:20:47.930097Z","shell.execute_reply":"2022-07-09T16:20:47.942207Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Lines below will give you a hint or solution code\n#step_1.hint()\n#step_1.solution()","metadata":{"execution":{"iopub.status.busy":"2022-07-09T16:20:47.944671Z","iopub.execute_input":"2022-07-09T16:20:47.9455Z","iopub.status.idle":"2022-07-09T16:20:47.95178Z","shell.execute_reply.started":"2022-07-09T16:20:47.945469Z","shell.execute_reply":"2022-07-09T16:20:47.950779Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Run the next code cell to get the MAE for this approach.","metadata":{}},{"cell_type":"code","source":"print(\"MAE from 1st Approach (drop categorical variables):\")\nprint(score_dataset(drop_X_train, drop_X_valid,y_train, y_valid))","metadata":{"execution":{"iopub.status.busy":"2022-07-09T16:20:47.95289Z","iopub.execute_input":"2022-07-09T16:20:47.953178Z","iopub.status.idle":"2022-07-09T16:20:49.326679Z","shell.execute_reply.started":"2022-07-09T16:20:47.953147Z","shell.execute_reply":"2022-07-09T16:20:49.325495Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Before jumping into ordinal encoding, we'll investigate the dataset.  Specifically, we'll look at the `'Condition2'` column.  The code cell below prints the unique entries in both the training and validation sets.","metadata":{}},{"cell_type":"code","source":"print(\"Unique values in 'Condition2' column (training dataset):\\n\", X_train[\"Condition2\"].unique())\nprint(\"\\nUnique values in 'Condition2' column (validation dataset):\\n\", X_valid[\"Condition2\"].unique())","metadata":{"execution":{"iopub.status.busy":"2022-07-09T16:20:49.327869Z","iopub.execute_input":"2022-07-09T16:20:49.328171Z","iopub.status.idle":"2022-07-09T16:20:49.335599Z","shell.execute_reply.started":"2022-07-09T16:20:49.328141Z","shell.execute_reply":"2022-07-09T16:20:49.334479Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Step 2: Ordinal encoding\n\n### Part A\n\nIf you now write code to: \n- fit an ordinal encoder to the training data, and then \n- use it to transform both the training and validation data, \n\nyou'll get an error.  Can you see why this is the case?  (_You'll need  to use the above output to answer this question._)","metadata":{}},{"cell_type":"code","source":"# Check your answer (Run this code cell to receive credit!)\nstep_2.a.check()","metadata":{"execution":{"iopub.status.busy":"2022-07-09T16:20:49.337213Z","iopub.execute_input":"2022-07-09T16:20:49.337705Z","iopub.status.idle":"2022-07-09T16:20:49.350474Z","shell.execute_reply.started":"2022-07-09T16:20:49.337675Z","shell.execute_reply":"2022-07-09T16:20:49.34957Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"#step_2.a.hint()","metadata":{"execution":{"iopub.status.busy":"2022-07-09T16:20:49.35568Z","iopub.execute_input":"2022-07-09T16:20:49.356342Z","iopub.status.idle":"2022-07-09T16:20:49.361304Z","shell.execute_reply.started":"2022-07-09T16:20:49.356305Z","shell.execute_reply":"2022-07-09T16:20:49.359693Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"This is a common problem that you'll encounter with real-world data, and there are many approaches to fixing this issue.  For instance, you can write a custom ordinal encoder to deal with new categories.  The simplest approach, however, is to drop the problematic categorical columns.  \n\nRun the code cell below to save the problematic columns to a Python list `bad_label_cols`.  Likewise, columns that can be safely ordinal encoded are stored in `good_label_cols`.","metadata":{}},{"cell_type":"code","source":"# Categorical columns in training data \nobject_cols = [col for col in X_train.columns if X_train[col].dtype == \"object\"]\n\n# Columns that can be safely ordinal encoded\ndo_label_cols = [col for col in object_cols if set(X_train[col]).issuperset(set(X_valid[col]))]\n\n# Challenging columns that will be dropped from the dataset\nno_label_cols = list(set(object_cols) - set(do_label_cols))\n\nprint(\"Columns to be ordinal encoded: \", do_label_cols)\nprint(\"\\nColumns to be dropped: \", no_label_cols)","metadata":{"execution":{"iopub.status.busy":"2022-07-09T16:20:49.36282Z","iopub.execute_input":"2022-07-09T16:20:49.363651Z","iopub.status.idle":"2022-07-09T16:20:49.388984Z","shell.execute_reply.started":"2022-07-09T16:20:49.363612Z","shell.execute_reply":"2022-07-09T16:20:49.387442Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### Part B\n\nUse the next code cell to ordinal encode the data in `X_train` and `X_valid`.  Set the preprocessed DataFrames to `label_X_train` and `label_X_valid`, respectively.  \n- We have provided code below to drop the categorical columns in `bad_label_cols` from the dataset. \n- You should ordinal encode the categorical columns in `good_label_cols`.  ","metadata":{}},{"cell_type":"markdown","source":"# 🔴","metadata":{}},{"cell_type":"code","source":"from sklearn.preprocessing import OrdinalEncoder\n\n# Drop categorical columns that will not be ordinal encoded \nlabel_X_train = X_train.drop(no_label_cols, axis = 1)\nlabel_X_valid = X_valid.drop(no_label_cols, axis = 1)\n\n# Ordinal encode categorical columns\nor_en = OrdinalEncoder()\nlabel_X_train[do_label_cols] = or_en.fit_transform(label_X_train[do_label_cols])\nlabel_X_valid[do_label_cols] = or_en.transform(label_X_valid[do_label_cols])\n\n# Check answer\nstep_2.b.check()","metadata":{"execution":{"iopub.status.busy":"2022-07-09T16:20:49.390768Z","iopub.execute_input":"2022-07-09T16:20:49.391411Z","iopub.status.idle":"2022-07-09T16:20:49.437085Z","shell.execute_reply.started":"2022-07-09T16:20:49.391373Z","shell.execute_reply":"2022-07-09T16:20:49.436337Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Lines below will give you a hint or solution code\n#step_2.b.hint()\n#step_2.b.solution()","metadata":{"execution":{"iopub.status.busy":"2022-07-09T16:20:49.438212Z","iopub.execute_input":"2022-07-09T16:20:49.439193Z","iopub.status.idle":"2022-07-09T16:20:49.443009Z","shell.execute_reply.started":"2022-07-09T16:20:49.43916Z","shell.execute_reply":"2022-07-09T16:20:49.442105Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Run the next code cell to get the MAE for this approach.","metadata":{}},{"cell_type":"code","source":"print(\"MAE from Approach 2 (Ordinal Encoding):\")\nprint(score_dataset(label_X_train, label_X_valid, y_train, y_valid)) ","metadata":{"execution":{"iopub.status.busy":"2022-07-09T16:20:49.445545Z","iopub.execute_input":"2022-07-09T16:20:49.446012Z","iopub.status.idle":"2022-07-09T16:20:50.817394Z","shell.execute_reply.started":"2022-07-09T16:20:49.445975Z","shell.execute_reply":"2022-07-09T16:20:50.816446Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"So far, you've tried two different approaches to dealing with categorical variables.  And, you've seen that encoding categorical data yields better results than removing columns from the dataset.\n\nSoon, you'll try one-hot encoding.  Before then, there's one additional topic we need to cover.  Begin by running the next code cell without changes.  ","metadata":{}},{"cell_type":"code","source":"# Get number of unique values in categorical columns \nobject_nunique = list(map(lambda col: X_train[col].nunique(), object_cols))\nd = dict(zip(object_cols, object_nunique))\n\n# Print number of unique values by column in ascending order\nsorted(d.items(), key = lambda x: x[1])","metadata":{"execution":{"iopub.status.busy":"2022-07-09T16:20:50.819046Z","iopub.execute_input":"2022-07-09T16:20:50.819362Z","iopub.status.idle":"2022-07-09T16:20:50.83529Z","shell.execute_reply.started":"2022-07-09T16:20:50.819338Z","shell.execute_reply":"2022-07-09T16:20:50.833824Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Step 3: Investigating cardinality\n\n### Part A\n\nThe output above shows, for each column with categorical data, the number of unique values in the column.  For instance, the `'Street'` column in the training data has two unique values: `'Grvl'` and `'Pave'`, corresponding to a gravel road and a paved road, respectively.\n\nWe refer to the number of unique entries of a categorical variable as the **cardinality** of that categorical variable.  For instance, the `'Street'` variable has cardinality 2.\n\nUse the output above to answer the questions below.","metadata":{}},{"cell_type":"code","source":"# How many categorical columns have more unique values than 10\nhigh_cardinality_numcols = len([col for col in X_train.columns if X_train[col].dtype == 'object' and X_train[col].nunique() > 10])\n\n# How many columns are needed to one-hot encode the 'Neighborhood' variable in the training data \nnum_cols_neighborhood = X_train.Neighborhood.nunique()\n\n# Check answers\nstep_3.a.check()","metadata":{"execution":{"iopub.status.busy":"2022-07-09T16:20:50.836468Z","iopub.execute_input":"2022-07-09T16:20:50.836737Z","iopub.status.idle":"2022-07-09T16:20:50.855684Z","shell.execute_reply.started":"2022-07-09T16:20:50.83671Z","shell.execute_reply":"2022-07-09T16:20:50.854832Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Lines below will give you a hint or solution code\n#step_3.a.hint()\n#step_3.a.solution()","metadata":{"execution":{"iopub.status.busy":"2022-07-09T16:20:50.858456Z","iopub.execute_input":"2022-07-09T16:20:50.859286Z","iopub.status.idle":"2022-07-09T16:20:50.86591Z","shell.execute_reply.started":"2022-07-09T16:20:50.859253Z","shell.execute_reply":"2022-07-09T16:20:50.86485Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### Part B\n\nFor large datasets with many rows, one-hot encoding can greatly expand the size of the dataset.  For this reason, we typically will only one-hot encode columns with relatively low cardinality.  Then, high cardinality columns can either be dropped from the dataset, or we can use ordinal encoding.\n\nAs an example, consider a dataset with 10,000 rows, and containing one categorical column with 100 unique entries.  \n- If this column is replaced with the corresponding one-hot encoding, how many entries are added to the dataset?  \n- If we instead replace the column with the ordinal encoding, how many entries are added?  \n\nUse your answers to fill in the lines below.","metadata":{}},{"cell_type":"code","source":"# How many entries are added to the dataset by replacing the column with one-hot encoding \nOH_entries_added = (100 *10000) - 10000\n\n# Entries added by replacing column with ordinal encoding\nlabel_entries_added = 10000 - 10000\nstep_3.b.check()","metadata":{"execution":{"iopub.status.busy":"2022-07-09T16:20:50.869184Z","iopub.execute_input":"2022-07-09T16:20:50.869805Z","iopub.status.idle":"2022-07-09T16:20:50.880078Z","shell.execute_reply.started":"2022-07-09T16:20:50.869764Z","shell.execute_reply":"2022-07-09T16:20:50.879254Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Lines below will give you a hint or solution code\n#step_3.b.hint()\n#step_3.b.solution()","metadata":{"execution":{"iopub.status.busy":"2022-07-09T16:20:50.881312Z","iopub.execute_input":"2022-07-09T16:20:50.882256Z","iopub.status.idle":"2022-07-09T16:20:50.888208Z","shell.execute_reply.started":"2022-07-09T16:20:50.882225Z","shell.execute_reply":"2022-07-09T16:20:50.887354Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Next, you'll experiment with one-hot encoding.  But, instead of encoding all of the categorical variables in the dataset, you'll only create a one-hot encoding for columns with cardinality less than 10.\n\nRun the code cell below without changes to set `low_cardinality_cols` to a Python list containing the columns that will be one-hot encoded.  Likewise, `high_cardinality_cols` contains a list of categorical columns that will be dropped from the dataset.","metadata":{}},{"cell_type":"code","source":"# Columns that will be one-hot encoded \nlow_cardinality_cols = list([col for col in object_cols if X_train[col].nunique() < 10])\n\n# Columns that will be dropped\nhigh_cardinality_cols = list(set(object_cols) - set(low_cardinality_cols))\n\nprint('Categorical columns that will be one-hot encoded: ', low_cardinality_cols)\nprint('\\nCategorical columns that will be dropped: ', high_cardinality_cols)","metadata":{"execution":{"iopub.status.busy":"2022-07-09T16:20:50.889649Z","iopub.execute_input":"2022-07-09T16:20:50.890247Z","iopub.status.idle":"2022-07-09T16:20:50.91276Z","shell.execute_reply.started":"2022-07-09T16:20:50.890204Z","shell.execute_reply":"2022-07-09T16:20:50.911328Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Step 4: One-hot encoding\n\nUse the next code cell to one-hot encode the data in `X_train` and `X_valid`.  Set the preprocessed DataFrames to `OH_X_train` and `OH_X_valid`, respectively.  \n- The full list of categorical columns in the dataset can be found in the Python list `object_cols`.\n- You should only one-hot encode the categorical columns in `low_cardinality_cols`.  All other categorical columns should be dropped from the dataset. ","metadata":{}},{"cell_type":"code","source":"from sklearn.preprocessing import OneHotEncoder\n\n# Remove categorical columns\nnum_cols_train = X_train.drop(object_cols, axis = 1)\nnum_cols_valid = X_valid.drop(object_cols, axis = 1)\n\nOH_en = OneHotEncoder(handle_unknown = 'ignore',  sparse = False)\n\n# One-hot encode low cardinality columns\nOH_cols_train = pd.DataFrame(OH_en.fit_transform(X_train[low_cardinality_cols]))\nOH_cols_valid = pd.DataFrame(OH_en.transform(X_valid[low_cardinality_cols]))\n\n# Return indexes removed by one-hot encoding \nOH_cols_train.index = X_train.index\nOH_cols_valid.index = X_valid.index\n\n# Add one-hot encoded columns to dataset\nOH_X_train = pd.concat([num_cols_train, OH_cols_train], axis = 1)\nOH_X_valid = pd.concat([num_cols_valid, OH_cols_valid], axis = 1)\n\n# Check answer\nstep_4.check()","metadata":{"execution":{"iopub.status.busy":"2022-07-09T16:20:50.914524Z","iopub.execute_input":"2022-07-09T16:20:50.914925Z","iopub.status.idle":"2022-07-09T16:20:50.970816Z","shell.execute_reply.started":"2022-07-09T16:20:50.914833Z","shell.execute_reply":"2022-07-09T16:20:50.969929Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Lines below will give you a hint or solution code\n#step_4.hint()\n#step_4.solution()","metadata":{"execution":{"iopub.status.busy":"2022-07-09T16:20:50.971952Z","iopub.execute_input":"2022-07-09T16:20:50.972251Z","iopub.status.idle":"2022-07-09T16:20:50.976513Z","shell.execute_reply.started":"2022-07-09T16:20:50.972225Z","shell.execute_reply":"2022-07-09T16:20:50.975393Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Run the next code cell to get the MAE for this approach.","metadata":{}},{"cell_type":"code","source":"print(\"MAE from Approach 3 (One-Hot Encoding):\")\nprint(score_dataset(OH_X_train, OH_X_valid, y_train, y_valid))","metadata":{"execution":{"iopub.status.busy":"2022-07-09T16:20:50.97773Z","iopub.execute_input":"2022-07-09T16:20:50.977988Z","iopub.status.idle":"2022-07-09T16:20:53.029201Z","shell.execute_reply.started":"2022-07-09T16:20:50.977955Z","shell.execute_reply":"2022-07-09T16:20:53.028055Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Generate test predictions and submit your results\n\nAfter you complete Step 4, if you'd like to use what you've learned to submit your results to the leaderboard, you'll need to preprocess the test data before generating predictions.\n\n**This step is completely optional, and you do not need to submit results to the leaderboard to successfully complete the exercise.**\n\nCheck out the previous exercise if you need help with remembering how to [join the competition](https://www.kaggle.com/c/home-data-for-ml-course) or save your results to CSV.  Once you have generated a file with your results, follow the instructions below:\n1. Begin by clicking on the **Save Version** button in the top right corner of the window.  This will generate a pop-up window.  \n2. Ensure that the **Save and Run All** option is selected, and then click on the **Save** button.\n3. This generates a window in the bottom left corner of the notebook.  After it has finished running, click on the number to the right of the **Save Version** button.  This pulls up a list of versions on the right of the screen.  Click on the ellipsis **(...)** to the right of the most recent version, and select **Open in Viewer**.  This brings you into view mode of the same page. You will need to scroll down to get back to these instructions.\n4. Click on the **Output** tab on the right of the screen.  Then, click on the file you would like to submit, and click on the **Submit** button to submit your results to the leaderboard.\n\nYou have now successfully submitted to the competition!\n\nIf you want to keep working to improve your performance, select the **Edit** button in the top right of the screen. Then you can change your code and repeat the process. There's a lot of room to improve, and you will climb up the leaderboard as you work.\n","metadata":{}},{"cell_type":"code","source":"print(X2.columns)\nX2.head()","metadata":{"execution":{"iopub.status.busy":"2022-07-09T16:20:53.030608Z","iopub.execute_input":"2022-07-09T16:20:53.031478Z","iopub.status.idle":"2022-07-09T16:20:53.058302Z","shell.execute_reply.started":"2022-07-09T16:20:53.031446Z","shell.execute_reply":"2022-07-09T16:20:53.056976Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"test_na_cols = [col for col in X2_test.columns if X2_test[col].isnull().any()]\n#print(test_na_cols)\n\nfor col in test_na_cols:\n    print({col: X2_test[col].dtype})","metadata":{"execution":{"iopub.status.busy":"2022-07-09T16:36:01.333323Z","iopub.execute_input":"2022-07-09T16:36:01.33371Z","iopub.status.idle":"2022-07-09T16:36:01.350728Z","shell.execute_reply.started":"2022-07-09T16:36:01.33368Z","shell.execute_reply":"2022-07-09T16:36:01.349605Z"},"collapsed":true,"jupyter":{"outputs_hidden":true},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Create Submission\n\npreds_object_cols = [col for col in X2.columns if X2[col].dtype == \"object\"]\n\n# Categorical columns that would be challenging to ordinal encode \npreds_do_label_cols = [col for col in preds_object_cols if set(X2_test[col]).issubset(set(X2[col]))]\npreds_no_label_cols = list(set(preds_object_cols)-set(preds_do_label_cols))\n\n# Drop columns that would be challenging to ordinal encode\nX2.drop(preds_no_label_cols, axis = 1)\nX2_test.drop(preds_no_label_cols, axis = 1)\n\n# Ordinal encode categorical variables \nX2[preds_do_label_cols] = or_en.fit_transform(X2[preds_do_label_cols])\nX2_test[preds_do_label_cols] = or_en.transform(X2_test[preds_do_label_cols])\n\n\n# Create model and make predictions\nfinal_model = RandomForestRegressor(n_estimators = 100)\nfinal_model.fit(X2, y)\npreds_test = final_model.predict(X2_test)\n\n# Create submission\noutput = pd.DataFrame({\n    \"Id\": X2_test.index,\n    \"SalePrice\": preds_test\n})\n\noutput.to_csv(\"submission.csv\", index = False)","metadata":{"execution":{"iopub.status.busy":"2022-07-09T16:39:51.124976Z","iopub.execute_input":"2022-07-09T16:39:51.125731Z","iopub.status.idle":"2022-07-09T16:39:52.354711Z","shell.execute_reply.started":"2022-07-09T16:39:51.125703Z","shell.execute_reply":"2022-07-09T16:39:52.353299Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Keep going\n\nWith missing value handling and categorical encoding, your modeling process is getting complex. This complexity gets worse when you want to save your model to use in the future. The key to managing this complexity is something called **pipelines**. \n\n**[Learn to use pipelines](https://www.kaggle.com/alexisbcook/pipelines)** to preprocess datasets with categorical variables, missing values and any other messiness your data throws at you.","metadata":{}},{"cell_type":"markdown","source":"---\n\n\n\n\n*Have questions or comments? Visit the [course discussion forum](https://www.kaggle.com/learn/intermediate-machine-learning/discussion) to chat with other learners.*","metadata":{}}]}