{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"pygments_lexer":"ipython3","nbconvert_exporter":"python","version":"3.6.4","file_extension":".py","codemirror_mode":{"name":"ipython","version":3},"name":"python","mimetype":"text/x-python"}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"# <p style=\"background-color:coral;font-family:newtimeroman;font-size:150%;color:white;text-align:center;border-radius:20px 20px;\"><b>Tree-Based Machine Learning Models</b></p>\n![](https://img.freepik.com/free-photo/cherry-tree-grassy-landscape_1048-7642.jpg?w=1380&t=st=1658943811~exp=1658944411~hmac=dcc4bebdb38fe88c3c78099f0d2aa02fd23b55d3449ba4c9bcf451e05ba82282)\n","metadata":{}},{"cell_type":"markdown","source":"<a id=\"toc\"></a>\n<b>Hi guys </b>😀\n\nIn this notebook, I'm going to talk about tree-based machine learning models.\n\n<b>Table of contents:</b>\n<ul>\n<li><a href=\"#Loading\">Loading the dataset</a></li>  \n<li><a href=\"#Understanding\">Understanding the dataset</a></li>         \n<li><a href=\"#Data-Preprocessing\">Data preprocessing</a></li>\n<li><a href=\"#Decision-Tree\">Decision tree</a></li>\n<li><a href=\"#Random-Forest\">Random forest</a></li>      \n<li><a href=\"#XGBoost\">XGBoost</a></li>    \n<li><a href=\"#Comet\">Optimizing hyperparameters with Comet</a></li> \n<li><a href=\"#Conclusion\">Conclusion</a></li>   \n</ul>","metadata":{}},{"cell_type":"markdown","source":"<a id=\"Loading\"></a>\n# <p style=\"background-color:coral;font-family:newtimeroman;font-size:150%;color:white;text-align:center;border-radius:20px 20px;\"><b>Loading the Dataset</b></p>\n<a id=\"Loading\"></a> \n\n<a id=\"0\"></a>\n<a href=\"#toc\" class=\"btn btn-primary btn-sm\" role=\"button\" aria-pressed=\"true\" \nstyle=\"color:blue; background-color:#dfa8e4\" data-toggle=\"popover\">Content</a>","metadata":{}},{"cell_type":"markdown","source":"Let's use the [house prices dataset](https://www.kaggle.com/c/home-data-for-ml-course) available on kaggle to show tree-based models. The data set includes the sales prices of the houses according to their various features.","metadata":{}},{"cell_type":"code","source":"import pandas as pd\ndf = pd.read_csv(\"../input/home-data-for-ml-course/train.csv\", index_col=\"Id\")\ndf_test = pd.read_csv(\"../input/home-data-for-ml-course/test.csv\", index_col=\"Id\")\ndf.head()","metadata":{"execution":{"iopub.status.busy":"2022-07-27T18:20:08.321851Z","iopub.execute_input":"2022-07-27T18:20:08.322210Z","iopub.status.idle":"2022-07-27T18:20:08.384776Z","shell.execute_reply.started":"2022-07-27T18:20:08.322178Z","shell.execute_reply":"2022-07-27T18:20:08.383701Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<a id=\"Understanding\"></a>\n# <p style=\"background-color:coral;font-family:newtimeroman;font-size:150%;color:white;text-align:center;border-radius:20px 20px;\"><b>Understanding the Dataset</b></p>\n\n<a href=\"#toc\" class=\"btn btn-primary btn-sm\" role=\"button\" aria-pressed=\"true\" \nstyle=\"color:blue; background-color:#dfa8e4\" data-toggle=\"popover\">Content</a>","metadata":{}},{"cell_type":"code","source":"df.shape","metadata":{"execution":{"iopub.status.busy":"2022-07-27T18:20:08.386643Z","iopub.execute_input":"2022-07-27T18:20:08.387112Z","iopub.status.idle":"2022-07-27T18:20:08.393995Z","shell.execute_reply.started":"2022-07-27T18:20:08.387054Z","shell.execute_reply":"2022-07-27T18:20:08.392891Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"df.dtypes","metadata":{"execution":{"iopub.status.busy":"2022-07-27T18:20:08.395991Z","iopub.execute_input":"2022-07-27T18:20:08.396798Z","iopub.status.idle":"2022-07-27T18:20:08.406965Z","shell.execute_reply.started":"2022-07-27T18:20:08.396752Z","shell.execute_reply":"2022-07-27T18:20:08.405401Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"df.info(verbose=False)","metadata":{"execution":{"iopub.status.busy":"2022-07-27T18:20:08.409045Z","iopub.execute_input":"2022-07-27T18:20:08.409739Z","iopub.status.idle":"2022-07-27T18:20:08.429398Z","shell.execute_reply.started":"2022-07-27T18:20:08.409704Z","shell.execute_reply":"2022-07-27T18:20:08.428234Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"df.describe()","metadata":{"execution":{"iopub.status.busy":"2022-07-27T18:20:08.431183Z","iopub.execute_input":"2022-07-27T18:20:08.431780Z","iopub.status.idle":"2022-07-27T18:20:08.520686Z","shell.execute_reply.started":"2022-07-27T18:20:08.431746Z","shell.execute_reply":"2022-07-27T18:20:08.519732Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"df.describe(include=[object])","metadata":{"execution":{"iopub.status.busy":"2022-07-27T18:20:08.522679Z","iopub.execute_input":"2022-07-27T18:20:08.523012Z","iopub.status.idle":"2022-07-27T18:20:08.609773Z","shell.execute_reply.started":"2022-07-27T18:20:08.522978Z","shell.execute_reply":"2022-07-27T18:20:08.608719Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## <span style=\"color:Orange\">Handling Missing Data</span>","metadata":{}},{"cell_type":"code","source":"df.isnull().sum()","metadata":{"execution":{"iopub.status.busy":"2022-07-27T18:20:08.613126Z","iopub.execute_input":"2022-07-27T18:20:08.614283Z","iopub.status.idle":"2022-07-27T18:20:08.634026Z","shell.execute_reply.started":"2022-07-27T18:20:08.614249Z","shell.execute_reply":"2022-07-27T18:20:08.633011Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Let me draw a bar plot to show the columns with missing data.","metadata":{}},{"cell_type":"code","source":"import matplotlib.pyplot as plt\nimport seaborn as sns\nimport warnings \nwarnings.filterwarnings(\"ignore\")\nsns.set_theme()\nsns.set(rc={'figure.figsize':(12,8), 'figure.dpi':300})\ncols_with_null=df.isnull().sum().sort_values(ascending=False).head(10)\nsns.barplot(x=cols_with_null.index,y=cols_with_null)\nplt.xticks(rotation=90)","metadata":{"scrolled":true,"execution":{"iopub.status.busy":"2022-07-27T18:20:08.642363Z","iopub.execute_input":"2022-07-27T18:20:08.642893Z","iopub.status.idle":"2022-07-27T18:20:10.067709Z","shell.execute_reply.started":"2022-07-27T18:20:08.642855Z","shell.execute_reply":"2022-07-27T18:20:10.066762Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"cols_to_drop=(cols_with_null.head(6).index).tolist()\ndf.drop(cols_to_drop,axis=1,inplace=True)","metadata":{"execution":{"iopub.status.busy":"2022-07-27T18:20:10.068747Z","iopub.execute_input":"2022-07-27T18:20:10.069093Z","iopub.status.idle":"2022-07-27T18:20:10.077487Z","shell.execute_reply.started":"2022-07-27T18:20:10.069048Z","shell.execute_reply":"2022-07-27T18:20:10.076522Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Let's create the target and feature variables.","metadata":{}},{"cell_type":"code","source":"y = df.SalePrice\nX = df.drop(['SalePrice'], axis=1)","metadata":{"execution":{"iopub.status.busy":"2022-07-27T18:20:10.078955Z","iopub.execute_input":"2022-07-27T18:20:10.079530Z","iopub.status.idle":"2022-07-27T18:20:10.090771Z","shell.execute_reply.started":"2022-07-27T18:20:10.079492Z","shell.execute_reply":"2022-07-27T18:20:10.089933Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Let's split the dataset into the train and test set.","metadata":{}},{"cell_type":"code","source":"from sklearn.model_selection import train_test_split\nX_train, X_val, y_train, y_val = train_test_split(X, y, train_size=0.8, random_state=0)","metadata":{"execution":{"iopub.status.busy":"2022-07-27T18:20:10.092040Z","iopub.execute_input":"2022-07-27T18:20:10.092642Z","iopub.status.idle":"2022-07-27T18:20:10.102362Z","shell.execute_reply.started":"2022-07-27T18:20:10.092607Z","shell.execute_reply":"2022-07-27T18:20:10.101456Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Let's see the shape of train and validation set.","metadata":{}},{"cell_type":"code","source":"print(X_train.shape)\nprint(X_val.shape)","metadata":{"execution":{"iopub.status.busy":"2022-07-27T18:20:10.105926Z","iopub.execute_input":"2022-07-27T18:20:10.106253Z","iopub.status.idle":"2022-07-27T18:20:10.115749Z","shell.execute_reply.started":"2022-07-27T18:20:10.106230Z","shell.execute_reply":"2022-07-27T18:20:10.112310Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Let's select the numerical and categorical columns.","metadata":{}},{"cell_type":"code","source":"categorical_cols = [cname for cname in X_train.columns \n                    if X_train[cname].nunique() < 10 and X_train[cname].dtype == \"object\"]\nnumerical_cols = numerical_cols = [cname for cname in X_train.columns \n                    if X_train[cname].dtype in ['int64', 'float64']]","metadata":{"execution":{"iopub.status.busy":"2022-07-27T18:20:10.117800Z","iopub.execute_input":"2022-07-27T18:20:10.118224Z","iopub.status.idle":"2022-07-27T18:20:10.139538Z","shell.execute_reply.started":"2022-07-27T18:20:10.118185Z","shell.execute_reply":"2022-07-27T18:20:10.138738Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Let me show you the total number of categorical and numerical columns.","metadata":{}},{"cell_type":"code","source":"print(\"The total number of categorical columns:\", len(categorical_cols))\nprint(\"The total number of numerical columns:\", len(numerical_cols))","metadata":{"execution":{"iopub.status.busy":"2022-07-27T18:20:10.141828Z","iopub.execute_input":"2022-07-27T18:20:10.142242Z","iopub.status.idle":"2022-07-27T18:20:10.149257Z","shell.execute_reply.started":"2022-07-27T18:20:10.142215Z","shell.execute_reply":"2022-07-27T18:20:10.147922Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Let's select these columns from the datasets.","metadata":{}},{"cell_type":"code","source":"my_cols = categorical_cols + numerical_cols\nX_train = X_train[my_cols].copy()\nX_val= X_val[my_cols].copy()\nX_test = df_test[my_cols].copy()","metadata":{"execution":{"iopub.status.busy":"2022-07-27T18:20:10.150744Z","iopub.execute_input":"2022-07-27T18:20:10.151220Z","iopub.status.idle":"2022-07-27T18:20:10.164652Z","shell.execute_reply.started":"2022-07-27T18:20:10.151185Z","shell.execute_reply":"2022-07-27T18:20:10.163382Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## <span style=\"color:Orange\">Pipelines for Data Preprocessing</span>","metadata":{}},{"cell_type":"code","source":"from sklearn.pipeline import Pipeline\nfrom sklearn.impute import SimpleImputer\nfrom sklearn.preprocessing import StandardScaler","metadata":{"execution":{"iopub.status.busy":"2022-07-27T18:20:10.166569Z","iopub.execute_input":"2022-07-27T18:20:10.167223Z","iopub.status.idle":"2022-07-27T18:20:10.267693Z","shell.execute_reply.started":"2022-07-27T18:20:10.167187Z","shell.execute_reply":"2022-07-27T18:20:10.266820Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Let's create a transformer for numerical columns.","metadata":{}},{"cell_type":"code","source":"numerical_transformer = Pipeline(steps=[\n    ('imputer_num', SimpleImputer(strategy='median')), \n    ('scaler', StandardScaler())\n])","metadata":{"execution":{"iopub.status.busy":"2022-07-27T18:20:10.268897Z","iopub.execute_input":"2022-07-27T18:20:10.269341Z","iopub.status.idle":"2022-07-27T18:20:10.274904Z","shell.execute_reply.started":"2022-07-27T18:20:10.269305Z","shell.execute_reply":"2022-07-27T18:20:10.273816Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Let's create a transformer for categorical columns.","metadata":{}},{"cell_type":"code","source":"from sklearn.preprocessing import OneHotEncoder\ncategorical_transformer = Pipeline(steps=[\n    ('imputer_cat', SimpleImputer(strategy='most_frequent')),\n    ('onehot', OneHotEncoder(handle_unknown='ignore'))\n])","metadata":{"execution":{"iopub.status.busy":"2022-07-27T18:20:10.276568Z","iopub.execute_input":"2022-07-27T18:20:10.276907Z","iopub.status.idle":"2022-07-27T18:20:10.285416Z","shell.execute_reply.started":"2022-07-27T18:20:10.276874Z","shell.execute_reply":"2022-07-27T18:20:10.284534Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Let's perform these transformers to the numerical and categorical columns.","metadata":{}},{"cell_type":"code","source":"from sklearn.compose import ColumnTransformer\npreprocessor = ColumnTransformer(transformers=[\n    ('num', numerical_transformer, numerical_cols),\n    ('cat', categorical_transformer, categorical_cols)])","metadata":{"execution":{"iopub.status.busy":"2022-07-27T18:20:10.286784Z","iopub.execute_input":"2022-07-27T18:20:10.287143Z","iopub.status.idle":"2022-07-27T18:20:10.297579Z","shell.execute_reply.started":"2022-07-27T18:20:10.287110Z","shell.execute_reply":"2022-07-27T18:20:10.296713Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<a id=\"Decision-Tree\"></a> \n# <p style=\"background-color:coral;font-family:newtimeroman;font-size:150%;color:white;text-align:center;border-radius:20px 20px;\"><b>Decision Tree</b></p>\n\n<a href=\"#toc\" class=\"btn btn-primary btn-sm\" role=\"button\" aria-pressed=\"true\" \nstyle=\"color:blue; background-color:#dfa8e4\" data-toggle=\"popover\">Content</a>","metadata":{}},{"cell_type":"markdown","source":"Decision Trees are a non-parametric supervised learning method used for classification and regression. It encodes a series of if-then-else rules. Each node in a tree contains a condition. If the condition is satisfied, we go to the right side of the tree; otherwise, we go to the left.","metadata":{}},{"cell_type":"code","source":"from sklearn.tree import DecisionTreeRegressor\nfrom sklearn.metrics import r2_score\nimport numpy as np\n\nmodel_dt = DecisionTreeRegressor(random_state=0)\nmy_pipeline_dt = Pipeline(steps=[('preprocessor', preprocessor), ('model_dt', model_dt)])\nmy_pipeline_dt.fit(X_train, y_train)\ny_val_pred_dt= my_pipeline_dt.predict(X_val)\nprint(\"Decision model performance on validation data\", np.round(r2_score(y_val, y_val_pred_dt),2))\ny_train_pred_dt= my_pipeline_dt.predict(X_train)\nprint(\"Decision model performance on the train data:\", np.round(r2_score(y_train, y_train_pred_dt),2))","metadata":{"scrolled":true,"execution":{"iopub.status.busy":"2022-07-27T18:21:28.913460Z","iopub.execute_input":"2022-07-27T18:21:28.913869Z","iopub.status.idle":"2022-07-27T18:21:29.026213Z","shell.execute_reply.started":"2022-07-27T18:21:28.913836Z","shell.execute_reply":"2022-07-27T18:21:29.025240Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## <span style=\"color:Orange\">Decision Tree with Grid Search</span>","metadata":{}},{"cell_type":"code","source":"from sklearn.model_selection import GridSearchCV\ndef ModelGS(model, param_grid, model_name):\n    my_pipeline = Pipeline(steps=[('preprocessor', preprocessor), ('model', model)])\n    GridCV = GridSearchCV(my_pipeline, param_grid, n_jobs= -1, verbose=1)\n    GridCV.fit(X_train,y_train)  \n    print(GridCV.best_params_)    \n    y_val_pred = GridCV.predict(X_val)\n    print(f\"The r2 score of the {model_name} model on the validation set: \", np.round(r2_score(y_val, y_val_pred),2))\n    y_train_pred= GridCV.predict(X_train)\n    print(f\"The r2 score of the {model_name} model on the train data:\", np.round(r2_score(y_train, y_train_pred),2))","metadata":{},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"from sklearn.model_selection import GridSearchCV\nparam_grid = {'model__max_depth': [1, 5, 7, 9, 11, 13, 15, 20],\n                 'model__min_samples_leaf': [1, 5, 10, 15, 20, 50, 100]}\nmodel = DecisionTreeRegressor(random_state=0)\nModelGS(model, param_grid,  \"decision tree\")","metadata":{},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<a id=\"Random-Forest\"></a> \n# <p style=\"background-color:coral;font-family:newtimeroman;font-size:150%;color:white;text-align:center;border-radius:20px 20px;\"><b>Random Forest</b></p>\n\n<a id=\"0\"></a>\n<a href=\"#toc\" class=\"btn btn-primary btn-sm\" role=\"button\" aria-pressed=\"true\" \nstyle=\"color:blue; background-color:#dfa8e4\" data-toggle=\"popover\">Content</a>","metadata":{}},{"cell_type":"markdown","source":"Random forest is a supervised machine learning algorithm that is used widely in classification and regression problems. You can think of a random forest as an ensemble of decision trees. The decision tree models tend to overfit the training data. You can overcome the overfitting problem using random forest.","metadata":{}},{"cell_type":"markdown","source":"## <span style=\"color:Orange\">Random Forest with Grid Search</span>","metadata":{}},{"cell_type":"code","source":"from sklearn.ensemble import RandomForestRegressor\nparam_grid = {'model__max_depth':[20,30,40],\n                 'model__n_estimators':[200,250,300],\n                 'model__min_samples_leaf':[1,2,3]}\nmodel = RandomForestRegressor(random_state=0)\nModelGS(model, param_grid, \"random forest\")","metadata":{},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## <span style=\"color:Orange\">Random Forest with RandomizedSearchCV</span>","metadata":{}},{"cell_type":"code","source":"from sklearn.model_selection import RandomizedSearchCV\n\ndef ModelRS(model, param_grid, model_name):\n    my_pipeline = Pipeline(steps=[('preprocessor', preprocessor), ('model', model)])\n    RanCV = RandomizedSearchCV(my_pipeline, param_grid, verbose=True)\n    RanCV.fit(X_train, y_train)\n    print(RanCV.best_params_)    \n    y_val_pred = RanCV.predict(X_val)\n    print(f\"The r2 score of the {model_name} model on the validation set: \", np.round(r2_score(y_val, y_val_pred),2))\n    y_train_pred= RanCV.predict(X_train)\n    print(f\"The r2 score of the {model_name} model on the train data:\", np.round(r2_score(y_train, y_train_pred),2))\n    \nparam_grid = {'model__max_depth':[20,30,40],\n              'model__n_estimators':[200,250,300],\n              'model__min_samples_leaf':[1,2,3]}\nmodel = RandomForestRegressor(random_state=0)\nModelRS(model, param_grid, \"random forest\")","metadata":{},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<a id=\"XGBoost\"></a> \n# <p style=\"background-color:coral;font-family:newtimeroman;font-size:150%;color:white;text-align:center;border-radius:20px 20px;\"><b>XGBoost</b></p>\n\n<a id=\"0\"></a>\n<a href=\"#toc\" class=\"btn btn-primary btn-sm\" role=\"button\" aria-pressed=\"true\" \nstyle=\"color:blue; background-color:#dfa8e4\" data-toggle=\"popover\">Content</a>","metadata":{}},{"cell_type":"markdown","source":"XGBoost is short for Extreme Gradient Boosting. You can use the XGBoost to implement gradient boosting. The key idea behind gradient boosting is to use gradient descent to minimize the errors of the residuals. XGBoost provides a parallel tree boosting that solves many data science problems in a fast and accurate way. Scikit-learn has another version of gradient boosting, but XGBoost has some technical advantages.","metadata":{}},{"cell_type":"markdown","source":"## <span style=\"color:Orange\">XGBoost with Grid Search</span>","metadata":{}},{"cell_type":"code","source":"from xgboost import XGBRegressor\nparam_grid = {'model__n_estimators':[110, 120, 130],\n              'model__max_depth':[3, 4, 5],\n              'model__learning_rate':[0.05, 0.1, 0.2],\n              'model__min_child_weight':[1, 2, 3],\n              'model__subsample':[0.8, 0.9, 1]}\nmodel = XGBRegressor(random_state=0)\nModelGS(model, param_grid, \"XGBoost\")","metadata":{},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## <span style=\"color:Orange\">XGBoost with RandomizedSearchCV</span>","metadata":{}},{"cell_type":"code","source":"param_grid = {'model__n_estimators':[110, 120, 130],\n              'model__max_depth':[3, 4, 5],\n              'model__learning_rate':[0.05, 0.1, 0.2],\n              'model__min_child_weight':[1, 2, 3],\n              'model__subsample':[0.8, 0.9, 1],\n              'model__colsample_bytree':[0.7, 0.8, 0.9],\n              'model__colsample_bylevel':[0.8, 0.9, 1],\n              'model__colsample_bynode':[0.6, 0.7, 0.8, 0.9]}\nmodel = XGBRegressor(random_state=0)\nModelRS(model, param_grid, \"XGBoost\")","metadata":{"scrolled":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## <span style=\"color:Orange\">Submission (Optional) </span>","metadata":{}},{"cell_type":"code","source":"params = {'subsample': 0.8, 'n_estimators': 120, \n          'min_child_weight': 2, 'max_depth': 4, \n          'learning_rate': 0.05, 'colsample_bytree': 0.9, \n          'colsample_bynode': 0.6, 'colsample_bylevel': 1, 'random_state':0}\nmodel = XGBRegressor(**params)\nmy_pipeline = Pipeline(steps=[('preprocessor', preprocessor), ('model', model)])\nmy_pipeline.fit(X_train,y_train)\npreds_test = my_pipeline.predict(X_test)\noutput = pd.DataFrame({'Id': X_test.index, 'SalePrice': preds_test})\nprint(output.head())\noutput = pd.DataFrame({'Id': X_test.index,'SalePrice': preds_test})\noutput.to_csv('submission.csv', index=False)","metadata":{"execution":{"iopub.status.busy":"2022-07-27T18:24:55.755930Z","iopub.execute_input":"2022-07-27T18:24:55.756294Z","iopub.status.idle":"2022-07-27T18:24:56.298169Z","shell.execute_reply.started":"2022-07-27T18:24:55.756264Z","shell.execute_reply":"2022-07-27T18:24:56.297405Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<a id=\"Conclusion\"></a> \n# <p style=\"background-color:coral;font-family:newtimeroman;font-size:150%;color:white;text-align:center;border-radius:20px 20px;\"><b>Conclusion</b></p>\n\n\n<a id=\"0\"></a>\n<a href=\"#toc\" class=\"btn btn-primary btn-sm\" role=\"button\" aria-pressed=\"true\" \nstyle=\"color:blue; background-color:#dfa8e4\" data-toggle=\"popover\">Content</a>","metadata":{}},{"cell_type":"markdown","source":"Tree-based models are very often used in machine learning projects. You can use these models for both your regression and classification problems. In this notebook, I showed how to implement the decision trees, random forest, and XGBoost algorithms using the housing prices dataset. It turned out that XGBoost outperformed other techniques for the housing price dataset. This is not surprising because XGBoost is the most used model for standard tabular data and has won most competitions on Kaggle.","metadata":{}},{"cell_type":"markdown","source":"<b>Thanks for reading 😀 If you like this notebook, please upvote it 😊<b/>\n\n\n<b> Don't forget to follow us on [YouTube](http://youtube.com/tirendazacademy) | [Medium](http://tirendazacademy.medium.com) | [Twitter](http://twitter.com/tirendazacademy) | [GitHub](http://github.com/tirendazacademy) | [Linkedin](https://www.linkedin.com/in/tirendaz-academy) | [Kaggle](https://www.kaggle.com/tirendazacademy)<b/>","metadata":{}}]}