{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"pygments_lexer":"ipython3","nbconvert_exporter":"python","version":"3.6.4","file_extension":".py","codemirror_mode":{"name":"ipython","version":3},"name":"python","mimetype":"text/x-python"}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"**This notebook is an exercise in the [Time Series](https://www.kaggle.com/learn/time-series) course.  You can reference the tutorial at [this link](https://www.kaggle.com/ryanholbrook/hybrid-models).**\n\n---\n","metadata":{}},{"cell_type":"markdown","source":"# Introduction #\n\nRun this cell to set everything up!","metadata":{}},{"cell_type":"code","source":"# Setup feedback system\nfrom learntools.core import binder\nbinder.bind(globals())\nfrom learntools.time_series.ex5 import *\n\n# Setup notebook\nfrom pathlib import Path\nfrom learntools.time_series.style import *  # plot style settings\n\nimport matplotlib.pyplot as plt\nimport pandas as pd\nfrom sklearn.linear_model import LinearRegression\nfrom sklearn.preprocessing import LabelEncoder\nfrom statsmodels.tsa.deterministic import DeterministicProcess\nfrom xgboost import XGBRegressor\n\n\ncomp_dir = Path('../input/store-sales-time-series-forecasting')\ndata_dir = Path(\"../input/ts-course-data\")\n\nstore_sales = pd.read_csv(\n    comp_dir / 'train.csv',\n    usecols=['store_nbr', 'family', 'date', 'sales', 'onpromotion'],\n    dtype={\n        'store_nbr': 'category',\n        'family': 'category',\n        'sales': 'float32',\n    },\n    parse_dates=['date'],\n    infer_datetime_format=True,\n)\nstore_sales['date'] = store_sales.date.dt.to_period('D')\nstore_sales = store_sales.set_index(['store_nbr', 'family', 'date']).sort_index()\n\nfamily_sales = (\n    store_sales\n    .groupby(['family', 'date'])\n    .mean()\n    .unstack('family')\n    .loc['2017']\n)","metadata":{"execution":{"iopub.status.busy":"2022-07-29T13:27:37.027303Z","iopub.execute_input":"2022-07-29T13:27:37.028012Z","iopub.status.idle":"2022-07-29T13:27:44.857465Z","shell.execute_reply.started":"2022-07-29T13:27:37.027916Z","shell.execute_reply":"2022-07-29T13:27:44.856147Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"-------------------------------------------------------------------------------\n\nIn the next two questions, you'll create a boosted hybrid for the *Store Sales* dataset by implementing a new Python class. Run this cell to create the initial class definition. You'll add `fit` and `predict` methods to give it a scikit-learn like interface.\n","metadata":{}},{"cell_type":"code","source":"# You'll add fit and predict methods to this minimal class\nclass BoostedHybrid:\n    def __init__(self, model_1, model_2):\n        self.model_1 = model_1\n        self.model_2 = model_2\n        self.y_columns = None  # store column names from fit method\n","metadata":{"execution":{"iopub.status.busy":"2022-07-29T13:53:26.180036Z","iopub.execute_input":"2022-07-29T13:53:26.180617Z","iopub.status.idle":"2022-07-29T13:53:26.186589Z","shell.execute_reply.started":"2022-07-29T13:53:26.180577Z","shell.execute_reply":"2022-07-29T13:53:26.185607Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# 1) Define fit method for boosted hybrid\n\nComplete the `fit` definition for the `BoostedHybrid` class. Refer back to steps 1 and 2 from the **Hybrid Forecasting with Residuals** section in the tutorial if you need.","metadata":{}},{"cell_type":"code","source":"def fit(self, X_1, X_2, y):\n    # YOUR CODE HERE: fit self.model_1\n    self.model_1.fit(X_1, y)\n\n    y_fit = pd.DataFrame(\n        # YOUR CODE HERE: make predictions with self.model_1\n         self.model_1.predict(X_1),\n        index=X_1.index, columns=y.columns,\n    )\n\n    # YOUR CODE HERE: compute residuals\n    y_resid = y - y_fit\n    y_resid = y_resid.stack().squeeze() # wide to long\n\n    # YOUR CODE HERE: fit self.model_2 on residuals\n    self.model_2.fit(X_2, y_resid)\n\n    # Save column names for predict method\n    self.y_columns = y.columns\n    # Save data for question checking\n    self.y_fit = y_fit\n    self.y_resid = y_resid\n\n\n# Add method to class\nBoostedHybrid.fit = fit\n\n\n# Check your answer\nq_1.check()","metadata":{"execution":{"iopub.status.busy":"2022-07-29T13:54:22.492626Z","iopub.execute_input":"2022-07-29T13:54:22.493053Z","iopub.status.idle":"2022-07-29T13:54:22.517471Z","shell.execute_reply.started":"2022-07-29T13:54:22.493019Z","shell.execute_reply":"2022-07-29T13:54:22.516541Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Lines below will give you a hint or solution code\n#q_1.hint()\n# q_1.solution()","metadata":{"execution":{"iopub.status.busy":"2022-07-29T13:54:26.336258Z","iopub.execute_input":"2022-07-29T13:54:26.336696Z","iopub.status.idle":"2022-07-29T13:54:26.341962Z","shell.execute_reply.started":"2022-07-29T13:54:26.336659Z","shell.execute_reply":"2022-07-29T13:54:26.341050Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"-------------------------------------------------------------------------------\n\n# 2) Define predict method for boosted hybrid\n\nNow define the `predict` method for the `BoostedHybrid` class. Refer back to step 3 from the **Hybrid Forecasting with Residuals** section in the tutorial if you need.","metadata":{}},{"cell_type":"code","source":"def predict(self, X_1, X_2):\n    y_pred = pd.DataFrame(\n        # YOUR CODE HERE: predict with self.model_1\n        self.model_1.predict(X_1), \n        index=X_1.index, columns=self.y_columns,\n    )\n    y_pred = y_pred.stack().squeeze()  # wide to long\n\n    # YOUR CODE HERE: add self.model_2 predictions to y_pred\n    y_pred +=  self.model_2.predict(X_2)\n    \n    return y_pred.unstack()  # long to wide\n\n\n# Add method to class\nBoostedHybrid.predict = predict\n\n\n# Check your answer\nq_2.check()","metadata":{"execution":{"iopub.status.busy":"2022-07-29T13:54:56.641012Z","iopub.execute_input":"2022-07-29T13:54:56.641445Z","iopub.status.idle":"2022-07-29T13:54:56.682963Z","shell.execute_reply.started":"2022-07-29T13:54:56.641409Z","shell.execute_reply":"2022-07-29T13:54:56.681771Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Lines below will give you a hint or solution code\n#q_2.hint()\n# q_2.solution()","metadata":{"execution":{"iopub.status.busy":"2022-07-29T13:55:00.446444Z","iopub.execute_input":"2022-07-29T13:55:00.446822Z","iopub.status.idle":"2022-07-29T13:55:00.452811Z","shell.execute_reply.started":"2022-07-29T13:55:00.446793Z","shell.execute_reply":"2022-07-29T13:55:00.451245Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"-------------------------------------------------------------------------------\n\nNow you're ready to use your new `BoostedHybrid` class to create a model for the *Store Sales* data. Run the next cell to set up the data for training.","metadata":{}},{"cell_type":"code","source":"# Target series\ny = family_sales.loc[:, 'sales']\n\n\n# X_1: Features for Linear Regression\ndp = DeterministicProcess(index=y.index, order=1)\nX_1 = dp.in_sample()\n\n\n# X_2: Features for XGBoost\nX_2 = family_sales.drop('sales', axis=1).stack()  # onpromotion feature\n\n# Label encoding for 'family'\nle = LabelEncoder()  # from sklearn.preprocessing\nX_2 = X_2.reset_index('family')\nX_2['family'] = le.fit_transform(X_2['family'])\n\n# Label encoding for seasonality\nX_2[\"day\"] = X_2.index.day  # values are day of the month","metadata":{"execution":{"iopub.status.busy":"2022-07-29T13:55:03.914677Z","iopub.execute_input":"2022-07-29T13:55:03.915110Z","iopub.status.idle":"2022-07-29T13:55:03.941619Z","shell.execute_reply.started":"2022-07-29T13:55:03.915067Z","shell.execute_reply":"2022-07-29T13:55:03.940446Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# 3) Train boosted hybrid\n\nCreate the hybrid model by initializing a `BoostedHybrid` class with `LinearRegression()` and `XGBRegressor()` instances.","metadata":{}},{"cell_type":"code","source":"# YOUR CODE HERE: Create LinearRegression + XGBRegressor hybrid with BoostedHybrid\nmodel = BoostedHybrid(\n    model_1=LinearRegression(),\n    model_2=XGBRegressor(),\n)\nmodel.fit(X_1, X_2, y)\n\n# YOUR CODE HERE: Fit and predict\ny_pred = model.predict(X_1, X_2)\ny_pred = y_pred.clip(0.0)\n\n# Check your answer\nq_3.check()","metadata":{"execution":{"iopub.status.busy":"2022-07-29T13:57:10.549476Z","iopub.execute_input":"2022-07-29T13:57:10.549896Z","iopub.status.idle":"2022-07-29T13:57:11.141196Z","shell.execute_reply.started":"2022-07-29T13:57:10.549865Z","shell.execute_reply":"2022-07-29T13:57:11.140168Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Lines below will give you a hint or solution code\n#q_3.hint()\n# q_3.solution()","metadata":{"execution":{"iopub.status.busy":"2022-07-29T13:57:06.602570Z","iopub.execute_input":"2022-07-29T13:57:06.602950Z","iopub.status.idle":"2022-07-29T13:57:06.607737Z","shell.execute_reply.started":"2022-07-29T13:57:06.602916Z","shell.execute_reply":"2022-07-29T13:57:06.606446Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"-------------------------------------------------------------------------------\n\nDepending on your problem, you might want to use other hybrid combinations than the linear regression + XGBoost hybrid you've created in the previous questions. Run the next cell to try other algorithms from scikit-learn.","metadata":{}},{"cell_type":"code","source":"# Model 1 (trend)\nfrom pyearth import Earth\nfrom sklearn.linear_model import ElasticNet, Lasso, Ridge\n\n# Model 2\nfrom sklearn.ensemble import ExtraTreesRegressor, RandomForestRegressor\nfrom sklearn.neighbors import KNeighborsRegressor\nfrom sklearn.neural_network import MLPRegressor\n\n# Boosted Hybrid\n\n# YOUR CODE HERE: Try different combinations of the algorithms above\nmodel = BoostedHybrid(\n    model_1=Ridge(),\n    model_2=KNeighborsRegressor(),\n)","metadata":{"execution":{"iopub.status.busy":"2022-07-29T13:57:15.026883Z","iopub.execute_input":"2022-07-29T13:57:15.027256Z","iopub.status.idle":"2022-07-29T13:57:15.177026Z","shell.execute_reply.started":"2022-07-29T13:57:15.027225Z","shell.execute_reply":"2022-07-29T13:57:15.175759Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"These are just some suggestions. You might discover other algorithms you like in the scikit-learn [User Guide](https://scikit-learn.org/stable/supervised_learning.html).\n\nUse the code in this cell to see the predictions your hybrid makes.","metadata":{}},{"cell_type":"code","source":"y_train, y_valid = y[:\"2017-07-01\"], y[\"2017-07-02\":]\nX1_train, X1_valid = X_1[: \"2017-07-01\"], X_1[\"2017-07-02\" :]\nX2_train, X2_valid = X_2.loc[:\"2017-07-01\"], X_2.loc[\"2017-07-02\":]\n\n# Some of the algorithms above do best with certain kinds of\n# preprocessing on the features (like standardization), but this is\n# just a demo.\nmodel.fit(X1_train, X2_train, y_train)\ny_fit = model.predict(X1_train, X2_train).clip(0.0)\ny_pred = model.predict(X1_valid, X2_valid).clip(0.0)\n\nfamilies = y.columns[0:6]\naxs = y.loc(axis=1)[families].plot(\n    subplots=True, sharex=True, figsize=(11, 9), **plot_params, alpha=0.5,\n)\n_ = y_fit.loc(axis=1)[families].plot(subplots=True, sharex=True, color='C0', ax=axs)\n_ = y_pred.loc(axis=1)[families].plot(subplots=True, sharex=True, color='C3', ax=axs)\nfor ax, family in zip(axs, families):\n    ax.legend([])\n    ax.set_ylabel(family)","metadata":{"execution":{"iopub.status.busy":"2022-07-29T13:57:32.875251Z","iopub.execute_input":"2022-07-29T13:57:32.875683Z","iopub.status.idle":"2022-07-29T13:57:34.616607Z","shell.execute_reply.started":"2022-07-29T13:57:32.875650Z","shell.execute_reply":"2022-07-29T13:57:34.615363Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# 4) Fit with different learning algorithms\n\nOnce you're ready to move on, run the next cell for credit on this question.","metadata":{}},{"cell_type":"code","source":"# View the solution (Run this cell to receive credit!)\nq_4.check()","metadata":{"lines_to_next_cell":2,"execution":{"iopub.status.busy":"2022-07-29T13:57:56.734386Z","iopub.execute_input":"2022-07-29T13:57:56.734767Z","iopub.status.idle":"2022-07-29T13:57:56.743336Z","shell.execute_reply.started":"2022-07-29T13:57:56.734737Z","shell.execute_reply":"2022-07-29T13:57:56.742464Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Keep Going #\n\n[**Convert any forecasting task**](https://www.kaggle.com/ryanholbrook/forecasting-with-machine-learning) to a machine learning problem with four ML forecasting strategies.","metadata":{}},{"cell_type":"markdown","source":"---\n\n\n\n\n*Have questions or comments? Visit the [course discussion forum](https://www.kaggle.com/learn/time-series/discussion) to chat with other learners.*","metadata":{}}]}