{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"pygments_lexer":"ipython3","nbconvert_exporter":"python","version":"3.6.4","file_extension":".py","codemirror_mode":{"name":"ipython","version":3},"name":"python","mimetype":"text/x-python"}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"# Using FBProphet to forecast the total price \nBased on [Sales Forecasting Using Facebook’s Prophet\n](https://medium.com/fritzheartbeat/sales-forecasting-using-facebooks-prophet-f9ae0214f196)","metadata":{}},{"cell_type":"markdown","source":"Sales forecasting is one the most common tasks in many sales-driven organizations. When done well, it enables organizations to adequately plan for the future with a degree of confidence. In this tutorial, we’ll use Prophet, a package developed by Facebook to show how one can achieve this. This package is available in both Python and R. We assume that the reader has a basic understanding of handling time series data in Python.","metadata":{}},{"cell_type":"code","source":"# This Python 3 environment comes with many helpful analytics libraries installed\n# It is defined by the kaggle/python Docker image: https://github.com/kaggle/docker-python\n# For example, here's several helpful packages to load\n\nimport numpy as np # linear algebra\nimport pandas as pd # data processing, CSV file I/O (e.g. pd.read_csv)\n","metadata":{"_uuid":"8f2839f25d086af736a60e9eeb907d3b93b6e0e5","_cell_guid":"b1076dfc-b9ad-4769-8c92-a6c4dae69d19","execution":{"iopub.status.busy":"2022-02-08T11:48:51.466452Z","iopub.execute_input":"2022-02-08T11:48:51.466859Z","iopub.status.idle":"2022-02-08T11:48:51.479319Z","shell.execute_reply.started":"2022-02-08T11:48:51.466753Z","shell.execute_reply":"2022-02-08T11:48:51.477591Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"transactions_train = pd.read_csv(\"/kaggle/input/h-and-m-personalized-fashion-recommendations/transactions_train.csv\")","metadata":{"execution":{"iopub.status.busy":"2022-02-08T11:49:06.343557Z","iopub.execute_input":"2022-02-08T11:49:06.343915Z","iopub.status.idle":"2022-02-08T11:50:15.941522Z","shell.execute_reply.started":"2022-02-08T11:49:06.34388Z","shell.execute_reply":"2022-02-08T11:50:15.939994Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"transactions_train.tail()","metadata":{"execution":{"iopub.status.busy":"2022-02-08T11:50:26.896372Z","iopub.execute_input":"2022-02-08T11:50:26.896792Z","iopub.status.idle":"2022-02-08T11:50:26.912421Z","shell.execute_reply.started":"2022-02-08T11:50:26.896715Z","shell.execute_reply":"2022-02-08T11:50:26.911124Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"transactions = transactions_train.groupby('t_dat')['price'].sum().reset_index()","metadata":{"execution":{"iopub.status.busy":"2022-02-08T11:54:08.761523Z","iopub.execute_input":"2022-02-08T11:54:08.761878Z","iopub.status.idle":"2022-02-08T11:54:11.7703Z","shell.execute_reply.started":"2022-02-08T11:54:08.761843Z","shell.execute_reply":"2022-02-08T11:54:11.769147Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"transactions['ds'] = transactions['t_dat']\ntransactions['y'] = transactions['price']\ntransactions = transactions[['ds','y']]","metadata":{"execution":{"iopub.status.busy":"2022-02-08T11:56:43.349927Z","iopub.execute_input":"2022-02-08T11:56:43.350528Z","iopub.status.idle":"2022-02-08T11:56:43.366086Z","shell.execute_reply.started":"2022-02-08T11:56:43.350479Z","shell.execute_reply":"2022-02-08T11:56:43.364681Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"transactions.tail()","metadata":{"execution":{"iopub.status.busy":"2022-02-08T12:02:09.6304Z","iopub.execute_input":"2022-02-08T12:02:09.630876Z","iopub.status.idle":"2022-02-08T12:02:09.646105Z","shell.execute_reply.started":"2022-02-08T12:02:09.630837Z","shell.execute_reply":"2022-02-08T12:02:09.64504Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Prophet works best with hourly and weekly data over several months. When working with Prophet, yearly data is most preferred","metadata":{}},{"cell_type":"code","source":"from fbprophet import Prophet","metadata":{"execution":{"iopub.status.busy":"2022-02-08T11:57:11.480665Z","iopub.execute_input":"2022-02-08T11:57:11.480983Z","iopub.status.idle":"2022-02-08T11:57:11.984992Z","shell.execute_reply.started":"2022-02-08T11:57:11.480948Z","shell.execute_reply":"2022-02-08T11:57:11.984077Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Start by creating an instance of the Prophet class and then fit it to our dataset.","metadata":{}},{"cell_type":"code","source":"model = Prophet()\nmodel.add_country_holidays(country_name='US')\nmodel.fit(transactions)","metadata":{"execution":{"iopub.status.busy":"2022-02-08T11:58:38.054404Z","iopub.execute_input":"2022-02-08T11:58:38.054726Z","iopub.status.idle":"2022-02-08T11:58:38.485586Z","shell.execute_reply.started":"2022-02-08T11:58:38.054694Z","shell.execute_reply":"2022-02-08T11:58:38.484374Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### Making Future Predictions\n","metadata":{}},{"cell_type":"markdown","source":"The next step is to prepare our model to make future predictions. This is achieved using the `Prophet.make_future_dataframe` method and passing the number of days we’d like to predict in the future. We use the periods attribute to specify this. This also include the historical dates. We’ll use these historical dates to compare the predictions with the actual values in the `ds` column.","metadata":{}},{"cell_type":"code","source":"future = model.make_future_dataframe(periods=7)\nfuture.tail()","metadata":{"execution":{"iopub.status.busy":"2022-02-08T11:58:39.565299Z","iopub.execute_input":"2022-02-08T11:58:39.56659Z","iopub.status.idle":"2022-02-08T11:58:39.58734Z","shell.execute_reply.started":"2022-02-08T11:58:39.566519Z","shell.execute_reply":"2022-02-08T11:58:39.585476Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### Obtaining the Forecasts","metadata":{}},{"cell_type":"markdown","source":"We use the `predict` method to make future predictions. This will generate a dataframe with a `yhat` column that will contain the predictions.","metadata":{}},{"cell_type":"code","source":"forecast = model.predict(future)","metadata":{"execution":{"iopub.status.busy":"2022-02-08T11:58:55.004326Z","iopub.execute_input":"2022-02-08T11:58:55.004633Z","iopub.status.idle":"2022-02-08T11:58:59.328535Z","shell.execute_reply.started":"2022-02-08T11:58:55.00459Z","shell.execute_reply":"2022-02-08T11:58:59.327239Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"If we check the head for our forecast dataframe, we’ll notice that it has a lot of columns. However, we are mainly interested in `ds`, `yhat`, `yhat_lower`, and `yhat_upper`. `yhat` is our predicted forecast, `yhat_lower` is the lower bound for our predictions, and `yhat_upper` is the upper bound.","metadata":{}},{"cell_type":"code","source":"forecast.sample()","metadata":{"execution":{"iopub.status.busy":"2022-02-08T11:59:04.646713Z","iopub.execute_input":"2022-02-08T11:59:04.646998Z","iopub.status.idle":"2022-02-08T11:59:04.677098Z","shell.execute_reply.started":"2022-02-08T11:59:04.64697Z","shell.execute_reply":"2022-02-08T11:59:04.676098Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"forecast[['ds', 'yhat', 'yhat_lower', 'yhat_upper']].tail()\n","metadata":{"execution":{"iopub.status.busy":"2022-02-08T11:59:16.675765Z","iopub.execute_input":"2022-02-08T11:59:16.676442Z","iopub.status.idle":"2022-02-08T11:59:16.691628Z","shell.execute_reply.started":"2022-02-08T11:59:16.6764Z","shell.execute_reply":"2022-02-08T11:59:16.690628Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### Plotting the Forecasts\n","metadata":{}},{"cell_type":"markdown","source":"Prophet has an inbuilt feature that enables us to plot the forecasts we just generated. This is achieved using `model.plot()` and passing in our forecasts as the argument. The blue line in the graph represents the predicted values while the black dots represents the data in our dataset.","metadata":{}},{"cell_type":"code","source":"plot = model.plot(forecast)\n# The blue line in the graph represents the predicted values while the black dots \n# represents the data in our dataset.","metadata":{"execution":{"iopub.status.busy":"2022-02-08T11:59:26.391146Z","iopub.execute_input":"2022-02-08T11:59:26.391502Z","iopub.status.idle":"2022-02-08T11:59:26.832701Z","shell.execute_reply.started":"2022-02-08T11:59:26.391465Z","shell.execute_reply":"2022-02-08T11:59:26.831714Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"#### Plotting the Forecast Components","metadata":{}},{"cell_type":"markdown","source":"The `plot_components` method plots the trend, yearly, and weekly seasonality of the time series data.","metadata":{}},{"cell_type":"code","source":"plot2 = model.plot_components(forecast)","metadata":{"execution":{"iopub.status.busy":"2022-02-08T11:59:41.730366Z","iopub.execute_input":"2022-02-08T11:59:41.731015Z","iopub.status.idle":"2022-02-08T11:59:42.825618Z","shell.execute_reply.started":"2022-02-08T11:59:41.730977Z","shell.execute_reply":"2022-02-08T11:59:42.824464Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"customers = pd.read_csv(\"/kaggle/input/h-and-m-personalized-fashion-recommendations/customers.csv\")","metadata":{"execution":{"iopub.status.busy":"2022-02-08T11:51:09.900315Z","iopub.execute_input":"2022-02-08T11:51:09.900664Z","iopub.status.idle":"2022-02-08T11:51:16.386621Z","shell.execute_reply.started":"2022-02-08T11:51:09.900631Z","shell.execute_reply":"2022-02-08T11:51:16.385285Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### Cross Validation\n","metadata":{}},{"cell_type":"markdown","source":"Next let’s measure the forecast error using the historical data. We’ll do this by comparing the predicted values with the actual values. In order to perform this operation, we select cut off points in the data history and fit the model with data up to that cut off point.\nAfterwards, we compare the actual values to the predicted values. The `cross_validation` method allows us to do this in Prophet. This method takes the following parameters, as explained below:\n\n- horizon the forecast horizon\n\n- initial the size of the initial training period\n\n- period the spacing between cutoff dates\n\nThe output of the cross_validation method is a dataframe containing `y` (the true values) and `yhat` (the predicted values). We’ll use this dataframe to compute the prediction errors.","metadata":{}},{"cell_type":"markdown","source":"### Time series cross validation to measure forecast error using historical data.\nSelect a cut off points in the past\nFit the model to the data up to that cut off point\nCompare the forecasted values to the actual values.","metadata":{}},{"cell_type":"code","source":"from fbprophet.diagnostics import cross_validation #measure forecast error using historical data\n# This is done by selecting cutoff points in the history, and for each of them fitting the model using \n# data only up to that cutoff point. We can then compare the forecasted values to the actual values\ndf_cv = cross_validation(model,horizon = '50 days') #  forecast horizon\n# By default, the initial training period is set to three times the horizon\n# cutoffs are made every half a horizon\ndf_cv.head()","metadata":{"execution":{"iopub.status.busy":"2022-02-08T12:00:23.870067Z","iopub.execute_input":"2022-02-08T12:00:23.870372Z","iopub.status.idle":"2022-02-08T12:00:53.378517Z","shell.execute_reply.started":"2022-02-08T12:00:23.870342Z","shell.execute_reply":"2022-02-08T12:00:53.377465Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"from fbprophet.diagnostics import performance_metrics\ndf_p = performance_metrics(df_cv)\ndf_p.tail()","metadata":{"execution":{"iopub.status.busy":"2022-02-08T12:00:55.795421Z","iopub.execute_input":"2022-02-08T12:00:55.795744Z","iopub.status.idle":"2022-02-08T12:00:55.961394Z","shell.execute_reply.started":"2022-02-08T12:00:55.795714Z","shell.execute_reply":"2022-02-08T12:00:55.96024Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"from fbprophet.plot import plot_cross_validation_metric\nfig = plot_cross_validation_metric(df_cv, metric='rmse')","metadata":{"execution":{"iopub.status.busy":"2022-02-08T12:01:06.639662Z","iopub.execute_input":"2022-02-08T12:01:06.639977Z","iopub.status.idle":"2022-02-08T12:01:07.176873Z","shell.execute_reply.started":"2022-02-08T12:01:06.639945Z","shell.execute_reply":"2022-02-08T12:01:07.175019Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### Obtaining the Performance Metrics\nWe use the `performance_metrics` utility to compute the Mean Squared Error(MSE), Root Mean Squared Error(RMSE), Mean Absolute Error(MAE), Mean Absolute Percentage Error(MAPE) and the coverage of the `yhat_lower` and `yhat_upper` estimates.","metadata":{}},{"cell_type":"code","source":"from fbprophet.diagnostics import performance_metrics\ndf_p = performance_metrics(df_cv)\ndf_p.head()","metadata":{"execution":{"iopub.status.busy":"2022-02-08T12:19:56.228735Z","iopub.execute_input":"2022-02-08T12:19:56.229808Z","iopub.status.idle":"2022-02-08T12:19:56.427053Z","shell.execute_reply.started":"2022-02-08T12:19:56.229744Z","shell.execute_reply":"2022-02-08T12:19:56.426106Z"},"trusted":true},"execution_count":null,"outputs":[]}]}