{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"pygments_lexer":"ipython3","nbconvert_exporter":"python","version":"3.6.4","file_extension":".py","codemirror_mode":{"name":"ipython","version":3},"name":"python","mimetype":"text/x-python"}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"# **<span style = 'color:blue'>Multivariate FBProphet Smart Homes Indoor Temperature Forecasting</span>**\n## **<span style='color:green'> Table of Contents:</span>**<a id=\"Table\"></a>\n\n* [1. Import the libraries](#Import)\n* [2. Load the data](#Load)\n* [3. Data Wrangling](#Wrangling)\n* [4. Statistical Data Analysis](#Statistical)\n* [5. Multivariate FBProphet](#FBProphet)\n    - [5.1 Data preprocessing for FBProphet model](#Preprocessing)\n    - [5.2 FBProphet model training](#Training)\n    - [5.3 Model Evaluation](#Evaluation)\n    - [5.4 Test Set Prediction & Kaggle Submission](#Prediction)","metadata":{"_uuid":"8f2839f25d086af736a60e9eeb907d3b93b6e0e5","_cell_guid":"b1076dfc-b9ad-4769-8c92-a6c4dae69d19"}},{"cell_type":"markdown","source":"## **<span style = 'color:green'>1. Import the required libraries</span>**<a id =\"Import\"></a>","metadata":{}},{"cell_type":"code","source":"# Import data handling & numerical libraries\nimport pandas as pd\nimport numpy as np\nfrom copy import copy\nimport datetime\n\n# Import Data Visualization libraries\nimport seaborn as sb\nimport matplotlib.pyplot as plt\n\n#import libraries for muting unnecessary warnings if needed\nimport warnings\nwarnings.filterwarnings('ignore')","metadata":{"execution":{"iopub.status.busy":"2022-07-06T16:52:25.466310Z","iopub.execute_input":"2022-07-06T16:52:25.466816Z","iopub.status.idle":"2022-07-06T16:52:26.605396Z","shell.execute_reply.started":"2022-07-06T16:52:25.466715Z","shell.execute_reply":"2022-07-06T16:52:26.604196Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## **<span style = 'color:green'>2. Load the dataset</span>**<a id ='Load'></a>\n[<div style=\"text-align: right\"> Back to Table of contents</div>](#Table)\n\nThe dataset collected from the monitor system mounted in a solar house corresponds to approximately 40 days of monitoring data. The Goal is to predict indoor temperature of a room (the Bedroom), in order to choose whether or not to activate the HVAC (Heating, Ventilation, and Air Conditioning) system. The data was sampled every minute, computing and uploading it smoothed with 15 minute means. The dataset includes dates, other sensor measurements, weather measurements and other information. It is a multivariate time-series dataset.\n\nIt has been established that the power consumption attributed to HVAC accounts for 53.9 % of total consumption, and the energy required to maintain the temperature is less than that required to drop or raise it.\n\nAs a result, a predictive model capable of predicting a room's indoor temperature (a short-term forecast of indoor temperature) would help in lowering overall energy consumption, by deciding whether or not to activate the HVAC system, at the appropriate time.\n","metadata":{}},{"cell_type":"code","source":"df = pd.read_csv('../input/smart-homes-temperature-time-series-forecasting/train.csv')\ndf_test1 = pd.read_csv('../input/smart-homes-temperature-time-series-forecasting/test.csv')\ndf.head(2)","metadata":{"execution":{"iopub.status.busy":"2022-07-06T16:52:26.608091Z","iopub.execute_input":"2022-07-06T16:52:26.608626Z","iopub.status.idle":"2022-07-06T16:52:26.683265Z","shell.execute_reply.started":"2022-07-06T16:52:26.608576Z","shell.execute_reply":"2022-07-06T16:52:26.682250Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Define a function to plot the entire dataframe\n# The function performs data visualization\ndef display_plot(data,x, y, fig_title):\n    plt.figure(figsize = (20,7))\n    plt.title(fig_title, loc='center', fontsize=20)\n    sb.barplot(x = x, y = y, palette = 'cool') \n    plt.tight_layout();\n\n# Plot the data\ndisplay_plot(df, df['Time'], df['Indoor_temperature_room'], \"Indoor room temperature\")","metadata":{"execution":{"iopub.status.busy":"2022-07-06T16:52:26.684889Z","iopub.execute_input":"2022-07-06T16:52:26.686186Z","iopub.status.idle":"2022-07-06T16:52:33.091442Z","shell.execute_reply.started":"2022-07-06T16:52:26.686131Z","shell.execute_reply":"2022-07-06T16:52:33.090181Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"display_plot(df, df['Date'], df['Indoor_temperature_room'], \"Indoor room temperature\")","metadata":{"execution":{"iopub.status.busy":"2022-07-06T16:52:33.094900Z","iopub.execute_input":"2022-07-06T16:52:33.095622Z","iopub.status.idle":"2022-07-06T16:52:34.922299Z","shell.execute_reply.started":"2022-07-06T16:52:33.095577Z","shell.execute_reply":"2022-07-06T16:52:34.920981Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"df.info()","metadata":{"execution":{"iopub.status.busy":"2022-07-06T16:52:34.924242Z","iopub.execute_input":"2022-07-06T16:52:34.925035Z","iopub.status.idle":"2022-07-06T16:52:34.947431Z","shell.execute_reply.started":"2022-07-06T16:52:34.924985Z","shell.execute_reply":"2022-07-06T16:52:34.946492Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"df.shape","metadata":{"execution":{"iopub.status.busy":"2022-07-06T16:52:34.948794Z","iopub.execute_input":"2022-07-06T16:52:34.949347Z","iopub.status.idle":"2022-07-06T16:52:34.955972Z","shell.execute_reply.started":"2022-07-06T16:52:34.949299Z","shell.execute_reply":"2022-07-06T16:52:34.954741Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## **<span style = 'color:green'>3. Data Wrangling</span>**<a id ='Wrangling'></a> \n[<div style=\"text-align: right\"> Back to Table of contents</div>](#Table) \n### **<span style = 'color:brown'>Checking for Missing Values</span>**","metadata":{}},{"cell_type":"code","source":"df[df.columns[df.isnull().sum()>0]].isnull().sum()","metadata":{"execution":{"iopub.status.busy":"2022-07-06T16:52:34.958091Z","iopub.execute_input":"2022-07-06T16:52:34.959488Z","iopub.status.idle":"2022-07-06T16:52:34.972792Z","shell.execute_reply.started":"2022-07-06T16:52:34.959373Z","shell.execute_reply":"2022-07-06T16:52:34.971482Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### **<span style=\"font-family: Segoe UI; font-size:1.0em;color:brown;\">Checking for duplicates</span>**","metadata":{}},{"cell_type":"code","source":"df.duplicated().sum()","metadata":{"execution":{"iopub.status.busy":"2022-07-06T16:52:34.974162Z","iopub.execute_input":"2022-07-06T16:52:34.974497Z","iopub.status.idle":"2022-07-06T16:52:34.992828Z","shell.execute_reply.started":"2022-07-06T16:52:34.974464Z","shell.execute_reply":"2022-07-06T16:52:34.991822Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"#creating datetime list with boundaries of raw data series, hourly frequency\ndatelist = pd.date_range(datetime.datetime(2012,3,12,11,45,0), datetime.datetime(2012,4,11,6,30,0), freq='15min').tolist()\n\n#extracting raw data series indices\nidx_list = df.index.to_list()\n\n#checking for anomalies by comparing the two\nprint(idx_list == datelist)\n#searching for anomalies\nprint(\"\\n No. of elements in full list:\", len(datelist), \"\\n No. of indices:\", len(idx_list), \"\\n No. of elements in set of indices:\", len(set(idx_list)))","metadata":{"execution":{"iopub.status.busy":"2022-07-06T16:52:34.993942Z","iopub.execute_input":"2022-07-06T16:52:34.994504Z","iopub.status.idle":"2022-07-06T16:52:35.020613Z","shell.execute_reply.started":"2022-07-06T16:52:34.994469Z","shell.execute_reply":"2022-07-06T16:52:35.019192Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## **<span style='color:green'>4. Statistical Data Analysis</span>**<a id ='Statistical'></a>\n[<div style=\"text-align: right\"> Back to Table of contents</div>](#Table) ","metadata":{}},{"cell_type":"code","source":"df.describe()","metadata":{"execution":{"iopub.status.busy":"2022-07-06T16:52:35.025054Z","iopub.execute_input":"2022-07-06T16:52:35.025605Z","iopub.status.idle":"2022-07-06T16:52:35.101084Z","shell.execute_reply.started":"2022-07-06T16:52:35.025562Z","shell.execute_reply":"2022-07-06T16:52:35.100122Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### **<span style = 'color:brown'>Data Summary with Tableone</span>**","metadata":{}},{"cell_type":"code","source":"pip install tableone --quiet","metadata":{"execution":{"iopub.status.busy":"2022-07-06T16:52:35.102694Z","iopub.execute_input":"2022-07-06T16:52:35.103227Z","iopub.status.idle":"2022-07-06T16:52:50.138464Z","shell.execute_reply.started":"2022-07-06T16:52:35.103193Z","shell.execute_reply":"2022-07-06T16:52:50.136991Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Import tableone\nfrom tableone import TableOne, load_dataset\n\n# Create a simple Table 1 with no grouping variable\n# Test for normality, multimodality (Hartigan's Dip Test), and far outliers (Tukey's test)\n#Create an instance of TableOne with the input arguments:\ntable1 = TableOne(df.drop(['Date','Time'], axis = 1), dip_test=True, normal_test=True, tukey_test=True)\n# View table1 (note the remarks below the table)\ntable1","metadata":{"execution":{"iopub.status.busy":"2022-07-06T16:52:50.140774Z","iopub.execute_input":"2022-07-06T16:52:50.141974Z","iopub.status.idle":"2022-07-06T16:52:50.698678Z","shell.execute_reply.started":"2022-07-06T16:52:50.141918Z","shell.execute_reply":"2022-07-06T16:52:50.697229Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"#### **<span style = 'color:brown'>Exploring the warning raised by Hartigan's Dip Test</span>**\nHartigan's Dip Test is a test for multimodality. The test has suggested that the CO2_(dinning-room), CO2_room, Relative_humidity_(dinning-room), Relative_humidity_room, Lighting_(dinning-room), Lighting_room, Meteo_Rain, Meteo_Sun_dusk, Meteo_Wind, Meteo_Sun_light_in_west_facade, Meteo_Sun_light_in_east_facade, Meteo_Sun_light_in_south_facade, Meteo_Sun_irradiance and Day_of_the_week may be multimodal. A distribution with one peak is called unimodal A distribution with two peaks is called bimodal A distribution with two peaks or more is multimodal A bimodal distribution is also multimodal, as there are multiple peaks. We'll plot the distributions here.","metadata":{}},{"cell_type":"code","source":"# Numerical features\nnum_feats=[col for col in df_test1.columns if df_test1[col].dtypes != 'object' and col != 'Id']\n\n# Plot distribution of numerical columns\nfig=plt.figure(figsize=(20,40))\nfor i, col in enumerate(num_feats):\n    plt.subplot(len(num_feats),2,2*i+1)\n    sb.distplot(df[col])\n    plt.title(f'Train Data')\n\nfor i, col in enumerate(num_feats):\n    plt.subplot(len(num_feats),2,2*i+2)\n    sb.distplot(df_test1[col], color = 'deeppink')\n    plt.title(f'Test Data')\n    \nfig.tight_layout()\nplt.show()","metadata":{"execution":{"iopub.status.busy":"2022-07-06T16:52:50.700056Z","iopub.execute_input":"2022-07-06T16:52:50.700411Z","iopub.status.idle":"2022-07-06T16:52:58.742483Z","shell.execute_reply.started":"2022-07-06T16:52:50.700378Z","shell.execute_reply":"2022-07-06T16:52:58.741522Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### **<span style = 'color:brown'>Exploring the warning raised by Tukey's rule</span>**\nTukey's rule has found far outliers in CO2_(dinning-room), CO2_room, Lighting_(dinning-room), Meteo_Rain, Meteo_Sun_light_in_west_facade, Meteo_Sun_light_in_east_facade. so we'll look at this in a boxplot","metadata":{}},{"cell_type":"code","source":"# Numerical features\nnum_feats=[col for col in df_test1.columns if df_test1[col].dtypes != 'object' and col != 'Id']\n\n# Plot distribution of numerical columns\nfig=plt.figure(figsize=(20,40))\nfor i, col in enumerate(num_feats):\n    plt.subplot(len(num_feats),2,2*i+1)\n    sb.boxplot(df[col], color = 'red')\n    plt.title(f'Train Data')\n\nfor i, col in enumerate(num_feats):\n    plt.subplot(len(num_feats),2,2*i+2)\n    sb.boxplot(df_test1[col], color = 'yellow')\n    plt.title(f'Test Data')\n    \nfig.tight_layout()\nplt.show()","metadata":{"execution":{"iopub.status.busy":"2022-07-06T16:52:58.743829Z","iopub.execute_input":"2022-07-06T16:52:58.744709Z","iopub.status.idle":"2022-07-06T16:53:02.965852Z","shell.execute_reply.started":"2022-07-06T16:52:58.744674Z","shell.execute_reply":"2022-07-06T16:53:02.964361Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Apply log or square root transformation technique columns contains the outliers to remove outliers. ","metadata":{}},{"cell_type":"code","source":"df_trans = df.copy()\ndf_trans['CO2_(dinning-room)'] = np.log(df_trans['CO2_(dinning-room)']) # right skewed data -> apply log transform\ndf_trans['CO2_room'] = np.log(df_trans['CO2_room'])  # right outliers data -> apply log transform\ndf_trans['Meteo_Rain'] = np.log(df_trans['Meteo_Rain']) # right skewed data -> apply log transform\ndf_trans['Lighting_(dinning-room)'] = np.log(df_trans['Lighting_(dinning-room)'])  # right outlier data -> apply square root transform\ndf_trans['Meteo_Sun_light_in_west_facade'] = df_trans['Meteo_Sun_light_in_west_facade']**0.5  # right outlier data -> apply square root transform\ndf_trans['Meteo_Sun_light_in_east_facade'] = df_trans['Meteo_Sun_light_in_east_facade']**0.5  # right outlier data -> apply square root transform","metadata":{"execution":{"iopub.status.busy":"2022-07-06T16:53:02.967150Z","iopub.execute_input":"2022-07-06T16:53:02.967760Z","iopub.status.idle":"2022-07-06T16:53:02.978410Z","shell.execute_reply.started":"2022-07-06T16:53:02.967727Z","shell.execute_reply":"2022-07-06T16:53:02.977152Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"After trying both log and square root combinations the best results are obtained for Lighting_(dinning-room), Meteo_Sun_light_in_west_facade and Meteo_Sun_light_in_east_facade. The other three CO2_(dinning-room), CO2_room, and Meteo_Rain exhibited no change at all. So its better to remove these three from the training set and retain first three after doing necessary transformation as above. \n\nPlotting them to visualize before and after transformation effects.  ","metadata":{}},{"cell_type":"code","source":"# Numerical features\nnum_feats=[col for col in df.columns if df[col].dtypes != 'object' and col != 'label']\n\n# Plot distribution of numerical columns\nfig=plt.figure(figsize=(20,40))\nfor i, col in enumerate(num_feats):\n    plt.subplot(len(num_feats),2,2*i+1)\n    sb.boxplot(df[col], color = 'magenta')\n    plt.title(f'Before Transform Train Data')\n\nfor i, col in enumerate(num_feats):\n    plt.subplot(len(num_feats),2,2*i+2)\n    sb.boxplot(df_trans[col], color = 'green')\n    plt.title(f'After Transform Train Data')\n    \nfig.tight_layout()\nplt.show()\n\ndf['Lighting_(dinning-room)'] = df_trans['Lighting_(dinning-room)']\ndf['Meteo_Sun_light_in_west_facade'] = df_trans['Meteo_Sun_light_in_west_facade'] \ndf['Meteo_Sun_light_in_east_facade'] = df_trans['Meteo_Sun_light_in_east_facade']","metadata":{"execution":{"iopub.status.busy":"2022-07-06T16:53:02.979788Z","iopub.execute_input":"2022-07-06T16:53:02.980410Z","iopub.status.idle":"2022-07-06T16:53:07.875315Z","shell.execute_reply.started":"2022-07-06T16:53:02.980374Z","shell.execute_reply":"2022-07-06T16:53:07.873942Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### **<span style = 'color:brown'>Shapiro-Wilk test for normality</span>**\nThe Shapiro-Wilk test is a test of normality. It is used to determine whether or not a sample comes from a normal distribution.","metadata":{}},{"cell_type":"code","source":"#perform Shapiro-Wilk test for normality\nfrom scipy.stats import shapiro\nshapiro(df['Indoor_temperature_room'])","metadata":{"execution":{"iopub.status.busy":"2022-07-06T16:53:07.877089Z","iopub.execute_input":"2022-07-06T16:53:07.877556Z","iopub.status.idle":"2022-07-06T16:53:07.888267Z","shell.execute_reply.started":"2022-07-06T16:53:07.877506Z","shell.execute_reply":"2022-07-06T16:53:07.886801Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Since the p-value is less than .05, the null hypothesis of normality is to be rejected. We do have sufficient evidence to say that the sample data does not come from a normal distribution.\n\n### **<span style = 'color:brown'>Stem-and-Leaf Plot</span>**\n\nA stem-and-leaf plot is a unique plot to visualize distribution  of data by splitting up each raw value in a dataset into a stem and a leaf. This chart provides a wealth of information like cumulative frequencies, maximum, minimum, distribution of data using raw values etc. For this, first install stemgraphic library.","metadata":{}},{"cell_type":"code","source":"pip install stemgraphic --quiet","metadata":{"execution":{"iopub.status.busy":"2022-07-06T16:53:07.889891Z","iopub.execute_input":"2022-07-06T16:53:07.890697Z","iopub.status.idle":"2022-07-06T16:53:21.399388Z","shell.execute_reply.started":"2022-07-06T16:53:07.890661Z","shell.execute_reply":"2022-07-06T16:53:21.397670Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"import stemgraphic\n\n#create stem-and-leaf plot\nfig, ax = stemgraphic.stem_graphic(df['Indoor_temperature_room'])","metadata":{"execution":{"iopub.status.busy":"2022-07-06T16:53:21.402012Z","iopub.execute_input":"2022-07-06T16:53:21.402591Z","iopub.status.idle":"2022-07-06T16:53:22.860617Z","shell.execute_reply.started":"2022-07-06T16:53:21.402519Z","shell.execute_reply":"2022-07-06T16:53:22.859335Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### **<span style = 'color:brown'>Distribution Plot</span>**","metadata":{}},{"cell_type":"code","source":"import holoviews as hv\nfrom holoviews import opts\nhv.extension('bokeh')\nhv.Distribution(df['Indoor_temperature_room']).opts(title=\"Indoor Temperature Distribution\", color=\"red\",\n                                                        xlabel=\"Indoor Temperature\", ylabel=\"Density\")\\\n.opts(opts.Distribution(width=700, height=300,tools=['hover'],show_grid=True))","metadata":{"execution":{"iopub.status.busy":"2022-07-06T16:53:22.862240Z","iopub.execute_input":"2022-07-06T16:53:22.862696Z","iopub.status.idle":"2022-07-06T16:53:26.003599Z","shell.execute_reply.started":"2022-07-06T16:53:22.862661Z","shell.execute_reply":"2022-07-06T16:53:26.002264Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### **<span style = 'color:brown'>Correlation matrix</span>**","metadata":{}},{"cell_type":"code","source":"plt.figure(figsize = (15,10))\nsb.heatmap(df.drop('Id', axis = 1).corr(), cmap ='BuGn', annot = True);","metadata":{"execution":{"iopub.status.busy":"2022-07-06T16:53:26.005887Z","iopub.execute_input":"2022-07-06T16:53:26.006418Z","iopub.status.idle":"2022-07-06T16:53:27.674643Z","shell.execute_reply.started":"2022-07-06T16:53:26.006368Z","shell.execute_reply":"2022-07-06T16:53:27.673356Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"corr = df.drop('Id', axis = 1).corr()\ncorr_top = corr[abs(corr)>=.3]['Indoor_temperature_room'].sort_values(ascending = False)[0:3]\ncorr_top","metadata":{"execution":{"iopub.status.busy":"2022-07-06T16:53:27.676343Z","iopub.execute_input":"2022-07-06T16:53:27.677637Z","iopub.status.idle":"2022-07-06T16:53:27.695145Z","shell.execute_reply.started":"2022-07-06T16:53:27.677591Z","shell.execute_reply":"2022-07-06T16:53:27.694144Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## **<span style = 'color:green'>5. FBProphet for multivariate regression analysis**<a id ='FBProphet'></a>\n[<div style=\"text-align: right\"> Back to Table of contents</div>](#Table) ","metadata":{}},{"cell_type":"markdown","source":"### **<span style = 'color:brown'>5.1 Data preprocessing for FBProphet model</span>**<a id =\"Preprocessing\"></a>\n\n**Processing of Date format for train set**","metadata":{}},{"cell_type":"code","source":"df1 = df.copy()\ndf1['Date'] = pd.to_datetime(df1['Date'], dayfirst=True)\ndf1['Date'] = df1['Date'].dt.strftime('%Y-%m-%d')\ndf_train = df1[['Date','Time', 'Indoor_temperature_room', 'Relative_humidity_room', 'Outdoor_relative_humidity_Sensor']]\ndf_train['DateTime'] = df_train['Date'] + ' ' + df_train['Time']\ndf_train['DateTime'] = pd.to_datetime(df_train['DateTime'])\ndf_train.drop(['Date','Time'], axis = 1, inplace = True)\ndf_train = df_train[['DateTime','Indoor_temperature_room', \n                     'Relative_humidity_room','Outdoor_relative_humidity_Sensor' ]]","metadata":{"execution":{"iopub.status.busy":"2022-07-06T16:53:27.696340Z","iopub.execute_input":"2022-07-06T16:53:27.697346Z","iopub.status.idle":"2022-07-06T16:53:27.736183Z","shell.execute_reply.started":"2022-07-06T16:53:27.697276Z","shell.execute_reply":"2022-07-06T16:53:27.735200Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"display_plot(df_train, df_train['DateTime'], df_train['Indoor_temperature_room'], \"Indoor room temperature\")","metadata":{"execution":{"iopub.status.busy":"2022-07-06T16:53:27.737551Z","iopub.execute_input":"2022-07-06T16:53:27.738144Z","iopub.status.idle":"2022-07-06T16:54:38.502073Z","shell.execute_reply.started":"2022-07-06T16:53:27.738107Z","shell.execute_reply":"2022-07-06T16:54:38.500714Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"**Processing of Date format for test set**","metadata":{}},{"cell_type":"code","source":"df_test = df_test1.copy()\ndf_test['Date'] = pd.to_datetime(df_test['Date'],  dayfirst=True)\ndf_test['Date'] = df_test['Date'].dt.strftime('%Y-%m-%d')\ndf_test = df_test[['Date','Time', 'Relative_humidity_room', 'Outdoor_relative_humidity_Sensor']]\ndf_test['DateTime'] = df_test['Date'] + ' ' + df_test['Time']\ndf_test['DateTime'] = pd.to_datetime(df_test['DateTime'])\ndf_test.drop(['Date','Time'], axis = 1, inplace = True)\ndf_test = df_test[['DateTime','Relative_humidity_room','Outdoor_relative_humidity_Sensor' ]]\ndf_test.columns = ['ds', 'Relative_humidity_room', 'Outdoor_relative_humidity_Sensor']\ndf_test","metadata":{"execution":{"iopub.status.busy":"2022-07-06T16:54:38.503443Z","iopub.execute_input":"2022-07-06T16:54:38.503762Z","iopub.status.idle":"2022-07-06T16:54:38.551868Z","shell.execute_reply.started":"2022-07-06T16:54:38.503730Z","shell.execute_reply":"2022-07-06T16:54:38.550629Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"**Train set split for validation**","metadata":{}},{"cell_type":"code","source":"df_train.columns = ['ds', 'y', 'Relative_humidity_room','Outdoor_relative_humidity_Sensor']\nsize = int(len(df_train) * 0.80)\ntrain, val = df_train[0:size], df_train[size:len(df)]\ntrain.shape, val.shape","metadata":{"execution":{"iopub.status.busy":"2022-07-06T16:54:38.553773Z","iopub.execute_input":"2022-07-06T16:54:38.554175Z","iopub.status.idle":"2022-07-06T16:54:38.564350Z","shell.execute_reply.started":"2022-07-06T16:54:38.554142Z","shell.execute_reply":"2022-07-06T16:54:38.563358Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### **<span style = 'color:brown'>5.2 FBProphet model training</span>**<a id =\"Training\"></a>\n[<div style=\"text-align: right\"> Back to Table of contents</div>](#Table) \n\nAdditional regressors can be added to the linear part of the model using the add_regressor method or function. The add_regressor function provides a more general interface for defining extra linear regressors. Note that regressors must be added prior to model fitting. The extra regressors must be known for both the history and for future dates. Columns with the regressor values will need to be present in both the fitting and prediction dataframes.","metadata":{}},{"cell_type":"code","source":"from prophet import Prophet\nm = Prophet(interval_width=0.90)\nm.add_regressor('Relative_humidity_room')\nm.add_regressor('Outdoor_relative_humidity_Sensor')\nm.add_seasonality(\n    name='15min', period=1, fourier_order = 2)\nm.fit(df_train)","metadata":{"execution":{"iopub.status.busy":"2022-07-06T16:54:38.566061Z","iopub.execute_input":"2022-07-06T16:54:38.566505Z","iopub.status.idle":"2022-07-06T16:54:46.318504Z","shell.execute_reply.started":"2022-07-06T16:54:38.566469Z","shell.execute_reply":"2022-07-06T16:54:46.317141Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"forecast = m.predict(val)\nplt2 = m.plot_components(forecast)","metadata":{"execution":{"iopub.status.busy":"2022-07-06T16:54:46.327177Z","iopub.execute_input":"2022-07-06T16:54:46.327610Z","iopub.status.idle":"2022-07-06T16:54:51.934062Z","shell.execute_reply.started":"2022-07-06T16:54:46.327576Z","shell.execute_reply":"2022-07-06T16:54:51.933023Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### **<span style = 'color:brown'>5.3 Model Evaluation</span>**<a id =\"Evaluation\"></a>\n[<div style=\"text-align: right\"> Back to Table of contents</div>](#Table) ","metadata":{}},{"cell_type":"code","source":"forecast = forecast[['ds','yhat']].set_index('ds')\ntest2 = val.copy()\ntest2 = test2.set_index('ds')\nfinal = pd.concat((forecast['yhat'], test2), axis =1)\nfinal = final[['y', 'yhat']]\nfinal","metadata":{"execution":{"iopub.status.busy":"2022-07-06T16:54:51.935496Z","iopub.execute_input":"2022-07-06T16:54:51.936016Z","iopub.status.idle":"2022-07-06T16:54:51.958240Z","shell.execute_reply.started":"2022-07-06T16:54:51.935983Z","shell.execute_reply":"2022-07-06T16:54:51.957027Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"def model_evaluation(final, y, ypred, model_name):\n    from sklearn.metrics import mean_squared_error,mean_absolute_error,explained_variance_score, r2_score\n    print(\"\\n \\n Model Evaluation Report: \")\n    print('Mean Absolute Error(MAE) of', model_name,':', mean_absolute_error(y, ypred))\n    print('Mean Squared Error(MSE) of', model_name,':', mean_squared_error(y, ypred))\n    print('Root Mean Squared Error (RMSE) of', model_name,':', np.sqrt(mean_squared_error(y, ypred)))\n    print('Explained Variance Score (EVS) of', model_name,':', explained_variance_score(y, ypred))\n    print('R2 of', model_name,':', (r2_score(y, ypred)).round(2))\n    print('\\n \\n')\n    \n    # Actual vs Predicted Plot\n    f, ax = plt.subplots(1,2, figsize=(12,6),dpi=100);\n    f.set_tight_layout('tight')\n    ax[0].plot(final)\n    ax[0].set_title('Predicted vs Original');\n    \n    plt.scatter(y, ypred, label=\"Actual vs Predicted\")\n    # Perfect predictions\n    ax[1].set_xlabel('Indoor Temperature in celsius')\n    ax[1].set_ylabel('Indoor Temperature in celsius')\n    ax[1].set_title('Expection vs Prediction')\n    ax[1].plot(y,y,'r', label=\"Perfect Expected Prediction\")\n    ax[1].legend()\n    f.text(0.95, 0.06, 'AUTHOR: RINI CHRISTY',\n         fontsize=12, color='green',\n         ha='left', va='bottom', alpha=0.5);","metadata":{"execution":{"iopub.status.busy":"2022-07-06T16:54:51.959834Z","iopub.execute_input":"2022-07-06T16:54:51.960160Z","iopub.status.idle":"2022-07-06T16:54:51.971871Z","shell.execute_reply.started":"2022-07-06T16:54:51.960128Z","shell.execute_reply":"2022-07-06T16:54:51.970979Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"model_evaluation(final, final['y'], final['yhat'], 'Multivariate FBProphet model')","metadata":{"execution":{"iopub.status.busy":"2022-07-06T16:54:51.973217Z","iopub.execute_input":"2022-07-06T16:54:51.973810Z","iopub.status.idle":"2022-07-06T16:54:52.667179Z","shell.execute_reply.started":"2022-07-06T16:54:51.973777Z","shell.execute_reply":"2022-07-06T16:54:52.665669Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### **<span style = 'color:brown'>5.4 Test Set Prediction & Kaggle Submission</span>**<a id =\"Prediction\"></a>\n[<div style=\"text-align: right\"> Back to Table of contents</div>](#Table) ","metadata":{}},{"cell_type":"code","source":"df_test","metadata":{"execution":{"iopub.status.busy":"2022-07-06T16:54:52.668922Z","iopub.execute_input":"2022-07-06T16:54:52.669279Z","iopub.status.idle":"2022-07-06T16:54:52.692666Z","shell.execute_reply.started":"2022-07-06T16:54:52.669245Z","shell.execute_reply":"2022-07-06T16:54:52.691344Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"forecast_test = m.predict(df_test)\nsubmission_preds = forecast_test['yhat']\ntest_ids = df_test1['Id']\ndf = pd.DataFrame({'Id': test_ids.values, 'Indoor_temperature_room': submission_preds})\ndf.to_csv('submission.csv', index = False)","metadata":{"execution":{"iopub.status.busy":"2022-07-06T16:54:52.696672Z","iopub.execute_input":"2022-07-06T16:54:52.697071Z","iopub.status.idle":"2022-07-06T16:54:58.011692Z","shell.execute_reply.started":"2022-07-06T16:54:52.697033Z","shell.execute_reply":"2022-07-06T16:54:58.010352Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"[<div style=\"text-align: right\"> Back to Table of contents</div>](#Table) ","metadata":{}}]}