{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"pygments_lexer":"ipython3","nbconvert_exporter":"python","version":"3.6.4","file_extension":".py","codemirror_mode":{"name":"ipython","version":3},"name":"python","mimetype":"text/x-python"}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"# **<span style = 'color:blue'>Part 2: Bi- and Uni-directional LSTM RNNs for Smart Homes Indoor Temperature Forecasting</span>**\nPart 1: [Multivariate FBProphet Smart Homes Indoor Temperature Forecasting]( https://www.kaggle.com/rinichristy/multivariate-fbprophet-smart-homes-temperature)\n\n## **<span style='color:green'> Contents:</span>**<a id=\"Table\"></a>\n\n* [1. Import the libraries](#Import)\n* [2. Dataset Information](#Dataset)\n* [3. Statistical Data Analysis](#Statistical)\n* [4. Data Wrangling](#Wrangling)\n    - [4.1 Choosing variables correlated with Indoor room temperature](#variables)\n    - [4.2 Feature engineering to extract Dateparts](#Dateparts)\n    - [4.3 Choosing number of lags](#Lags)\n* [5. Build Multivariate Bidirectional & Unidirectional RNN LSTM Time Series Models](#RNN)\n    - [5.1 Recurrent neural networks (RNN): Introduction](#Introduction)\n    - [5.2 Model: Bidirectional LSTM Architecture](#Bidirectional)\n    - [5.3 Model: Unidirectional LSTM Architecture](#Unidirectional)\n* [6. Model Train, Evaluation & Final Report](#Report)\n    - [6.1 Train & evaluate the model](#Train-Evaluate)\n    - [6.2 Model Forecast](#Forecast)\n    - [6.3 Kaggle Submission](#Submission)","metadata":{"_uuid":"8f2839f25d086af736a60e9eeb907d3b93b6e0e5","_cell_guid":"b1076dfc-b9ad-4769-8c92-a6c4dae69d19"}},{"cell_type":"markdown","source":"## **<span style = 'color:green'>1. Import the required libraries</span>**<a id =\"Import\"></a>","metadata":{}},{"cell_type":"code","source":"# Import data handling & numerical libraries\nimport pandas as pd\nimport numpy as np\nfrom copy import copy\nimport datetime\n\n# Import Data Visualization libraries\nimport seaborn as sb\nimport matplotlib.pyplot as plt\n\n#import libraries for muting unnecessary warnings if needed\nimport warnings\nwarnings.filterwarnings('ignore')","metadata":{"execution":{"iopub.status.busy":"2022-07-08T09:15:51.778924Z","iopub.execute_input":"2022-07-08T09:15:51.780855Z","iopub.status.idle":"2022-07-08T09:15:52.343273Z","shell.execute_reply.started":"2022-07-08T09:15:51.780727Z","shell.execute_reply":"2022-07-08T09:15:52.342339Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## **<span style = 'color:green'>2. Dataset information</span>**<a id ='Dataset'></a>\n[<div style=\"text-align: right\"> Back to Table of contents</div>](#Table)\n\nThe dataset collected from the monitor system mounted in a solar house corresponds to approximately 40 days of monitoring data. The Goal is to predict indoor temperature of a room (the Bedroom), in order to choose whether or not to activate the HVAC (Heating, Ventilation, and Air Conditioning) system. The data was sampled every minute, computing and uploading it smoothed with 15 minute means. The dataset includes dates, other sensor measurements, weather measurements and other information. It is a multivariate time-series dataset.\n\nIt has been established that the power consumption attributed to HVAC accounts for 53.9 % of total consumption, and the energy required to maintain the temperature is less than that required to drop or raise it.\n\nAs a result, a predictive model capable of predicting a room's indoor temperature (a short-term forecast of indoor temperature) would help in lowering overall energy consumption, by deciding whether or not to activate the HVAC system, at the appropriate time.\n\nBelow displayed map provides an idea of locations of solar house sensors and actuators. ","metadata":{}},{"cell_type":"code","source":"from IPython.display import Image\nurl = '../input/smart-homes-temperature-time-series-forecasting/Solar house sensors and actuators map.png'\nImage(url,width=700, height=700)","metadata":{"execution":{"iopub.status.busy":"2022-07-08T09:15:52.348321Z","iopub.execute_input":"2022-07-08T09:15:52.348877Z","iopub.status.idle":"2022-07-08T09:15:52.369711Z","shell.execute_reply.started":"2022-07-08T09:15:52.348844Z","shell.execute_reply":"2022-07-08T09:15:52.368340Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"df = pd.read_csv('../input/smart-homes-temperature-time-series-forecasting/train.csv')\ndf_test1 = pd.read_csv('../input/smart-homes-temperature-time-series-forecasting/test.csv')\ndf.head(2)","metadata":{"execution":{"iopub.status.busy":"2022-07-08T09:15:52.371241Z","iopub.execute_input":"2022-07-08T09:15:52.372220Z","iopub.status.idle":"2022-07-08T09:15:52.418543Z","shell.execute_reply.started":"2022-07-08T09:15:52.372178Z","shell.execute_reply":"2022-07-08T09:15:52.417614Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## **<span style='color:green'>3. Statistical Data Analysis</span>**<a id ='Statistical'></a>\nSee [Part1](https://www.kaggle.com/rinichristy/multivariate-fbprophet-smart-homes-temperature) for this section","metadata":{}},{"cell_type":"markdown","source":"## **<span style = 'color:green'>4. Data Wrangling</span>**<a id ='Wrangling'></a>","metadata":{}},{"cell_type":"code","source":"# Numerical features\nnum_feats=[col for col in df.columns if df[col].dtypes != 'object'  and col !='Day_of_the_week' and col != 'Id']\n\n# Plot distribution of numerical columns\nfig=plt.figure(figsize=(20,40))\nfor i, col in enumerate(num_feats):\n    plt.subplot(len(num_feats),2,1*i+1)\n    sb.lineplot(df['Date'], df[col], color ='green')\n    plt.title(col.upper())  \nfig.tight_layout()\nplt.show()","metadata":{"execution":{"iopub.status.busy":"2022-07-08T09:15:52.421750Z","iopub.execute_input":"2022-07-08T09:15:52.422416Z","iopub.status.idle":"2022-07-08T09:16:09.814400Z","shell.execute_reply.started":"2022-07-08T09:15:52.422376Z","shell.execute_reply":"2022-07-08T09:16:09.813094Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### **<span style = 'color:brown'>4.1 Choosing variables correlated with Indoor room temperature</span>**<a id = 'variables'></a>\n[<div style=\"text-align: right\"> Back to Table of contents</div>](#Table)\n\n**Cross-correlation plot**\n\nTo examine the cross-correlation between for example, Sun Irradiance hours and indoor temperature, which is a process subject to seasonal lag - whereby maximum indoor temperature will lag the period of maximum sunlight irradiance.\n\nThe matplotlib xcorr() function can be used to plot the cross correlation between Indoor_temperature_room and other variable's time series to determine if a lag is present, and if so by how many time periods. The cross-correlation is plotted for the data as follows:","metadata":{}},{"cell_type":"code","source":"plt.xcorr(df['Indoor_temperature_room'], df['Meteo_Sun_irradiance'], normed=True, usevlines=True, maxlags=100)\nplt.title(\"Sun irradiance versus Indoor Temperature\");","metadata":{"execution":{"iopub.status.busy":"2022-07-08T09:16:09.816225Z","iopub.execute_input":"2022-07-08T09:16:09.817267Z","iopub.status.idle":"2022-07-08T09:16:10.036728Z","shell.execute_reply.started":"2022-07-08T09:16:09.817225Z","shell.execute_reply":"2022-07-08T09:16:10.035427Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"In this instance, the strongest correlation between Sun irradiance and Indoor temperature comes lags by approximately 20 days, i.e. this is when the strongest correlation between the two time series is observed. The below xcorrelation plots display the individual influence of parameters such as relative humidity, wind, solar irradiation, lighting etc on indoor temperature. ","metadata":{}},{"cell_type":"code","source":"# Numerical features\nnum_feats=[col for col in df.columns if df[col].dtypes != 'object' and col != 'Id']\n\n# Plot distribution of numerical columns\nfig=plt.figure(figsize=(15,40))\nfor i, col in enumerate(num_feats):\n    plt.subplot(len(num_feats),2,1*i+1)\n    plt.xcorr(df['Indoor_temperature_room'], df[col], normed=True, usevlines=True, maxlags=100, color = 'red')\n    plt.title(col.upper())\n    \nfig.tight_layout()\nplt.show()","metadata":{"execution":{"iopub.status.busy":"2022-07-08T09:16:10.038324Z","iopub.execute_input":"2022-07-08T09:16:10.038680Z","iopub.status.idle":"2022-07-08T09:16:12.933886Z","shell.execute_reply.started":"2022-07-08T09:16:10.038648Z","shell.execute_reply":"2022-07-08T09:16:12.932524Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Data with a sine or cosine-like wave indicate seasonality, which repeat itself periodically. To know more about this, define a function to make lags using shift method and then call the function on each variable to make any desirable number of lags.","metadata":{}},{"cell_type":"code","source":"def make_lags(ts, lags, start):\n    return pd.concat(\n        {\n            f'y_lag_{i}': ts.shift(i)\n            for i in range(start, lags + 1)\n        },\n        axis=1)\nX = make_lags(df['Meteo_Sun_irradiance'], lags=25, start = 15)\nX","metadata":{"execution":{"iopub.status.busy":"2022-07-08T09:16:12.936069Z","iopub.execute_input":"2022-07-08T09:16:12.936552Z","iopub.status.idle":"2022-07-08T09:16:12.967565Z","shell.execute_reply.started":"2022-07-08T09:16:12.936505Z","shell.execute_reply":"2022-07-08T09:16:12.966331Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"X['Indoor_temperature_room'] = df['Indoor_temperature_room']\nX.loc[25:, :].corr()['Indoor_temperature_room'].sort_values(ascending = False)[0:5]","metadata":{"execution":{"iopub.status.busy":"2022-07-08T09:16:12.969277Z","iopub.execute_input":"2022-07-08T09:16:12.969783Z","iopub.status.idle":"2022-07-08T09:16:12.983824Z","shell.execute_reply.started":"2022-07-08T09:16:12.969736Z","shell.execute_reply":"2022-07-08T09:16:12.982858Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"z = make_lags(df['Lighting_room'], lags=25, start = 15)\nz = z.fillna(0)\nz['Indoor_temperature_room'] = df['Indoor_temperature_room']\nz.loc[25:, :].corr()['Indoor_temperature_room'].sort_values(ascending = False)[0:5]","metadata":{"execution":{"iopub.status.busy":"2022-07-08T09:16:12.985118Z","iopub.execute_input":"2022-07-08T09:16:12.985979Z","iopub.status.idle":"2022-07-08T09:16:13.004289Z","shell.execute_reply.started":"2022-07-08T09:16:12.985943Z","shell.execute_reply":"2022-07-08T09:16:13.003388Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"z = make_lags(df['Meteo_Wind'], lags=50, start = 1)\nz = z.fillna(0)\nz['Indoor_temperature_room'] = df['Indoor_temperature_room']\nabs(z.loc[25:, :].corr()['Indoor_temperature_room']).sort_values(ascending = False)[0:5]","metadata":{"execution":{"iopub.status.busy":"2022-07-08T09:16:13.005520Z","iopub.execute_input":"2022-07-08T09:16:13.006011Z","iopub.status.idle":"2022-07-08T09:16:13.046248Z","shell.execute_reply.started":"2022-07-08T09:16:13.005976Z","shell.execute_reply":"2022-07-08T09:16:13.045372Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"df_train1 = df.copy()\ndf_test = df_test1.copy()\ndf_train_test=pd.concat([df.drop(['Indoor_temperature_room'],axis=1),df_test1],ignore_index=True)\ndf_train_test","metadata":{"execution":{"iopub.status.busy":"2022-07-08T09:16:13.047317Z","iopub.execute_input":"2022-07-08T09:16:13.048025Z","iopub.status.idle":"2022-07-08T09:16:13.086051Z","shell.execute_reply.started":"2022-07-08T09:16:13.047992Z","shell.execute_reply":"2022-07-08T09:16:13.084613Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### **<span style = 'color:brown'>4.2 Feature engineering to extract Dateparts</span>**<a id = 'Dateparts'></a>\n[<div style=\"text-align: right\"> Back to Table of contents</div>](#Table)","metadata":{}},{"cell_type":"code","source":"def preprocss_data(df):\n    df1 = df.copy()\n    df1['Date'] = pd.to_datetime(df1['Date'], dayfirst=True)\n    df1['Date'] = df1['Date'].dt.strftime('%Y-%m-%d')\n\n    df1['DateTime'] = df1['Date'] + ' ' + df1['Time']\n    df1['DateTime'] = pd.to_datetime(df1['DateTime'])\n    df1.drop(['Date','Time'], axis = 1, inplace = True)\n    \n    from fastai.tabular.core import add_datepart\n    # make a Date copy because \"add_datepart\" do delete the orginal formatted Date.\n    #In other words, the function transfer the column to Date Parts\n    df_formatted = df1.copy()\n    df_formatted['Formatted Date'] = df_formatted[\"DateTime\"]\n    df_formatted = add_datepart(df_formatted, 'Formatted Date')\n    df_formatted.head()\n    \n    def hours2timing(x):\n        if x in [22,23,0,1,2,3]:\n            timing = 'Night'\n        elif x in range(4, 12):\n            timing = 'Morning'\n        elif x in range(12, 17):\n            timing = 'Afternoon'\n        elif x in range(17, 22):\n            timing = 'Evening'\n        else:\n            timing = 'X'\n        return timing\n\n    df_formatted['hour'] = df_formatted['DateTime'].apply(lambda x : x.hour)\n    df_formatted['timing'] = df_formatted['hour'].apply(hours2timing)\n    timing = df_formatted[['timing']]\n    \n    \n    #OrdinalEncoder from scikit learn, which allows multi-column encoding \n    # It can be used to convert categorical features in numerical data type.\n    from sklearn.preprocessing import OrdinalEncoder\n    for i in df_formatted.select_dtypes(['object', 'bool']).columns:\n        Oe = OrdinalEncoder().fit(df_formatted[[i]])\n        df_formatted[i] = Oe.transform(df_formatted[[i]]).astype(int)\n    return  df_formatted\n\ndf_formatted_train_test = preprocss_data(df_train_test)\ndf_formatted_train = preprocss_data(df)\ndf_formatted_train","metadata":{"execution":{"iopub.status.busy":"2022-07-08T09:16:13.087736Z","iopub.execute_input":"2022-07-08T09:16:13.088908Z","iopub.status.idle":"2022-07-08T09:16:13.868593Z","shell.execute_reply.started":"2022-07-08T09:16:13.088855Z","shell.execute_reply":"2022-07-08T09:16:13.867324Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Plot the xcorr plot of the entire formatted dataset once again to hunt for new added features, if any. ","metadata":{}},{"cell_type":"code","source":"# Numerical features\nnum_feats=[col for col in df_formatted_train.columns if df_formatted_train[col].dtypes != 'object' and col != 'Id']\n\n# Plot distribution of numerical columns\nfig=plt.figure(figsize=(15,40))\nfor i, col in enumerate(num_feats):\n    plt.subplot(len(num_feats),2,1*i+1)\n    plt.xcorr(df_formatted_train['Indoor_temperature_room'], df_formatted_train[col], normed=True, usevlines=True, maxlags=100, color = 'purple')\n    plt.title(col.upper())\n    \nfig.tight_layout()\nplt.show()","metadata":{"execution":{"iopub.status.busy":"2022-07-08T09:16:13.875247Z","iopub.execute_input":"2022-07-08T09:16:13.876283Z","iopub.status.idle":"2022-07-08T09:16:19.891738Z","shell.execute_reply.started":"2022-07-08T09:16:13.876224Z","shell.execute_reply":"2022-07-08T09:16:19.890851Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Make lags of some of the numerical columns and check their corresponding correlation with 'Indoor_temperature_room' to find the lag that has the most impact on temperature.","metadata":{}},{"cell_type":"code","source":"hr = make_lags(df_formatted_train['hour'], lags=25, start=1)\nhr = hr.fillna(0)\nhr['Indoor_temperature_room'] = df['Indoor_temperature_room']\nhr.corr()['Indoor_temperature_room'].sort_values(ascending = False)[0:5]","metadata":{"execution":{"iopub.status.busy":"2022-07-08T09:16:19.893211Z","iopub.execute_input":"2022-07-08T09:16:19.893581Z","iopub.status.idle":"2022-07-08T09:16:19.918749Z","shell.execute_reply.started":"2022-07-08T09:16:19.893544Z","shell.execute_reply":"2022-07-08T09:16:19.917887Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"timing = make_lags(df_formatted_train['timing'], lags=10, start = 1)\ntiming = timing.fillna(0)\ntiming['Indoor_temperature_room'] = df['Indoor_temperature_room']\nabs(timing.corr()['Indoor_temperature_room']).sort_values(ascending = False)[0:5]","metadata":{"execution":{"iopub.status.busy":"2022-07-08T09:16:19.919835Z","iopub.execute_input":"2022-07-08T09:16:19.920909Z","iopub.status.idle":"2022-07-08T09:16:19.937821Z","shell.execute_reply.started":"2022-07-08T09:16:19.920872Z","shell.execute_reply":"2022-07-08T09:16:19.936659Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"humidity = make_lags(df_formatted_train['Outdoor_relative_humidity_Sensor'], lags=25, start = 1)\nhumidity = humidity.fillna(0)\nhumidity['Indoor_temperature_room'] = df['Indoor_temperature_room']\nabs(humidity.corr()['Indoor_temperature_room']).sort_values(ascending = False)[0:5]","metadata":{"execution":{"iopub.status.busy":"2022-07-08T09:16:19.939069Z","iopub.execute_input":"2022-07-08T09:16:19.940114Z","iopub.status.idle":"2022-07-08T09:16:19.963394Z","shell.execute_reply.started":"2022-07-08T09:16:19.940051Z","shell.execute_reply":"2022-07-08T09:16:19.962237Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Now, start building the training dataset to fit the model. ","metadata":{}},{"cell_type":"code","source":"df_formatted_train_test1 = df_formatted_train_test.copy()\ndf_train=pd.merge(df_formatted_train_test1,df[['Id','Indoor_temperature_room']],on='Id')\ndf_train = df_train.fillna(method = 'bfill')\ndf_train","metadata":{"execution":{"iopub.status.busy":"2022-07-08T09:16:19.964952Z","iopub.execute_input":"2022-07-08T09:16:19.965352Z","iopub.status.idle":"2022-07-08T09:16:20.221063Z","shell.execute_reply.started":"2022-07-08T09:16:19.965318Z","shell.execute_reply":"2022-07-08T09:16:20.219563Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"df_train_lag = df_train.copy()\ndf_train_lag['Meteo_Sun_irradiance'] = df_train_lag['Meteo_Sun_irradiance'].shift(19)\ndf_train_lag['Lighting_room'] = df_train_lag['Lighting_room'].shift(21)\ndf_train_lag['Outdoor_relative_humidity_Sensor'] = df_train_lag['Outdoor_relative_humidity_Sensor'].shift(6)\ndf_train_lag['Meteo_Wind'] = df_train_lag['Meteo_Wind'].shift(12)","metadata":{"execution":{"iopub.status.busy":"2022-07-08T09:16:20.223557Z","iopub.execute_input":"2022-07-08T09:16:20.224094Z","iopub.status.idle":"2022-07-08T09:16:20.235477Z","shell.execute_reply.started":"2022-07-08T09:16:20.224022Z","shell.execute_reply":"2022-07-08T09:16:20.234122Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"corr = df_train_lag.drop(['Id', 'DateTime'], axis = 1).corr()\nabs(corr['Indoor_temperature_room']).sort_values(ascending = False)[0:10]","metadata":{"execution":{"iopub.status.busy":"2022-07-08T09:16:20.236906Z","iopub.execute_input":"2022-07-08T09:16:20.237274Z","iopub.status.idle":"2022-07-08T09:16:20.261905Z","shell.execute_reply.started":"2022-07-08T09:16:20.237239Z","shell.execute_reply":"2022-07-08T09:16:20.261022Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"print(df_formatted_train_test[['Lighting_room', 'Meteo_Sun_irradiance']].corr())","metadata":{"execution":{"iopub.status.busy":"2022-07-08T09:16:20.263528Z","iopub.execute_input":"2022-07-08T09:16:20.263855Z","iopub.status.idle":"2022-07-08T09:16:20.275424Z","shell.execute_reply.started":"2022-07-08T09:16:20.263825Z","shell.execute_reply":"2022-07-08T09:16:20.274430Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"* After applying lag, it is clearly visible that Meteo_Sun_irradiance is the most correlated with Indoor temperature. \n* This is then followed by Outdoor_relative_humidity_Sensor, \n* and thenby 'Lighting_room', which is almost perfectly collinear to sun irradiance (0.92). So either one should be selected and it is better to drop Lighting from the list because it is the direct effect of the irradiance. \n* Next in the list is  'Relative_humidity_room' \n* and 'hour' also have moderate correlation with Indoor temperature.  \n\n### **<span style = 'color:brown'>4.3 Choosing number of lags</span>**<a id =\"Lags\"></a>\n[<div style=\"text-align: right\"> Back to Table of contents</div>](#Table)\n\n#### **Lag plot**\n\nA lag plot to compare indoor room temperature autocorrelationfrom each observation in the dataset against the one from a previous observation. ","metadata":{}},{"cell_type":"code","source":"#from pandas import plotting\n#plotting.lag_plot(df['Indoor_temperature_room'], lag=1, marker='2', c='blue')\n\nfrom pandas.plotting import lag_plot\nplt.figure()\nlag_plot(df['Indoor_temperature_room'], lag=1, marker='2', c='blue')\nplt.title('lag = 1')\nplt.figure()\nlag_plot(df['Indoor_temperature_room'], lag=5, marker='2', c='green')\nplt.title('lag = 5');\nplt.figure()\nlag_plot(df['Indoor_temperature_room'], lag=10, marker='2', c='red')\nplt.title('lag = 10');\n","metadata":{"execution":{"iopub.status.busy":"2022-07-08T09:16:20.277009Z","iopub.execute_input":"2022-07-08T09:16:20.277757Z","iopub.status.idle":"2022-07-08T09:16:20.867356Z","shell.execute_reply.started":"2022-07-08T09:16:20.277718Z","shell.execute_reply":"2022-07-08T09:16:20.866203Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"If the plot shows a linear pattern, it suggests autocorrelation is present. A positive linear trend (i.e. going upwards from left to right) is suggestive of positive autocorrelation and the tighter the data is clustered around the diagonal, the more autocorrelation is present; perfectly autocorrelated data will cluster in a single diagonal line.\nSince the lag plot is linear, the underlying structure is of the autoregressive in nature. The linear shape to the plot suggests that an autoregressive model with 1 lag is probably the best choice for this data. Almost the entire values are concentrated on the diagonal in lag plot, it suggests a strong autocorrelation.\n\nThe plots overlapped on top of each other shows the difference in dispersion of points on increasing the lags. ","metadata":{}},{"cell_type":"code","source":"from pandas import plotting\nplt.figure()\nplotting.lag_plot(df['Indoor_temperature_room'], lag=1)\nplotting.lag_plot(df['Indoor_temperature_room'], lag=10, c= 'red', alpha =0.05)\nplt.title('lag 1 (blue) vs lag 10 (red)');","metadata":{"execution":{"iopub.status.busy":"2022-07-08T09:16:20.868870Z","iopub.execute_input":"2022-07-08T09:16:20.869253Z","iopub.status.idle":"2022-07-08T09:16:21.051968Z","shell.execute_reply.started":"2022-07-08T09:16:20.869220Z","shell.execute_reply":"2022-07-08T09:16:21.051112Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"from pandas.plotting import autocorrelation_plot\n\nautocorrelation_plot(df['Indoor_temperature_room'].head(50));","metadata":{"execution":{"iopub.status.busy":"2022-07-08T09:16:21.053184Z","iopub.execute_input":"2022-07-08T09:16:21.054124Z","iopub.status.idle":"2022-07-08T09:16:21.208623Z","shell.execute_reply.started":"2022-07-08T09:16:21.054055Z","shell.execute_reply":"2022-07-08T09:16:21.207443Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"#### **Partial Autocorrelation Plot of the Indoor Temperatures**","metadata":{}},{"cell_type":"code","source":"from statsmodels.graphics.tsaplots import plot_pacf\nplot_pacf(df['Indoor_temperature_room'], title = \"Partial Autocorrelation Plot of the Indoor Temperatures\");","metadata":{"execution":{"iopub.status.busy":"2022-07-08T09:16:21.211151Z","iopub.execute_input":"2022-07-08T09:16:21.212270Z","iopub.status.idle":"2022-07-08T09:16:21.483374Z","shell.execute_reply.started":"2022-07-08T09:16:21.212232Z","shell.execute_reply":"2022-07-08T09:16:21.482162Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"The lag plots indicate that the relationship of Indoor temperature to its lags is mostly linear, while the partial autocorrelation suggests the dependence can be captured efficiently using lags 1, 2, 3, 7, 8 etc. Since xcorr plot shows a correlation at 4 lags for 'Outdoor_relative_humidity_Sensor', 1 for 'hour' and 19 for 'Meteo_Sun_irradiance', it is better to settle with a lag of 8 for building the memory of the model. So this is going to be a multivariate time series with Meteo_Sun_irradiance','hour','Outdoor_relative_humidity_Sensor', and 'Indoor_temperature_room' and their lag of 8 to be the independent variables for predicting the dependent variable 'Indoor_temperature_room'.\n\nA sample test with Linear regression is run as follows:\n\nFirst, a univariate autocorrelation run.","metadata":{}},{"cell_type":"code","source":"# Create target series and data splits\nX = make_lags(df.Indoor_temperature_room, lags=8, start = 1)\nX = X.fillna(0.0)\n\ny = df.Indoor_temperature_room.copy()\nfrom sklearn.linear_model import LinearRegression\nfrom sklearn.model_selection import train_test_split\nX_train, X_test, y_train, y_test = train_test_split(X, y, test_size=500, shuffle=False)\n\n# Fit and predict\nmodel = LinearRegression()  # `fit_intercept=True` since we didn't use DeterministicProcess\nmodel.fit(X_train, y_train)\ny_pred = pd.Series(model.predict(X_train), index=y_train.index)\ny_fore = pd.Series(model.predict(X_test), index=y_test.index)\n\nplt.rc(\"figure\", autolayout=True, figsize=(11, 4))\n\nplot_params = dict(\n    color=\"0.75\",\n    style=\".-\",\n    markeredgecolor=\"0.25\",\n    markerfacecolor=\"0.25\",\n)\nax = y_train.plot(**plot_params)\nax = y_test.plot(**plot_params)\nax = y_pred.plot(ax=ax)\n_ = y_fore.plot(ax=ax, color='C3')\n","metadata":{"execution":{"iopub.status.busy":"2022-07-08T09:16:21.485493Z","iopub.execute_input":"2022-07-08T09:16:21.485968Z","iopub.status.idle":"2022-07-08T09:16:21.878152Z","shell.execute_reply.started":"2022-07-08T09:16:21.485919Z","shell.execute_reply":"2022-07-08T09:16:21.876645Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"To improve the forecast add some leading indicators, time series that could provide an \"early warning\" for changes in temperature changes. For this approach add those above mentioned variables to the training data. ","metadata":{}},{"cell_type":"code","source":"ts = df_train[['Meteo_Sun_irradiance', 'hour', 'Outdoor_relative_humidity_Sensor']]\n# Create three lags for each search term\nX0 = make_lags(ts, lags=8, start=1)\n\n# Create four lags for the target, as before\nX1 = make_lags(df_train['Indoor_temperature_room'], lags=8, start=1)\n\n# Combine to create the training data\nX = pd.concat([X0, X1], axis=1).fillna(0.0)\n\nX_train, X_test, y_train, y_test = train_test_split(X, y, test_size=500, shuffle=False)\n\nmodel = LinearRegression()\nmodel.fit(X_train, y_train)\ny_pred = pd.Series(model.predict(X_train), index=y_train.index)\ny_fore = pd.Series(model.predict(X_test), index=y_test.index)\n\nax = y_test.plot(**plot_params)\n_ = y_fore.plot(ax=ax, color='C3')\n","metadata":{"execution":{"iopub.status.busy":"2022-07-08T09:16:21.880114Z","iopub.execute_input":"2022-07-08T09:16:21.880484Z","iopub.status.idle":"2022-07-08T09:16:22.240373Z","shell.execute_reply.started":"2022-07-08T09:16:21.880451Z","shell.execute_reply":"2022-07-08T09:16:22.239148Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Even though it looks simple yet promising, with this approach past known values required to be used to predict the unknown values get exhausted at the end of the series. Only the very next value can be predicted. The forecast comes to a halt beyond the first value of the test set where the dependent variable is unknown to the model. To predict a very long series  another model is required which can retain the memory and carry it forward to next in line till the end.  Here comes the advantage of RNN over these models. \n\n## **<span style = 'color:green'>5. Build Multivariate Bidirectional & Unidirectional RNN LSTM Time Series Models**<a id ='RNN'></a>\n[<div style=\"text-align: right\"> Back to Table of contents</div>](#Table)\n### **<span style = 'color:brown'>5.1 Recurrent neural networks (RNN): Introduction</span>**<a id = 'Introduction'></a>\nThe main advantage of using Recurrent neural networks (RNN) is its ability to learn from the immediately preceding data in a sequence like Time series.  It remembers the sequence of the data and uses any patterns involved to make the prediction with the help of feedback loops. By gathering information from the previous loop, it converts both dependent and independent variable to its memory for its prediction of next future dependent value. However it suffers from vanishing or exploding gradient issues, because of which the network doesn’t learn much from the data which is far away from the current position. To overcome this problem Long Short Term Memory Networks, (LSTM) can be used. An LSTM is capable of learning long-term dependencies using short term memory by means of input, output, and forget gates. These gates help to remember the crucial information and forgets the unnecessary information that it learns throughout the network.\n    \nThe chosen variables are: 'Meteo_Sun_irradiance','hour','Outdoor_relative_humidity_Sensor', 'Indoor_temperature_room'","metadata":{}},{"cell_type":"code","source":"df_train=df_train[['Meteo_Sun_irradiance','hour','Outdoor_relative_humidity_Sensor', 'Indoor_temperature_room']]","metadata":{"execution":{"iopub.status.busy":"2022-07-08T09:16:22.242466Z","iopub.execute_input":"2022-07-08T09:16:22.243257Z","iopub.status.idle":"2022-07-08T09:16:22.250500Z","shell.execute_reply.started":"2022-07-08T09:16:22.243204Z","shell.execute_reply":"2022-07-08T09:16:22.249235Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"#### **<span style = 'color:purple'>Scaling the data</span>**\nThis is required for faster processing of data and also to remove any bias that may arise due to difference in range of variables, if any.","metadata":{}},{"cell_type":"code","source":"from sklearn.preprocessing import StandardScaler\nscaler= StandardScaler()\nscaler.fit(df_train)\ndf_train_scaled =scaler.transform(df_train)\nprint(df_train_scaled[0:3], '\\n, \\n')\ndf_train_scaled.shape","metadata":{"execution":{"iopub.status.busy":"2022-07-08T09:16:22.251990Z","iopub.execute_input":"2022-07-08T09:16:22.252850Z","iopub.status.idle":"2022-07-08T09:16:22.271994Z","shell.execute_reply.started":"2022-07-08T09:16:22.252812Z","shell.execute_reply":"2022-07-08T09:16:22.270563Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"#### **<span style = 'color:purple'>Create a new multivariate timeseies dataset with lag features suitable for forecast</span>**\nLoop through the data to create multivariate, multistep predictors (X) and target (y) arrays that will allow the model to see past 8 instances of data (lags) and forecast the next 3 instances (forecasts). ","metadata":{}},{"cell_type":"code","source":"lags=8\nforecasts=3\nX, y = [], []\nfor i in range(len(df_train_scaled) - forecasts-lags):\n  X.append(df_train_scaled[i:(i + lags)])\n  y.append(df_train_scaled[:,-1][i+lags: i+lags+ forecasts])\nX,y =  np.array(X), np.array(y)\ndf_train_scaled.shape, X.shape, y.shape,","metadata":{"execution":{"iopub.status.busy":"2022-07-08T09:16:22.273385Z","iopub.execute_input":"2022-07-08T09:16:22.274370Z","iopub.status.idle":"2022-07-08T09:16:22.294652Z","shell.execute_reply.started":"2022-07-08T09:16:22.274334Z","shell.execute_reply":"2022-07-08T09:16:22.293175Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"#### **<span style = 'color:purple'>Train- Validation Split</span>**\nThe training will occur on 8/10th of the data, reserving the last 2/10th for validation.","metadata":{}},{"cell_type":"code","source":"size = int(len(X) * 0.80)\nX_train = X[0:size]\ny_train = y[0:size]\nX_val = X[size:len(X)]\ny_val = y[size:len(X)]\nX_train.shape, y_train.shape, X_val.shape, y_val.shape","metadata":{"execution":{"iopub.status.busy":"2022-07-08T09:16:22.296362Z","iopub.execute_input":"2022-07-08T09:16:22.297511Z","iopub.status.idle":"2022-07-08T09:16:22.308367Z","shell.execute_reply.started":"2022-07-08T09:16:22.297461Z","shell.execute_reply":"2022-07-08T09:16:22.307123Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"X_train, y_train, X_val, y_val[0:5]","metadata":{"execution":{"iopub.status.busy":"2022-07-08T09:16:22.310061Z","iopub.execute_input":"2022-07-08T09:16:22.311066Z","iopub.status.idle":"2022-07-08T09:16:22.325378Z","shell.execute_reply.started":"2022-07-08T09:16:22.311029Z","shell.execute_reply":"2022-07-08T09:16:22.324014Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"\n### **<span style = 'color:brown'>5.2 Model: Bidirectional LSTM Architecture</span>**<a id ='Bidirectional'></a>\n[<div style=\"text-align: right\"> Back to Table of contents</div>](#Table)\nWhile Uni directional LSTM RNN calculates the results of forwarding LSTM, Bi-Directional LSTM RNN deals with both forward and backward LSTM. A Bidirectional RNN is a combination of two RNNs training the network in opposite directions, one from the beginning to the end of a sequence, and the other, from the end to the beginning of a sequence, ie., past to future and future to past. This is achieved by adding one more LSTM layer, which reverses the direction of information flow by which the input sequence flows backward in the additional LSTM layer. Here in the below codes a bi-LSTM layer using keras is added in to a regular neural network. This is done by the following way: \n\n1. import the Bidirectional class and LSTM class provided by keras to implement Bi-LSTM in keras.\n2. Next step is to wrap the LSTM layer inside the Bidirectional class.","metadata":{}},{"cell_type":"code","source":"from keras.models import Sequential\nfrom keras.layers import Dense\nfrom keras.layers import SimpleRNN\n#from keras.layers import GRU\nfrom keras.layers import LSTM\nfrom keras.layers import Dropout\nimport tensorflow as tf\nfrom tensorflow.keras.layers import Bidirectional\n\nn_steps = X_train.shape[-2]\nn_features =X_train.shape[-1]\ninput_shape=(n_steps, n_features)\nmodel_bi = Sequential([Bidirectional(LSTM(256, return_sequences=True), input_shape= input_shape),\n    Dense(20, activation='tanh'),\n    Bidirectional(LSTM(128,return_sequences=True, activation = 'tanh')),\n    Dense(20, activation='tanh'),\n    Bidirectional(LSTM(128,return_sequences=False, activation = 'tanh')),\n    Dense(20, activation='tanh'),\n    tf.keras.layers.Dropout(0.20),\n    Dense(units=3, activation = 'linear'),\n])\nmodel_bi.compile(optimizer='adam', loss='mse')\nmodel_bi.summary()\ntf.keras.utils.plot_model(model_bi)","metadata":{"execution":{"iopub.status.busy":"2022-07-08T09:16:22.326950Z","iopub.execute_input":"2022-07-08T09:16:22.327676Z","iopub.status.idle":"2022-07-08T09:16:25.903402Z","shell.execute_reply.started":"2022-07-08T09:16:22.327624Z","shell.execute_reply":"2022-07-08T09:16:25.901780Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### **<span style = 'color:brown'>5.2 Model: Unidirectional LSTM Architecture</span>**<a id ='Unidirectional'></a>\n[<div style=\"text-align: right\"> Back to Table of contents</div>](#Table)\nThe model consists of input, 2 hidden layers and 1 output layer with 3 dense units corresponding to 3 forecasts. The input to the hidden layer with 64 units comes from the input layer with defined shape and 128 units. All the nodes are fully connected. The output of the hidden layer will go to the second hidden layer with 64 units layer in the network, which is then going to the final and output layer with 3 nodes. All layers except the final one will have the same tanh activation. The activation function for the final dense output layer will be linear. The code for adding this layer is as follows:","metadata":{}},{"cell_type":"code","source":"from keras.models import Sequential\nfrom keras.layers import Dense\nfrom keras.layers import SimpleRNN\nfrom keras.layers import GRU\nfrom keras.layers import LSTM\nfrom keras.layers import Dropout\nimport tensorflow as tf\nfrom tensorflow.keras.optimizers import Adam\n\nn_steps = X_train.shape[-2]\nn_features =X_train.shape[-1]\n#input_shape=(n_steps, n_features)\nmodel=Sequential()\nmodel.add(LSTM(128,activation='tanh',return_sequences=True, input_shape=(n_steps, n_features)))\nmodel.add(Dropout(0.2))\nmodel.add(LSTM(64,activation='tanh',return_sequences=True))\nmodel.add(Dropout(0.2))\nmodel.add(LSTM(64,activation='tanh'))\nmodel.add(Dense(3, activation='linear'))\nmodel.compile(optimizer=Adam(learning_rate=0.0005), loss='mse')\nmodel.summary()\ntf.keras.utils.plot_model(model)","metadata":{"execution":{"iopub.status.busy":"2022-07-08T09:16:25.906073Z","iopub.execute_input":"2022-07-08T09:16:25.907020Z","iopub.status.idle":"2022-07-08T09:16:26.719781Z","shell.execute_reply.started":"2022-07-08T09:16:25.906964Z","shell.execute_reply":"2022-07-08T09:16:26.718339Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## **<span style = 'color:green'>6. Model Train, Evaluation & Final Report</span>**<a id ='Report'></a>\n[<div style=\"text-align: right\"> Back to Table of contents</div>](#Table) \n\n### **<span style = 'color:brown'>6.1 Train & evaluate the model</span>**<a id = 'Train-Evaluate'></a>\nDefine a function to train & evaluate the model.","metadata":{}},{"cell_type":"code","source":"def model_train_evaluation(y, X, model, model_name):\n    #Model run\n    from tensorflow.keras.callbacks import EarlyStopping\n    early_stop = EarlyStopping(monitor='val_loss', mode='min', verbose=1, patience=3)\n    #history = model.fit(X_train,y_train, epochs=6,validation_data=(X_val, y_val),callbacks=early_stop)\n    history = model.fit(X_train,y_train, epochs=6,validation_split=0.01, callbacks=early_stop)\n    history_frame = pd.DataFrame(history.history)\n    \n    # model performance plot\n    plt.figure(figsize=(20,5))\n    plt.plot(history.history['loss'], label='training loss')\n    plt.plot(history.history['val_loss'], label='validation loss')\n    plt.ylabel('loss')\n    plt.xlabel('epoch')\n    plt.legend(loc='best')\n    plt.title(model_name)\n    plt.show()\n    \n    # Model Evaluation metrics\n    from sklearn.metrics import mean_squared_error,mean_absolute_error,explained_variance_score, r2_score, mean_absolute_percentage_error\n    ypred =model.predict(X)\n    print(\"\\n \\n Model Evaluation Report: \")\n    print('Mean Absolute Error(MAE) of', model_name,':', mean_absolute_error(y, ypred))\n    print('Mean Squared Error(MSE) of', model_name,':', mean_squared_error(y, ypred))\n    print('Root Mean Squared Error (RMSE) of', model_name,':', mean_squared_error(y, ypred, squared = False))\n    print('Mean absolute percentage error (MAPE) of', model_name,':', mean_absolute_percentage_error(y, ypred))\n    print('Explained Variance Score (EVS) of', model_name,':', explained_variance_score(y, ypred))\n    print('R2 of', model_name,':', (r2_score(y, ypred)).round(2))\n    print('\\n \\n')\n    \n    # Actual vs Predicted Plot\n    f, ax = plt.subplots(figsize=(12,6),dpi=100);\n    plt.scatter(y, ypred, label=\"Actual vs Predicted\")\n    # Perfect predictions\n    plt.xlabel('Indoor Temperature in celsius')\n    plt.ylabel('Indoor Temperature in celsius')\n    plt.title('Expection vs Prediction')\n    plt.plot(y,y,'r', label=\"Perfect Expected Prediction\")\n    plt.legend()\n    f.text(0.95, 0.06, 'AUTHOR: RINI CHRISTY',\n         fontsize=12, color='green',\n         ha='left', va='bottom', alpha=0.5);\n\nmodel_train_evaluation(y_val, X_val,model_bi, 'Bidirectional LSTM model')","metadata":{"execution":{"iopub.status.busy":"2022-07-08T09:16:26.722381Z","iopub.execute_input":"2022-07-08T09:16:26.722882Z","iopub.status.idle":"2022-07-08T09:17:14.798475Z","shell.execute_reply.started":"2022-07-08T09:16:26.722831Z","shell.execute_reply":"2022-07-08T09:17:14.797202Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"model_train_evaluation(y_val, X_val,model, 'Unidirectional LSTM model')","metadata":{"execution":{"iopub.status.busy":"2022-07-08T09:17:14.800399Z","iopub.execute_input":"2022-07-08T09:17:14.800775Z","iopub.status.idle":"2022-07-08T09:17:36.999703Z","shell.execute_reply.started":"2022-07-08T09:17:14.800739Z","shell.execute_reply":"2022-07-08T09:17:36.998173Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### **<span style = 'color:brown'>6.2 Model Forecast</span>**<a id = 'Forecast'></a>\n[<div style=\"text-align: right\"> Back to Table of contents</div>](#Table)","metadata":{}},{"cell_type":"code","source":"df_test=pd.merge(df_test1['Id'],df_formatted_train_test,on='Id')\ndf_test=df_test[['Meteo_Sun_irradiance','hour','Outdoor_relative_humidity_Sensor']]\ndf_test","metadata":{"execution":{"iopub.status.busy":"2022-07-08T09:17:37.001572Z","iopub.execute_input":"2022-07-08T09:17:37.002108Z","iopub.status.idle":"2022-07-08T09:17:37.030772Z","shell.execute_reply.started":"2022-07-08T09:17:37.002041Z","shell.execute_reply":"2022-07-08T09:17:37.029192Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"train_test=df_train.append(df_test,ignore_index=True).fillna(0)\ntrain_test","metadata":{"execution":{"iopub.status.busy":"2022-07-08T09:17:37.033270Z","iopub.execute_input":"2022-07-08T09:17:37.033709Z","iopub.status.idle":"2022-07-08T09:17:37.060141Z","shell.execute_reply.started":"2022-07-08T09:17:37.033672Z","shell.execute_reply":"2022-07-08T09:17:37.058685Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"**Inverse Scale Transform**\n\nThe standard score of a sample x is calculated as:\n\n$ z = \\frac{(x - u)} s $\n\nSo, x can be inverse scale transformed as, \n\n$x = (z * s)+u$\n\nFor that, first find the mean and standard deviation of dependent column. \n","metadata":{}},{"cell_type":"code","source":"y_mean=df_train['Indoor_temperature_room'].mean()\ny_std=scaler.scale_[df_train.shape[1]-1]\ny_mean, y_std","metadata":{"execution":{"iopub.status.busy":"2022-07-08T09:17:37.061549Z","iopub.execute_input":"2022-07-08T09:17:37.061901Z","iopub.status.idle":"2022-07-08T09:17:37.072614Z","shell.execute_reply.started":"2022-07-08T09:17:37.061867Z","shell.execute_reply":"2022-07-08T09:17:37.071146Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Now, scale transform x to predict the unknown indoor temperature, which is then applied to the above formula to obtain the final forecasted result.\n\nStart with bidirectional model. ","metadata":{}},{"cell_type":"code","source":"for i in range(len(df_train),len(train_test)):\n    X_test=train_test[i-lags:i]\n    X_test=scaler.transform(X_test)\n    val=model_bi.predict(X_test.reshape(1,lags,X_test.shape[1]))\n    val2=y_std*val[:,0]+y_mean\n    train_test.loc[i,'Indoor_temperature_room']=val2\n    \nfinal_test = train_test[len(df_train):][['Indoor_temperature_room']]\nfinal_test","metadata":{"execution":{"iopub.status.busy":"2022-07-08T09:17:37.075019Z","iopub.execute_input":"2022-07-08T09:17:37.075949Z","iopub.status.idle":"2022-07-08T09:19:09.254711Z","shell.execute_reply.started":"2022-07-08T09:17:37.075890Z","shell.execute_reply":"2022-07-08T09:19:09.253467Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"#Plot the Graph\nplt.figure(figsize=(20,5))\nplt.plot(df_train['Indoor_temperature_room'], label='training values', color = 'blue')\nplt.plot(final_test['Indoor_temperature_room'], label='forcasting', color = 'red')\nplt.legend(loc='best')\nplt.title('Bidirectional LSTM Model')\nplt.show()","metadata":{"execution":{"iopub.status.busy":"2022-07-08T09:19:09.262964Z","iopub.execute_input":"2022-07-08T09:19:09.264203Z","iopub.status.idle":"2022-07-08T09:19:09.566768Z","shell.execute_reply.started":"2022-07-08T09:19:09.264160Z","shell.execute_reply":"2022-07-08T09:19:09.565437Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Now, its the turn of Unidirectional model.","metadata":{}},{"cell_type":"code","source":"for i in range(len(df_train),len(train_test)):\n    X_test=train_test[i-lags:i]\n    X_test=scaler.transform(X_test)\n    val=model.predict(X_test.reshape(1,lags,X_test.shape[1]))\n    val2=y_std*val[:,0]+y_mean\n    train_test.loc[i,'Indoor_temperature_room']=val2\n    \nfinal_test = train_test[len(df_train):][['Indoor_temperature_room']]\nfinal_test","metadata":{"execution":{"iopub.status.busy":"2022-07-08T09:19:09.568656Z","iopub.execute_input":"2022-07-08T09:19:09.569502Z","iopub.status.idle":"2022-07-08T09:20:37.367403Z","shell.execute_reply.started":"2022-07-08T09:19:09.569457Z","shell.execute_reply":"2022-07-08T09:20:37.365914Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"#Plot the Graph\nplt.figure(figsize=(20,5))\nplt.plot(df_train['Indoor_temperature_room'], label='training values', color = 'blue')\nplt.plot(final_test['Indoor_temperature_room'], label='forcasting', color = 'red')\nplt.title('Unidirectional LSTM Model')\nplt.legend(loc='best')\nplt.show()","metadata":{"execution":{"iopub.status.busy":"2022-07-08T09:20:37.369538Z","iopub.execute_input":"2022-07-08T09:20:37.369910Z","iopub.status.idle":"2022-07-08T09:20:37.682924Z","shell.execute_reply.started":"2022-07-08T09:20:37.369874Z","shell.execute_reply":"2022-07-08T09:20:37.681943Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### **<span style = 'color:brown'>6.3 Kaggle Submission</span>**<a id =\"Submission\"></a>\n[<div style=\"text-align: right\"> Back to Table of contents</div>](#Table)\n","metadata":{}},{"cell_type":"code","source":"final=pd.DataFrame({'Id':df_test1['Id'],'Indoor_temperature_room':np.array(train_test[len(df_train):]['Indoor_temperature_room'])})\nfinal.to_csv('submission.csv',index=False)","metadata":{"execution":{"iopub.status.busy":"2022-07-08T09:20:37.684393Z","iopub.execute_input":"2022-07-08T09:20:37.685493Z","iopub.status.idle":"2022-07-08T09:20:37.701050Z","shell.execute_reply.started":"2022-07-08T09:20:37.685446Z","shell.execute_reply":"2022-07-08T09:20:37.699403Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### **Concluding remarks:**\nIt seems Univariate model has worked comparitively well with the chosen variables 'Meteo_Sun_irradiance','hour','Outdoor_relative_humidity_Sensor', 'Indoor_temperature_room', with a lag of 8. This is confirmed by the error metrics such as MAE, MSE, RMSE, MAPE etc. Bidirectional model eventhough it achieved pretty good accuracy, it seems overfitting the model, with a slightly higher error than unidirectional.  \n\n### **References:**\n1. [Hands-on Time Series Analysis with Python By   B V Vishwas,  Ashish Patel](https://link.springer.com/book/10.1007/978-1-4842-5992-4)\n\n\n2. [Temperature Forecasting Analysis-part2 By PAVANKUMAR20](https://www.kaggle.com/code/pavankumar20/temperature-forecasting-analysis-part2)\n\nThanks for reading.","metadata":{}}]}