{"cells":[{"metadata":{},"cell_type":"markdown","source":"# Simple Pytorch Model\n\nThis kernel is a simple basemodel to predict the unavailability of cars in a car rental agency using Pytorch.\nWhen a car is unavailable, its status can be either “in maintenance” or “being washed”. The goal is to predict the number of cars entering and leaving each of the two status, for each of the four shifts in a day, for one week ahead.\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F210371%2Fa725239bd0154225ea57089b917ea9ae%2Flocaliza_exemplo.png?generation=1594062202781499&alt=media)\n\n* Prepara Dataset\n* Model\n* Train\n* Prepare Submission","execution_count":null},{"metadata":{"_uuid":"8f2839f25d086af736a60e9eeb907d3b93b6e0e5","_cell_guid":"b1076dfc-b9ad-4769-8c92-a6c4dae69d19","trusted":true},"cell_type":"code","source":"import numpy as np # linear algebra\nimport pandas as pd # data processing, CSV file I/O (e.g. pd.read_csv)\nimport missingno as msno\nimport os\nimport torch\nimport gc\n","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"# determine the supported device\ndef get_device():\n    if torch.cuda.is_available():\n        device = torch.device('cuda:0')\n    else:\n        device = torch.device('cpu') # don't have GPU \n    return device\n\ndevice = get_device()","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"!pip install torchbearer\n!pip install livelossplot","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"## Prepara Dataset","execution_count":null},{"metadata":{},"cell_type":"markdown","source":"We are using only 2018 data for train, it's because of the limitation of RAM memory here, but it's ok.","execution_count":null},{"metadata":{"trusted":true},"cell_type":"code","source":"import glob\nimport pandas as pd\nfrom typing import List\n\ndef read_csv(path, re, dtype = None):\n    all_files = glob.glob(os.path.join(path, re))     # advisable to use os.path.join as this makes concatenation OS independent\n\n    df = (pd.read_csv(f, dtype=dtype) for f in all_files)\n    df = pd.concat(df, ignore_index=True)\n    return df\n\ndf = read_csv('/kaggle/input/kddbr-2020/', \"2018*.csv\")\ndf.head()","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"df.info()","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"df.shape","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"The problem contains multiples inputs and output...\n\nThe input variables are 557 features that may be correlated to the target variables. They are labeled from 0 up to 556, with the first four being categorical variables, and the rest being numerical. For the numerical variables, all the previous fourteen (14) values are provided (one for each day in two weeks before the tuple date). The input column names follow the syntax below:\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F210371%2Faf05e6e353863f45c6ce70e4618be43f%2Fdatakdd1.png?generation=1594062355235890&alt=media)\n\nThe output variables consist of a set of 16 (sixteen) columns, representing each of the 7 (seven) days ahead of the tuple date, 4 (four) shifts a day and 4 (four) status transitions a day, two for each status (entering or leaving). The column names follow the syntax detailed below:\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F210371%2F9cf037d978a6ff49919c3c1e6cadafbd%2Fdatakdd2.png?generation=1594062372260850&alt=media)","execution_count":null},{"metadata":{"trusted":true},"cell_type":"code","source":"input_columns  = df.columns[df.columns.str.contains(\"input\")]\noutput_columns = df.columns[df.columns.str.contains(\"output\")]\n\nprint(input_columns[:10], output_columns[:10])","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"Add timestamp information","execution_count":null},{"metadata":{"trusted":true},"cell_type":"code","source":"## Extract date information\n\ndef get_date_features(df, column):\n    df['input_dt_sin_quarter']     = np.sin(2*np.pi*df[column].dt.quarter/4)\n    df['input_dt_sin_day_of_week'] = np.sin(2*np.pi*df[column].dt.dayofweek/6)\n    df['input_dt_sin_day_of_year'] = np.sin(2*np.pi*df[column].dt.dayofyear/365)\n    df['input_dt_sin_day']         = np.sin(2*np.pi*df[column].dt.day/30)\n    df['input_dt_sin_month']       = np.sin(2*np.pi*df[column].dt.month/12)\n    \n    \ndf['date']  = pd.to_datetime(df['date'])\nget_date_features(df, 'date')\ndf.head()","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"input_columns  = df.columns[df.columns.str.contains(\"input\")]\ninput_columns","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"We need observe if have null columns:","execution_count":null},{"metadata":{"trusted":true},"cell_type":"code","source":"# Filter onlu nissing values\nnull_columns=df.columns[df.isnull().any()]\nmsno.bar(df.sample(1000)[null_columns])","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"null_columns","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"We will fill with zero values.","execution_count":null},{"metadata":{"trusted":true},"cell_type":"code","source":"df = df.fillna(0)","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"df[input_columns].shape","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"gc.collect()","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"import torch\nfrom torch.utils.data import DataLoader\nfrom torch.utils.data.dataset import Dataset\n\nclass CustomDataset(Dataset):\n    def __init__(self, data, targets):\n        self.data  = data\n        self.targets = targets\n\n    def __len__(self):\n        return len(self.data)\n    \n    def __getitem__(self, idx):\n        if torch.is_tensor(idx):\n            idx = idx.tolist()\n            \n        input  = self.data[idx]\n        target = self.targets[idx]\n        \n        return torch.Tensor(input.astype(float)).to(device), torch.Tensor(target.astype(float)).to(device)\n    \ninput, output = CustomDataset(df[input_columns].values, df[output_columns].values).__getitem__(1)\n[input, output]","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"## Model\n\nWe are using an architecture encoder-decoder to predict multiple outputs","execution_count":null},{"metadata":{"trusted":true},"cell_type":"code","source":"import torch\nimport torch.nn as nn\nimport torch.nn.functional as F\n\nclass EncoderDecoder(nn.Module):\n    def __init__(self, input_size, n_factors, output_size):\n        super(EncoderDecoder,self).__init__()\n        \n        self.encoder = nn.Sequential(\n            nn.Linear(input_size, int(input_size/2)),\n            nn.ReLU(True),\n            nn.Linear(int(input_size/2), int(input_size/4)),\n            nn.ReLU(True),\n            nn.Linear(int(input_size/4), n_factors),\n            nn.ReLU(True))\n        \n        self.decoder = nn.Sequential(\n            nn.Linear(n_factors, int(output_size/2)),\n            nn.ReLU(True),\n            nn.Linear(int(output_size/2), output_size),\n            nn.ReLU(True),            \n            nn.Linear(output_size, output_size))\n        \n        self.dropout = torch.nn.Dropout(0.2)\n        \n    def forward(self,x):\n        x = self.encoder(x)\n        x = self.dropout(x)\n        x = self.decoder(x)\n        return x","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"# Params\ninput_size  = len(input_columns) \noutput_size = len(output_columns)\nn_factors   = 50\n\n\n# Model\nmodel       = EncoderDecoder(input_size, n_factors, output_size).to(device)\nmodel","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"## Train","execution_count":null},{"metadata":{"trusted":true},"cell_type":"code","source":"#device      = torch.device(\"cuda\" if torch.cuda.is_available() else \"cpu\")\n\n# Train Params\nbatch_size  = 128\nnum_epochs  = 30\n\n\n# Loss and optimizer\n\n# create a nn class (just-for-fun choice :-) \nclass RMSELoss(nn.Module):\n    def __init__(self):\n        super().__init__()\n        self.mse = nn.MSELoss()\n        \n    def forward(self,yhat,y):\n        return torch.sqrt(self.mse(yhat,y))\n\ncriterion   = RMSELoss()\noptimizer   = torch.optim.Adam(model.parameters(),weight_decay=1e-5)\n\nmodel","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"### Split Dataset Train/test","execution_count":null},{"metadata":{"trusted":true},"cell_type":"code","source":"from sklearn.model_selection import train_test_split\n# Data Loader\nclass NoAutoCollationDataLoader(DataLoader):\n    @property\n    def _auto_collation(self):\n        return False\n\n    @property\n    def _index_sampler(self):\n        return self.batch_sampler\n    \n# Split dataset\n#input_mean[np.isnan(input_mean)] = 0\nX = df[input_columns].values\nY = df[output_columns].fillna(0).values\n\nx_train, x_val, y_train, y_val = train_test_split(X, Y, test_size=0.20, random_state=42)\nlen(x_train), len(x_val)","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"# Pytorch Data Loader\ntrain_loader = NoAutoCollationDataLoader(CustomDataset(x_train, y_train), batch_size=batch_size, shuffle=True)\nval_loader   = NoAutoCollationDataLoader(CustomDataset(x_val, y_val), batch_size=batch_size, shuffle=False)","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"### Train \n\nTrain Pytorch Model ","execution_count":null},{"metadata":{"trusted":true},"cell_type":"code","source":"import torchbearer\nfrom torchbearer import Trial\nfrom torchbearer.callbacks import LiveLossPlot\nfrom torchbearer.callbacks.checkpointers import ModelCheckpoint\nfrom torchbearer.callbacks.early_stopping import EarlyStopping\n\nweigth_path = \"/kaggle/working/weights.pt\"\ncallbacks = [\n                LiveLossPlot(), \n                EarlyStopping(min_delta=1e-3, patience=10, monitor='val_loss', mode='min'),\n                ModelCheckpoint(weigth_path, save_best_only=True, monitor='val_loss', mode='min')\n            ]","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"%matplotlib inline\n\ntrial = Trial(model, optimizer, criterion, metrics=['loss'], callbacks=callbacks).to(device)\ntrial.with_generators(train_generator=train_loader, val_generator=val_loader)\nhist = trial.run(epochs=num_epochs)","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"!ls /kaggle/working/","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"### Evaluation","execution_count":null},{"metadata":{"trusted":true},"cell_type":"code","source":"## Load Model\n\nstate_dict = torch.load(weigth_path, map_location=device)\nmodel.load_state_dict(state_dict[\"model\"])\n#model.eval()","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"# Predict\nval_predictions = model(torch.Tensor(x_val).to(device)).detach().cpu().numpy()","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"# Load trained model\n\n#val_loader   = NoAutoCollationDataLoader(CustomDataset(x_val, y_val), batch_size=batch_size, shuffle=False)\ntrial           = Trial(model, optimizer, criterion).with_test_generator(val_loader).to(device)\nval_predictions = trial.predict().detach().cpu().numpy()#.reshape(-1)\n\nval_predictions.shape, y_val.shape","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"#  Weighted Root Mean Squared Error (WRMSE),\n# 0\t1.00\n# 1\t0.75\n# 2\t0.60\n# 3\t0.50\n# 4\t0.43\n# 5\t0.38\n# 6\t0.33\n\nweigths = [1]*16 + [0.75]*16 + [0.6]*16 + [0.5]*16 + [0.43]*16 + [0.38]*16 + [0.33]*16\n\ndef wrmse(predictions, targets, weigths):\n    #Is it???\n    return np.sqrt((((predictions - targets) ** 2)*weigths).mean())\n\ndef rmse(predictions, targets):\n    return np.sqrt(((predictions - targets) ** 2).mean())\n\nerror_wrmse = wrmse(val_predictions, y_val, weigths)\nerror_rmse  = rmse(val_predictions, y_val)\n\nprint(\"WRMSE:\", error_wrmse)\nprint(\"RMSE:\", error_rmse)","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"Weighted Root Mean Squared Error (WRMSE) - Is it??? ","execution_count":null},{"metadata":{"trusted":true},"cell_type":"code","source":"# Clean Memory\ndel x_train\ndel x_val\ndel df\n\ndel train_loader\ndel val_loader\n\ngc.collect()","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"## Prepare Submission\n\nWe will use the model in 2019 dataset for submission...","execution_count":null},{"metadata":{"trusted":true},"cell_type":"code","source":"## Load Model\n\nstate_dict = torch.load(weigth_path, map_location=device)\nmodel.load_state_dict(state_dict[\"model\"])\nmodel.eval()","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"# Load and Prepare data\ndf_pred = read_csv('/kaggle/input/kddbr-2020/', \"public2019*.csv\").fillna(0)\n\n# transform\ndf_pred['date']  = pd.to_datetime(df_pred['date'])\nget_date_features(df_pred, 'date')\n\ndf_pred.head()","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"# Predict\ninputs = torch.Tensor(df_pred[input_columns].values).to(device)\npred   = model(inputs).cpu().detach().numpy()\npred","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"df_pred_sub = pd.DataFrame(pred)\ndf_pred_sub.columns = output_columns\ndf_pred_sub['id']   = df_pred['id']\ndf_pred_sub.head()","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"###","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"### Build Submission","execution_count":null},{"metadata":{"trusted":true},"cell_type":"code","source":"## submission\ndf_sub = []\nfor i, row in df_pred_sub.iterrows():\n    for column, value in zip(output_columns, row.values):\n        id = \"{}_{}\".format(int(row.id), column)\n        df_sub.append([id, value])\n\ndf_sub = pd.DataFrame(df_sub)\ndf_sub.columns = ['id', 'value']\ndf_sub.to_csv('/kaggle/working/submission.csv', index=False)\ndf_sub","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"","execution_count":null,"outputs":[]}],"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"pygments_lexer":"ipython3","nbconvert_exporter":"python","version":"3.6.4","file_extension":".py","codemirror_mode":{"name":"ipython","version":3},"name":"python","mimetype":"text/x-python"}},"nbformat":4,"nbformat_minor":4}