{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"pygments_lexer":"ipython3","nbconvert_exporter":"python","version":"3.6.4","file_extension":".py","codemirror_mode":{"name":"ipython","version":3},"name":"python","mimetype":"text/x-python"}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"## Understanding Root Mean Squared Logarithmic Error (RMSLE)","metadata":{}},{"cell_type":"markdown","source":"In this kernel we take a deep dive into Root Mean Squared Logarithmic Error (RMSLE). This is the metric used for the [ASHRAE Energy Prediction competition](https://www.kaggle.com/c/ashrae-energy-prediction) and a common metric for regression problems. It is an extension on [Mean Squared Error (MSE)](https://en.wikipedia.org/wiki/Root-mean-square_deviation) that is mainly used when predictions have large deviations, which is the case with this energy prediction competition. Values range from 0 up to millions and we don't want to punish deviations in prediction as much as with MSE.\n\nWe will explore the metric itself as well as some of the best naive predictions that are possible.\n\nP.S. This kernel is the first in a series of kernels on competition metrics. Feel free to check out [\"Episode Two\" of Understanding the metric on Quadratic Weighted Kappa](https://www.kaggle.com/carlolepelaars/understanding-the-metric-quadratic-weighted-kappa) and [\"Episode Three\" of Understanding the metric on Spearman's Rho](https://www.kaggle.com/carlolepelaars/understanding-the-metric-spearman-s-rho).","metadata":{}},{"cell_type":"code","source":"# Update scikit-learn to use squared=False for mean_squared_log_error\n!pip install -U scikit-learn","metadata":{"execution":{"iopub.status.busy":"2022-07-27T15:50:32.637544Z","iopub.execute_input":"2022-07-27T15:50:32.638029Z","iopub.status.idle":"2022-07-27T15:50:47.595721Z","shell.execute_reply.started":"2022-07-27T15:50:32.637948Z","shell.execute_reply":"2022-07-27T15:50:47.594585Z"},"_kg_hide-output":true,"_kg_hide-input":true,"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## Table Of Contents","metadata":{}},{"cell_type":"markdown","source":"- [Preparation](#1)\n- [The Metric (RMSLE)](#2)\n- [Best Baselines](#3)\n- [Submission with best constant](#4)","metadata":{}},{"cell_type":"markdown","source":"## Preparation <a id=\"1\"></a>","metadata":{}},{"cell_type":"code","source":"# Standard libraries\nimport os\nimport numpy as np\nimport pandas as pd\nimport random as rn\nimport tensorflow as tf\n\n# Visualization\n%matplotlib inline\nimport matplotlib.pyplot as plt\n\n# For calculating the metric\nfrom sklearn.metrics import mean_squared_log_error\n\n# Path specifications\nBASE_PATH = \"../input/ashrae-energy-prediction/\"\nTRAIN_PATH = BASE_PATH + \"train.csv\"\nSAMP_SUB_PATH = BASE_PATH + \"sample_submission.csv\"\n\n# Seed for reproducability\nseed = 1234\nrn.seed(seed)\nnp.random.seed(seed)\nos.environ['PYTHONHASHSEED'] = str(seed)","metadata":{"_kg_hide-output":true,"_kg_hide-input":true,"execution":{"iopub.status.busy":"2022-07-27T15:52:04.803635Z","iopub.execute_input":"2022-07-27T15:52:04.804139Z","iopub.status.idle":"2022-07-27T15:52:12.534015Z","shell.execute_reply.started":"2022-07-27T15:52:04.804030Z","shell.execute_reply":"2022-07-27T15:52:12.532753Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Read in data\ndf = pd.read_csv(TRAIN_PATH)\n# Remove outliers\ndf = df[df['meter_reading'] < 250000]","metadata":{"_kg_hide-input":true,"_kg_hide-output":true,"execution":{"iopub.status.busy":"2022-07-27T15:52:12.536425Z","iopub.execute_input":"2022-07-27T15:52:12.537403Z","iopub.status.idle":"2022-07-27T15:52:29.444629Z","shell.execute_reply.started":"2022-07-27T15:52:12.537365Z","shell.execute_reply":"2022-07-27T15:52:29.443069Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## The metric (RMSLE) <a id=\"2\"></a>","metadata":{}},{"cell_type":"markdown","source":"The Root Mean Squared Log Error (RMSLE) can be defined using a slight modification on sklearn's `mean_squared_log_error` function, which itself a modification on the familiar [Mean Squared Error (MSE)](https://scikit-learn.org/stable/modules/generated/sklearn.metrics.mean_squared_error.html) metric.\n\nThe formula for RMSLE is represented as follows:\n\nRMSLE = $\\sqrt{\\frac{1}{n} \\sum_{i=1}^n (\\log(p_i + 1) - \\log(a_i+1))^2 }$\n\nWhere:\n\n$n$ is the total number of observations in the (public/private) data set,\n\n$p_i$ is your prediction of target, and\n\n$a_i$ is the actual target for $i$.\n\n$log(x)$ is the natural logarithm of $x$ ($log_e(x)$.\n\n[Formula source: Kaggle Evaluation page for the ASHRAE 2019 competition](https://www.kaggle.com/c/ashrae-energy-prediction/overview/evaluation)","metadata":{}},{"cell_type":"markdown","source":"2022 UPDATE: Credits to [David Gilbertson](https://www.kaggle.com/davidg707) for suggesting some optimizations since this notebook was created:\n1. `squared=False` is available for `mean_squared_log_error` in sklearn since 2021. This makes the need for wrapping `np.sqrt` unnecessary.\n2. The NumPy solution can be computed in a completely vectorized way. This makes the NumPy implementation a lot faster than the scikit-learn implementation.","metadata":{}},{"cell_type":"code","source":"def RMSLE(y_true: np.array, y_pred: np.array) -> np.float64:\n    \"\"\"\n    The Root Mean Squared Log Error (RMSLE) metric \n        \n    :param y_true: The ground truth labels given in the dataset\n    :param y_pred: Our predictions\n    :return: The RMSLE score\n    \"\"\"\n    return mean_squared_log_error(y_true, y_pred, squared=False)","metadata":{"_uuid":"d629ff2d2480ee46fbb7e2d37f6b5fab8052498a","_cell_guid":"79c7e3d0-c299-4dcb-8224-4455121ee9b0","execution":{"iopub.status.busy":"2022-07-27T15:52:29.445960Z","iopub.execute_input":"2022-07-27T15:52:29.447054Z","iopub.status.idle":"2022-07-27T15:52:29.453237Z","shell.execute_reply.started":"2022-07-27T15:52:29.447014Z","shell.execute_reply":"2022-07-27T15:52:29.452406Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"def NumPyRMSLE(y_true: list, y_pred: list) -> float:\n    \"\"\"\n    The Root Mean Squared Log Error (RMSLE) metric using only NumPy\n    \n    :param y_true: The ground truth labels given in the dataset\n    :param y_pred: Our predictions\n    :return: The RMSLE score\n    \"\"\"\n    n = len(y_true)\n    msle = np.sqrt(np.mean(np.square(np.log1p(y_pred) - np.log1p(y_true))))\n    return msle","metadata":{"execution":{"iopub.status.busy":"2022-07-27T15:52:29.455223Z","iopub.execute_input":"2022-07-27T15:52:29.456264Z","iopub.status.idle":"2022-07-27T15:52:29.467509Z","shell.execute_reply.started":"2022-07-27T15:52:29.456225Z","shell.execute_reply":"2022-07-27T15:52:29.466236Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"RMSLE For Tensorflow (From [Shashi Prakash Tripathi](https://www.kaggle.com/shishu1421) in the comments)","metadata":{}},{"cell_type":"code","source":"def RMSLETF(y_pred:tf.Tensor, y_true:tf.Tensor) -> tf.float64:\n    \"\"\"\n    The Root Mean Squared Log Error (RMSLE) metric for TensorFlow / Keras\n        \n    :param y_true: The ground truth labels given in the dataset\n    :param y_pred: Predicted values\n    :return: The RMSLE score\n    \"\"\"\n    y_pred = tf.cast(y_pred, tf.float64)\n    y_true = tf.cast(y_true, tf.float64) \n    y_pred = tf.nn.relu(y_pred) \n    return tf.sqrt(tf.reduce_mean(tf.squared_difference(tf.log1p(y_pred), tf.log1p(y_true))))","metadata":{"execution":{"iopub.status.busy":"2022-07-27T15:52:29.469302Z","iopub.execute_input":"2022-07-27T15:52:29.469701Z","iopub.status.idle":"2022-07-27T15:52:29.479989Z","shell.execute_reply.started":"2022-07-27T15:52:29.469668Z","shell.execute_reply":"2022-07-27T15:52:29.478531Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## Best baselines <a id=\"3\"></a>","metadata":{}},{"cell_type":"markdown","source":"### 1. Constant predictions","metadata":{}},{"cell_type":"markdown","source":"Making a constant prediction is a good way to get a sense of what it means to have good performance on model. You can build a complex model that can get 95% accuracy, but if a constant prediction will give you that as well than 95% doesn't look so good anymore. A metric like [AUC (Area under ROC Curve)](https://scikit-learn.org/stable/modules/generated/sklearn.metrics.roc_auc_score.html) will for example give a more honest representation for constant predictions, which is 0.5 and the same for random predictions.\n\nThe best constant score for RMSLE is the exponential of the mean of the log target values. It can be expressed as a formula in the following way.\n\n#### Best constant for RMSLE = $\\mathrm{e}^{mean(log_e(targets))}$\n\nWe will derive this conclusion for ourselves by iteratively increasing the constant value and checking where the RMSLE score is the lowest.","metadata":{}},{"cell_type":"code","source":"mean = np.mean(df['meter_reading'])\n\nprint(f\"RMSLE for predicting only 0: {round(RMSLE(df['meter_reading'], np.zeros(len(df))), 5)}\")\nprint(f\"RMSLE for predicting only 1: {round(RMSLE(df['meter_reading'], np.ones(len(df))), 5)}\")\nprint(f\"RMSLE for predicting only 50: {round(RMSLE(df['meter_reading'], np.full(len(df), 50)), 5)}\")\nprint(f\"RMSLE for predicting the mean ({round(mean, 2)}): {round(RMSLE(df['meter_reading'], np.full(len(df), mean)), 5)}\")","metadata":{"_kg_hide-input":true,"execution":{"iopub.status.busy":"2022-07-27T15:52:29.481897Z","iopub.execute_input":"2022-07-27T15:52:29.483026Z","iopub.status.idle":"2022-07-27T15:52:33.089317Z","shell.execute_reply.started":"2022-07-27T15:52:29.482990Z","shell.execute_reply":"2022-07-27T15:52:33.088220Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"const_rmsles = dict()\nfor i in range(75):\n    const = i * 2\n    rmsle = round(RMSLE(df['meter_reading'], np.full(len(df), const)), 5)\n    const_rmsles[const] = rmsle\n\nxs = list(const_rmsles.keys())\nys = list(const_rmsles.values())\n\npd.DataFrame(ys, index=xs).plot(figsize=(15, 10), legend=None)\nplt.scatter(min(const_rmsles, key=const_rmsles.get), sorted(ys)[0], color='red')\nplt.title(\"RMSLE scores for constant predictions\", fontsize=18, weight='bold')\nplt.xticks(fontsize=14)\nplt.xlabel(\"Constant\", fontsize=14)\nplt.ylabel(\"RMSLE\", rotation=0, fontsize=14);","metadata":{"_kg_hide-input":true,"execution":{"iopub.status.busy":"2022-07-27T15:52:33.090629Z","iopub.execute_input":"2022-07-27T15:52:33.091051Z","iopub.status.idle":"2022-07-27T15:53:45.488942Z","shell.execute_reply.started":"2022-07-27T15:52:33.091020Z","shell.execute_reply":"2022-07-27T15:53:45.487564Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"In the graph above we can see that the RMSLE score is the lowest around 60. However, you probably would like to know the exact value. This can easily be calculate using familiar NumPy functions.","metadata":{}},{"cell_type":"code","source":"# Formulate the best constant for this metric\nbest_const = np.expm1(np.mean(np.log1p(df['meter_reading'])))","metadata":{"_kg_hide-output":true,"execution":{"iopub.status.busy":"2022-07-27T15:53:45.490997Z","iopub.execute_input":"2022-07-27T15:53:45.491453Z","iopub.status.idle":"2022-07-27T15:53:45.943702Z","shell.execute_reply.started":"2022-07-27T15:53:45.491394Z","shell.execute_reply":"2022-07-27T15:53:45.942535Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"print(f\"The best constant for our data is: {best_const}...\")\nprint(f\"RMSLE for predicting the best possible constant on our data: {round(RMSLE(df['meter_reading'], np.full(len(df), best_const)), 5)}\\n\")\n\nprint(\"This is the optimal RMSLE score that we can get with only a constant prediction and using all data available.\\n\\\nWe therefore call it the best 'Naive baseline'\\n\\\nA model should at least perform better than this RMSLE score.\")","metadata":{"_kg_hide-input":true,"execution":{"iopub.status.busy":"2022-07-27T15:53:45.945543Z","iopub.execute_input":"2022-07-27T15:53:45.945966Z","iopub.status.idle":"2022-07-27T15:53:46.916544Z","shell.execute_reply.started":"2022-07-27T15:53:45.945923Z","shell.execute_reply":"2022-07-27T15:53:46.915394Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### 2. Random Predictions","metadata":{}},{"cell_type":"markdown","source":"It can be interesting to check the score that random predictions will give you. This will for example help you identify when your predictions are shuffled accidentally or when there is something wrong with your model. For this competition there is no clear \"maximum value\" that would make sense to predict. We therefore will try out different maximum thresholds for random predictions. The minimum value will stay 0 for our experiment.","metadata":{}},{"cell_type":"code","source":"# Random predictions\nrand_rmsles = dict()\nfor i in range(15):\n    magn = 10**(0.2 * (i + 1))\n    rand_preds = np.random.randint(0, magn, len(df))\n    rmsle = round(RMSLE(df['meter_reading'], rand_preds), 5)\n    rand_rmsles[magn] = rmsle\n\nxs = list(rand_rmsles.keys())\nys = list(rand_rmsles.values())  \n    \npd.DataFrame(ys, index=xs).plot(figsize=(15, 10), legend=None)\nplt.scatter(min(rand_rmsles, key=rand_rmsles.get),sorted(ys)[0],color='red')\nplt.title(\"RMSLE scores for random predictions\", fontsize=18, weight='bold')\nplt.xticks(fontsize=14)\nplt.xlabel(\"Maximum value\", fontsize=14)\nplt.ylabel(\"RMSLE\", rotation=0, fontsize=14);","metadata":{"_kg_hide-input":true,"execution":{"iopub.status.busy":"2022-07-27T15:51:09.174768Z","iopub.status.idle":"2022-07-27T15:51:09.175491Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"We can conclude that selecting random values between 0 and 160 will yield close to optimal performance regarding random predictions. Note that the best RMSLE score for random predictions (around 2.34) is not better than the best constant prediction. It can nonetheless still be interesting to analyze what we can expect from predicting random values.","metadata":{}},{"cell_type":"markdown","source":"## Submission with best constant <a id=\"4\"></a>","metadata":{}},{"cell_type":"code","source":"# Read in sample submission and fill all predictions with the best constant\nsamp_sub = pd.read_csv(SAMP_SUB_PATH)\nsamp_sub['meter_reading'] = best_const\nsamp_sub.to_csv(\"best_constant_submission.csv\", index=False)","metadata":{"execution":{"iopub.status.busy":"2022-07-27T15:51:09.176983Z","iopub.status.idle":"2022-07-27T15:51:09.177564Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Check Final Submission\nprint(\"Final Submission:\")\nsamp_sub.head(2)","metadata":{"_kg_hide-input":true,"execution":{"iopub.status.busy":"2022-07-27T15:51:09.179292Z","iopub.status.idle":"2022-07-27T15:51:09.179799Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"**That's it! If you like this Kaggle kernel, feel free to give an upvote and leave a comment! Your feedback is also very welcome! I will try to implement your suggestions in this kernel!**","metadata":{}}]}