{
  "cells": [
    {
      "cell_type": "markdown",
      "metadata": {
        "_cell_guid": "e04e8c91-580f-f64a-6ca7-772fb06b7a29"
      },
      "source": [
        "This notebooks looks at the published multi-log-loss evaluation function that is used for this competition."
      ]
    },
    {
      "cell_type": "code",
      "execution_count": null,
      "metadata": {
        "_cell_guid": "bf46132c-cab0-88f2-c961-4065420c9de4"
      },
      "outputs": [],
      "source": [
        "# This R environment comes with all of CRAN preinstalled, as well as many other helpful packages\n",
        "# The environment is defined by the kaggle/rstats docker image: https://github.com/kaggle/docker-rstats\n",
        "# For example, here's several helpful packages to load in \n",
        "\n",
        "library(ggplot2) # Data visualization\n",
        "library(readr) # CSV file I/O, e.g. the read_csv function\n",
        "library(dplyr)"
      ]
    },
    {
      "cell_type": "markdown",
      "metadata": {
        "_cell_guid": "736b0377-51c6-4241-12ef-605310f67762"
      },
      "source": [
        "There is a link in the overview to the multi log loss function. [https://www.kaggle.com/wiki/MultiClassLogLoss][1]\n",
        "\n",
        "Note I updated device_id to listing_id for this dataset.\n",
        "\n",
        "  [1]: https://www.kaggle.com/wiki/MultiClassLogLoss"
      ]
    },
    {
      "cell_type": "code",
      "execution_count": null,
      "metadata": {
        "_cell_guid": "c0975fdf-ebda-1860-eb91-1e8f58a80b65"
      },
      "outputs": [],
      "source": [
        "multiloss <- function(predicted, actual){\n",
        "  #to add: reorder the rows\n",
        "\n",
        "  predicted_m <- as.matrix(select(predicted, -listing_id))\n",
        "  # bound predicted by max value\n",
        "  predicted_m <- apply(predicted_m, c(1,2), function(x) max(min(x, 1-10^(-15)), 10^(-15)))\n",
        "\n",
        "  actual_m <- as.matrix(select(actual, -listing_id))\n",
        "  score <- -sum(actual_m*log(predicted_m))/nrow(predicted_m)\n",
        "\n",
        "  return(score)\n",
        "}"
      ]
    },
    {
      "cell_type": "markdown",
      "metadata": {
        "_cell_guid": "feb46bf4-a944-0d90-8762-ddd69e4927a9"
      },
      "source": [
        "If the actual output is going to be a 1.0 for the true class (e.g. 'medium') and then a 0.0 for the other 2 classes then the above function does not seem to be concerned with costing false class probabilities - it only seems to be concerned with costing the accuracy of the true class. Multiplying actual_m which will consist of zeros and ones will ignore anything on the prediction side that is multiplied by zero.\n",
        "\n",
        "That seems wrong so I built some data to test. The first example shows a fully correct prediction with a tiny error."
      ]
    },
    {
      "cell_type": "code",
      "execution_count": null,
      "metadata": {
        "_cell_guid": "a5593ac8-0c10-9ff5-ce50-9d329da80d6a"
      },
      "outputs": [],
      "source": [
        "listing_id = c(1001)\n",
        "low = c(0)\n",
        "medium = c(0)\n",
        "high = c(1)\n",
        "\n",
        "actual = data.frame(listing_id, high, medium, low)\n",
        "\n",
        "actual\n",
        "\n",
        "multiloss(actual, actual)"
      ]
    },
    {
      "cell_type": "markdown",
      "metadata": {
        "_cell_guid": "b05a1ea9-9541-e269-2709-19dc1aa51fec"
      },
      "source": [
        "This example shows a fully incorrect prediction - we have allocated a probability of 1.0 to the wrong class. There is a high cost to this misclassification."
      ]
    },
    {
      "cell_type": "code",
      "execution_count": null,
      "metadata": {
        "_cell_guid": "c2d7a181-a288-4617-0b92-9021ea061d10"
      },
      "outputs": [],
      "source": [
        "low = c(0)\n",
        "medium = c(1)\n",
        "high = c(0)\n",
        "\n",
        "pred_binary_wrong = data.frame(listing_id, high, medium, low)\n",
        "\n",
        "pred_binary_wrong\n",
        "multiloss(pred_binary_wrong, actual)"
      ]
    },
    {
      "cell_type": "markdown",
      "metadata": {
        "_cell_guid": "8804229b-dc3c-56bc-a3fd-58f9d64fe9de"
      },
      "source": [
        "The 3rd example simply submits probability of 1.0 for each class. It scores the same as the correct prediction."
      ]
    },
    {
      "cell_type": "code",
      "execution_count": null,
      "metadata": {
        "_cell_guid": "a638c0df-0759-409a-c53c-fb4de4d402d7"
      },
      "outputs": [],
      "source": [
        "low = c(1)\n",
        "medium = c(1)\n",
        "high = c(1)\n",
        "\n",
        "pred_all_1s = data.frame(listing_id, high, medium, low)\n",
        "\n",
        "pred_all_1s\n",
        "multiloss(pred_all_1s, actual)"
      ]
    },
    {
      "cell_type": "markdown",
      "metadata": {
        "_cell_guid": "fadfcdba-c1b8-f85b-e3e8-d2730ce97e92"
      },
      "source": [
        "I then submitted a file with a probability of 1.0 in every position to see what that would do. It got a public score of 1.09861 which means it is not behaving like the above.\n",
        "\n",
        "For paranoia I extend to 3 test cases below. Same behaviour."
      ]
    },
    {
      "cell_type": "code",
      "execution_count": null,
      "metadata": {
        "_cell_guid": "a711a51a-7548-b187-dc75-91b412a3af0a"
      },
      "outputs": [],
      "source": [
        "listing_id = c(1001, 1002, 1003)\n",
        "low = c(0, 0, 1)\n",
        "medium = c(0, 1, 0)\n",
        "high = c(1, 0, 0)\n",
        "\n",
        "actual = data.frame(listing_id, high, medium, low)\n",
        "\n",
        "actual\n",
        "\n",
        "multiloss(actual, actual)"
      ]
    },
    {
      "cell_type": "code",
      "execution_count": null,
      "metadata": {
        "_cell_guid": "d32a08a1-3a84-0e2b-8388-65c3f2f8ac82"
      },
      "outputs": [],
      "source": [
        "low = c(1, 1, 1)\n",
        "medium = c(1, 1, 1)\n",
        "high = c(1, 1, 1)\n",
        "\n",
        "all_ones = data.frame(listing_id, high, medium, low)\n",
        "\n",
        "all_ones\n",
        "\n",
        "multiloss(all_ones, actual)"
      ]
    },
    {
      "cell_type": "markdown",
      "metadata": {
        "_cell_guid": "80c84383-6bbf-6c1a-7140-64ee4bc17b16"
      },
      "source": [
        "Does this make sense? I imagine that knowing the sensitivities of the evaluation function could influence how predictions could be thresholded in order to achieve a better score?"
      ]
    }
  ],
  "metadata": {
    "_change_revision": 0,
    "_is_fork": false,
    "kernelspec": {
      "display_name": "R",
      "language": "R",
      "name": "ir"
    },
    "language_info": {
      "codemirror_mode": "r",
      "file_extension": ".r",
      "mimetype": "text/x-r-source",
      "name": "R",
      "pygments_lexer": "r",
      "version": "3.3.3"
    }
  },
  "nbformat": 4,
  "nbformat_minor": 0
}