{
  "cells": [
    {
      "cell_type": "markdown",
      "metadata": {
        "_cell_guid": "bed0985f-a709-02da-e343-aa0b01b9d08a"
      },
      "source": [
        "NOAA Sea Lions - Simple Feature Prediction in R\n",
        "===============================\n",
        "\n",
        "This code will get you started in R. It extracts some basic features from the images.\n",
        "\n",
        "Step 1 - Packages and Variables\n",
        "----------------------------------\n",
        "\n",
        "We start by loading packages and setting some important variables.\n",
        "\n",
        "Note that the Kaggle kernel environment only has 10 images. I trust there are more images in the full data set."
      ]
    },
    {
      "cell_type": "code",
      "execution_count": null,
      "metadata": {
        "_cell_guid": "78d951e9-1c06-5d1b-deb9-867753197794"
      },
      "outputs": [],
      "source": [
        "# Load packages and set important constants\n",
        "packages <- c(\"readr\", \"dplyr\", \"purrr\", \"stringr\", \"xgboost\")\n",
        "purrr::walk(packages, library, character.only = TRUE, warn.conflicts = FALSE)\n",
        "\n",
        "\n",
        "\n",
        "train_filename <- \"../input/Train/train.csv\"\n",
        "#test_filename <- \"../input/test.csv\"\n",
        "train_image_directory <- \"../input/Train/\"\n",
        "train_image_marked_directory <- \"../input/TrainDotted/\"\n",
        "\n",
        "sample_submission_filename <- \"../input/sample_submission.csv\"\n",
        "results_path <- \"./\"\n",
        "sub_name <- paste0(results_path, \"median_submission.csv\")\n",
        "\n",
        "submission <- read_csv(sample_submission_filename)\n",
        "\n",
        "target_variables <- setdiff(names(submission), c('test_id'))\n",
        "train_variables <- c(target_variables)\n",
        "\n",
        "SEED <- 100\n",
        "\n"
      ]
    },
    {
      "cell_type": "markdown",
      "metadata": {
        "_cell_guid": "d8a10966-da07-31df-4d16-462f7124637e"
      },
      "source": [
        "Step 2 - Functions\n",
        "------------------\n",
        "\n",
        "Functions do all the work. Let's start with some basic functions to calculate features. The important function here is create_features. Add your favorite features and rise to the top of the leaderboard.\n",
        "\n",
        "This test example simply calculates the file size."
      ]
    },
    {
      "cell_type": "code",
      "execution_count": null,
      "metadata": {
        "_cell_guid": "0cdbbc19-5330-40ea-acb5-62a72f8e78f9"
      },
      "outputs": [],
      "source": [
        "get_feature_names <- function(train, test, train_variables = train_variables){\n",
        "  ## Get list of variables used in training\n",
        "  var_names <- intersect(names(as.data.frame(train)), names(as.data.frame(test)))\n",
        "  var_names <- setdiff(var_names, train_variables)\n",
        "  return(var_names)\n",
        "}\n",
        "\n",
        "create_features <- function(image_directory, train_file, target_vars, train = TRUE){\n",
        "    flist <- list.files(image_directory, pattern = \".jpg\", full.names = TRUE)\n",
        "    file_feature <- file.size(flist)\n",
        "    \n",
        "    train_data <- read_csv(train_file)\n",
        "    train_median <- apply(train_data[target_variables],2,median)\n",
        "    train_results <- rep(train_median, length(flist))\n",
        "    train_results_matrix <- matrix(train_results, nrow = length(flist), \n",
        "                                   ncol = length(train_median), \n",
        "                                   byrow = TRUE)\n",
        "    \n",
        "    results <- cbind(file_feature, train_results_matrix)\n",
        "    colnames(results) <- c('fsize', names(train_median))\n",
        "    \n",
        "    return(results)\n",
        "}\n",
        "\n",
        "build_median_model <- function(train_data, var_names, target_variables){\n",
        "    median <- train_data[1, target_variables]\n",
        "    return(median)\n",
        "}\n",
        "\n",
        "make_predictions <- function(model, target_variables, sample_submission){\n",
        "    preds <- sample_submission\n",
        "    for(var in target_variables){\n",
        "        preds[var] <- model[var]\n",
        "    }\n",
        "    return(preds)\n",
        "}"
      ]
    },
    {
      "cell_type": "markdown",
      "metadata": {
        "_cell_guid": "4c5a6304-a082-08d8-c328-172671fe5bc2"
      },
      "source": [
        "Step 3 Run the model and make predictions\n",
        "-----------------------------------------\n",
        "\n"
      ]
    },
    {
      "cell_type": "code",
      "execution_count": null,
      "metadata": {
        "_cell_guid": "4df829b0-3701-910d-a2a6-bc98b89d0031"
      },
      "outputs": [],
      "source": [
        "features <- create_features(train_image_directory, train_filename, target_variables)\n",
        "var_names <- get_feature_names(features, features, train_variables)\n",
        "model <- build_median_model(features, var_names, target_variables)\n",
        "sub <- make_predictions(model, target_variables, submission )\n",
        "write_csv(sub, sub_name)"
      ]
    }
  ],
  "metadata": {
    "_change_revision": 0,
    "_is_fork": false,
    "kernelspec": {
      "display_name": "R",
      "language": "R",
      "name": "ir"
    },
    "language_info": {
      "codemirror_mode": "r",
      "file_extension": ".r",
      "mimetype": "text/x-r-source",
      "name": "R",
      "pygments_lexer": "r",
      "version": "3.4.0"
    }
  },
  "nbformat": 4,
  "nbformat_minor": 0
}