{
  "cells": [
    {
      "cell_type": "markdown",
      "metadata": {
        "_cell_guid": "100c9ac5-9885-55ef-e081-78daa303aa17"
      },
      "source": [
        "# Predicting MNIST labels with a CNN in Keras\n",
        "\n",
        "In this notebook we will analyze a bit the MNIST dataset, build a CNN with Keras, train it and finally test the accuracy of our predictions.\n",
        "\n",
        "But first, lets see how the training and testing data is organized:\n",
        "> The data files `train.csv` and `test.csv` contain gray-scale images of hand-drawn digits, from zero through nine.\n",
        "\n",
        ">Each image is 28 pixels in height and 28 pixels in width, for a total of 784 pixels in total. Each pixel has a single pixel-value associated with it, indicating the lightness or darkness of that pixel, with higher numbers meaning darker. This pixel-value is an integer between 0 and 255, inclusive.\n",
        "\n",
        ">The training data set, (`train.csv`), has 785 columns. The first column, called \"label\", is the digit that was drawn by the user. The rest of the columns contain the pixel-values of the associated image.\n",
        "\n",
        "\n",
        "So, lets import our dependencies:"
      ]
    },
    {
      "cell_type": "code",
      "execution_count": null,
      "metadata": {
        "_cell_guid": "37468289-7fb9-1784-95c7-2eae6019073f"
      },
      "outputs": [],
      "source": [
        "import numpy as np # linear algebra\n",
        "import pandas as pd # data processing, CSV file I/O (e.g. pd.read_csv)\n",
        "\n",
        "import matplotlib.pyplot as plt\n",
        "import warnings # current version of seaborn generates a bunch of warnings that we'll ignore\n",
        "warnings.filterwarnings(\"ignore\")\n",
        "import seaborn as sns\n",
        "%matplotlib inline\n",
        "\n",
        "from keras.models import Sequential\n",
        "from keras.layers import Conv2D, MaxPooling2D, Dropout, Activation, Flatten, Dense\n",
        "from keras.optimizers import Adam\n",
        "from sklearn.model_selection import train_test_split\n",
        "\n",
        "# Input data files are available in the \"../input/\" directory.\n",
        "# For example, running this (by clicking run or pressing Shift+Enter) will list the files in the input directory\n",
        "\n",
        "from subprocess import check_output\n",
        "print(\"Files in Input Directory:\")\n",
        "print(check_output([\"ls\", \"../input\"]).decode(\"utf8\")) \n",
        "\n",
        "# Any results you write to the current directory are saved as output."
      ]
    },
    {
      "cell_type": "code",
      "execution_count": null,
      "metadata": {
        "_cell_guid": "d809d2dd-b597-25e2-bc7f-569d508ce32f"
      },
      "outputs": [],
      "source": [
        "train = pd.read_csv(\"../input/train.csv\")\n",
        "test = pd.read_csv(\"../input/test.csv\")"
      ]
    },
    {
      "cell_type": "markdown",
      "metadata": {
        "_cell_guid": "6f992968-1758-7524-c7c2-a562fb3596db"
      },
      "source": [
        "Now that we have our raw data lets see how it looks like"
      ]
    },
    {
      "cell_type": "code",
      "execution_count": null,
      "metadata": {
        "_cell_guid": "6a52c6ee-a6c7-35b4-81b3-62256b9e5e7e"
      },
      "outputs": [],
      "source": [
        "train.head()"
      ]
    },
    {
      "cell_type": "markdown",
      "metadata": {
        "_cell_guid": "e6930817-6666-6bf8-cb0a-cb2907cb977d"
      },
      "source": [
        "We can see that our training data is simply a series of rows, where each row contains the true label of the image followed by a succession of 784 (= 28 * 28) values that indicate the darkness of each pixel.\n",
        "\n",
        "Now lets take a look at our testing data:"
      ]
    },
    {
      "cell_type": "code",
      "execution_count": null,
      "metadata": {
        "_cell_guid": "253353f3-ed24-7975-cdd4-eee741c74a60"
      },
      "outputs": [],
      "source": [
        "test.head()"
      ]
    },
    {
      "cell_type": "markdown",
      "metadata": {
        "_cell_guid": "1a2310c9-cd7d-ec30-b58d-41c907119827"
      },
      "source": [
        "The testing data looks the same as training data but it has no labels. This is ok since the whole idea of this is to predict which labels our testing data should have.\n",
        "\n",
        "Now, how many testing and training points do we have?"
      ]
    },
    {
      "cell_type": "code",
      "execution_count": null,
      "metadata": {
        "_cell_guid": "182eba00-3db5-a06c-d46b-0453e76a039c"
      },
      "outputs": [],
      "source": [
        "num_training = len(train.values)\n",
        "num_testing = len(test.values)\n",
        "\n",
        "print(\"Amount of training data:\", num_training, \"pairs of images and labels.\")\n",
        "print(\"Amount of testing data:\", num_testing, \"images.\")"
      ]
    },
    {
      "cell_type": "code",
      "execution_count": null,
      "metadata": {
        "_cell_guid": "da76bce5-9f5d-ab6f-b28d-17e04bd87390"
      },
      "outputs": [],
      "source": [
        "# Here we are defining the x & y variables for training and testing\n",
        "\n",
        "y_train = np.array(train.pop(\"label\").values) # array containing correct labels | Shape -> (42000,)\n",
        "x_train = np.array(train.values) # array of images for training | Shape -> (42000, 784)\n",
        "x_test = np.array(test.values) # array of images for testing | Shape -> (28000, 784)"
      ]
    },
    {
      "cell_type": "markdown",
      "metadata": {
        "_cell_guid": "26431226-84c6-9dec-1502-a6900f7898f2"
      },
      "source": [
        "Now we want to convert the pixeles of our `x_train` and `x_test` into 28x28 matrices, so we will end up with arrays of matrices. We could also proceed without doing this but I find it easier to think about 2D images than 1D images."
      ]
    },
    {
      "cell_type": "code",
      "execution_count": null,
      "metadata": {
        "_cell_guid": "b0b48198-0382-db17-17d3-bdd9ef7b5243"
      },
      "outputs": [],
      "source": [
        "x_train = x_train.reshape(num_training, 28, 28) # resulting shape => (42000, 28, 28)\n",
        "x_test = x_test.reshape(num_testing, 28, 28) # resulting shape => (28000, 28, 28)"
      ]
    },
    {
      "cell_type": "code",
      "execution_count": null,
      "metadata": {
        "_cell_guid": "6ebe89fa-2ac3-e803-f277-1bb602e840f1"
      },
      "outputs": [],
      "source": [
        "# Now lets take a look at some of our triaining images!\n",
        "\n",
        "plt.figure(figsize=(11,6))\n",
        "for i in range(66): \n",
        "    plt.subplot(6,11,i+1)\n",
        "    plt.imshow(x_train[i])\n",
        "    plt.xticks([])\n",
        "    plt.yticks([])\n",
        "    \n",
        "plt.tight_layout()"
      ]
    },
    {
      "cell_type": "markdown",
      "metadata": {
        "_cell_guid": "22185fe1-6a33-630a-ad14-c217feefbfc1"
      },
      "source": [
        "Great! It seems everything looks good, our initial data is in place and now we are ready to start. Lets begin by visualizing our training dataset with t-SNE."
      ]
    },
    {
      "cell_type": "code",
      "execution_count": null,
      "metadata": {
        "_cell_guid": "0db40a18-c1eb-3d73-5e7b-a2915c1c36a4"
      },
      "outputs": [],
      "source": [
        "# First of, lets standardize our features by removing the mean and scaling to unit variance\n",
        "rawx = train.values[:1000] # I don't have enough memory to run the tsne on all the data so here we'll limit it\n",
        "from sklearn.preprocessing import StandardScaler\n",
        "xscaled = StandardScaler().fit_transform(rawx)\n",
        "\n",
        "from sklearn.manifold import TSNE\n",
        "tsne = TSNE()\n",
        "vis = tsne.fit_transform(xscaled) \n",
        "vis = [{'X': vis[i][0], 'Y': vis[i][1], 'K': y_train[i]} for i in range(len(vis))] # transform to dict for plotting"
      ]
    },
    {
      "cell_type": "code",
      "execution_count": null,
      "metadata": {
        "_cell_guid": "05ef38ce-d22d-e0ae-9ebd-2fca321585fb"
      },
      "outputs": [],
      "source": [
        "sns.FacetGrid(pd.DataFrame.from_dict(vis), hue=\"K\", size=8) \\\n",
        "   .map(plt.scatter, \"X\", \"Y\") \\\n",
        "   .add_legend()"
      ]
    },
    {
      "cell_type": "markdown",
      "metadata": {
        "_cell_guid": "59937a5a-6442-f584-71bf-8af996952ebc"
      },
      "source": [
        "Looks pretty nice! We can see how t-sne separates our data points into individual clusters (kind of).\n",
        "\n",
        "Now lets plot a PCA visualization of the data:"
      ]
    },
    {
      "cell_type": "code",
      "execution_count": null,
      "metadata": {
        "_cell_guid": "a69a6cc6-17b2-a4ff-f635-ae6ba18b962d"
      },
      "outputs": [],
      "source": [
        "from sklearn.decomposition import PCA\n",
        "\n",
        "pca = PCA(n_components=5)\n",
        "pca_ani = pca.fit_transform(xscaled)\n",
        "\n",
        "xpca = pca_ani[:, 0]\n",
        "ypca = pca_ani[:, 1]\n",
        "vispca = [{'X': xpca[i], 'Y': ypca[i], 'K': y_train[i]} for i in range(len(xpca))]"
      ]
    },
    {
      "cell_type": "code",
      "execution_count": null,
      "metadata": {
        "_cell_guid": "c05c5a08-0e75-da0c-ee06-43f48c30a091"
      },
      "outputs": [],
      "source": [
        "sns.FacetGrid(pd.DataFrame.from_dict(vispca), hue=\"K\", size=8) \\\n",
        "   .map(plt.scatter, \"X\", \"Y\") \\\n",
        "   .add_legend()"
      ]
    },
    {
      "cell_type": "markdown",
      "metadata": {
        "_cell_guid": "48122ba8-cc29-d5f1-b626-b1f79e3b6688"
      },
      "source": [
        "## Preparing our data for the network.\n",
        "\n",
        "Now that we've meddled a bit with our data and gotten to know it pretty well we can begin to build our network. We will start by normalizing our X datapoints, then we'll proceed to hot-encode our `y_labels`.\n",
        "\n",
        "### Normalization"
      ]
    },
    {
      "cell_type": "code",
      "execution_count": null,
      "metadata": {
        "_cell_guid": "764ee8ed-7330-3e60-9e28-4658eb30a5f9"
      },
      "outputs": [],
      "source": [
        "# normalize X datapoints. We want them to be between 0 and 1\n",
        "x_train = x_train / 255. # Perform an elementwise division by 255.0 (the maximum possible value for each pixel)\n",
        "x_test = x_test / 255."
      ]
    },
    {
      "cell_type": "markdown",
      "metadata": {
        "_cell_guid": "a2cc7567-d676-6a21-3ec7-9fcca5dd6a4f"
      },
      "source": [
        "### One Hot Encode\n",
        "Now we will hot-encode all our `y_train` labels using sklearn's LabelBinarizer"
      ]
    },
    {
      "cell_type": "code",
      "execution_count": null,
      "metadata": {
        "_cell_guid": "e48bed33-5247-1b31-03dc-d39ad7ec2e34"
      },
      "outputs": [],
      "source": [
        "from sklearn.preprocessing import LabelBinarizer\n",
        "y_train_hot = LabelBinarizer().fit_transform(y_train)"
      ]
    },
    {
      "cell_type": "markdown",
      "metadata": {
        "_cell_guid": "b9eff15a-a1be-9388-8de0-8343d773aeda"
      },
      "source": [
        "## Neural Net with Keras\n",
        "\n",
        "Here is where we built our neural network. We'll be basing the model on LeNet-5 which is a simply yet effective CNN Yann LeCun exposed in his paper [Gradient-Based Learning Applied to Document Recognition](http://yann.lecun.com/exdb/publis/pdf/lecun-01a.pdf). We'll also be including [dropout](https://www.cs.toronto.edu/~hinton/absps/JMLRdropout.pdf) to prevent overfitting.\n",
        "\n",
        "Architecture of LeNet-5:\n",
        "<img src=\"http://eblearn.sourceforge.net/lib/exe/lenet5.png\" width=\"100%\"/>"
      ]
    },
    {
      "cell_type": "code",
      "execution_count": null,
      "metadata": {
        "_cell_guid": "045c6461-6f05-d7d3-0d22-c5792841fea0"
      },
      "outputs": [],
      "source": [
        "model = Sequential()\n",
        "\n",
        "# Convolutional Layer 1\n",
        "model.add(Conv2D(6, 5, 5, input_shape=(28,28,1), bias=True,\n",
        "                 border_mode=\"same\"))\n",
        "model.add(Activation('relu'))\n",
        "model.add(MaxPooling2D(pool_size=(2, 2), strides=(2,2), border_mode=\"same\"))\n",
        "\n",
        "# Dropout Layer 1\n",
        "model.add(Dropout(p=0.12))\n",
        "\n",
        "# Convolutional Layer 2\n",
        "model.add(Conv2D(16, 5, 5, bias=True,\n",
        "                 border_mode=\"same\"))\n",
        "model.add(Activation('relu'))\n",
        "model.add(MaxPooling2D(pool_size=(2, 2), strides=(2,2), border_mode=\"same\"))\n",
        "\n",
        "# Convolutional Layer 3\n",
        "model.add(Conv2D(35, 5, 5, bias=True,\n",
        "                 border_mode=\"same\"))\n",
        "model.add(Activation('relu'))\n",
        "model.add(MaxPooling2D(pool_size=(2, 2), strides=(2,2), border_mode=\"same\"))\n",
        "\n",
        "# Flatten convolutional result so we can feed data to fully connected layers\n",
        "model.add(Flatten())\n",
        "\n",
        "# Fully connected 1\n",
        "model.add(Dense(120, bias=True))\n",
        "model.add(Activation('relu'))\n",
        "\n",
        "# Dropout Layer 2\n",
        "model.add(Dropout(p=0.5))\n",
        "\n",
        "# Fully connected 1\n",
        "model.add(Dense(84, bias=True))\n",
        "model.add(Activation('relu'))\n",
        "\n",
        "model.add(Dense(10, bias=True))\n",
        "model.add(Activation('softmax'))\n",
        "\n",
        "model.compile(optimizer=Adam(lr=0.001),\n",
        "              loss=\"categorical_crossentropy\",\n",
        "              metrics=['accuracy'],\n",
        "              decay=1)"
      ]
    },
    {
      "cell_type": "markdown",
      "metadata": {
        "_cell_guid": "be6f4647-93ec-88b2-c88c-52e1a7b9d294"
      },
      "source": [
        "### Training the network\n",
        "Now that we have our model defined we'll train it! But before this we need to add an extra dimension to our input."
      ]
    },
    {
      "cell_type": "code",
      "execution_count": null,
      "metadata": {
        "_cell_guid": "e9a9e5a8-212d-eafb-399b-7a3da6757123"
      },
      "outputs": [],
      "source": [
        "# Add extra dimension: Keras' is expecting an input of shape (num_x_train, 28, 28, 1) \n",
        "# but for the moment the shape of x_train is (num_x_train, 28, 28). We can easily add this extra \n",
        "# dimension using np.expand_dims(axis=4).\n",
        "x_train = np.expand_dims(x_train, axis=4)\n",
        "\n",
        "# We'll also take advantage of the occassion to do the same modification to x_test\n",
        "x_test = np.expand_dims(x_test, axis=4)"
      ]
    },
    {
      "cell_type": "code",
      "execution_count": null,
      "metadata": {
        "_cell_guid": "9928103e-500d-1d4a-a28f-98e0de5cff1d"
      },
      "outputs": [],
      "source": [
        "# Now: Train the model! Remember we one-hot encoded y_train so we shold use y_train_hot as our labels.\n",
        "\n",
        "# Note: verbose=2 gives us one log line per epoch.\n",
        "training_hist = model.fit(x_train, y_train_hot, nb_epoch=16, batch_size=64, verbose=2, validation_split=0.23)"
      ]
    },
    {
      "cell_type": "markdown",
      "metadata": {
        "_cell_guid": "33f52545-18fb-96a4-a47f-f041fb60306f"
      },
      "source": [
        "The `training_hist` variable we assigned above contains the following for each epoch:\n",
        "- val_loss \n",
        "- val_acc\n",
        "- loss \n",
        "- acc\n",
        "\n",
        "Lets plot those to see how our network behaved during training:"
      ]
    },
    {
      "cell_type": "code",
      "execution_count": null,
      "metadata": {
        "_cell_guid": "9bba2205-6e9d-e1e9-401c-820ee0553dd6"
      },
      "outputs": [],
      "source": [
        "hdf = pd.DataFrame.from_dict(training_hist.history)\n",
        "hdf['epochs'] = list(range(16)) # 16 is the number of epochs we used for training"
      ]
    },
    {
      "cell_type": "markdown",
      "metadata": {
        "_cell_guid": "dc26e0e5-1043-6125-74cf-a88c241bd768"
      },
      "source": [
        "### Validation Accuracy (Yellow) and Test Accuracy (Red) VS Epochs"
      ]
    },
    {
      "cell_type": "code",
      "execution_count": null,
      "metadata": {
        "_cell_guid": "bd3a6dd2-f0f0-b9c1-301a-4d81ad9ebbfc"
      },
      "outputs": [],
      "source": [
        "sns.FacetGrid(hdf, size=6) \\\n",
        "   .map(sns.pointplot, \"epochs\", \"val_acc\", color=\"y\") \\\n",
        "   .map(sns.pointplot, \"epochs\", \"acc\", color=\"r\") \\\n",
        "   .set(xlabel='Epochs', ylabel='Validation Accuracy (Yellow) and Test Accuracy (Red)') "
      ]
    },
    {
      "cell_type": "markdown",
      "metadata": {
        "_cell_guid": "c0bdefda-0037-7ef5-db8a-a63cd661e230"
      },
      "source": [
        "### Validation Loss (Magenta) and Training Loss (Green) VS Epochs"
      ]
    },
    {
      "cell_type": "code",
      "execution_count": null,
      "metadata": {
        "_cell_guid": "bfb60afb-1fb4-1e8d-fbe5-c58f45d762ac"
      },
      "outputs": [],
      "source": [
        "sns.FacetGrid(hdf, size=6) \\\n",
        "   .map(sns.pointplot, \"epochs\", \"loss\", color=\"g\") \\\n",
        "   .map(sns.pointplot, \"epochs\", \"val_loss\", color=\"m\") \\\n",
        "   .set(xlabel='Epochs', ylabel='Validation Loss (Magenta) and Training Loss (Green)') "
      ]
    },
    {
      "cell_type": "markdown",
      "metadata": {
        "_cell_guid": "f979d215-2d84-fca2-3228-1771a64779ec"
      },
      "source": [
        "# Predictions\n",
        "First we'll create a numpy array of predictions using `model.predict_classes`, we'll save this model to a file and finally show a bit of info on the predictions we generated."
      ]
    },
    {
      "cell_type": "code",
      "execution_count": null,
      "metadata": {
        "_cell_guid": "c61a0a66-e916-99a0-66ff-0e8966dd3ada"
      },
      "outputs": [],
      "source": [
        "pred = model.predict_classes(x_test, verbose=2) # Shape -> (28000,)\n",
        "\n",
        "sub_df = pd.DataFrame()\n",
        "sub_df[\"ImageId\"] = list(range(1, num_testing + 1))\n",
        "sub_df[\"Label\"] = pred\n",
        "\n",
        "# sub_df.to_csv(\"mnist_predictions.csv\", header=True, index=False)"
      ]
    },
    {
      "cell_type": "code",
      "execution_count": null,
      "metadata": {
        "_cell_guid": "fb72c789-33ac-fe31-3173-310e87980519"
      },
      "outputs": [],
      "source": [
        "# Display head and lenght of our predictions:\n",
        "print(\"Amount of test points:\", num_testing)\n",
        "print(\"Amount of predictions:\", len(pred))\n",
        "\n",
        "sub_df.head()"
      ]
    },
    {
      "cell_type": "markdown",
      "metadata": {
        "_cell_guid": "457ae358-99b0-d0a6-cefe-7e288ae03fe8"
      },
      "source": [
        "And that's it! Let me know if you have any questions and or suggestions :)"
      ]
    }
  ],
  "metadata": {
    "_change_revision": 0,
    "_is_fork": false,
    "kernelspec": {
      "display_name": "Python 3",
      "language": "python",
      "name": "python3"
    },
    "language_info": {
      "codemirror_mode": {
        "name": "ipython",
        "version": 3
      },
      "file_extension": ".py",
      "mimetype": "text/x-python",
      "name": "python",
      "nbconvert_exporter": "python",
      "pygments_lexer": "ipython3",
      "version": "3.6.0"
    }
  },
  "nbformat": 4,
  "nbformat_minor": 0
}