{
  "cells": [
    {
      "cell_type": "markdown",
      "metadata": {
        "_cell_guid": "89234f43-d25b-6f0c-9d3e-9feb2a90d551"
      },
      "source": [
        "# Introduction\n",
        "Contests have become a very popular way to develop new algorithms. The basic idea is to give a large group of people (participants) access to exactly the same training data and testing data.\n",
        "### Training Data\n",
        "The training data has some input and the desired output. In the case of this example the input is an image and the output is the identification of the number which is inside the image. In the case of the 2017 Kaggle Data Science Bowl this is a Chest CT of a person and the output is a 1 or 0 if they have cancer or not. The training data is then used to build a number of different tools to solve the problem.\n",
        "### Testing Data\n",
        "The testing data are also made up of inputs and desired outputs, but the participants are just given the inputs (in this case just the images) upon which the different algorithms can be tested. Each participant submits their predictions or guesses for the ouputs to the organizer who can then compare the guesses with the known outputs. The algorithm which _matches_ the most wins (there are many different criteria for choosing the best, but in this case we will just count matches).\n",
        "\n",
        "### Standard Imports\n",
        "Here we just import some of the standard functions we will need later"
      ]
    },
    {
      "cell_type": "code",
      "execution_count": null,
      "metadata": {
        "_cell_guid": "64087669-6c03-2cfa-4ed2-ca7e7d561110"
      },
      "outputs": [],
      "source": [
        "from __future__ import print_function, division\n",
        "import numpy as np\n",
        "import matplotlib.pyplot as plt\n",
        "import os # for operating system commands like dealing with paths\n",
        "DATA_PATH = '../input' # where are the test.csv and train.csv files located\n",
        "test_data_path = os.path.join(DATA_PATH, 'test.csv')\n",
        "train_data_path = os.path.join(DATA_PATH, 'train.csv')"
      ]
    },
    {
      "cell_type": "markdown",
      "metadata": {
        "_cell_guid": "4f1783cb-09df-079c-7ef8-3c7782b04396"
      },
      "source": [
        "Load the training data from the file. Since we know the first column is the 'label' of the number we extract that from the rest and reform the remaining columns as a 28x28 array"
      ]
    },
    {
      "cell_type": "code",
      "execution_count": null,
      "metadata": {
        "_cell_guid": "6d09d6c3-d058-1583-425c-12dde6be6344"
      },
      "outputs": [],
      "source": [
        "%%time\n",
        "# this takes around 30 seconds so be patient\n",
        "train_data = np.loadtxt(train_data_path, delimiter = ',', skiprows = 1)\n",
        "numb_id = train_data[:,0] # just the number id\n",
        "numb_vec = train_data[:,1:] # the array of the images"
      ]
    },
    {
      "cell_type": "code",
      "execution_count": null,
      "metadata": {
        "_cell_guid": "a6c9721f-d0f4-d702-d069-a75b3a105566"
      },
      "outputs": [],
      "source": [
        "print('Input Data:', train_data.shape)\n",
        "print('Number ID:', numb_id.shape)\n",
        "print('Number Vector:', numb_vec.shape)"
      ]
    },
    {
      "cell_type": "markdown",
      "metadata": {
        "_cell_guid": "b0fd78c7-585e-8168-c4b4-5b6190361e9e"
      },
      "source": [
        "The image is formatted as a vector (a single long array), we want to reshape it to look like an image"
      ]
    },
    {
      "cell_type": "code",
      "execution_count": null,
      "metadata": {
        "_cell_guid": "b85981a4-5f57-c3a2-6a27-93e1e833fc75"
      },
      "outputs": [],
      "source": [
        "numb_image = numb_vec.reshape(-1, 28, 28)\n",
        "print('Number Image', numb_image.shape)"
      ]
    },
    {
      "cell_type": "markdown",
      "metadata": {
        "_cell_guid": "31ce3aff-b74a-6e38-8421-93c801a86549"
      },
      "source": [
        "Now we can show the image to make sure it is reasonable"
      ]
    },
    {
      "cell_type": "code",
      "execution_count": null,
      "metadata": {
        "_cell_guid": "07835ecd-e63b-dd43-d3da-679e65745c1d"
      },
      "outputs": [],
      "source": [
        "%matplotlib inline\n",
        "fig, ax1 = plt.subplots(1,1)\n",
        "ax1.matshow(numb_image[0], cmap = 'gray')\n",
        "ax1.set_title('Current Digit {}'.format(numb_id[0]))\n",
        "ax1.axis('off')"
      ]
    },
    {
      "cell_type": "markdown",
      "metadata": {
        "_cell_guid": "f7214785-f7c0-c78e-4399-e9f34f34ef4f"
      },
      "source": [
        "# Building a Classifier\n",
        "Here we use the full machine learning approach with a tool called TPOT which automatically builds and tests models.  You can read more about the state of fully automated machine learning [here](http://blog.yhat.com/posts/state-of-automl.html)\n",
        "\n",
        "![TPOP Pipeline][1]\n",
        "\n",
        "\n",
        "  [1]: http://blog.yhat.com/static/img/tpot-ml-pipeline.png"
      ]
    },
    {
      "cell_type": "code",
      "execution_count": null,
      "metadata": {
        "_cell_guid": "a826fc96-f86f-6f4b-45d4-8f12c0ad1d6a"
      },
      "outputs": [],
      "source": [
        "from tpot import TPOTClassifier\n",
        "auto_classifier = TPOTClassifier(generations=1, population_size=5, verbosity=2)"
      ]
    },
    {
      "cell_type": "code",
      "execution_count": null,
      "metadata": {
        "_cell_guid": "554822c7-0d23-c08d-0b8d-9ef7b194efaf"
      },
      "outputs": [],
      "source": [
        "%%time\n",
        "np.random.seed(1234)\n",
        "test_idx = np.random.choice(range(len(numb_vec)), 8000) # since the whole dataset takes up too much memory\n",
        "auto_classifier.fit(numb_vec[test_idx], numb_id[test_idx])"
      ]
    },
    {
      "cell_type": "markdown",
      "metadata": {
        "_cell_guid": "7fc22146-f70b-7fc7-b419-e3f199538f20"
      },
      "source": [
        "### Build the Classifier\n",
        "Using the dictionary of example digits we can build a very simple classifier using the MSE to find the best match"
      ]
    },
    {
      "cell_type": "code",
      "execution_count": null,
      "metadata": {
        "_cell_guid": "2cfc5917-e35c-9c67-ce35-7f11b1035dee"
      },
      "outputs": [],
      "source": [
        "rand_digit = np.random.choice(range(len(numb_image))) # just picks a random digit\n",
        "rand_digit_vec = numb_image[rand_digit].reshape(1,-1)\n",
        "guess_digit = auto_classifier.predict(rand_digit_vec)\n",
        "# show the probabilities for each class\n",
        "guess_digit_prob = auto_classifier._fitted_pipeline.predict_proba(rand_digit_vec)\n",
        "guess_dict = dict(enumerate(guess_digit_prob[0]))\n",
        "print('Guessed {}, actual result was {}'.format(guess_digit, numb_id[rand_digit]))\n",
        "print('Score for other numbers:', guess_dict)\n",
        "# show the results\n",
        "fig, (ax_img, ax_score) = plt.subplots(1,2, figsize = (10, 5))\n",
        "ax_img.imshow(numb_image[rand_digit], cmap = 'gray', interpolation = 'none')\n",
        "ax_img.set_title('Guessed {}, actual result was {}'.format(guess_digit, numb_id[rand_digit]))\n",
        "ax_score.bar(list(guess_dict.keys()), list(guess_dict.values()))\n",
        "ax_score.set_xlabel('Digit')\n",
        "ax_score.set_ylabel('Probability')\n",
        "ax_score.set_title('Probability for each digit')"
      ]
    },
    {
      "cell_type": "markdown",
      "metadata": {
        "_cell_guid": "a7862a40-77ff-fa7b-0cf3-4c68a052bef2"
      },
      "source": [
        "The classified doesn't work perfectly but it does a fairly good job of at least coming close with most of the guesses"
      ]
    },
    {
      "cell_type": "markdown",
      "metadata": {
        "_cell_guid": "8df3caaf-439d-00a0-de3c-2355b6b35894"
      },
      "source": [
        "# Submitting an Entry\n",
        "We can now submit an entry to the contest by running our digit classifier on all of the test images. The data for the test images is the same as the training data but there are no labels."
      ]
    },
    {
      "cell_type": "code",
      "execution_count": null,
      "metadata": {
        "_cell_guid": "3ba4676d-206c-4777-0f3f-59bc09a781da"
      },
      "outputs": [],
      "source": [
        "test_vec = np.loadtxt(test_data_path, delimiter = ',', skiprows = 1)\n",
        "print('Test Vec', test_vec.shape)\n",
        "test_image = test_vec.reshape(-1, 28, 28)\n",
        "print('Test Image', test_image.shape)"
      ]
    },
    {
      "cell_type": "markdown",
      "metadata": {
        "_cell_guid": "b5a93c64-feba-26f1-147c-a49307ccdbb6"
      },
      "source": [
        "## Make a guess for each image\n",
        "Here we make a guess for each image by using our ```classify_image``` function and just keeping the first result (the digit) and not the entire dictionary"
      ]
    },
    {
      "cell_type": "code",
      "execution_count": null,
      "metadata": {
        "_cell_guid": "7dda61d8-6662-9a6a-fa71-9beffc66e4bd"
      },
      "outputs": [],
      "source": [
        "%%time\n",
        "guess_test_data = auto_classifier.predict(test_vec)"
      ]
    },
    {
      "cell_type": "markdown",
      "metadata": {
        "_cell_guid": "ca06d9d3-9916-c294-670a-0857e10d67c8"
      },
      "source": [
        "### Save the results\n",
        "Here we save the results to a file called submission that we can upload to Kaggle and see how we perform"
      ]
    },
    {
      "cell_type": "code",
      "execution_count": null,
      "metadata": {
        "_cell_guid": "a8ba545f-665b-48b5-f166-104443c63e5d"
      },
      "outputs": [],
      "source": [
        "with open('submission.csv', 'w') as out_file:\n",
        "    out_file.write('ImageId,Label\\n')\n",
        "    for img_id, guess_label in enumerate(guess_test_data):\n",
        "        out_file.write('%d,%d\\n' % (img_id+1, guess_label))"
      ]
    },
    {
      "cell_type": "markdown",
      "metadata": {
        "_cell_guid": "5ff8a338-c49f-3089-b97f-4a95543ccfb1"
      },
      "source": [
        "# Basic Tasks\n",
        "- Submit the output to the competition, where does it come in? Is this better or worse than expected? What would random guessing be?\n",
        "\n",
        "# Idea Questions\n",
        "- What causes the classifier to fail on certain images?\n",
        "- How could the classifier be improved?\n",
        "- Could filtering or image enhancement improve the result, if so how\n",
        "\n",
        "# Challenge Tasks\n",
        "- Try improving the classifier by using SSIM or another similarity metric.\n",
        "- Try using more examples (or better examples)\n",
        "- Implement a version using image enhancement and compare the results"
      ]
    }
  ],
  "metadata": {
    "_change_revision": 0,
    "_is_fork": false,
    "kernelspec": {
      "display_name": "Python 3",
      "language": "python",
      "name": "python3"
    },
    "language_info": {
      "codemirror_mode": {
        "name": "ipython",
        "version": 3
      },
      "file_extension": ".py",
      "mimetype": "text/x-python",
      "name": "python",
      "nbconvert_exporter": "python",
      "pygments_lexer": "ipython3",
      "version": "3.6.0"
    }
  },
  "nbformat": 4,
  "nbformat_minor": 0
}