{
  "cells": [
    {
      "cell_type": "markdown",
      "metadata": {
        "_cell_guid": "70c44336-351f-b3f2-7590-a0f49b23b615"
      },
      "source": [
        "Introduction\n",
        "------------\n",
        "\n",
        "The [scikit-learn article][1] on Decision Trees give a good introduction to this non-parametric supervised learning method. They are used for classification and regression. The goal is to create a model that predicts the value of a target variable by learning simple decision rules inferred from the data features. \n",
        "\n",
        "A decision tree is a flowchart-like structure in which each internal node represents a \"test\" on an attribute (e.g. whether a coin flip comes up heads or tails), each branch represents the outcome of the test and each leaf node represents a class label (decision taken after computing all attributes). The paths from root to leaf represents classification rules.\n",
        "\n",
        "**Objectives:** \n",
        "\n",
        "- Understand how Decision Trees work  \n",
        "- Build decision tree using Python package\n",
        "- Understand the problem of overfitting \n",
        "- Understand Cross Validation & Holdout Validation \n",
        "\n",
        "Hence, features transformations , features selection, cross-validation, parameters tuning etc. are outside the scope of this notebook :)  \n",
        "\n",
        "**Bibliography**\n",
        "\n",
        "- [Python for Data Analysis Part 29: Decision Trees][2] - from hamelg.blogspot\n",
        "\n",
        "\n",
        "  [1]: http://scikit-learn.org/stable/modules/tree.html#tree\n",
        "  [2]: http://hamelg.blogspot.co.uk/2015/11/python-for-data-analysis-part-29.html"
      ]
    },
    {
      "cell_type": "markdown",
      "metadata": {
        "_cell_guid": "5ac6b48f-916e-dcc2-5c88-a5b2e68bdc7b"
      },
      "source": [
        "Data manipulations\n",
        "------------------\n",
        "\n",
        "Visit my [first notebook][1] for more details :) \n",
        "\n",
        "\n",
        "  [1]: https://www.kaggle.com/malymandalay/titanic/titanic-from-cleaning-to-visualisation"
      ]
    },
    {
      "cell_type": "code",
      "execution_count": null,
      "metadata": {
        "_cell_guid": "5c97aeba-8f1d-5232-4887-e78318292f3d"
      },
      "outputs": [],
      "source": [
        "# Load in our libraries\n",
        "import pandas as pd\n",
        "import numpy as np\n",
        "import re\n",
        "import sklearn\n",
        "\n",
        "from sklearn import tree\n",
        "from sklearn import preprocessing\n",
        "\n",
        "import seaborn as sns\n",
        "import matplotlib.pyplot as plt\n",
        "%matplotlib inline"
      ]
    },
    {
      "cell_type": "code",
      "execution_count": null,
      "metadata": {
        "_cell_guid": "759ecc4a-5c6a-5d37-722a-8678af33631f"
      },
      "outputs": [],
      "source": [
        "train = pd.read_csv('../input/train.csv')\n",
        "test = pd.read_csv('../input/test.csv')\n",
        "\n",
        "full_data = [train , test]\n",
        "\n",
        "# Missing values \n",
        "\n",
        "for dataset in full_data: \n",
        "    mean_age = dataset['Age'].mean()\n",
        "    std_age = dataset['Age'].std() \n",
        "    missing_age = dataset['Age'].isnull().sum()\n",
        "    \n",
        "    age_null_random_list = np.random.randint(mean_age - std_age, mean_age + std_age, size=missing_age)\n",
        "    dataset['Age'][np.isnan(dataset['Age'])] = age_null_random_list\n",
        "    dataset['Age'] = dataset['Age'].astype(int)\n",
        "    \n",
        "for dataset in full_data:\n",
        "    dataset['Fare'] = dataset['Fare'].fillna(train['Fare'].median())  \n",
        "    \n",
        "for dataset in full_data:\n",
        "    dataset['Embarked'] = dataset['Embarked'].fillna('S') #most common value \n",
        "\n",
        "# Create bands  \n",
        "    \n",
        "for dataset in full_data:\n",
        "    dataset.loc[ dataset['Fare'] <= 7.91, 'Fare']                               = 0\n",
        "    dataset.loc[(dataset['Fare'] > 7.91) & (dataset['Fare'] <= 14.454), 'Fare'] = 1\n",
        "    dataset.loc[(dataset['Fare'] > 14.454) & (dataset['Fare'] <= 31), 'Fare']   = 2\n",
        "    dataset.loc[ dataset['Fare'] > 31, 'Fare']                                  = 3\n",
        "    dataset['Fare'] = dataset['Fare'].astype(int)\n",
        "     \n",
        "bins = [0, 18, 34, 50, 200] \n",
        "group_names = [1, 2, 3 , 4]\n",
        "\n",
        "for dataset in full_data:\n",
        "    categories = pd.cut(dataset['Age'], bins, labels=group_names)\n",
        "    dataset['Age_cat'] = pd.cut(dataset['Age'], bins, labels=group_names)\n",
        "\n",
        "# Create new feature FamilySize as a combination of SibSp and Parch\n",
        "for dataset in full_data:\n",
        "    dataset['FamilySize'] = 0\n",
        "    dataset['FamilySize'] = dataset['SibSp'] + dataset['Parch'] + 1\n",
        "    dataset.loc[dataset['FamilySize'] > 8,\"FamilySize\"]=8 #meaning 8+ \n",
        "    \n",
        "# Create new feature IsAlone from FamilySize\n",
        "for dataset in full_data:\n",
        "    dataset['IsAlone'] = 0\n",
        "    dataset.loc[dataset['FamilySize'] == 1, 'IsAlone'] = 1\n",
        "\n",
        "# Create Title variable \n",
        "def get_title(name):\n",
        "    title_search = re.search(' ([A-Za-z]+)\\.', name)\n",
        "    # If the title exists, extract and return it.\n",
        "    if title_search:\n",
        "        return title_search.group(1)\n",
        "    return \"\"\n",
        "# Create a new feature Title, containing the titles of passenger names\n",
        "for dataset in full_data:\n",
        "    dataset['Title'] = dataset['Name'].apply(get_title)\n",
        "# Group all non-common titles into one single grouping \"Rare\"\n",
        "for dataset in full_data:\n",
        "    dataset['Title'] = dataset['Title'].replace(['Lady', 'Countess','Capt', 'Col','Don', 'Dr', 'Major', 'Rev', 'Sir', 'Jonkheer', 'Dona'], 'Rare')\n",
        "\n",
        "    dataset['Title'] = dataset['Title'].replace('Mlle', 'Miss')\n",
        "    dataset['Title'] = dataset['Title'].replace('Ms', 'Miss')\n",
        "    dataset['Title'] = dataset['Title'].replace('Mme', 'Mrs')\n",
        "\n",
        "for dataset in full_data:\n",
        "    dataset['Mother'] = 0\n",
        "    dataset.loc[(dataset['Age'] > 18) & (dataset['Parch'] > 1) & (dataset['Title'] != 'Miss'),\"Mother\"]=1\n",
        "    #dataset['Mother'][(dataset['Age'] > 18) & (dataset['Parch'] > 1) & (dataset['Title'] != 'Miss')] = 1\n",
        "\n",
        "# So, we can classify passengers as males, females, and child\n",
        "def get_person(passenger):\n",
        "    age,sex = passenger\n",
        "    return 'child' if age < 16 else sex\n",
        "\n",
        "for dataset in full_data:\n",
        "    dataset['Person'] = dataset[['Age','Sex']].apply(get_person,axis=1)"
      ]
    },
    {
      "cell_type": "code",
      "execution_count": null,
      "metadata": {
        "_cell_guid": "f4119ccf-b3cf-4846-7d18-832253f05da7"
      },
      "outputs": [],
      "source": [
        "test.head(3)"
      ]
    },
    {
      "cell_type": "markdown",
      "metadata": {
        "_cell_guid": "e96a5b74-cb62-1be1-5142-d1d60ffa96c3"
      },
      "source": [
        "Decision trees\n",
        "--------------\n",
        "\n",
        "**Understand how it works**\n",
        "\n",
        "Let's only use 1  variable Person and create a model based on the training set using decision trees in Python. \n",
        "\n",
        "In the previous post, we wrote a code to change our numeric variables into dummy variables. \n",
        "Here, we will use [LabelEncoder][1], which is a utility class to help normalize labels such that they contain only values between 0 and n_classes-1. \n",
        "\n",
        "\n",
        "  [1]: http://scikit-learn.org/stable/modules/preprocessing_targets.html#preprocessing-targets"
      ]
    },
    {
      "cell_type": "code",
      "execution_count": null,
      "metadata": {
        "_cell_guid": "03fe01aa-6b49-8e7d-c74b-b89ae7361346"
      },
      "outputs": [],
      "source": [
        "# Initialize label encoder\n",
        "label_encoder = preprocessing.LabelEncoder()\n",
        "\n",
        "# Convert Sex variable to numeric\n",
        "encoded_Person = label_encoder.fit_transform(train[\"Person\"])\n",
        "\n",
        "# Initialize model\n",
        "tree_model = tree.DecisionTreeClassifier()\n",
        "\n",
        "# Train the model\n",
        "tree_model.fit(X = pd.DataFrame(encoded_Person), \n",
        "               y = train[\"Survived\"])"
      ]
    },
    {
      "cell_type": "code",
      "execution_count": null,
      "metadata": {
        "_cell_guid": "57e38697-61bd-78ee-f851-732741287455"
      },
      "outputs": [],
      "source": [
        "## I wanted to see an image of the decision tree but this doesn't seem to be working on Kaggle :) \n",
        "\n",
        "#import pydotplus\n",
        "#from pydotplus import graphviz\n",
        "#from IPython.display import Image\n",
        "       \n",
        "#dot_data = tree.export_graphviz(tree_model, feature_names=[\"Person\"])\n",
        "#graph = pydotplus.graphviz.graph_from_dot_file(dot_data)\n",
        "\n",
        "#Image(graph.create_png())             # Display image*"
      ]
    },
    {
      "cell_type": "code",
      "execution_count": null,
      "metadata": {
        "_cell_guid": "e1356145-0c65-ceb6-4a95-6aebc11a006d"
      },
      "outputs": [],
      "source": [
        "# Get survival probability\n",
        "preds = tree_model.predict_proba(X = pd.DataFrame(encoded_Person))\n",
        "\n",
        "pd.crosstab(preds[:,0], train[\"Person\"])"
      ]
    },
    {
      "cell_type": "markdown",
      "metadata": {
        "_cell_guid": "1dfc3025-d4e4-7402-593d-1277a0eead35"
      },
      "source": [
        "You can imagine this table like a tree with 1 node (the \"gender\") and each category (child / woman / man) are in 3 different leaves. \n",
        "\n",
        "Childen have a 45% death probability and men a 83% probability to die. Women have the lowest probability to die (24%). \n",
        "\n",
        "Let's check model accuracy : "
      ]
    },
    {
      "cell_type": "code",
      "execution_count": null,
      "metadata": {
        "_cell_guid": "3a23ed25-62a2-1360-b66d-b823b1b47eb0"
      },
      "outputs": [],
      "source": [
        "tree_model.score(X =  pd.DataFrame(encoded_Person), \n",
        "                 y = train[\"Survived\"])"
      ]
    },
    {
      "cell_type": "markdown",
      "metadata": {
        "_cell_guid": "38003a80-0f05-b20f-daf2-bb4e466a50b6"
      },
      "source": [
        "The model is 778% accurate on the training data"
      ]
    },
    {
      "cell_type": "markdown",
      "metadata": {
        "_cell_guid": "9a7f1658-e85a-23b7-3de8-a5f1802f8577"
      },
      "source": [
        "**Let's input more data ! , or the danger of overfitting**\n",
        "\n",
        "Decision-tree learners can create over-complex trees that do not generalise the data well. The default arguments for max_depth and min_samples_split were set to None: this means that no limit on the depth of your tree was set and we might be overfitting.\n",
        "\n",
        "This means that while your model describes the training data extremely well, it doesn't generalize to new data, which is frankly the point of prediction. \n",
        "\n",
        "Maybe we can improve the overfit model by making a less complex model? In DecisionTreeRegressor, the depth of our model is defined by two parameters: \n",
        "- the max_depth parameter determines when the splitting up of the decision tree stops. \n",
        "- the min_samples_split parameter monitors the amount of observations in a bucket. If a certain threshold is not reached (e.g minimum 10 passengers) no further splitting can be done"
      ]
    },
    {
      "cell_type": "code",
      "execution_count": null,
      "metadata": {
        "_cell_guid": "3dc15ff1-064a-6494-3a0b-94029a46dcb9"
      },
      "outputs": [],
      "source": [
        "train.head(3)"
      ]
    },
    {
      "cell_type": "code",
      "execution_count": null,
      "metadata": {
        "_cell_guid": "b3692155-d2c8-2cbb-e4e2-0b94361104fa"
      },
      "outputs": [],
      "source": [
        "# Convert categorical variables variable to numeric\n",
        "\n",
        "encoded_Person = label_encoder.fit_transform(train[\"Person\"])\n",
        "encoded_Embarked = label_encoder.fit_transform(train[\"Embarked\"])\n",
        "encoded_Age = label_encoder.fit_transform(train[\"Age_cat\"])\n",
        "encoded_Title = label_encoder.fit_transform(train[\"Title\"])\n",
        "\n",
        "# define predictors \n",
        "#variables = [encoded_Person, encoded_Embarked, encoded_Age, encoded_Title, train[\"Fare\"] , train[\"FamilySize\"] ]\n",
        "variables = [encoded_Person,  train[\"Age\"] ,  train[\"Pclass\"] , encoded_Title,  train[\"Fare\"] , train[\"FamilySize\"] ]\n",
        "# Make data frame of predictors\n",
        "predictors = pd.DataFrame(variables).T\n",
        "\n",
        "max_depth = 8\n",
        "min_samples_split = 10\n",
        "\n",
        "# Initialize model with options \n",
        "tree_model = tree.DecisionTreeClassifier(max_depth = max_depth, min_samples_split = min_samples_split, random_state = 1)\n",
        "\n",
        "# Train the model\n",
        "tree_model.fit(X = predictors, \n",
        "               y = train[\"Survived\"])\n",
        "\n",
        "# Get survival probability\n",
        "preds = tree_model.predict_proba(X = predictors)\n",
        "\n",
        "# Get cross table (only with 2 variables )\n",
        "#pd.crosstab(preds[:,0], columns = [train[\"Age_cat\"], \n",
        " #                                  train[\"Sex\"]])\n",
        "\n",
        "tree_model.score(X =  predictors, \n",
        "                y = train[\"Survived\"])"
      ]
    },
    {
      "cell_type": "markdown",
      "metadata": {
        "_cell_guid": "b0247268-d4a0-b54f-5c6f-3770f6dc979e"
      },
      "source": [
        "Cross Validation & Holdout Validation\n",
        "-------------------------------------\n",
        "\n",
        "We have created our model on training data, and we saw that the model fits training data quite closely - but how is it going to perform on new data ? \n",
        "\n",
        "(and also, when submitting the test sample to Kaggle it was a bit annoying to see that the model didn't perform quite as well on the test set - and I realised that only when Kaggle came back with a lower accuracy - I would like to get an idea of how the model perform on new data BEFORE submitting :) . \n",
        "\n",
        "We can get a better sense of a model's expected performance on unseen data by setting a portion of our training data aside when creating a model, and then using that set aside data to evaluate the model's performance.\n",
        "\n",
        "I like Hemelg's [blog explanation][1]  on hold-out sample - and let's follow his approach on creating holdout and cross validation samples \n",
        "\n",
        "**Cross validation** involves splitting the training data into two or more partitions and creating a model for each partition where the partition acts as the validation set and the remaining data acts as the training data.\n",
        "\n",
        "A common form of cross validation is \"k-fold\" cross validation, which randomly splits data into some number k (a user specified parameter) partitions and then creates k models, each tested against one of the partitions. Each of the k models are then combined into one aggregate final model.\n",
        "\n",
        "The primary advantage of cross validation is it uses all the training data to build and assess the final model. The main drawback is that building and testing several models can be computationally expensive, so it tends to take much longer than holdout validation.\n",
        "\n",
        "We are going to follow Hemelg approach (with less variables) , to give an example . For our submission, we will use a Holdout validation. \n",
        "\n",
        "\n",
        "  [1]: http://hamelg.blogspot.co.uk/2015/11/python-for-data-analysis-part-29.htmlHoldout"
      ]
    },
    {
      "cell_type": "code",
      "execution_count": null,
      "metadata": {
        "_cell_guid": "2b689d1d-3a10-71e4-76a7-b275f9ecdb5b"
      },
      "outputs": [],
      "source": [
        "from sklearn.cross_validation import KFold\n",
        "\n",
        "cv = KFold(n=len(train),  # Number of elements\n",
        "           n_folds=10,            # Desired number of cv folds\n",
        "           random_state=12)       # Set a random seed\n",
        "\n",
        "#After creating a cross validation object, you can loop over each fold and train and evaluate a your model on each one:\n",
        "\n",
        "fold_accuracy = []\n",
        "\n",
        "# Convert Sex variable to numeric\n",
        "encoded_sex = label_encoder.fit_transform(train[\"Sex\"])\n",
        "train[\"Sex\"] = encoded_sex\n",
        "\n",
        "\n",
        "for train_fold, valid_fold in cv:\n",
        "    train_k = train.loc[train_fold] # Extract train data with cv indices\n",
        "    valid_k = train.loc[valid_fold] # Extract valid data with cv indices\n",
        "    \n",
        "    model = tree_model.fit(X = train_k[[\"Sex\",\"Pclass\",\"Age\",\"Fare\"]], \n",
        "                           y = train_k[\"Survived\"])\n",
        "    valid_acc = model.score(X = valid_k[[\"Sex\",\"Pclass\",\"Age\",\"Fare\"]], \n",
        "                            y = valid_k[\"Survived\"])\n",
        "    fold_accuracy.append(valid_acc)    \n",
        "\n",
        "print(\"Accuracy per fold: \", fold_accuracy, \"\\n\")\n",
        "print(\"Average accuracy: \", sum(fold_accuracy)/len(fold_accuracy))"
      ]
    },
    {
      "cell_type": "markdown",
      "metadata": {
        "_cell_guid": "c3560363-0a17-00cf-afb1-85ac0dba4769"
      },
      "source": [
        "We'd like the target variable's classes to have roughly the same proportion across each fold when performing cross validation for a classification problem. To perform stratified cross validation, use the StratifiedKFold() function instead of KFold().\n",
        "\n",
        "Notice that the average accuracy across each fold is higher than the non-stratified K-fold example. The cross_val_score function is useful for testing models and tuning model parameters (finding optimal values for arguments like maximum tree depth that affect model performance.)."
      ]
    },
    {
      "cell_type": "code",
      "execution_count": null,
      "metadata": {
        "_cell_guid": "543ff3d2-6bff-fed9-360a-ed2d7282724d"
      },
      "outputs": [],
      "source": [
        "from sklearn.cross_validation import cross_val_score\n",
        "\n",
        "scores = cross_val_score(estimator= tree_model,     # Model to test\n",
        "                X= train[[\"Sex\",\"Pclass\",   # Train Data\n",
        "                                  \"Age\",\"Fare\"]],  \n",
        "                y = train[\"Survived\"],      # Target variable\n",
        "                scoring = \"accuracy\",               # Scoring metric    \n",
        "                cv=10)                              # Cross validation folds\n",
        "\n",
        "print(\"Accuracy per fold: \")\n",
        "print(scores)\n",
        "print(\"Average accuracy: \", scores.mean())"
      ]
    },
    {
      "cell_type": "markdown",
      "metadata": {
        "_cell_guid": "091951d9-9cf6-89ac-11f8-918823cfa8f1"
      },
      "source": [
        "**Holdout validation** involves splitting the training data into two parts, a training set and a validation set, building a model with the training set and then assessing performance with the validation set. In theory, model performance on the hold-out validation set should roughly mirror the performance you'd expect to see on unseen test data. In practice, holdout validation is fast and it can work well, especially on large data sets, but it has some pitfalls.\n",
        "\n",
        "Reserving a portion of the training data for a holdout set means you aren't using all the data at your disposal to build your model in the validation phase. This can lead to suboptimal performance, especially in situations where you don't have much data to work with. In addition, if you use the same holdout validation set to assess too many different models, you may end up finding a model that fits the validation set well due to chance that won't necessarily generalize well to unseen data. Despite these shortcomings, it is worth learning how to use a holdout validation set in Python. "
      ]
    },
    {
      "cell_type": "code",
      "execution_count": null,
      "metadata": {
        "_cell_guid": "2c49d4fa-fca1-1348-9292-013f4bc722ff"
      },
      "outputs": [],
      "source": [
        "from sklearn.cross_validation import train_test_split\n",
        "\n",
        "v_train, v_test = train_test_split(train,     # Data set to split\n",
        "                                   test_size = 0.25,  # Split ratio\n",
        "                                   random_state=1,    # Set random seed\n",
        "                                   stratify = train[\"Survived\"]) #*\n",
        "\n",
        "# Training set size for validation\n",
        "print(v_train.shape)\n",
        "# Test set size for validation\n",
        "print(v_test.shape)"
      ]
    },
    {
      "cell_type": "code",
      "execution_count": null,
      "metadata": {
        "_cell_guid": "e5233872-0ebd-c905-4d78-9b7cdf639bee"
      },
      "outputs": [],
      "source": [
        "full_data = [v_train , v_test]\n",
        "\n",
        "max_depth = 8\n",
        "min_samples_split = 10\n",
        "tree_model = tree.DecisionTreeClassifier(max_depth = max_depth, min_samples_split = min_samples_split, random_state = 1)\n",
        "\n",
        "def model_validation(dataset):\n",
        "    # Convert categorical variables variable to numeric\n",
        "    encoded_Person = label_encoder.fit_transform(dataset[\"Person\"])\n",
        "    encoded_Embarked = label_encoder.fit_transform(dataset[\"Embarked\"])\n",
        "    encoded_Age = label_encoder.fit_transform(dataset[\"Age_cat\"])\n",
        "    encoded_Title = label_encoder.fit_transform(dataset[\"Title\"])\n",
        "\n",
        "    variables = [encoded_Person,  dataset[\"Age\"] ,  dataset[\"Pclass\"] , encoded_Title,  dataset[\"Fare\"] , dataset[\"FamilySize\"] ]\n",
        "    # Make data frame of predictors\n",
        "    predictors = pd.DataFrame(variables).T\n",
        "\n",
        "    tree_model.fit(X = predictors, \n",
        "                   y = dataset[\"Survived\"])\n",
        "\n",
        "    # Get survival probability\n",
        "    preds = tree_model.predict_proba(X = predictors)\n",
        "\n",
        "    return tree_model.score(X =  predictors, \n",
        "                        y = dataset[\"Survived\"])\n",
        "    \n",
        "x = model_validation(v_train)\n",
        "y = model_validation(v_test)\n",
        "\n",
        "print (\"Train Accuracy  = \", x) \n",
        "print (\"Test Accuracy  = \", y) "
      ]
    },
    {
      "cell_type": "markdown",
      "metadata": {
        "_cell_guid": "7c633ffd-e423-276a-4de6-9068d1a85f2e"
      },
      "source": [
        "**Submission**"
      ]
    },
    {
      "cell_type": "code",
      "execution_count": null,
      "metadata": {
        "_cell_guid": "1cd2625e-71eb-9448-c482-6806c625a3d8"
      },
      "outputs": [],
      "source": [
        "# Convert categorical variables variable to numeric\n",
        "encoded_Person = label_encoder.fit_transform(test[\"Person\"])\n",
        "encoded_Embarked = label_encoder.fit_transform(test[\"Embarked\"])\n",
        "encoded_Age = label_encoder.fit_transform(test[\"Age_cat\"])\n",
        "encoded_Title = label_encoder.fit_transform(test[\"Title\"])\n",
        "\n",
        "variables_test = [encoded_Person,  test[\"Age\"] ,  test[\"Pclass\"] , encoded_Title,  test[\"Fare\"] , test[\"FamilySize\"] ]\n",
        "# Make data frame of predictors\n",
        "test_features = pd.DataFrame(variables_test).T\n",
        "\n",
        "# Make test set predictions\n",
        "test_preds = tree_model.predict(X=test_features)\n",
        "\n",
        "# Create a submission for Kaggle\n",
        "submission = pd.DataFrame({\"PassengerId\":test[\"PassengerId\"],\n",
        "                           \"Survived\":test_preds})\n",
        "\n",
        "# Save submission to CSV\n",
        "submission.to_csv(\"tutorial_dectree_submission.csv\", \n",
        "                  index=False)        # Do not save index values"
      ]
    }
  ],
  "metadata": {
    "_change_revision": 0,
    "_is_fork": false,
    "kernelspec": {
      "display_name": "Python 3",
      "language": "python",
      "name": "python3"
    },
    "language_info": {
      "codemirror_mode": {
        "name": "ipython",
        "version": 3
      },
      "file_extension": ".py",
      "mimetype": "text/x-python",
      "name": "python",
      "nbconvert_exporter": "python",
      "pygments_lexer": "ipython3",
      "version": "3.6.0"
    }
  },
  "nbformat": 4,
  "nbformat_minor": 0
}