{
  "cells": [
    {
      "cell_type": "markdown",
      "metadata": {
        "_cell_guid": "092a7182-adba-a40d-f4f9-eb98446025b5"
      },
      "source": [
        "## First Steps ##"
      ]
    },
    {
      "cell_type": "markdown",
      "metadata": {
        "_cell_guid": "e00640c0-c56c-e34f-c268-c262dd40ae53"
      },
      "source": [
        "These are my first steps in the Kaggle environment. \n",
        "My plan is to invest some time every day and develop (a) a exploratory analysis of the data and (b) set up a prediction model using models that I'm not sure of yet. The goal is to get a foot into the world of Kaggle and get myself \u00e0 jour with the methods commonly used these days."
      ]
    },
    {
      "cell_type": "code",
      "execution_count": null,
      "metadata": {
        "_cell_guid": "e4f3b4bd-6fd3-6fdd-abf4-c98785d4feed"
      },
      "outputs": [],
      "source": [
        "# This Python 3 environment comes with many helpful analytics libraries installed\n",
        "# It is defined by the kaggle/python docker image: https://github.com/kaggle/docker-python\n",
        "# For example, here's several helpful packages to load in \n",
        "\n",
        "import numpy as np # linear algebra\n",
        "import pandas as pd # data processing, CSV file I/O (e.g. pd.read_csv)\n",
        "\n",
        "# Input data files are available in the \"../input/\" directory.\n",
        "# For example, running this (by clicking run or pressing Shift+Enter) will list the files in the input directory\n",
        "\n",
        "from subprocess import check_output\n",
        "print(check_output([\"ls\", \"../input\"]).decode(\"utf8\"))\n",
        "\n",
        "# Any results you write to the current directory are saved as output."
      ]
    },
    {
      "cell_type": "code",
      "execution_count": null,
      "metadata": {
        "_cell_guid": "ac0e7bce-9a84-47b8-341d-ddc3e158e2f4"
      },
      "outputs": [],
      "source": [
        "training = pd.read_csv(\"../input/train.csv\")\n",
        "test = pd.read_csv(\"../input/test.csv\")"
      ]
    },
    {
      "cell_type": "markdown",
      "metadata": {
        "_cell_guid": "c0d12160-9f90-18ee-272c-3d68ded15a2e"
      },
      "source": [
        "Now that the data is loaded, let's see what we're dealing with"
      ]
    },
    {
      "cell_type": "code",
      "execution_count": null,
      "metadata": {
        "_cell_guid": "157f39fa-c0f3-2b1a-805a-b15e40d5612f"
      },
      "outputs": [],
      "source": [
        "training.head()"
      ]
    },
    {
      "cell_type": "markdown",
      "metadata": {
        "_cell_guid": "ca25eed5-6535-e28c-cabb-8c03d9b4d962"
      },
      "source": [
        "First things first: since we're interested in predicting the sales price of an object, it's wise to get acquainted with that variable first. For starters, a histogram:"
      ]
    },
    {
      "cell_type": "code",
      "execution_count": null,
      "metadata": {
        "_cell_guid": "f3d19737-c7b3-9138-db5a-62182b1bc1fb"
      },
      "outputs": [],
      "source": [
        "import matplotlib.pyplot as plt\n",
        "price_hist = plt.hist(training['SalePrice'], bins=40)"
      ]
    },
    {
      "cell_type": "markdown",
      "metadata": {
        "_cell_guid": "05c1ae2c-ee31-9774-2ef6-d9c58285beb7"
      },
      "source": [
        "That tells us that the sales price is not very dispersed. Most of the objects lie within the 100,000 - 200,000 price range with a prominent peak at around 140.000 or so.\n",
        "For good measure, I'll also take a look at the first four empirical moments of the random variable. Just to show that I was paying attention in Statistics 101."
      ]
    },
    {
      "cell_type": "code",
      "execution_count": null,
      "metadata": {
        "_cell_guid": "dd21b6b6-f5e9-3f81-5d52-9b2dd5f8b687"
      },
      "outputs": [],
      "source": [
        "import scipy.stats as scipy\n",
        "sales_mean = training['SalePrice'].mean()\n",
        "sales_var = training['SalePrice'].var()\n",
        "sales_skew = scipy.skew(training['SalePrice'])\n",
        "sales_kurt = scipy.kurtosis(training['SalePrice'])\n",
        "\n",
        "print(\"The first four moments of SalePrice are\")\n",
        "print(\"------------\")\n",
        "print(\"1. Mean:\") \n",
        "print(sales_mean)\n",
        "print(\"------------\")\n",
        "print(\"2. Variance:\") \n",
        "print(sales_var)\n",
        "print(\"------------\")\n",
        "print(\"3. Skewness:\") \n",
        "print(sales_skew)\n",
        "print(\"------------\")\n",
        "print(\"4. Kurtosis:\") \n",
        "print(sales_kurt)"
      ]
    },
    {
      "cell_type": "markdown",
      "metadata": {
        "_cell_guid": "bc81c6dc-d3fd-48fc-e7f0-1a07c965d108"
      },
      "source": [
        "So the mean is somewhat expected, the variance has no helpful interpretation here.\n",
        "Skewness is positive, which quantifies the eye-ball inspection of having a right-skewed distribution.\n",
        "The kurtosis was slightly more suprising to me. On the histogram the tails look quite flat, but with a kurtosis of significantly greater than 3, we seem to have a leptokurtic specimen. Didn't see that coming, but in the end I don't think that this will jeopardize the results of the predictive analysis."
      ]
    },
    {
      "cell_type": "markdown",
      "metadata": {
        "_cell_guid": "fdf575e8-475e-6dbc-4a61-77bc43d68da0"
      },
      "source": [
        "## Quick look at the features ##"
      ]
    },
    {
      "cell_type": "markdown",
      "metadata": {
        "_cell_guid": "a4fe503e-251c-9279-94d5-b6f6094ac205"
      },
      "source": [
        "In order to make the process of feature examination a little leaner, the respective correlations with the target variable are of interest."
      ]
    },
    {
      "cell_type": "code",
      "execution_count": null,
      "metadata": {
        "_cell_guid": "4843a492-d8b9-8647-d2f6-c5927f941151"
      },
      "outputs": [],
      "source": [
        "import seaborn as sns\n",
        "corr = training.corr()\n",
        "f, ax = plt.subplots(figsize=(12, 9))\n",
        "sns.heatmap(corr, vmax=.8, square=True)"
      ]
    },
    {
      "cell_type": "markdown",
      "metadata": {
        "_cell_guid": "4ad57506-c8f9-6a2d-0dd3-96e01a508c86"
      },
      "source": [
        "The last row has the most informational content for the moment. \n",
        "Apparently these variables (colored in a slightly deeper shade of red) have explainatory power for SalePrice:\n",
        "\n",
        " - OverallQual\n",
        " - YearBuild (and only this one, since it's highly correlated with YearRemodAdd)\n",
        " - TotalBsmtSF\n",
        " - GrLivArea\n",
        " - FullBath\n",
        " - TotRmsAbvGrd\n",
        " - GarageArea (as the only Garage variable due to high correlation all over the place)"
      ]
    },
    {
      "cell_type": "markdown",
      "metadata": {
        "_cell_guid": "b81c47b5-2a9a-17f9-0c2f-bb60eeb30905"
      },
      "source": [
        "### Missing Data ###"
      ]
    },
    {
      "cell_type": "markdown",
      "metadata": {
        "_cell_guid": "9d8be0ac-4548-3864-0afd-280c6a076a89"
      },
      "source": [
        "In a last exploratory step, I'm going to check if any of these features have missing data that needs to be taken care of."
      ]
    },
    {
      "cell_type": "code",
      "execution_count": null,
      "metadata": {
        "_cell_guid": "53828d03-e403-b5e8-c315-a16001908aef"
      },
      "outputs": [],
      "source": [
        "selected_features = ['OverallQual','YearBuilt', 'TotalBsmtSF', 'GrLivArea', 'FullBath', 'TotRmsAbvGrd', 'GarageArea']\n",
        "training[selected_features].isnull().sum()"
      ]
    },
    {
      "cell_type": "markdown",
      "metadata": {
        "_cell_guid": "2e3ee624-5795-c656-070f-4f8e4ab8be13"
      },
      "source": [
        "Neat, no missing values in my selected variables! So I guess it's time to go ahead and get acquainted with the ubiquitous XGBoost algorithm."
      ]
    },
    {
      "cell_type": "markdown",
      "metadata": {
        "_cell_guid": "8102f6ae-edad-ae39-a4c4-6126eaafb5cd"
      },
      "source": [
        "## Getting to know XGBoost ##"
      ]
    },
    {
      "cell_type": "markdown",
      "metadata": {
        "_cell_guid": "4f9f943f-3f33-ae89-d830-138c1c412dd7"
      },
      "source": [
        "Since I'm such a noob, I'll stick to the cooking recipe quite closely for now. The instructions I'm following come from machinelearningmastery.com ([http://machinelearningmastery.com/develop-first-xgboost-model-python-scikit-learn/][1])\n",
        "\n",
        "\n",
        "  [1]: http://machinelearningmastery.com/develop-first-xgboost-model-python-scikit-learn/ \"XGBoost for complete Beginners\""
      ]
    },
    {
      "cell_type": "markdown",
      "metadata": {
        "_cell_guid": "4707b309-f933-dddb-651e-38a62496b7ee"
      },
      "source": [
        "Let's start with the equivalent of heating up the oven:"
      ]
    },
    {
      "cell_type": "code",
      "execution_count": null,
      "metadata": {
        "_cell_guid": "5a8495cd-4330-6548-3211-5de42dab596f"
      },
      "outputs": [],
      "source": [
        "import xgboost\n",
        "from sklearn import model_selection\n",
        "from sklearn.metrics import accuracy_score"
      ]
    },
    {
      "cell_type": "markdown",
      "metadata": {
        "_cell_guid": "835cec13-40ee-b80c-fe0f-1d263160a109"
      },
      "source": [
        "Next, we're separating the yolk from the egg:"
      ]
    },
    {
      "cell_type": "code",
      "execution_count": null,
      "metadata": {
        "_cell_guid": "cea60a20-7d04-4d8a-27cb-2dedea3dfe1d"
      },
      "outputs": [],
      "source": [
        "train_X = training[selected_features]\n",
        "train_y = training['SalePrice']\n",
        "test_X = test[selected_features]"
      ]
    },
    {
      "cell_type": "markdown",
      "metadata": {
        "_cell_guid": "6c3b9a01-c985-6110-03df-e67880675617"
      },
      "source": [
        "After stirring it around, the eggs will now be poured into the frying pan:"
      ]
    },
    {
      "cell_type": "code",
      "execution_count": null,
      "metadata": {
        "_cell_guid": "e7340e32-4da4-b9cd-21b2-ca8f5e5e095d"
      },
      "outputs": [],
      "source": [
        "xgb_model = xgboost.XGBRegressor()\n",
        "xgb_model.fit(train_X, train_y)"
      ]
    },
    {
      "cell_type": "markdown",
      "metadata": {
        "_cell_guid": "6fc7bcdb-547e-55d4-5d05-202f076d5e75"
      },
      "source": [
        "Flip it, and it's done!"
      ]
    },
    {
      "cell_type": "code",
      "execution_count": null,
      "metadata": {
        "_cell_guid": "37f52665-9b1a-e398-c2f9-4c4f9547996b"
      },
      "outputs": [],
      "source": [
        "prediction = xgb_model.predict(test_X)\n",
        "print(prediction)\n",
        "print(len(prediction))"
      ]
    },
    {
      "cell_type": "markdown",
      "metadata": {
        "_cell_guid": "81484881-213b-1669-e39d-427f87df73e6"
      },
      "source": [
        "Of course this omelette needs some seasoning (i.e. parameter tuning), but I'll leave that exercise for another day."
      ]
    },
    {
      "cell_type": "code",
      "execution_count": null,
      "metadata": {
        "_cell_guid": "44b5f8fa-a047-7f7d-34a3-fac647335e5d"
      },
      "outputs": [],
      "source": [
        "np.savetxt('house_prices_prediction.csv',prediction, delimiter=',')"
      ]
    }
  ],
  "metadata": {
    "_change_revision": 0,
    "_is_fork": false,
    "kernelspec": {
      "display_name": "Python 3",
      "language": "python",
      "name": "python3"
    },
    "language_info": {
      "codemirror_mode": {
        "name": "ipython",
        "version": 3
      },
      "file_extension": ".py",
      "mimetype": "text/x-python",
      "name": "python",
      "nbconvert_exporter": "python",
      "pygments_lexer": "ipython3",
      "version": "3.6.0"
    }
  },
  "nbformat": 4,
  "nbformat_minor": 0
}