{
  "cells": [
    {
      "cell_type": "markdown",
      "metadata": {
        "_cell_guid": "69815d51-d630-11a3-33b0-3366daacc683"
      },
      "source": [
        "# Import the data set"
      ]
    },
    {
      "cell_type": "code",
      "execution_count": null,
      "metadata": {
        "_cell_guid": "33031c47-9a07-f354-92e1-bce872547a2b"
      },
      "outputs": [],
      "source": [
        "# This Python 3 environment comes with many helpful analytics libraries installed\n",
        "# It is defined by the kaggle/python docker image: https://github.com/kaggle/docker-python\n",
        "# For example, here's several helpful packages to load in \n",
        "\n",
        "import numpy as np # linear algebra\n",
        "import pandas as pd # data processing, CSV file I/O (e.g. pd.read_csv)\n",
        "import matplotlib.pyplot as plt\n",
        "import seaborn as sb\n",
        "from scipy.stats import norm\n",
        "from scipy import stats\n",
        "from scipy.optimize import minimize\n",
        "\n",
        "HousesSold = pd.read_csv('../input/train.csv')\n",
        "print(HousesSold['SalePrice'].describe())"
      ]
    },
    {
      "cell_type": "markdown",
      "metadata": {
        "_cell_guid": "19544ea6-4177-5e36-6e67-c7017633a113"
      },
      "source": [
        "The data set is imported. The mean is higher than median implying that SalePrice is right skewed. We need to see how may variables are there in this data set."
      ]
    },
    {
      "cell_type": "code",
      "execution_count": null,
      "metadata": {
        "_cell_guid": "b1b10c48-3159-98d6-c15c-51292a9991de"
      },
      "outputs": [],
      "source": [
        "print(HousesSold.shape)\n",
        "print(HousesSold.columns)"
      ]
    },
    {
      "cell_type": "markdown",
      "metadata": {
        "_cell_guid": "1fbd029f-534a-1c28-70ad-a9e05e042d28"
      },
      "source": [
        "There are 81 variables in total in this data set and a total of 1460 observations"
      ]
    },
    {
      "cell_type": "code",
      "execution_count": null,
      "metadata": {
        "_cell_guid": "d4b020fc-a2ef-886a-e434-0092687d111e"
      },
      "outputs": [],
      "source": [
        " #*SalePrice*"
      ]
    },
    {
      "cell_type": "code",
      "execution_count": null,
      "metadata": {
        "_cell_guid": "045438e0-84b6-e567-4052-e5e720074230"
      },
      "outputs": [],
      "source": [
        "sb.distplot(HousesSold['SalePrice'])\n",
        "print(\"Skewness: %f\" % HousesSold['SalePrice'].skew())\n",
        "print(\"Kurtosis: %f\" % HousesSold['SalePrice'].kurt())"
      ]
    },
    {
      "cell_type": "markdown",
      "metadata": {
        "_cell_guid": "fd72d682-9635-4525-ac17-01eafd797187"
      },
      "source": [
        "# Heath Map\n",
        "\n",
        "Variables that are strongly correlated with SalePrice variable can be spotted from the heat map of their correlations with each other."
      ]
    },
    {
      "cell_type": "code",
      "execution_count": null,
      "metadata": {
        "_cell_guid": "6ad0f9f7-d13c-200e-97ec-9852d53aeae9"
      },
      "outputs": [],
      "source": [
        "corrmat = HousesSold.corr()\n",
        "plt.subplots(1, figsize=(12, 9))\n",
        "\n",
        "# Draw the heatmap using seaborn\n",
        "sb.heatmap(corrmat, vmax=1.0, square=True)\n",
        "plt.yticks(rotation=0)\n",
        "plt.xticks(rotation=90)\n",
        "plt.subplots_adjust(bottom=0.2)"
      ]
    },
    {
      "cell_type": "markdown",
      "metadata": {
        "_cell_guid": "121e78fb-34e2-5c07-e12e-a6a056f0ff0c"
      },
      "source": [
        "# Better Correlations with the *SalePrice Variable*\n",
        "\n",
        "We should spot the variables that have stronger correlations with the *SalePrice*."
      ]
    },
    {
      "cell_type": "code",
      "execution_count": null,
      "metadata": {
        "_cell_guid": "bf147d4b-ddea-bf46-3d12-2b25f30128e7"
      },
      "outputs": [],
      "source": [
        "k = 11\n",
        "cols = corrmat.nlargest(k, 'SalePrice')['SalePrice'].index\n",
        "cm = np.corrcoef(HousesSold[cols].values.T)\n",
        "plt.subplots(1, figsize=(12, 9))\n",
        "sb.set(font_scale=1)\n",
        "sb.heatmap(cm, cbar=True, annot = True, square=True, fmt='.2f', annot_kws={'size': 10}, yticklabels=cols.values, xticklabels=cols.values)\n",
        "plt.xticks(rotation=90)\n",
        "plt.yticks(rotation=0)\n",
        "plt.subplots_adjust(bottom=0.2)"
      ]
    },
    {
      "cell_type": "markdown",
      "metadata": {
        "_cell_guid": "e749e466-64ec-286e-bf90-19cfaae5f8de"
      },
      "source": [
        "# *SalePrice*\n",
        "\n",
        "Thus, we observe that the variable *SalePrice* has considerable relationship (r > 0.5) with the following variables:\n",
        "\n",
        " 1. *OverallQual*\n",
        " 2. *GrLivArea*\n",
        " 3. *GarageCars*\n",
        " 4. *GarageArea*\n",
        " 5. *TotalBsmtSF*\n",
        " 6. *1stFlrSF*\n",
        " 7.  *FullBath*\n",
        " 8. *TotRmsAbvGrd*\n",
        " 9. *YearBuilt*\n",
        " 10. *YearRemodAdd*\n",
        "\n",
        "A further investigation reveals that the variables *GrLivArea* and *TotRmsAbvGrd* are strongly correlated with each other, and that absolutely makes sense. We can, therefore, happily drop one of these variables. Let's keep *GrLivArea* and drop *TotRmsAbvGrd*.\n",
        "\n",
        "Similarly *GarageCars* and *GarageArea* are redundant, we can drop *GarageArea*. We will also keep *TotalBsmtSF* and drop *1stFlrSF*.\n",
        "\n",
        "We would like to keep other variables as they are strong candidates of being the predictors of *SalePrice* variable. The new list of variables that we would like to use as the predictors of *SalePrice* are:\n",
        "\n",
        " 1. *OverallQual*\n",
        " 2. *GrLivArea*\n",
        " 3. *GarageCars*\n",
        " 4. *TotalBsmtSF*\n",
        " 5.  *FullBath*\n",
        " 6. *YearBuilt*\n",
        " 7. *YearRemodAdd*"
      ]
    },
    {
      "cell_type": "markdown",
      "metadata": {
        "_cell_guid": "862d2453-0529-8466-b3a5-9b65e2667a5e"
      },
      "source": [
        "# Scatterplots of strong candidates with the *SalePrice*"
      ]
    },
    {
      "cell_type": "code",
      "execution_count": null,
      "metadata": {
        "_cell_guid": "26ea7f53-3634-b15f-4e32-abc58760cab9"
      },
      "outputs": [],
      "source": [
        "sb.set()\n",
        "cols = ['SalePrice', 'OverallQual', 'GrLivArea', 'GarageCars', 'TotalBsmtSF', 'FullBath', 'YearBuilt', 'YearRemodAdd']\n",
        "sb.pairplot(HousesSold[cols], size = 2.2)"
      ]
    },
    {
      "cell_type": "markdown",
      "metadata": {
        "_cell_guid": "19230544-dac1-19d9-18a4-2d9060666965"
      },
      "source": [
        "A quick observation says that the response variable *SalePrice* is right skewed. The two response variables *GrLivArea* and * TotalBsmtSF* are left skewed, the other three variables *OverallQual*, *GarageCars*, and *FullBath* seem normal. The variable *YearBuilt* is left skewed, whereas *YearRemodAdd* is bimodal.\n",
        "\n",
        "The scatter plots also show a quick overview of the behaviour of the response variable with the explanatory variables. All these relations are positive (week - moderately strong) and linear/non-linear.\n",
        "\n",
        "# Missing values\n",
        "\n",
        "It is important that we remove the variables with the significant number of missing values."
      ]
    },
    {
      "cell_type": "code",
      "execution_count": null,
      "metadata": {
        "_cell_guid": "acc45505-70a6-0fae-af45-8077bf5bd1dd"
      },
      "outputs": [],
      "source": [
        "#missing data\n",
        "allMissing = HousesSold.isnull().sum().sort_values(ascending=False)\n",
        "percentage = (HousesSold.isnull().sum()/HousesSold.isnull().count()).sort_values(ascending=False)\n",
        "missingData = pd.concat([allMissing, percentage*100], axis=1, keys=['TotalMissing', 'Percentage'])\n",
        "missingData.head(20)"
      ]
    },
    {
      "cell_type": "markdown",
      "metadata": {
        "_cell_guid": "5f9253e1-1798-a108-0dac-92fe8409ce75"
      },
      "source": [
        "A barplot of the variables with missing values is helpful here. We are okay to drop all these variables except *Electrical* because none of these variables seems to have a serious impact on the response variable. We can drop the observation in the data set *HousesSold* that has missing observation in *Electrical* variable."
      ]
    },
    {
      "cell_type": "code",
      "execution_count": null,
      "metadata": {
        "_cell_guid": "f07779a8-63e4-7af7-e43d-a1a71b3934a4"
      },
      "outputs": [],
      "source": [
        "missingData.head(19)[['Percentage']].plot(kind='bar', title =\"Missing Values\", figsize=(15, 10), legend=False, fontsize=13)\n",
        "plt.subplots_adjust(bottom=0.2)\n",
        "plt.ylabel(\"Percentage of Missing Values\")"
      ]
    },
    {
      "cell_type": "code",
      "execution_count": null,
      "metadata": {
        "_cell_guid": "6409bbe7-9088-29e1-7bfb-f17f7e5f64e3"
      },
      "outputs": [],
      "source": [
        "HousesSold1 = HousesSold.drop((missingData[missingData['TotalMissing'] > 1]).index,1)\n",
        "HousesSold2 = HousesSold1.drop(HousesSold1.loc[HousesSold1['Electrical'].isnull()].index)\n",
        "print(HousesSold2.shape)"
      ]
    },
    {
      "cell_type": "markdown",
      "metadata": {
        "_cell_guid": "921b5446-8dd6-212a-9839-f1248ebb5f5b"
      },
      "source": [
        "# Transformations\n",
        "\n",
        "## *SalePrice*\n",
        "\n",
        "We will apply logarithmic transformation on this variable."
      ]
    },
    {
      "cell_type": "code",
      "execution_count": null,
      "metadata": {
        "_cell_guid": "4a745cc8-98a4-f9eb-ecfa-bcca4c530c3c"
      },
      "outputs": [],
      "source": [
        "#histogram and normal probability plot\n",
        "sb.distplot(HousesSold2['SalePrice'], fit=norm);\n",
        "fig = plt.figure()\n",
        "res = stats.probplot(HousesSold2['SalePrice'], plot=plt)"
      ]
    },
    {
      "cell_type": "markdown",
      "metadata": {
        "_cell_guid": "734b08fc-d009-b35c-c16f-f16b76d7d8e2"
      },
      "source": [
        "As agreed before this is a right skewed distribution. The logarithmic distribution might help..."
      ]
    },
    {
      "cell_type": "code",
      "execution_count": null,
      "metadata": {
        "_cell_guid": "48030d8f-8a66-9afc-d240-594d97c42e49"
      },
      "outputs": [],
      "source": [
        "HousesSold2['SalePrice'] = np.log(HousesSold2['SalePrice'])"
      ]
    },
    {
      "cell_type": "code",
      "execution_count": null,
      "metadata": {
        "_cell_guid": "bf933fd7-9380-6a83-8c2f-ea26e567e1cf"
      },
      "outputs": [],
      "source": [
        "sb.distplot(HousesSold2['SalePrice'], fit=norm);\n",
        "fig = plt.figure()\n",
        "res = stats.probplot(HousesSold2['SalePrice'], plot=plt)"
      ]
    },
    {
      "cell_type": "markdown",
      "metadata": {
        "_cell_guid": "902a76e4-dfe3-1a76-d1b0-93e190fe8e99"
      },
      "source": [
        "This looks better.\n",
        "\n",
        "## *GrLivArea*"
      ]
    },
    {
      "cell_type": "code",
      "execution_count": null,
      "metadata": {
        "_cell_guid": "7ba08b33-25b2-3a72-c65d-c58db2df2ca4"
      },
      "outputs": [],
      "source": [
        "#histogram and normal probability plot\n",
        "sb.distplot(HousesSold2['GrLivArea'], fit=norm);\n",
        "fig = plt.figure()\n",
        "res = stats.probplot(HousesSold2['GrLivArea'], plot=plt)"
      ]
    },
    {
      "cell_type": "markdown",
      "metadata": {
        "_cell_guid": "b61913a6-a476-368c-5233-f0e1d25cef8c"
      },
      "source": [
        "We can delete the two extreme values on the right hand side here. Then using log transformation..."
      ]
    },
    {
      "cell_type": "code",
      "execution_count": null,
      "metadata": {
        "_cell_guid": "311afdbc-7e9d-b35a-479c-47fcc22a3947"
      },
      "outputs": [],
      "source": [
        "HousesSold2.sort_values(by = 'GrLivArea', ascending = False)[:2]"
      ]
    },
    {
      "cell_type": "code",
      "execution_count": null,
      "metadata": {
        "_cell_guid": "cc1480ca-a43a-92bd-ffe5-2d333c812716"
      },
      "outputs": [],
      "source": [
        "HousesSold2 = HousesSold2.drop(HousesSold2[HousesSold2['Id'] == 1299].index)\n",
        "HousesSold2 = HousesSold2.drop(HousesSold2[HousesSold2['Id'] == 524].index)\n",
        "HousesSold2['GrLivArea'] = np.log(HousesSold2['GrLivArea'])"
      ]
    },
    {
      "cell_type": "code",
      "execution_count": null,
      "metadata": {
        "_cell_guid": "c66a6b76-8347-dc41-b9e8-026a7c33c7be"
      },
      "outputs": [],
      "source": [
        "sb.distplot(HousesSold2['GrLivArea'], fit=norm);\n",
        "fig = plt.figure()\n",
        "res = stats.probplot(HousesSold2['GrLivArea'], plot=plt)"
      ]
    },
    {
      "cell_type": "markdown",
      "metadata": {
        "_cell_guid": "8dfc5fff-72ae-3423-b205-79b89f0223e6"
      },
      "source": [
        "Its also good...\n",
        "\n",
        "## *TotalBsmtSF*"
      ]
    },
    {
      "cell_type": "code",
      "execution_count": null,
      "metadata": {
        "_cell_guid": "e40e6173-4a3b-e2d1-9169-544c8d774f42"
      },
      "outputs": [],
      "source": [
        "#histogram and normal probability plot\n",
        "sb.distplot(HousesSold2['TotalBsmtSF'], fit=norm);\n",
        "fig = plt.figure()\n",
        "res = stats.probplot(HousesSold2['TotalBsmtSF'], plot=plt)"
      ]
    },
    {
      "cell_type": "markdown",
      "metadata": {
        "_cell_guid": "f808fb88-c12e-dc0b-f4fe-e2029158433d"
      },
      "source": [
        "We see a slight problem here.... Some observations with zero basement area. We do not want to remove this. For the moment let's not transform this.\n",
        "\n",
        "## Other variables\n",
        "\n",
        "We will for the time being keep the other variables as they are and try to look at their relationships with the response variable.\n",
        "\n",
        "# Relationships with the *SalePrice*\n",
        "## *OverallQual*"
      ]
    },
    {
      "cell_type": "code",
      "execution_count": null,
      "metadata": {
        "_cell_guid": "9213b5d0-c331-2e12-8906-824dd235198a"
      },
      "outputs": [],
      "source": [
        "plt.scatter(HousesSold2['OverallQual'], HousesSold2['SalePrice']);"
      ]
    },
    {
      "cell_type": "markdown",
      "metadata": {
        "_cell_guid": "88859f2c-9a30-bb55-8a77-333581942746"
      },
      "source": [
        "or"
      ]
    },
    {
      "cell_type": "code",
      "execution_count": null,
      "metadata": {
        "_cell_guid": "995ec704-6851-1bc9-db02-94a9a9aeebb7"
      },
      "outputs": [],
      "source": [
        "f, ax = plt.subplots(figsize=(8, 6))\n",
        "fig = sb.boxplot(x=\"OverallQual\", y=\"SalePrice\", data=HousesSold2)"
      ]
    },
    {
      "cell_type": "markdown",
      "metadata": {
        "_cell_guid": "5bd14976-2e6a-1621-e68a-8b45bbe397dd"
      },
      "source": [
        "That seems nice and linear relationship.\n",
        "\n",
        "## *GrLivArea*"
      ]
    },
    {
      "cell_type": "code",
      "execution_count": null,
      "metadata": {
        "_cell_guid": "1b39ee9e-d666-700d-0117-17d851efc36e"
      },
      "outputs": [],
      "source": [
        "plt.scatter(HousesSold2['GrLivArea'], HousesSold2['SalePrice']);"
      ]
    },
    {
      "cell_type": "markdown",
      "metadata": {
        "_cell_guid": "417b79f8-dd51-9af2-c225-a7c47bd1a81e"
      },
      "source": [
        "this is also good.\n",
        "\n",
        "## *GarageCars*"
      ]
    },
    {
      "cell_type": "code",
      "execution_count": null,
      "metadata": {
        "_cell_guid": "ad1620f3-c232-5ce7-9bd7-9fe7af117191"
      },
      "outputs": [],
      "source": [
        "f, ax = plt.subplots(figsize=(8, 6))\n",
        "fig = sb.boxplot(x=\"GarageCars\", y=\"SalePrice\", data=HousesSold2)"
      ]
    },
    {
      "cell_type": "markdown",
      "metadata": {
        "_cell_guid": "85594b2e-6999-97d3-6d31-69ab7c99c847"
      },
      "source": [
        "or"
      ]
    },
    {
      "cell_type": "code",
      "execution_count": null,
      "metadata": {
        "_cell_guid": "130c419d-c196-febe-a1fe-79f78793ae2c"
      },
      "outputs": [],
      "source": [
        "plt.scatter(HousesSold2['GarageCars'], HousesSold2['SalePrice']);"
      ]
    },
    {
      "cell_type": "markdown",
      "metadata": {
        "_cell_guid": "1274b2b3-1764-7ce6-7218-b896309a33a3"
      },
      "source": [
        "The variable *SalePrice* is related with the *GarageCars* linearly. The drop in house price for a garage of 4 cars space could be for any reason but these are just a few entries which makes this category less significant anyway.\n",
        "\n",
        "## *TotalBsmtSF*"
      ]
    },
    {
      "cell_type": "code",
      "execution_count": null,
      "metadata": {
        "_cell_guid": "0c55458e-e0b5-81e4-0c50-3dcee287ffad"
      },
      "outputs": [],
      "source": [
        "plt.scatter(HousesSold2['TotalBsmtSF'], HousesSold2['SalePrice']);"
      ]
    },
    {
      "cell_type": "markdown",
      "metadata": {
        "_cell_guid": "f4b607b8-ff47-4c7e-123a-8067dbae6223"
      },
      "source": [
        "This looks okay (except for the zeros basement entries). We can, however, remove two outliers with large basement area..."
      ]
    },
    {
      "cell_type": "code",
      "execution_count": null,
      "metadata": {
        "_cell_guid": "706ad199-676c-476e-ee62-f283c111b91e"
      },
      "outputs": [],
      "source": [
        "HousesSold2.sort_values(by = 'TotalBsmtSF', ascending = False)[:2]"
      ]
    },
    {
      "cell_type": "code",
      "execution_count": null,
      "metadata": {
        "_cell_guid": "1dee5541-dbb5-cad7-3036-b29261b75900"
      },
      "outputs": [],
      "source": [
        "HousesSold2 = HousesSold2.drop(HousesSold2[HousesSold2['Id'] == 333].index)\n",
        "HousesSold2 = HousesSold2.drop(HousesSold2[HousesSold2['Id'] == 497].index)\n",
        "plt.scatter(HousesSold2['TotalBsmtSF'], HousesSold2['SalePrice']);"
      ]
    },
    {
      "cell_type": "markdown",
      "metadata": {
        "_cell_guid": "32011820-a531-d56c-74d1-b6c419c9c421"
      },
      "source": [
        "## *FullBath*"
      ]
    },
    {
      "cell_type": "code",
      "execution_count": null,
      "metadata": {
        "_cell_guid": "e29d39e6-cca1-8033-15a2-3a8f39e7101e"
      },
      "outputs": [],
      "source": [
        "plt.scatter(HousesSold2['FullBath'], HousesSold2['SalePrice']);"
      ]
    },
    {
      "cell_type": "markdown",
      "metadata": {
        "_cell_guid": "b1c83889-3960-8d36-bb20-6ba4204612fa"
      },
      "source": [
        "or"
      ]
    },
    {
      "cell_type": "code",
      "execution_count": null,
      "metadata": {
        "_cell_guid": "4c81f53b-5996-40cb-9483-78c7d0c1becb"
      },
      "outputs": [],
      "source": [
        "f, ax = plt.subplots(figsize=(8, 6))\n",
        "fig = sb.boxplot(x=\"FullBath\", y=\"SalePrice\", data=HousesSold2)"
      ]
    },
    {
      "cell_type": "markdown",
      "metadata": {
        "_cell_guid": "9c312592-adfe-f1ad-c764-fa7229b2adc9"
      },
      "source": [
        "Again, there are only a few entries with zero number of full baths, therefore, it is okay that its average position does not lie on the line of relationship.\n",
        "\n",
        "## *YearBuilt*"
      ]
    },
    {
      "cell_type": "code",
      "execution_count": null,
      "metadata": {
        "_cell_guid": "3028417e-9945-a14a-5f86-55fcc791a797"
      },
      "outputs": [],
      "source": [
        "plt.scatter(HousesSold2['YearBuilt'], HousesSold2['SalePrice']);"
      ]
    },
    {
      "cell_type": "markdown",
      "metadata": {
        "_cell_guid": "96e5ce98-ca8a-ed9f-80cf-83fddf230d6c"
      },
      "source": [
        "I will count this as a nonlinear or quadratic relationship with the *SalePrice* variable. Some really old houses are being sold on decent prices probably for their historic nature.\n",
        "\n",
        "## *YearRemodAdd*"
      ]
    },
    {
      "cell_type": "code",
      "execution_count": null,
      "metadata": {
        "_cell_guid": "b3a770d6-860b-49c7-dffe-c1788ee66928"
      },
      "outputs": [],
      "source": [
        "plt.scatter(HousesSold2['YearRemodAdd'], HousesSold2['SalePrice']);"
      ]
    },
    {
      "cell_type": "markdown",
      "metadata": {
        "_cell_guid": "8c84b4a1-204f-e397-74fc-5768d1313fc7"
      },
      "source": [
        "Probably a linear relationship?\n",
        "\n",
        "# Final Model\n",
        "\n",
        "The model, we propose to predict the *SalePrice* is:\n",
        "\n",
        "*log(SP)* = *a OQ+ b log(GLA) + c GC + d TB + e FB + (f YB^2 + g YB + h) + i YM*\n",
        "\n",
        "Here:\n",
        "*a, b, c, d, e, f, g, h, i* are unknown constants and\n",
        "\n",
        "*SP = SalePrice*\n",
        "\n",
        "*OQ = OverallQual*\n",
        "\n",
        "and so on.."
      ]
    },
    {
      "cell_type": "code",
      "execution_count": null,
      "metadata": {
        "_cell_guid": "ce5b5456-ecc0-3e79-2476-f80e7a46cebd"
      },
      "outputs": [],
      "source": [
        "def f(x):   # The function that returns sum of squares of deviations from the true SalePrice\n",
        "    Yp = np.array([0]*len(HousesSold2.index));\n",
        "    devofThis = np.array([0]*len(HousesSold2.index));\n",
        "    sqofThis = np.array([0]*len(HousesSold2.index));\n",
        "    for i in range(len(HousesSold2.index)):\n",
        "        Yp[i] = x[0] * float(HousesSold2['OverallQual'].iloc[i]) + \\\n",
        "        x[1] * HousesSold2['GrLivArea'].iloc[i] + \\\n",
        "        x[2] * float(HousesSold2['GarageCars'].iloc[i]) + \\\n",
        "        x[3] * HousesSold2['TotalBsmtSF'].iloc[i] + \\\n",
        "        x[4] * float(HousesSold2['FullBath'].iloc[i]) + \\\n",
        "        x[5] * float(HousesSold2['YearBuilt'].iloc[i]) ** 2 +\\\n",
        "        x[6] * float(HousesSold2['YearBuilt'].iloc[i]) + x[7] + \\\n",
        "        x[8] * float(HousesSold2['YearRemodAdd'].iloc[i]);\n",
        "        devofThis[i]=Yp[i] - HousesSold2['SalePrice'].iloc[i];\n",
        "        sqofThis[i] = devofThis[i] ** 2;\n",
        "    return sqofThis.sum()\n",
        "def f_der(x):\n",
        "    der = np.zeros_like(x)\n",
        "    Yp = np.array([0]*len(HousesSold2.index));\n",
        "    devofThis = np.array([0]*len(HousesSold2.index));\n",
        "    devofThisOQ = np.array([0]*len(HousesSold2.index));\n",
        "    devofThisGLA = np.array([0]*len(HousesSold2.index));\n",
        "    devofThisGC = np.array([0]*len(HousesSold2.index));\n",
        "    devofThisTB = np.array([0]*len(HousesSold2.index));\n",
        "    devofThisFB = np.array([0]*len(HousesSold2.index));\n",
        "    devofThisYB2 = np.array([0]*len(HousesSold2.index));\n",
        "    devofThisYB = np.array([0]*len(HousesSold2.index));\n",
        "    devofThisYM = np.array([0]*len(HousesSold2.index));\n",
        "    for i in range(len(HousesSold2.index)):\n",
        "        Yp[i] = x[0] * float(HousesSold2['OverallQual'].iloc[i]) +\\\n",
        "        x[1] * float(HousesSold2['GrLivArea'].iloc[i]) +\\\n",
        "        x[2] * float(HousesSold2['GarageCars'].iloc[i]) + \\\n",
        "        x[3] * float(HousesSold2['TotalBsmtSF'].iloc[i]) + \\\n",
        "        x[4] * float(HousesSold2['FullBath'].iloc[i]) + \\\n",
        "        x[5] * float(HousesSold2['YearBuilt'].iloc[i]) ** 2 +\\\n",
        "        x[6] * float(HousesSold2['YearBuilt'].iloc[i]) + x[7] + \\\n",
        "        x[8] * float(HousesSold2['YearRemodAdd'].iloc[i]);\n",
        "        devofThis[i] = Yp[i] - float(HousesSold2['SalePrice'].iloc[i]);\n",
        "        devofThisOQ[i] = 2 * devofThis[i] * float(HousesSold2['OverallQual'].iloc[i]);\n",
        "        devofThisGLA[i] = 2 * devofThis[i] * float(HousesSold2['GrLivArea'].iloc[i]);\n",
        "        devofThisGC[i] = 2 * devofThis[i] * float(HousesSold2['GarageCars'].iloc[i]);\n",
        "        devofThisTB[i] = 2 * devofThis[i] * float(HousesSold2['TotalBsmtSF'].iloc[i]);\n",
        "        devofThisFB[i] = 2 * devofThis[i] * float(HousesSold2['FullBath'].iloc[i]);\n",
        "        devofThisYB2[i] = 2 * devofThis[i] * float(HousesSold2['YearBuilt'].iloc[i]) ** 2;\n",
        "        devofThisYB[i] = 2 * devofThis[i] * float(HousesSold2['YearBuilt'].iloc[i]);\n",
        "        devofThisYM[i] = 2 * devofThis[i] * float(HousesSold2['YearRemodAdd'].iloc[i])\n",
        "    der[0] = devofThisOQ.sum();\n",
        "    der[1] = devofThisGLA.sum();\n",
        "    der[2] = devofThisGC.sum();\n",
        "    der[3] = devofThisTB.sum();\n",
        "    der[4] = devofThisFB.sum();\n",
        "    der[5] = devofThisYB2.sum();\n",
        "    der[6] = devofThisYB.sum();\n",
        "    der[7] = devofThis.sum();\n",
        "    der[8] = devofThisYM.sum();\n",
        "    return der\n",
        "    \n",
        "bnds = ((0.0, None), (0.0, None), (0.0, None), (0.0, None), (0.0, None), (0.0, None), (0.0, None), (0.0,None), (0.0, None))\n",
        "res=minimize(f, [0.5, 3.1, 6.01, 0.01, 6.0, 0.0000001, 0.001, 8.0, 0.001], method='BFGS',jac=f_der,options={'xtol': 1e-8,'disp': True})\n",
        "print(res.x)"
      ]
    },
    {
      "cell_type": "code",
      "execution_count": null,
      "metadata": {
        "_cell_guid": "08607a33-d869-19e0-e906-fdea0d86d477"
      },
      "outputs": [],
      "source": [
        "Ypred = np.array([0]*len(HousesSold2.index));\n",
        "yM = HousesSold2[\"SalePrice\"].mean();\n",
        "sstot = np.array([0]*len(HousesSold2.index));\n",
        "ssres = np.array([0]*len(HousesSold2.index));\n",
        "for i in range(len(HousesSold2.index)):\n",
        "    Ypred[i] = res.x[0] * float(HousesSold2['OverallQual'].iloc[i]) +\\\n",
        "    res.x[1] * float(HousesSold2['GrLivArea'].iloc[i]) + \\\n",
        "    res.x[2] * float(HousesSold2['GarageCars'].iloc[i]) + \\\n",
        "    res.x[3] * float(HousesSold2['TotalBsmtSF'].iloc[i]) + \\\n",
        "    res.x[4] * float(HousesSold2['FullBath'].iloc[i]) + \\\n",
        "    res.x[5] * float(HousesSold2['YearBuilt'].iloc[i]) ** 2 +\\\n",
        "    res.x[6] * float(HousesSold2['YearBuilt'].iloc[i]) + res.x[7] + \\\n",
        "    res.x[8] * float(HousesSold2['YearRemodAdd'].iloc[i])\n",
        "    sstot[i] = (float(HousesSold2['SalePrice'].iloc[i]) - yM) ** 2;\n",
        "    ssres[i] = (float(HousesSold2['SalePrice'].iloc[i]) - Ypred[i]) ** 2;\n",
        "\n",
        "SStot = sstot.sum();\n",
        "SSres = ssres.sum();\n",
        "\n",
        "print(1- SSres/SStot)\n",
        "sb.distplot(ssres)"
      ]
    },
    {
      "cell_type": "code",
      "execution_count": null,
      "metadata": {
        "_cell_guid": "055241f7-975f-dd72-6aeb-4ae34fa05926"
      },
      "outputs": [],
      "source": [
        "print(Ypred[150])\n",
        "print(float(HousesSold2['SalePrice'].iloc[150]))"
      ]
    }
  ],
  "metadata": {
    "_change_revision": 0,
    "_is_fork": false,
    "kernelspec": {
      "display_name": "R",
      "language": "R",
      "name": "ir"
    },
    "language_info": {
      "codemirror_mode": "r",
      "file_extension": ".r",
      "mimetype": "text/x-r-source",
      "name": "R",
      "pygments_lexer": "r",
      "version": "3.3.3"
    }
  },
  "nbformat": 4,
  "nbformat_minor": 0
}