{
  "cells": [
    {
      "cell_type": "markdown",
      "metadata": {
        "_cell_guid": "f9eea58c-6938-7b57-5f4a-9243dcf1309b"
      },
      "source": [
        "## Overview \n",
        "\n",
        "Now I am in the Top 4% of this task. \n",
        "My notebook is very simple and understandable, it doesn't consist from difficult formulas,  plots and lots of words  :)\n",
        "\n",
        "I want to improve my decision if you have any advice or comments I will be glad to know about this.\n",
        "\n",
        "I slightly changed parameters of functions. You can select the appropriate parameters if you want"
      ]
    },
    {
      "cell_type": "code",
      "execution_count": null,
      "metadata": {
        "_cell_guid": "c9eb7bfa-48e0-4639-f322-293fb25ba437"
      },
      "outputs": [],
      "source": [
        "import pandas as pd\n",
        "import numpy as np\n",
        "import scipy \n",
        "import seaborn as sns\n",
        "import matplotlib.pyplot as plt\n",
        "%matplotlib inline\n",
        "pd.options.display.max_rows = 999\n",
        "verbose = False # param for debugging"
      ]
    },
    {
      "cell_type": "code",
      "execution_count": null,
      "metadata": {
        "_cell_guid": "afd782c6-e28c-ff81-09c5-04581cd33579"
      },
      "outputs": [],
      "source": [
        "# Load data\n",
        "train = pd.read_csv('../input/train.csv', header=0,sep=',')\n",
        "test = pd.read_csv('../input/test.csv', header=0)"
      ]
    },
    {
      "cell_type": "code",
      "execution_count": null,
      "metadata": {
        "_cell_guid": "af64a38f-29c1-9c70-4ed5-6f883c8f89b3"
      },
      "outputs": [],
      "source": [
        "# Let's look at the general information about the training set\n",
        "train.info()"
      ]
    },
    {
      "cell_type": "markdown",
      "metadata": {
        "_cell_guid": "92ceace4-2329-029b-0358-8ad81422fb7e"
      },
      "source": [
        "## 1 Processing of the dataset\n",
        "\n",
        "### 1.1 Data Dictionary\n",
        "\n",
        "    **Variable**\t**Definition**\t                            **Key**\n",
        "\n",
        "    survival\t       Survival\t                                  0 = No, 1 = Yes\n",
        "    pclass\t         Ticket class\t                              1 = 1st, 2 = 2nd, 3 = 3rd\n",
        "    sex\t            Sex\t\n",
        "    Age\t            Age in years\t\n",
        "    sibsp\t          # of siblings / spouses aboard the Titanic\t\n",
        "    parch              # of parents / children aboard the Titanic\t\n",
        "    ticket\t         Ticket number\t\n",
        "    fare\t           Passenger fare\t\n",
        "    cabin\t          Cabin number\t\n",
        "    embarked\t       Port of Embarkation\t                      C = Cherbourg, Q = Queenstown, S = Southampton\n",
        "\n",
        "### 1.2 NaN values"
      ]
    },
    {
      "cell_type": "code",
      "execution_count": null,
      "metadata": {
        "_cell_guid": "e3c3b62b-2a28-9306-d2f2-1f74b6ae3b46"
      },
      "outputs": [],
      "source": [
        "# Create a function for processing NaN values\n",
        "\n",
        "def fillNaN(df=train,drop=True):\n",
        "    if drop:\n",
        "        df.dropna(subset=['Embarked','Fare'],inplace=True,axis=0)  # Drop  NaN for Embarked and Fare\n",
        "    else:\n",
        "        df['Fare']=df['Fare'].fillna(0.)                           # Fill NaN as 0 for Fare\n",
        "        df['Embarked']=df['Embarked'].fillna('S')                  # Fill NaN as S for Embarked\n",
        "    df['Age']=df['Age'].fillna(95.)                                # Fill NaN as 95 for Age. I chose 95 because max Age is 80\n",
        "    df['Cabin']=df['Cabin'].fillna('-1')                           # Fill NaN as -1 for Cabin \n",
        "fillNaN()\n",
        "fillNaN(test,False)\n",
        "\n",
        "# Check empty values\n",
        "train.info()"
      ]
    },
    {
      "cell_type": "markdown",
      "metadata": {
        "_cell_guid": "a14f95dc-b9c4-5b72-ce43-94ff1cfdbba8"
      },
      "source": [
        "### 1.3 Mapping object (string) values to  numeric type & creating groups from values\n",
        "\n",
        "1. As we saw above, we have 5 object columns. There are  **Name, Sex, Ticket, Cabin, Embarked**. For future calculation we should map this values to numbers. Let's try to map string values to numbers for **Sex, Cabin, Embarked**\n",
        "2. Also we have 2 columns with large range of values. There are **Cabin and Age**. For future calculation we should group this values\n",
        "3. We have some columns useless for predictions, there are  **Name and Ticket**\n",
        "\n",
        " **Cabin**"
      ]
    },
    {
      "cell_type": "code",
      "execution_count": null,
      "metadata": {
        "_cell_guid": "a96dab60-5d8e-c37e-3236-c46e96850796"
      },
      "outputs": [],
      "source": [
        "train.groupby('Cabin').count()\n",
        "# We will see 'T' as noise"
      ]
    },
    {
      "cell_type": "code",
      "execution_count": null,
      "metadata": {
        "_cell_guid": "50005f08-40b9-ed01-5f95-99abfed601e5"
      },
      "outputs": [],
      "source": [
        "# I want to get code for Cabine, \n",
        "# I suppose that the first letter is ship's deck ( depends on location and class of passenger)\n",
        "\n",
        "data_clear = train.copy()  \n",
        "\n",
        "# T is noise in the field\n",
        "data_clear = data_clear[train.Cabin!= 'T']    \n",
        "\n",
        "\n",
        "cabin,cab_map = [],[]\n",
        "\n",
        "def modCabin(df):\n",
        "    global cabin,cab_map\n",
        "    df['cab'] = df['Cabin']                  # Create a new column\n",
        "    df['cab'] = [ x[0] for x in df.Cabin]         # get the first letter from Cabin\n",
        "    if len(cabin) == 0:\n",
        "        cabin = df['cab'].unique()                         # get unique values\n",
        "        cabin.sort(axis=0)                                         # sort values\n",
        "        print(cabin)                                                # note NA = '-'\n",
        "        cab_map =np.arange(len(cabin))                             # List of numbers for mapping \n",
        "    df['cab'].replace(cabin,cab_map,inplace=True)      # Map letters to numbers \n",
        "    df['cab'].astype(dtype='int64')                  # Convert data type to  int64"
      ]
    },
    {
      "cell_type": "code",
      "execution_count": null,
      "metadata": {
        "_cell_guid": "18c32339-372a-268e-ec22-bfd2d9cb08a9"
      },
      "outputs": [],
      "source": [
        "modCabin(data_clear)\n",
        "modCabin(test)"
      ]
    },
    {
      "cell_type": "markdown",
      "metadata": {
        "_cell_guid": "1f0c6e50-8d06-03dd-772a-e70bfe2716c3"
      },
      "source": [
        "**Sex & Embarked**"
      ]
    },
    {
      "cell_type": "code",
      "execution_count": null,
      "metadata": {
        "_cell_guid": "99292be2-bcdb-59f2-7fbf-6aa20d52e89c"
      },
      "outputs": [],
      "source": [
        "# Create a function to mapping string to int \n",
        "def modSexEmb(df):\n",
        "    df['Sex'].replace(['male','female'],[0,1],inplace=True) # Map male = 0, female = 1\n",
        "    df['Sex'].astype(dtype='int64')                         # Convert data type to  int64\n",
        "\n",
        "    df['Embarked'].replace(['C','Q','S'],[0,1,2],inplace=True) # Map C = 0, Q = 1, S = 2, N = 3\n",
        "    df['Embarked'].astype(dtype='int64')                    # Convert data type to  int64\n",
        "modSexEmb(data_clear)\n",
        "modSexEmb(test)"
      ]
    },
    {
      "cell_type": "markdown",
      "metadata": {
        "_cell_guid": "5e11ce73-32c4-f0b0-5897-ba3b3d1f88c0"
      },
      "source": [
        "**Drop useless fields**"
      ]
    },
    {
      "cell_type": "code",
      "execution_count": null,
      "metadata": {
        "_cell_guid": "089d3801-2077-81c4-9d9c-074bf5c382be"
      },
      "outputs": [],
      "source": [
        "# Create a function for dropping useless fields \n",
        "def drop(df, delPas=True):\n",
        "    df.drop('Cabin',axis=1,inplace=True)             # Drop 'Cabin' field\n",
        "    if delPas == True:\n",
        "        df.drop('PassengerId',axis=1,inplace=True)       # Drop 'PassengerId' field\n",
        "    df.drop('Name',axis=1,inplace=True)              # Drop 'Name' field\n",
        "    df.drop('Ticket',axis=1,inplace=True)            # Drop 'Ticket' field\n",
        "    df.info()                                        # Check  \n",
        "drop(data_clear)\n",
        "drop(test,False)"
      ]
    },
    {
      "cell_type": "markdown",
      "metadata": {
        "_cell_guid": "4129dca4-8209-5bef-899a-d3fa7b5da800"
      },
      "source": [
        " **Grouping  values**\n",
        "\n",
        "As we saw above, Age and Fare have a lot of values and they must be grouped"
      ]
    },
    {
      "cell_type": "code",
      "execution_count": null,
      "metadata": {
        "_cell_guid": "89fb2baf-237c-a2f7-02ce-e57984d43530"
      },
      "outputs": [],
      "source": [
        "# Create a function for groupping\n",
        "def split(col_range, df,col):\n",
        "    n=0\n",
        "    for i in col_range:\n",
        "        \n",
        "        if i == col_range[len(col_range)-1]:\n",
        "            df[col].replace(df[(df[col]>=i)][col],i,inplace=True)\n",
        "        else:\n",
        "            alpha = col_range[n+1]-col_range[n]\n",
        "            df[col].replace( df[(df[col]>=i) & (df[col]<i+alpha)][col],i,inplace=True)\n",
        "        n+=1\n",
        "    return df[col]\n",
        "\n",
        "\"\"\"Age\"\"\"\n",
        "ages_range = [0, 9.5, 15, 20, 35, 40, 50, 60, 90]\n",
        "data_clear['Age'] = split(ages_range, data_clear.copy(), 'Age')\n",
        "test['Age'] = split(ages_range, test.copy(), 'Age')\n",
        "\"\"\"Fare\"\"\"\n",
        "fare_range = [0,20,40,60,80,100,150,200,250,500]\n",
        "data_clear['Fare'] = split(fare_range, data_clear.copy(), 'Fare')\n",
        "test['Fare'] = split(fare_range, test.copy(), 'Fare')"
      ]
    },
    {
      "cell_type": "markdown",
      "metadata": {
        "_cell_guid": "afbaa3b3-9dc4-8dac-b091-4fc4b5a47136"
      },
      "source": [
        "## 2 Draw graphics\n",
        "\n",
        "Now we will draw some plots for processed values"
      ]
    },
    {
      "cell_type": "code",
      "execution_count": null,
      "metadata": {
        "_cell_guid": "4014b309-9f57-9cc5-5abf-1c1b362e201e"
      },
      "outputs": [],
      "source": [
        "params = {'Sex':[0,1],       # Who is Survived (Male vs Female)\n",
        "         'Pclass':[1,2,3],               # What class\n",
        "         'Embarked':[0,1,2,3],  # Embarked\n",
        "         'SibSp':[0,1,2,3,4,5,6,7,8],    # \u2116 of siblings\n",
        "         'Parch':[0,1,2,3,4,5,6],        # \u2116 parents / children\n",
        "         'cab':np.arange(len(cabin)),\n",
        "         'Age':ages_range,\n",
        "         'Fare':fare_range\n",
        "         }             \n",
        "plt.figure(figsize=(20, 40))\n",
        "plot_number=0\n",
        "for key, value in params.items():\n",
        "    vals =[]\n",
        "    if verbose:\n",
        "        print( key)\n",
        "    for i in value:\n",
        "        c = data_clear[data_clear[key]==i][key].count()\n",
        "        v = data_clear[data_clear[key]==i]['Survived']\n",
        "        if verbose:\n",
        "            print  (i , c)\n",
        "            print ('Survived', data_clear[(data_clear[key]==i)&(data_clear['Survived']==1)]['Survived'].count())\n",
        "            print ('Died', data_clear[(data_clear[key]==i)&(data_clear['Survived']==0)]['Survived'].count())\n",
        "        vals.append(v)\n",
        "    plot_number+=1\n",
        "    ax = plt.subplot(5, 2, plot_number)\n",
        "    \n",
        "    plt.title(key)\n",
        "    plt.xlabel(key)\n",
        "    plt.ylabel('Numbers of people')\n",
        "    plt.hist((vals), histtype='bar', bins=5,label=value,cumulative =False, normed=True)\n",
        "    ax.legend(prop={'size': 10})\n",
        "plt.show()"
      ]
    },
    {
      "cell_type": "markdown",
      "metadata": {
        "_cell_guid": "ead19fc0-eed7-7bf6-3962-d5af9c327e70"
      },
      "source": [
        "And very interesting graph  Age / Survived / Sex"
      ]
    },
    {
      "cell_type": "code",
      "execution_count": null,
      "metadata": {
        "_cell_guid": "160dff5c-c288-6410-c947-a5c28c66b115"
      },
      "outputs": [],
      "source": [
        "def plot_distribution( df , var , target , **kwargs ):\n",
        "    row = kwargs.get( 'row' , None )\n",
        "    col = kwargs.get( 'col' , None )\n",
        "    facet = sns.FacetGrid( df , hue=target , aspect=4 , row = row , col = col )\n",
        "    facet.map( sns.kdeplot , var , shade= True )\n",
        "    facet.set( xlim=( 0 , df[ var ].max() ) )\n",
        "    facet.add_legend()\n",
        "plot_distribution(train , var = 'Age' , target = 'Survived' , row = 'Sex' )   "
      ]
    },
    {
      "cell_type": "markdown",
      "metadata": {
        "_cell_guid": "0a9c9add-220f-fe61-0125-7c08268f55df"
      },
      "source": [
        "### 3 Processing"
      ]
    },
    {
      "cell_type": "code",
      "execution_count": null,
      "metadata": {
        "_cell_guid": "f7b64d01-0b2d-9812-8344-2c8333006c4f"
      },
      "outputs": [],
      "source": [
        "# Separarte numeric, categorial and target(y) parameters\n",
        "\n",
        "numeric_cols=['SibSp','Parch','Fare']\n",
        "\n",
        "y = data_clear['Survived'].copy()\n",
        "data_clear.drop(['Survived'],axis=1,inplace=True) \n",
        "\n",
        "categorical_cols = list(set(data_clear.columns.values.tolist()) - set(numeric_cols))"
      ]
    },
    {
      "cell_type": "code",
      "execution_count": null,
      "metadata": {
        "_cell_guid": "b77b5ddb-de09-b8ab-4bbc-4ba13517f78f"
      },
      "outputs": [],
      "source": [
        "from sklearn.linear_model import LogisticRegression as LR\n",
        "from sklearn.feature_extraction import DictVectorizer as DV\n",
        "from sklearn.cross_validation import train_test_split\n",
        "from sklearn.svm import SVC, LinearSVC\n",
        "from sklearn.grid_search import GridSearchCV\n",
        "from sklearn.metrics import roc_auc_score\n",
        "from sklearn import cross_validation, datasets, metrics, tree"
      ]
    },
    {
      "cell_type": "code",
      "execution_count": null,
      "metadata": {
        "_cell_guid": "0919288f-b155-763a-c3f0-37dca6a0b5ff"
      },
      "outputs": [],
      "source": [
        "# Create copies of data for numeric and categorial columns\n",
        "X_cat = data_clear[categorical_cols].copy()\n",
        "for i in X_cat.columns.values:\n",
        "    X_cat[i] = X_cat[i].astype(str)\n",
        "\n",
        "X_num = data_clear[numeric_cols].copy()"
      ]
    },
    {
      "cell_type": "code",
      "execution_count": null,
      "metadata": {
        "_cell_guid": "080076d7-6dc3-224e-e97a-2bc9d4f154ea"
      },
      "outputs": [],
      "source": [
        "# Create function for transform categorial data to matrix\n",
        "def catTransform(X_cat):\n",
        "    encoder = DV(sparse = False)\n",
        "    X_cat_oh = encoder.fit_transform(X_cat.T.to_dict().values())\n",
        "    np.set_printoptions(threshold=np.nan)\n",
        "    print (X_cat_oh.shape)\n",
        "    return X_cat_oh\n",
        "X_cat_oh = catTransform(X_cat)"
      ]
    },
    {
      "cell_type": "code",
      "execution_count": null,
      "metadata": {
        "_cell_guid": "dc16eaa0-1fbb-71bd-c4e7-55e5499cc27f"
      },
      "outputs": [],
      "source": [
        "# Devide data  to train and test as 70/30\n",
        "\n",
        "(X_train, \n",
        " X_test, \n",
        " y_train, y_test) = train_test_split(X_num, y, \n",
        "                                     test_size=0.3, \n",
        "                                     random_state=0,\n",
        "                                    stratify=y)\n",
        "(X_train_cat_oh,\n",
        " X_test_cat_oh) = train_test_split(X_cat_oh, \n",
        "                                   test_size=0.3, \n",
        "                                   random_state=0,\n",
        "                                  stratify=y)\n",
        "(X_train_cat,\n",
        " X_test_cat) = train_test_split(X_cat, \n",
        "                                   test_size=0.3, \n",
        "                                   random_state=0,\n",
        "                                  stratify=y)"
      ]
    },
    {
      "cell_type": "markdown",
      "metadata": {
        "_cell_guid": "0c7c93ff-05cf-53cd-69c8-c1fb95fd75bf"
      },
      "source": [
        "**LogisticRegression**"
      ]
    },
    {
      "cell_type": "code",
      "execution_count": null,
      "metadata": {
        "_cell_guid": "99c990c5-aa28-fefb-92af-8f9709104997"
      },
      "outputs": [],
      "source": [
        "# LogisticRegression\n",
        "param_grid = {'C': np.linspace(1,2,50),'class_weight':['balanced']}\n",
        "cv = 3\n",
        "\n",
        "estimator = LR('l1')\n",
        "grid = GridSearchCV(estimator, param_grid,cv=cv)\n",
        "grid.fit(np.hstack([X_train,X_train_cat_oh]), y_train)\n",
        "\n",
        "lr_pred = grid.predict(np.hstack([X_test,X_test_cat_oh]))\n",
        "cv_score_lr = cross_validation.cross_val_score(grid, np.hstack([X_test,X_test_cat_oh]), y_test, cv = 10).mean()\n",
        "print (cv_score_lr)"
      ]
    },
    {
      "cell_type": "markdown",
      "metadata": {
        "_cell_guid": "438ca057-8eb5-ab3a-134a-d82efd49ec53"
      },
      "source": [
        "**SVC**"
      ]
    },
    {
      "cell_type": "code",
      "execution_count": null,
      "metadata": {
        "_cell_guid": "c4ea8e9f-1a82-cdcd-00fe-ae30884216b5"
      },
      "outputs": [],
      "source": [
        "# SVC\n",
        "\n",
        "svc = SVC()\n",
        "svc.fit(np.hstack([X_train,X_train_cat_oh]), y_train)\n",
        "svc_pred = svc.predict(np.hstack([X_test,X_test_cat_oh]))\n",
        "cv_score_svc = cross_validation.cross_val_score(svc, np.hstack([X_test,X_test_cat_oh]), y_test, cv = 10).mean()\n",
        "print (cv_score_svc)"
      ]
    },
    {
      "cell_type": "markdown",
      "metadata": {
        "_cell_guid": "278f0ae2-2aa9-b60e-4839-7518c01beae4"
      },
      "source": [
        "**Decision Tree**\n",
        "\n",
        "Belowe  for train, predict and crossval I use all set of data because these are features of algorithms"
      ]
    },
    {
      "cell_type": "code",
      "execution_count": null,
      "metadata": {
        "_cell_guid": "227ea01b-3a64-879e-9ce6-8ce0d409b669"
      },
      "outputs": [],
      "source": [
        "# Decision Tree\n",
        "clf = tree.DecisionTreeClassifier(random_state = 1, min_samples_leaf = 5, max_depth = 6)\n",
        "clf.fit(np.hstack([X_num,X_cat]), y)\n",
        "clf_pred = clf.predict(np.hstack([X_num,X_cat]))\n",
        "\n",
        "cv_score_clf = cross_validation.cross_val_score(clf, np.hstack([X_num,X_cat]), y, cv = 10).mean()  \n",
        "print (cv_score_clf)"
      ]
    },
    {
      "cell_type": "markdown",
      "metadata": {
        "_cell_guid": "40913fb4-46d0-39d8-4606-8e41667bf7a4"
      },
      "source": [
        "**Random Forest**\n",
        "\n",
        "I will add a graphic to see results of the algoritmh "
      ]
    },
    {
      "cell_type": "code",
      "execution_count": null,
      "metadata": {
        "_cell_guid": "bf097746-3f38-6150-f56f-31ed2f13deef"
      },
      "outputs": [],
      "source": [
        "from sklearn import ensemble, learning_curve \n",
        "\n",
        "# Random Forest\n",
        "\n",
        "\n",
        "rf_classifier = ensemble.RandomForestClassifier(n_estimators = 200, random_state = 1, min_samples_leaf = 5)\n",
        "train_sizes, train_scores, test_scores = learning_curve.learning_curve(rf_classifier, np.hstack([X_num,X_cat]), y, \n",
        "                                                                       train_sizes=np.arange(0.1,1., 0.2), \n",
        "                                                                       cv=10, scoring='accuracy')\n",
        "plt.grid(True)\n",
        "plt.plot(train_sizes, train_scores.mean(axis = 1), 'g-', marker='o', label='train')\n",
        "plt.plot(train_sizes, test_scores.mean(axis = 1), 'r-', marker='o', label='test')\n",
        "plt.ylim((0.0, 1.05))\n",
        "plt.legend(loc='lower right')\n",
        "rf_classifier.fit(np.hstack([X_num,X_cat]), y)\n",
        "rf_pred = rf_classifier.predict(np.hstack([X_num,X_cat]))\n",
        "\n",
        "cv_score_rt = cross_validation.cross_val_score(rf_classifier, np.hstack([X_num,X_cat]), y, cv = 15).mean()\n",
        "print (cv_score_rt)"
      ]
    },
    {
      "cell_type": "markdown",
      "metadata": {
        "_cell_guid": "8160e3e0-30af-14ba-db8f-0f339168741f"
      },
      "source": [
        "**Bagging**"
      ]
    },
    {
      "cell_type": "code",
      "execution_count": null,
      "metadata": {
        "_cell_guid": "c225172b-aec7-dff2-b573-21ac062d35db"
      },
      "outputs": [],
      "source": [
        "# Bagging\n",
        "clf2 = tree.DecisionTreeClassifier(random_state = 1, min_samples_leaf = 5)\n",
        "bagging = ensemble.BaggingClassifier(clf2,n_estimators =100)\n",
        "bagging.fit(np.hstack([X_num,X_cat]), y)\n",
        "bag_pred = bagging.predict(np.hstack([X_num,X_cat]))\n",
        "cv_score_bag = cross_validation.cross_val_score(bagging, np.hstack([X_num,X_cat]), y, cv = 10).mean()\n",
        "print (cv_score_bag)"
      ]
    },
    {
      "cell_type": "markdown",
      "metadata": {
        "_cell_guid": "08c2452b-4fed-bea0-912f-ac0b3bbdec9b"
      },
      "source": [
        "**Bagging with features**"
      ]
    },
    {
      "cell_type": "code",
      "execution_count": null,
      "metadata": {
        "_cell_guid": "e8b2753e-9a6a-1a18-d3a8-18ffc388320b"
      },
      "outputs": [],
      "source": [
        "# Bagging with features\n",
        "\n",
        "d =  np.hstack([X_train,X_train_cat]).shape[1]\n",
        "bagging = ensemble.BaggingClassifier(clf2,n_estimators =100,max_features=d)\n",
        "bagging.fit(np.hstack([X_num,X_cat]), y)\n",
        "bagf_pred = bagging.predict(np.hstack([X_num,X_cat]))\n",
        "cv_score_bagf = cross_validation.cross_val_score(bagging, np.hstack([X_num,X_cat]), y, cv = 10).mean()\n",
        "print (cv_score_bagf)"
      ]
    },
    {
      "cell_type": "markdown",
      "metadata": {
        "_cell_guid": "6d769004-5b83-6ece-4c5c-32598fe41434"
      },
      "source": [
        "**AdaBoost**"
      ]
    },
    {
      "cell_type": "code",
      "execution_count": null,
      "metadata": {
        "_cell_guid": "975ed22a-6b85-0cb8-7417-a69262942eda"
      },
      "outputs": [],
      "source": [
        "abc = ensemble.AdaBoostClassifier(n_estimators = 50)\n",
        "abc.fit(np.hstack([X_num,X_cat]), y)\n",
        "abc_pred = abc.predict(np.hstack([X_num,X_cat]))\n",
        "cv_score_abc = cross_validation.cross_val_score(abc, np.hstack([X_num,X_cat]), y, cv = 10).mean()\n",
        "print (cv_score_abc)"
      ]
    },
    {
      "cell_type": "markdown",
      "metadata": {
        "_cell_guid": "5a3ea606-c32b-f8e3-1ead-5b7a6bab4fcd"
      },
      "source": [
        "**Gradient Boosting**"
      ]
    },
    {
      "cell_type": "code",
      "execution_count": null,
      "metadata": {
        "_cell_guid": "6219b8b8-fae9-6a4c-c52d-7bbd7136e658"
      },
      "outputs": [],
      "source": [
        "bc = ensemble.GradientBoostingClassifier(n_estimators = 100,learning_rate=1,max_depth=1, random_state=0)\n",
        "bc.fit(np.hstack([X_num,X_cat]), y)\n",
        "bc_pred = bc.predict(np.hstack([X_num,X_cat]))\n",
        "cv_score_bc = cross_validation.cross_val_score(bc, np.hstack([X_num,X_cat]), y, cv = 10).mean()\n",
        "print (cv_score_bc)"
      ]
    },
    {
      "cell_type": "code",
      "execution_count": null,
      "metadata": {
        "_cell_guid": "9d088153-6717-d2dc-8682-ba59d21dab89"
      },
      "outputs": [],
      "source": [
        "res = {'algorithm':['Logistic Regression','SVC','Decision Tree', 'Random Forest','Bagging','Bagging with features','AdaBoost','Gradient Boosting'],\n",
        "      'accuracy':[cv_score_lr,cv_score_svc,cv_score_clf,cv_score_rt,cv_score_bag,cv_score_bagf,cv_score_abc,cv_score_bc]}"
      ]
    },
    {
      "cell_type": "code",
      "execution_count": null,
      "metadata": {
        "_cell_guid": "ee6b3c4e-7eaf-5d2c-2b14-05be7b93fe5d"
      },
      "outputs": [],
      "source": [
        "import pandas as pd\n",
        "df = pd.DataFrame(res)\n",
        "df.sort_values('accuracy').head(10)"
      ]
    },
    {
      "cell_type": "markdown",
      "metadata": {
        "_cell_guid": "639958c7-9458-76c4-598f-2568583a25fb"
      },
      "source": [
        "**Now, the most appropriate algorithm is Bagging with features**"
      ]
    },
    {
      "cell_type": "code",
      "execution_count": null,
      "metadata": {
        "_cell_guid": "a168e1fd-0954-2ffb-1154-fa51aedc7418"
      },
      "outputs": [],
      "source": [
        ""
      ]
    }
  ],
  "metadata": {
    "_change_revision": 0,
    "_is_fork": false,
    "kernelspec": {
      "display_name": "Python 3",
      "language": "python",
      "name": "python3"
    },
    "language_info": {
      "codemirror_mode": {
        "name": "ipython",
        "version": 3
      },
      "file_extension": ".py",
      "mimetype": "text/x-python",
      "name": "python",
      "nbconvert_exporter": "python",
      "pygments_lexer": "ipython3",
      "version": "3.6.0"
    }
  },
  "nbformat": 4,
  "nbformat_minor": 0
}