{"cells":[{"metadata":{"_uuid":"c12e32b4731b2b4eeca13051f3fa3703bc9c75b9"},"cell_type":"markdown","source":"# Titanic Data Science Solution(2) \n## Introduction\nThis notebook is the second part of [titanic-data-science-solution-jimmy-modified](https://www.kaggle.com/icwangjimmy/titanic-data-science-solution-jimmy-modified). In this notebook, I reuse other experts' kernels and build a very basic and simple introductory primer to the method of ensembling (combining) base learning models, in particular the variant of ensembling known as Stacking. In a nutshell stacking uses as a first-level (base), the predictions of a few basic classifiers and then uses another model at the second-level (meta) to predict the output from the earlier first-level predictions.\n\n"},{"metadata":{"_uuid":"d520194bd554c9ff3e2c1949dd71849291fd47df"},"cell_type":"markdown","source":"## Prepare Python Packages"},{"metadata":{"trusted":true,"_uuid":"26d3d9538a99f10944cf4b1d99ee810f6c3167eb"},"cell_type":"code","source":"import pandas as pd\nimport numpy as np\n# import re\nimport sklearn\nimport xgboost as xgb\nimport seaborn as sns\nimport matplotlib.pyplot as plt\n%matplotlib inline\n\nimport plotly.offline as py\npy.init_notebook_mode(connected=True)\nimport plotly.graph_objs as go\nimport plotly.tools as tls\n\nimport warnings\nwarnings.filterwarnings('ignore')\n\n# Going to use these 5 base models for the stacking\nfrom sklearn.ensemble import (RandomForestClassifier, AdaBoostClassifier, \n                              GradientBoostingClassifier, ExtraTreesClassifier)\n\n# from sklearn.cross_validation import KFold\nfrom sklearn.model_selection import KFold\nfrom sklearn.linear_model import LogisticRegression\nfrom sklearn.svm import SVC, LinearSVC\n# from sklearn.ensemble import RandomForestClassifier\nfrom sklearn.neighbors import KNeighborsClassifier\nfrom sklearn.linear_model import Perceptron\nfrom sklearn.linear_model import SGDClassifier\nfrom sklearn.tree import DecisionTreeClassifier\nfrom sklearn.neural_network import MLPClassifier\n\n# for hyperparameter search\nfrom sklearn import model_selection\nfrom sklearn.metrics import accuracy_score\nfrom sklearn.model_selection import GridSearchCV","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"2d307b99ee3d19da3c1cddf509ed179c21dec94a","_cell_guid":"6b5dc743-15b1-aac6-405e-081def6ecca1"},"cell_type":"markdown","source":"<a id=\"step2\"></a>\n# 重复一遍前一个notebook的数据操作，为新的model做准备！\n# Acquire training and testing data\n"},{"metadata":{"_uuid":"13f38775c12ad6f914254a08f0d1ef948a2bd453","_cell_guid":"e7319668-86fe-8adc-438d-0eef3fd0a982","trusted":true},"cell_type":"code","source":"train_df = pd.read_csv('../input/train.csv')\ntest_df = pd.read_csv('../input/test.csv')\ncombine = [train_df, test_df]","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"8fd4f208a6e725bb5c478f358b88c6391a82acef"},"cell_type":"code","source":"train_df.info()","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"fd440648da1551ad70821c13563247ea0377f334"},"cell_type":"code","source":"train_df.head()","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"195d47e0187f7beeb598de5d36f294be1a37db72"},"cell_type":"code","source":"test_df.head()","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"73a9111a8dc2a6b8b6c78ef628b6cae2a63fc33f","_cell_guid":"cfac6291-33cc-506e-e548-6cad9408623d"},"cell_type":"markdown","source":"<a id=\"step4\"></a>\n# Wrangle, prepare, cleanse the data\n- Wrangle data (Feature Engineering)"},{"metadata":{"_uuid":"e328d9882affedcfc4c167aa5bb1ac132547558c","_cell_guid":"da057efe-88f0-bf49-917b-bb2fec418ed9","trusted":true},"cell_type":"code","source":"print(\"Before\", train_df.shape, test_df.shape, combine[0].shape, combine[1].shape)\n\ntrain_df = train_df.drop(['Ticket', 'Cabin'], axis=1)\ntest_df = test_df.drop(['Ticket', 'Cabin'], axis=1)\ncombine = [train_df, test_df]\n\nprint(\"After\", train_df.shape, test_df.shape, combine[0].shape, combine[1].shape)","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"2e0331b4b1c58ed3d310df504c39775e245b77a7"},"cell_type":"code","source":"for dataset in combine:\n    dataset['Title'] = dataset.Name.str.extract(' ([A-Za-z]+)\\.', expand=False)\n\npd.crosstab(train_df['Title'], train_df['Sex'])","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"b8cd938fba61fb4e226c77521b012f4bb8aa01d0","_cell_guid":"553f56d7-002a-ee63-21a4-c0efad10cfe9","trusted":true},"cell_type":"code","source":"for dataset in combine:\n    dataset['Title'] = dataset['Title'].replace(['Lady', 'Countess','Capt', 'Col', \n                                                 'Don', 'Dr', 'Major', 'Rev', 'Sir', 'Jonkheer', 'Dona'], 'Rare')\n\n    dataset['Title'] = dataset['Title'].replace('Mlle', 'Miss')\n    dataset['Title'] = dataset['Title'].replace('Ms', 'Miss')\n    dataset['Title'] = dataset['Title'].replace('Mme', 'Mrs')\n    \ntrain_df[['Title', 'Survived']].groupby(['Title'], as_index=False).mean()","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"de245fe76474d46995a5acc31b905b8aaa5893f6","_cell_guid":"6d46be9a-812a-f334-73b9-56ed912c9eca"},"cell_type":"markdown","source":"We can convert the categorical titles to ordinal."},{"metadata":{"_uuid":"e805ad52f0514497b67c3726104ba46d361eb92c","_cell_guid":"67444ebc-4d11-bac1-74a6-059133b6e2e8","trusted":true},"cell_type":"code","source":"title_mapping = {\"Mr\": 1, \"Miss\": 2, \"Mrs\": 3, \"Master\": 4, \"Rare\": 5}\nfor dataset in combine:\n    dataset['Title'] = dataset['Title'].map(title_mapping)\n    dataset['Title'] = dataset['Title'].fillna(0)\n\ntrain_df.head()","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"5fefaa1b37c537dda164c87a757fe705a99815d9","_cell_guid":"f27bb974-a3d7-07a1-f7e4-876f6da87e62"},"cell_type":"markdown","source":"Now we can safely drop the Name feature from training and testing datasets. We also do not need the PassengerId feature in the training dataset."},{"metadata":{"_uuid":"1da299cf2ffd399fd5b37d74fb40665d16ba5347","_cell_guid":"9d61dded-5ff0-5018-7580-aecb4ea17506","trusted":true},"cell_type":"code","source":"train_df = train_df.drop(['Name', 'PassengerId'], axis=1)\ntest_df = test_df.drop(['Name'], axis=1)\ncombine = [train_df, test_df]\ntrain_df.shape, test_df.shape","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"48f2278363295957905f71798ae265af7a7fae0d"},"cell_type":"code","source":"train_df.head()","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"a1ac66c79b279d94860e66996d3d8dba801a6d9a","_cell_guid":"2c8e84bb-196d-bd4a-4df9-f5213561b5d3"},"cell_type":"markdown","source":"### Converting a categorical feature\n\nNow we can convert features which contain strings to numerical values. This is required by most model algorithms. Doing so will also help us in achieving the feature completing goal.\n\nLet us start by converting Sex feature to a new feature called Gender where female=1 and male=0."},{"metadata":{"_uuid":"840498eaee7baaca228499b0a5652da9d4edaf37","_cell_guid":"c20c1df2-157c-e5a0-3e24-15a828095c96","trusted":true},"cell_type":"code","source":"for dataset in combine:\n    dataset['Sex'] = dataset['Sex'].map( {'female': 1, 'male': 0} ).astype(int)\n\ntrain_df.head()","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"6da8bfe6c832f4bd2aa1312bdd6b8b4af48a012e","_cell_guid":"d72cb29e-5034-1597-b459-83a9640d3d3a"},"cell_type":"markdown","source":"### Completing a numerical continuous feature\n\nNow we should start estimating and completing features with missing or null values. We will first do this for the Age feature.\n\nWe can consider three methods to complete a numerical continuous feature.\n\n1. A simple way is to generate random numbers between mean and [standard deviation](https://en.wikipedia.org/wiki/Standard_deviation).\n\n**2. More accurate way of guessing missing values is to use other correlated features. In our case we note correlation among Age, Gender, and Pclass. Guess Age values using [median](https://en.wikipedia.org/wiki/Median) values for Age across sets of Pclass and Gender feature combinations. So, median Age for Pclass=1 and Gender=0, Pclass=1 and Gender=1, and so on...**\n\n3. Combine methods 1 and 2. So instead of guessing age values based on median, use random numbers between mean and standard deviation, based on sets of Pclass and Gender combinations.\n\nMethod 1 and 3 will introduce random noise into our models. The results from multiple executions might vary. We will prefer method 2."},{"metadata":{"_uuid":"6b22ac53d95c7979d5f4580bd5fd29d27155c347","_cell_guid":"a4f166f9-f5f9-1819-66c3-d89dd5b0d8ff"},"cell_type":"markdown","source":"Let us start by preparing an empty array to contain guessed Age values based on Pclass x Gender combinations."},{"metadata":{"_uuid":"24a0971daa4cbc3aa700bae42e68c17ce9f3a6e2","_cell_guid":"9299523c-dcf1-fb00-e52f-e2fb860a3920","trusted":true},"cell_type":"code","source":"guess_ages = np.zeros((2,3))\nguess_ages","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"8acd90569767b544f055d573bbbb8f6012853385","_cell_guid":"ec9fed37-16b1-5518-4fa8-0a7f579dbc82"},"cell_type":"markdown","source":"Now we iterate over Sex (0 or 1) and Pclass (1, 2, 3) to calculate guessed values of Age for the six combinations."},{"metadata":{"_uuid":"31198f0ad0dbbb74290ebe135abffa994b8f58f3","_cell_guid":"a4015dfa-a0ab-65bc-0cbe-efecf1eb2569","trusted":true},"cell_type":"code","source":"for dataset in combine:\n    for i in range(0, 2):\n        for j in range(0, 3):\n            guess_df = dataset[(dataset['Sex'] == i) & (dataset['Pclass'] == j+1)]['Age'].dropna()\n\n            # age_mean = guess_df.mean()\n            # age_std = guess_df.std()\n            # age_guess = rnd.uniform(age_mean - age_std, age_mean + age_std)\n\n            age_guess = guess_df.median()\n\n            # Convert random age float to nearest .5 age\n            guess_ages[i,j] = int( age_guess/0.5 + 0.5 ) * 0.5\n            \n    for i in range(0, 2):\n        for j in range(0, 3):\n            dataset.loc[(dataset.Age.isnull()) & (dataset.Sex == i) & (dataset.Pclass == j+1),'Age'] = guess_ages[i,j]\n\n    dataset['Age'] = dataset['Age'].astype(int)\n\ntrain_df.head()","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"e7c52b44b703f28e4b6f4ddba67ab65f40274550","_cell_guid":"dbe0a8bf-40bc-c581-e10e-76f07b3b71d4"},"cell_type":"markdown","source":"Let us create Age bands and determine correlations with Survived."},{"metadata":{"_uuid":"5c8b4cbb302f439ef0d6278dcfbdafd952675353","_cell_guid":"725d1c84-6323-9d70-5812-baf9994d3aa1","trusted":true},"cell_type":"code","source":"train_df['AgeBand'] = pd.cut(train_df['Age'], 5)\ntrain_df[['AgeBand', 'Survived']].groupby(['AgeBand'], as_index=False).mean().sort_values(by='AgeBand', ascending=True)","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"856392dd415ac14ab74a885a37d068fc7a58f3a5","_cell_guid":"ba4be3a0-e524-9c57-fbec-c8ecc5cde5c6"},"cell_type":"markdown","source":"Let us replace Age with ordinals based on these bands."},{"metadata":{"_uuid":"ee13831345f389db407c178f66c19cc8331445b0","_cell_guid":"797b986d-2c45-a9ee-e5b5-088de817c8b2","trusted":true},"cell_type":"code","source":"for dataset in combine:    \n    dataset.loc[ dataset['Age'] <= 16, 'Age'] = 0\n    dataset.loc[(dataset['Age'] > 16) & (dataset['Age'] <= 32), 'Age'] = 1\n    dataset.loc[(dataset['Age'] > 32) & (dataset['Age'] <= 48), 'Age'] = 2\n    dataset.loc[(dataset['Age'] > 48) & (dataset['Age'] <= 64), 'Age'] = 3\n    dataset.loc[ dataset['Age'] > 64, 'Age']\ntrain_df.head()","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"8e3fbc95e0fd6600e28347567416d3f0d77a24cc","_cell_guid":"004568b6-dd9a-ff89-43d5-13d4e9370b1d"},"cell_type":"markdown","source":"We can now remove the AgeBand feature."},{"metadata":{"_uuid":"1ea01ccc4a24e8951556d97c990aa0136da19721","_cell_guid":"875e55d4-51b0-5061-b72c-8a23946133a3","trusted":true},"cell_type":"code","source":"train_df = train_df.drop(['AgeBand'], axis=1)\ncombine = [train_df, test_df]\ntrain_df.head()","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"e3d4a2040c053fbd0486c8cfc4fec3224bd3ebb3","_cell_guid":"1c237b76-d7ac-098f-0156-480a838a64a9"},"cell_type":"markdown","source":"### Create new feature combining existing features\n\nWe can create a new feature for FamilySize which combines Parch and SibSp. This will enable us to drop Parch and SibSp from our datasets."},{"metadata":{"_uuid":"33d1236ce4a8ab888b9fac2d5af1c78d174b32c7","_cell_guid":"7e6c04ed-cfaa-3139-4378-574fd095d6ba","trusted":true},"cell_type":"code","source":"for dataset in combine:\n    dataset['FamilySize'] = dataset['SibSp'] + dataset['Parch'] + 1\n\ntrain_df[['FamilySize', 'Survived']].groupby(['FamilySize'], as_index=False).mean().sort_values(by='Survived', ascending=False)","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"67f8e4474cd1ecf4261c153ce8b40ea23cf659e4","_cell_guid":"842188e6-acf8-2476-ccec-9e3451e4fa86"},"cell_type":"markdown","source":"We can create another feature called IsAlone."},{"metadata":{"_uuid":"3b8db81cc3513b088c6bcd9cd1938156fe77992f","_cell_guid":"5c778c69-a9ae-1b6b-44fe-a0898d07be7a","trusted":true},"cell_type":"code","source":"for dataset in combine:\n    dataset['IsAlone'] = 0\n    dataset.loc[dataset['FamilySize'] == 1, 'IsAlone'] = 1\n\ntrain_df[['IsAlone', 'Survived']].groupby(['IsAlone'], as_index=False).mean()","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"3da4204b2c78faa54a94bbad78a8aa85fbf90c87","_cell_guid":"e6b87c09-e7b2-f098-5b04-4360080d26bc"},"cell_type":"markdown","source":"Let us drop Parch, SibSp, and FamilySize features in favor of IsAlone."},{"metadata":{"_uuid":"1e3479690ef7cd8ee10538d4f39d7117246887f0","_cell_guid":"74ee56a6-7357-f3bc-b605-6c41f8aa6566","trusted":true},"cell_type":"code","source":"# train_df = train_df.drop(['Parch', 'SibSp', 'FamilySize'], axis=1)\n# test_df = test_df.drop(['Parch', 'SibSp', 'FamilySize'], axis=1)\ntrain_df = train_df.drop(['Parch', 'SibSp'], axis=1)\ntest_df = test_df.drop(['Parch', 'SibSp'], axis=1)\n\ncombine = [train_df, test_df]\n\ntrain_df.head()","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"71b800ed96407eba05220f76a1288366a22ec887","_cell_guid":"f890b730-b1fe-919e-fb07-352fbd7edd44"},"cell_type":"markdown","source":"We can also create an artificial feature combining Pclass and Age."},{"metadata":{"_uuid":"aac2c5340c06210a8b0199e15461e9049fbf2cff","_cell_guid":"305402aa-1ea1-c245-c367-056eef8fe453","trusted":true},"cell_type":"code","source":"for dataset in combine:\n    dataset['Age*Class'] = dataset.Age * dataset.Pclass\n\ntrain_df.loc[:, ['Age*Class', 'Age', 'Pclass']].head(10)","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"8264cc5676db8cd3e0b3e3f078cbaa74fd585a3c","_cell_guid":"13292c1b-020d-d9aa-525c-941331bb996a"},"cell_type":"markdown","source":"### Completing a categorical feature\n\nEmbarked feature takes S, Q, C values based on port of embarkation. Our training dataset has two missing values. We simply fill these with the most common occurance."},{"metadata":{"_uuid":"1e3f8af166f60a1b3125a6b046eff5fff02d63cf","_cell_guid":"bf351113-9b7f-ef56-7211-e8dd00665b18","trusted":true},"cell_type":"code","source":"freq_port = train_df.Embarked.dropna().mode()[0]\nfreq_port","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"d85b5575fb45f25749298641f6a0a38803e1ff22","_cell_guid":"51c21fcc-f066-cd80-18c8-3d140be6cbae","trusted":true},"cell_type":"code","source":"for dataset in combine:\n    dataset['Embarked'] = dataset['Embarked'].fillna(freq_port)\n    \ntrain_df[['Embarked', 'Survived']].groupby(['Embarked'], as_index=False).mean().sort_values(by='Survived', ascending=False)","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"d8830e997995145314328b6218b5606df04499b0","_cell_guid":"f6acf7b2-0db3-e583-de50-7e14b495de34"},"cell_type":"markdown","source":"### Converting categorical feature to numeric\n\nWe can now convert the EmbarkedFill feature by creating a new numeric Port feature."},{"metadata":{"_uuid":"e480a1ef145de0b023821134896391d568a6f4f9","_cell_guid":"89a91d76-2cc0-9bbb-c5c5-3c9ecae33c66","trusted":true},"cell_type":"code","source":"for dataset in combine:\n    dataset['Embarked'] = dataset['Embarked'].map( {'S': 0, 'C': 1, 'Q': 2} ).astype(int)\n\ntrain_df.head()","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"d79834ebc4ab9d48ed404584711475dbf8611b91","_cell_guid":"e3dfc817-e1c1-a274-a111-62c1c814cecf"},"cell_type":"markdown","source":"### Quick completing and converting a numeric feature\n\nWe can now complete the Fare feature for single missing value in test dataset using mode to get the value that occurs most frequently for this feature. We do this in a single line of code.\n\nNote that we are not creating an intermediate new feature or doing any further analysis for correlation to guess missing feature as we are replacing only a single value. The completion goal achieves desired requirement for model algorithm to operate on non-null values.\n\nWe may also want round off the fare to two decimals as it represents currency."},{"metadata":{"_uuid":"aacb62f3526072a84795a178bd59222378bab180","_cell_guid":"3600cb86-cf5f-d87b-1b33-638dc8db1564","trusted":true},"cell_type":"code","source":"# test_df['Fare'].fillna(test_df['Fare'].dropna().median(), inplace=True)\ntest_df['Fare'].fillna(test_df['Fare'].dropna().mode()[0], inplace=True)\ntest_df.head()","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"3466d98e83899d8b38a36ede794c68c5656f48e6","_cell_guid":"4b816bc7-d1fb-c02b-ed1d-ee34b819497d"},"cell_type":"markdown","source":"We can now create FareBand."},{"metadata":{"_uuid":"b9a78f6b4c72520d4ad99d2c89c84c591216098d","_cell_guid":"0e9018b1-ced5-9999-8ce1-258a0952cbf2","trusted":true},"cell_type":"code","source":"train_df['FareBand'] = pd.qcut(train_df['Fare'], 4)\ntrain_df[['FareBand', 'Survived']].groupby(['FareBand'], as_index=False).mean().sort_values(by='FareBand', ascending=True)","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"89400fba71af02d09ff07adf399fb36ac4913db6","_cell_guid":"d65901a5-3684-6869-e904-5f1a7cce8a6d"},"cell_type":"markdown","source":"Convert the Fare feature to ordinal values based on the FareBand."},{"metadata":{"_uuid":"640f305061ec4221a45ba250f8d54bb391035a57","_cell_guid":"385f217a-4e00-76dc-1570-1de4eec0c29c","trusted":true},"cell_type":"code","source":"for dataset in combine:\n    dataset.loc[ dataset['Fare'] <= 7.91, 'Fare'] = 0\n    dataset.loc[(dataset['Fare'] > 7.91) & (dataset['Fare'] <= 14.454), 'Fare'] = 1\n    dataset.loc[(dataset['Fare'] > 14.454) & (dataset['Fare'] <= 31), 'Fare']   = 2\n    dataset.loc[ dataset['Fare'] > 31, 'Fare'] = 3\n    dataset['Fare'] = dataset['Fare'].astype(int)\n\ntrain_df = train_df.drop(['FareBand'], axis=1)\ncombine = [train_df, test_df]\n    ","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"640f305061ec4221a45ba250f8d54bb391035a57","_cell_guid":"385f217a-4e00-76dc-1570-1de4eec0c29c","trusted":true},"cell_type":"code","source":"train_df.head(10)","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"531994ed95a3002d1759ceb74d9396db706a41e2","_cell_guid":"27272bb9-3c64-4f9a-4a3b-54f02e1c8289"},"cell_type":"markdown","source":"And the test dataset."},{"metadata":{"_uuid":"8453cecad81fcc44de3f4e4e4c3ce6afa977740d","_cell_guid":"d2334d33-4fe5-964d-beac-6aa620066e15","trusted":true},"cell_type":"code","source":"test_df.head(10)","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"9412bbaf0609c04b1fc7995eece6d5a1b84b230b"},"cell_type":"code","source":"# Store our passenger ID for easy access\nPassengerId = test_df['PassengerId']","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"13501e742fba8caf023818fb0a21620bc9a765f4"},"cell_type":"markdown","source":"# We are ready to model again!!!"},{"metadata":{"_uuid":"966d0db2bf6738cef4730b87d6c36d4626a25b93"},"cell_type":"markdown","source":"## Some key points about model selection and construction 模型选择与构建的几个关键问题"},{"metadata":{"_uuid":"8e0f39a48e491629982e38156a2a6d34fc37f73a"},"cell_type":"markdown","source":"## Overfitting Discussion 过度拟合讨论\n![](https://github.com/icwangjimmy/machine_learning_and_python_in_finance/raw/master/pic/overfitting.png)\nAs you see in the chart above. Underfitting is when the model fails to capture important aspects of the data and therefore introduces more bias and performs poorly. On the other hand, Overfitting is when the model performs too well on the training data but does poorly in the validation set or test sets. This situation is also known as having less bias but more variation and perform poorly as well. Ideally we want to configure a model that performs well not only in the training data, but also in the test data. This is where bias-variance tradeoff comes in. When we have a model that overfits meaning more less biased and more possible chance of variance, we introduce some bais in exchance of having much less variance. "},{"metadata":{"_uuid":"55185649e7edee2fce22c2a34ea447eb22d92bb5"},"cell_type":"markdown","source":"## The holdout method 保持法\nA classic and popular approach for estimating the generalization performance of machine learning models is holdout cross-validation. Using the holdout method, we **split our initial dataset into a separate training and test dataset—the former is used for model training, and the latter is used to estimate its performance.** However, in typical machine learning applications, we are also interested in tuning and comparing different parameter settings to further improve the performance for making predictions on unseen data. This process is called model selection, where the term model selection refers to a given classification problem for which we want to select the optimal values of tuning parameters (also called hyperparameters). However, if we reuse the same test dataset over and over again during model selection, it will become part of our training data and thus the model will be more likely to overfit. Despite this issue, many people still use the test set for model selection, which is not a good machine learning practice.\n\nA better way of using the holdout method for model selection is to **separate the data into three parts: a training set, a validation set, and a test set.** The training set is used to fit the different models, and the performance on the validation set is then used for the model selection. The advantage of having a test set that the model hasn't seen before during the training and model selection steps is that we can obtain a less biased estimate of its ability to generalize to new data. The following figure illustrates the concept of holdout cross-validation where we use a validation set to repeatedly evaluate the performance of the model after training using different parameter values. Once we are satisfied with the tuning of parameter values, we estimate the models' generalization error on the test dataset:\n![](https://github.com/icwangjimmy/machine_learning_and_python_in_finance/raw/master/pic/holdout.png)\nA disadvantage of the holdout method is that the performance estimate is sensitive to how we partition the training set into the training and validation subsets; the estimate will vary for different samples of the data. In the next subsection, we will take a look at a more robust technique for performance estimation, k-fold cross-validation, where we repeat the holdout method k times on k subsets of the training data."},{"metadata":{"_uuid":"19b749ace35c308f0a625b7d6ba99728c3369e8c"},"cell_type":"markdown","source":"## K-fold cross-validation K折交叉确认法\nIn k-fold cross-validation, we randomly split the training dataset into k folds without replacement, where k-1 folds are used for the model training and one fold is used for testing. This procedure is repeated k times so that we obtain k models and performance estimates.\n\nWe then calculate the average performance of the models based on the different, independent folds to obtain a performance estimate that is less sensitive to the subpartitioning of the training data compared to the holdout method. Typically, we use k-fold cross-validation for model tuning, that is, finding the optimal hyperparameter values that yield a satisfying generalization performance. Once we have found satisfactory hyperparameter values, we can retrain the model on the complete training set and obtain a final performance estimate using the independent test set.\n\n**Since k-fold cross-validation is a resampling technique without replacement, the advantage of this approach is that each sample point will be part of a training and test dataset exactly once, which yields a lower-variance estimate of the model performance than the holdout method.**\n![](https://github.com/icwangjimmy/machine_learning_and_python_in_finance/raw/master/pic/kfold1.png)"},{"metadata":{"_uuid":"3bfbfe06f855e62d8887e8b810f4a474b9da8ea0"},"cell_type":"markdown","source":"## Model Stacking 模型堆叠\nEnsemble Learning is a learning mechanism that uses multiple learning algorithms to obtain better predictive performance than could be obtained from any of the constituent learning algorithms. Many of the popular modern machine learning algorithms are actually ensembles.In short ,this is a technique that take a collection of **weak learners** and **form a single, strong learner**.For example Random Forests(Bagging) and Gradient Boosting(Boosting)are both ensemble learners.Stacking is a technique that belongs to these family of learning.\n\nStacking-a Meta Modeling Technique is introduced by Wolpert in the year 1992.In Stacking there are two types of learners called **Base Learners** and a **Meta Learner**.Base Learners and Meta Learners are the normal machine learning algorithms like Random Forests, SVM, Perceptron etc. Base Learners try to fit the normal data sets where as Meta learner fit on the predictions of the base Learner.\n![](https://github.com/icwangjimmy/machine_learning_and_python_in_finance/raw/master/pic/stack1.png)\n\n### Stacking Technique involves the following Steps:-\n\n- 1.Split the training data into 2 disjoint sets\n- 2.Train several Base Learners on the first part\n- 3.Test the Base Learners on the second part and make predictions\n- 4.Using the predictions from (3) as inputs,the correct responses from the output, train the higher level learner or meta level Learner\n\nMeta Learner is kind of trying to find the optimal combination of base learners. Let us take an example of classification problem where we are trying to classify 4 classes and as a part of traditional paradigm we are testing various models and we find out that Logistic Regression is making better predictions on class 1 data and SVM is making better on class 2 and class 4 and KNN is doing better on class 3 and class 2.This performance is predicted because in general no model is perfect and has its own advantages and disadvantages. So, if we train a model on the predictions of these model can we get better results? This is the idea on which this entire concept is built upon. So if we train a Random Forest Classifier on these predictions of LR, SVM, KNN we get better results.\n![](https://github.com/icwangjimmy/machine_learning_and_python_in_finance/raw/master/pic/stack2.gif)\n\nLets Try to give shape to this technique:-\n\n## Set up the ensemble.\n- Specify a list of L base algorithms (with a specific set of model parameters).\n- Specify a meta learning algorithm.\n\n## Train the ensemble.\n- Train each of the L base algorithms on the training set.\n- Perform k-fold cross-validation on each of these learners and collect the cross-validated predicted values from each of the L algorithms.\n- Train the meta learning algorithm on the level-one data. The “ensemble model” consists of the L base learning models and the meta learning model, which can then be used to generate predictions on a test set.\n\n## Predict on new data.\n- To generate ensemble predictions, first generate predictions from the base learners.\n- Feed those predictions into the meta learner to generate the ensemble prediction."},{"metadata":{"trusted":true,"_uuid":"3aeca26aa45c8c8e1485711b60357d5f24d3ee80"},"cell_type":"markdown","source":"## N-Level Stacking 可继续扩展到多层堆叠\nThe concept of Stacking can be extended to many Levels.\n![](https://github.com/icwangjimmy/machine_learning_and_python_in_finance/raw/master/pic/stack3.png)\n\n"},{"metadata":{"_uuid":"a7f401c24d24c01ae53056b26218622d49983755"},"cell_type":"markdown","source":"\n## Ensembling & Stacking models\nFinally after repeating the feature engineering we have done before and preparing all the basic concepts of model selection and construction, we arrive at the meat and gist of the this notebook.\n\nCreating a Stacking ensemble! 现在用上面提到的概念来做个堆叠模型吧！"},{"metadata":{"_uuid":"6d3c6a35ee0bdef934ba9569eca24f5326e0e9a3"},"cell_type":"markdown","source":"### Helpers via Python Classes 面向对象编程一瞥\nHere we invoke the use of Python's classes to help make it more convenient for us. For any newcomers to programming, one normally hears Classes being used in conjunction with Object-Oriented Programming (OOP). In short, a class helps to extend some code/program for creating objects (variables for old-school peeps) as well as to implement functions and methods specific to that class.\n![](https://github.com/icwangjimmy/machine_learning_and_python_in_finance/raw/master/pic/python_class.png)\n\nIn the section of code below, we essentially write a class SklearnHelper that allows one to extend the inbuilt methods (such as train, predict and fit) common to all the Sklearn classifiers. Therefore this cuts out redundancy as won't need to write the same methods five times if we wanted to invoke five different classifiers."},{"metadata":{"trusted":true,"_uuid":"e2cc197a2707782aec7a39ef8dc980349aeabae3"},"cell_type":"code","source":"# Some useful parameters which will come in handy later on\ntrain = train_df\ntest = test_df\n\nntrain = train.shape[0]\nntest = test.shape[0]\nSEED = 0 # for reproducibility\nNFOLDS = 5 # set folds for out-of-fold prediction\n# kf = KFold(ntrain, n_folds= NFOLDS, random_state=SEED)\nkf = KFold(n_splits=NFOLDS, random_state=SEED)\n# Class to extend the Sklearn classifier\nclass SklearnHelper(object):\n    def __init__(self, clf, seed=0, has_random_state=1, params=None):\n        if(has_random_state==1):\n            params['random_state'] = seed\n        print(params)\n        self.clf = clf(**params)\n\n    def train(self, x_train, y_train):\n        self.clf.fit(x_train, y_train)\n\n    def predict(self, x):\n        return self.clf.predict(x)\n    \n    def predict_proba(self, x):\n        return self.clf.predict_proba(x)\n    \n    def fit(self,x,y):\n        return self.clf.fit(x,y)\n    \n    def feature_importances(self,x,y):\n        print(self.clf.fit(x,y).feature_importances_)","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"a4e39a386e4ce79067d61087f67321db6538ade9"},"cell_type":"markdown","source":"Bear with me for those who already know this but for people who have not created classes or objects in Python before, let me explain what the code given above does. In creating my base classifiers, I will only use the models already present in the Sklearn library and therefore only extend the class for that.\n\ndef init : Python standard for invoking the default constructor for the class. This means that when you want to create an object (classifier), you have to give it the parameters of clf (what sklearn classifier you want), seed (random seed) and params (parameters for the classifiers).\n\nThe rest of the code are simply methods of the class which simply call the corresponding methods already existing within the sklearn classifiers. Essentially, we have created a wrapper class to extend the various Sklearn classifiers so that this should help us reduce having to write the same code over and over when we implement multiple learners to our stacker."},{"metadata":{"_uuid":"c02f0449cd5402b388b6a2ab1ff3c5be19936d73"},"cell_type":"markdown","source":"### Out-of-Fold Predictions\n\nNow as alluded to above in the introductory section, stacking uses predictions of base classifiers as input for training to a second-level model. However one cannot simply train the base models on the full training data, generate predictions on the full test set and then output these for the second-level training. This runs the risk of your base model predictions already having \"seen\" the test set and therefore overfitting when feeding these predictions.\n\nBasically, you need to cut training data into two parts (train and validation), and train the base model with train data, then predict with validation data, so it can pass the validation data prediction as meta features to the meta model for second-level model training.\n![](https://github.com/icwangjimmy/machine_learning_and_python_in_finance/raw/master/pic/kfold.png)"},{"metadata":{"trusted":true,"_uuid":"f3e6df250d5a7fb7b74204b735d4bde158ddc988"},"cell_type":"code","source":"def get_oof(clf, x_train, y_train, x_test):\n    oof_train = np.zeros((ntrain,)) #record train set prediction result for each test_index\n    oof_test = np.zeros((ntest,)) #record test set prediction result for final prediction\n    oof_test_skf = np.empty((NFOLDS, ntest)) #record prediciton from each clf for all the x_test set\n\n#     for i, (train_index, test_index) in enumerate(kf):\n    for i, (train_index, test_index) in enumerate(kf.split(train)):    \n#         print(\"TRAIN:\", train_index, \"TEST:\", test_index)\n        x_tr = x_train[train_index]\n        y_tr = y_train[train_index]\n        x_te = x_train[test_index]\n\n        clf.train(x_tr, y_tr)\n\n#         oof_train[test_index] = clf.predict(x_te) #each clf has a validation result, so totally 5 meta features\n#         oof_test_skf[i, :] = clf.predict(x_test) #test prediction for different clf from different training set under CV\n#         print(clf.predict_proba(x_te)[:,1].shape)\n        oof_train[test_index] = clf.predict_proba(x_te)[:,1] #each clf has a validation result, so totally 5 meta features\n        oof_test_skf[i, :] = clf.predict_proba(x_test)[:,1] #test prediction for different clf from different training set under CV\n\n    oof_test[:] = oof_test_skf.mean(axis=0)\n    return oof_train.reshape(-1, 1), oof_test.reshape(-1, 1)","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"10eeb6ae81cbc64696ef0d993795878bf2b89669"},"cell_type":"markdown","source":"# Generating our Base First-Level Models \n\nSo now let us prepare seven learning models as our first level classification. These models can all be conveniently invoked via the Sklearn library and are listed as follows:\n\n 1. Random Forest classifier\n 2. Extra Trees classifier\n 3. AdaBoost classifer\n 4. Gradient Boosting classifer\n 5. Support Vector Machine\n 6. Logistic Regression\n 7. K-Nearest Neighbors"},{"metadata":{"_uuid":"a43be976ada2f01d171c9faf44b3858815e5ebbe"},"cell_type":"markdown","source":"**Parameters**\n\nJust a quick summary of the parameters that we will be listing here for completeness,\n\n**n_jobs** : Number of cores used for the training process. If set to -1, all cores are used.\n\n**n_estimators** : Number of classification trees in your learning model ( set to 10 per default)\n\n**max_depth** : Maximum depth of tree, or how much a node should be expanded. Beware if set to too high  a number would run the risk of overfitting as one would be growing the tree too deep\n\n**verbose** : Controls whether you want to output any text during the learning process. A value of 0 suppresses all text while a value of 3 outputs the tree learning process at every iteration.\n\n Please check out the full description via the official Sklearn website. There you will find that there are a whole host of other useful parameters that you can play around with. "},{"metadata":{"trusted":true,"_uuid":"11eb2d159377da1272bfc2cf351b7e5de6b205bc"},"cell_type":"code","source":"# Put in our parameters for said classifiers\n# Random Forest parameters\n# rf_params = {\n#     'n_jobs': -1,\n#     'n_estimators': 100,\n#      'warm_start': True, \n#     'max_depth': 6,\n#     'min_samples_leaf': 2,\n#     'max_features' : 'sqrt',\n#     'verbose': 0\n# }\n\n# RandomForest parameters we got from titanic-data-science-solution-jimmy-modified notebook\nrf_params = {'bootstrap': True, \n             'criterion': 'entropy',\n             'max_depth': 5, \n             'max_features': 'sqrt', \n             'min_samples_leaf': 1, \n             'min_samples_split': 5,\n             'min_weight_fraction_leaf': 0.0, \n             'n_estimators': 100, \n             'n_jobs': -1,\n             'verbose': 0,\n             'warm_start': False\n             }\n\n# Extra Trees Parameters\net_params = {\n    'n_jobs': -1,\n    'n_estimators':500,\n    #'max_features': 0.5,\n    'max_depth': 8,\n    'min_samples_leaf': 2,\n    'verbose': 0\n}\n\n# AdaBoost parameters\nada_params = {\n    'n_estimators': 500,\n    'learning_rate' : 0.75\n}\n\n# Gradient Boosting parameters\ngb_params = {\n    'n_estimators': 500,\n     #'max_features': 0.2,\n    'max_depth': 5,\n    'min_samples_leaf': 2,\n    'verbose': 0\n}\n\n# Support Vector Classifier parameters \nsvc_params = {\n    'kernel' : 'linear',\n    'C' : 0.025,\n    'probability': True\n}\n\n# Logistic Regression\nlogreg_params = {    \n}\n\n# KNN\nknn_params ={\n    'n_neighbors': 3\n}\n","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"5c252286963b6d41bac7c5f43f77c3b32adca013"},"cell_type":"markdown","source":"Furthermore, since having mentioned about Objects and classes within the OOP framework, let us now create 5 objects that represent our 5 learning models via our Helper Sklearn Class we defined earlier."},{"metadata":{"trusted":true,"_uuid":"421ee8fe51dfab171fe986ffa57853a98322fd03"},"cell_type":"code","source":"# Create 5 objects that represent our 5 models\nrf = SklearnHelper(clf=RandomForestClassifier, seed=SEED, params=rf_params)\net = SklearnHelper(clf=ExtraTreesClassifier, seed=SEED, params=et_params)\nada = SklearnHelper(clf=AdaBoostClassifier, seed=SEED, params=ada_params)\ngb = SklearnHelper(clf=GradientBoostingClassifier, seed=SEED, params=gb_params)\nsvc = SklearnHelper(clf=SVC, seed=SEED, params=svc_params)\nlog = SklearnHelper(clf=LogisticRegression, seed=SEED, params=logreg_params)\nknn = SklearnHelper(clf=KNeighborsClassifier, seed=SEED, has_random_state=0, params=knn_params)","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"a4aa23c24d59ac3788d174a965e37a920ae7564d"},"cell_type":"markdown","source":"**Creating NumPy arrays out of our train and test sets**\n\nGreat. Having prepared our first layer base models as such, we can now ready the training and test test data for input into our classifiers by generating NumPy arrays out of their original dataframes as follows:"},{"metadata":{"trusted":true,"_uuid":"d96de94f01e32fb6627755b94d83614fcf25995e"},"cell_type":"code","source":"# Create Numpy arrays of train, test and target ( Survived) dataframes to feed into our models\ny_train = train['Survived'].ravel()\ntrain = train.drop(['Survived'], axis=1)\nx_train = train.values # Creates an array of the train data\ndrop_elements = ['PassengerId']\ntest  = test.drop(drop_elements, axis = 1)\nx_test = test.values # Creats an array of the test data","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"37755b58b8b2e181bcfdc64e0bd5503f1b51d364"},"cell_type":"code","source":"# Feature Scaling for non-tree based classifiers\n## We will be using standardscaler to transform\nfrom sklearn.preprocessing import StandardScaler\nsc = StandardScaler()\n\n## transforming \"train_x\"\nx_train = sc.fit_transform(x_train)\n\n## transforming \"The testset\"\nx_test = sc.transform(x_test)","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"517c96abc64a9fc51592a299331e291e5e040c23"},"cell_type":"markdown","source":"**Output of the First level Predictions** \n\nWe now feed the training and test data into our 5 base classifiers and use the Out-of-Fold prediction function we defined earlier to generate our first level predictions. Allow a handful of minutes for the chunk of code below to run."},{"metadata":{"trusted":true,"_uuid":"6e13868125a2ef0932865e605cdc7a986b67413f"},"cell_type":"code","source":"%%time\n# Create our OOF train and test predictions. These base results will be used as new features\net_oof_train, et_oof_test = get_oof(et, x_train, y_train, x_test) # Extra Trees\nrf_oof_train, rf_oof_test = get_oof(rf,x_train, y_train, x_test) # Random Forest\nada_oof_train, ada_oof_test = get_oof(ada, x_train, y_train, x_test) # AdaBoost \ngb_oof_train, gb_oof_test = get_oof(gb,x_train, y_train, x_test) # Gradient Boost\nsvc_oof_train, svc_oof_test = get_oof(svc,x_train, y_train, x_test) # Support Vector Classifier\nlog_oof_train, log_oof_test = get_oof(log,x_train, y_train, x_test) # Logistic Regression\nknn_oof_train, knn_oof_test = get_oof(knn,x_train, y_train, x_test) # KNN\nprint(\"Training is complete\")","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"c7c3cfd240d8729d8551ebc08ec5971fd88a80c4"},"cell_type":"markdown","source":"**Feature importances generated from the different classifiers**\n\nNow having learned our the first-level classifiers, we can utilise a very nifty feature of the Sklearn models and that is to output the importances of the various features in the training and test sets with one very simple line of code.\n\nAs per the Sklearn documentation, most of the classifiers are built in with an attribute which returns feature importances by simply typing in **.feature_importances_**. Therefore we will invoke this very useful attribute via our function earliand plot the feature importances as such"},{"metadata":{"trusted":true,"_uuid":"67ed2eff3d93169f6df498798c7b4bcae76252e3"},"cell_type":"code","source":"rf_feature = rf.feature_importances(x_train,y_train)\net_feature = et.feature_importances(x_train, y_train)\nada_feature = ada.feature_importances(x_train, y_train)\ngb_feature = gb.feature_importances(x_train,y_train)","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"bbd71aa2d14eb725410f1438875fcb79b5992b48"},"cell_type":"code","source":"rf_features = [0.13796186, 0.28025103, 0.0400734,  0.06107588, 0.02110676, 0.26240436,\n 0.09555471, 0.01240719, 0.08916479]\net_features = [0.16660931, 0.4105351,  0.02365337, 0.06920249, 0.03212467, 0.18273988,\n 0.05908173, 0.02362692, 0.03242655]\nada_features = [0.058, 0.214, 0.064, 0.044, 0.016, 0.386, 0.06,  0.008, 0.15 ]\ngb_features = [0.15048707, 0.01816345, 0.02917081, 0.05736517, 0.02881883, 0.54593532,\n 0.12270411, 0.00780863, 0.03954661]","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"9a408ec7ad13b929df304bd39cbe40965bd91570"},"cell_type":"markdown","source":"Create a dataframe from the lists containing the feature importance data for easy plotting via the Plotly package."},{"metadata":{"trusted":true,"_uuid":"58a92870f360b6e8f0d487e9832d149ccbf5a202"},"cell_type":"code","source":"cols = train.columns.values\n# Create a dataframe with features\nfeature_dataframe = pd.DataFrame( {'features': cols,\n     'Random Forest feature importances': rf_features,\n     'Extra Trees  feature importances': et_features,\n      'AdaBoost feature importances': ada_features,\n    'Gradient Boost feature importances': gb_features\n    })","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"96cb5140fc07f9bb6e0ab61b48b00f0ac9a4b101"},"cell_type":"code","source":"feature_dataframe","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"901145f6b8fd0207747fd23fe470a25da1e98b1b"},"cell_type":"markdown","source":"**Interactive feature importances via Plotly scatterplots**\n\nI'll use the interactive Plotly package at this juncture to visualise the feature importances values of the different classifiers  via a plotly scatter plot by calling \"Scatter\" as follows:"},{"metadata":{"trusted":true,"scrolled":false,"_uuid":"ac3cbbd8d9c24ba3de4bd390e1c90098fb12277e"},"cell_type":"code","source":"# Scatter plot \ntrace = go.Scatter(\n    y = feature_dataframe['Random Forest feature importances'].values,\n    x = feature_dataframe['features'].values,\n    mode='markers',\n    marker=dict(\n        sizemode = 'diameter',\n        sizeref = 1,\n        size = 25,\n#       size= feature_dataframe['AdaBoost feature importances'].values,\n        #color = np.random.randn(500), #set color equal to a variable\n        color = feature_dataframe['Random Forest feature importances'].values,\n        colorscale='Portland',\n        showscale=True\n    ),\n    text = feature_dataframe['features'].values\n)\ndata = [trace]\n\nlayout= go.Layout(\n    autosize= True,\n    title= 'Random Forest Feature Importance',\n    hovermode= 'closest',\n#     xaxis= dict(\n#         title= 'Pop',\n#         ticklen= 5,\n#         zeroline= False,\n#         gridwidth= 2,\n#     ),\n    yaxis=dict(\n        title= 'Feature Importance',\n        ticklen= 5,\n        gridwidth= 2\n    ),\n    showlegend= False\n)\nfig = go.Figure(data=data, layout=layout)\npy.iplot(fig,filename='scatter2010')\n\n# Scatter plot \ntrace = go.Scatter(\n    y = feature_dataframe['Extra Trees  feature importances'].values,\n    x = feature_dataframe['features'].values,\n    mode='markers',\n    marker=dict(\n        sizemode = 'diameter',\n        sizeref = 1,\n        size = 25,\n#       size= feature_dataframe['AdaBoost feature importances'].values,\n        #color = np.random.randn(500), #set color equal to a variable\n        color = feature_dataframe['Extra Trees  feature importances'].values,\n        colorscale='Portland',\n        showscale=True\n    ),\n    text = feature_dataframe['features'].values\n)\ndata = [trace]\n\nlayout= go.Layout(\n    autosize= True,\n    title= 'Extra Trees Feature Importance',\n    hovermode= 'closest',\n#     xaxis= dict(\n#         title= 'Pop',\n#         ticklen= 5,\n#         zeroline= False,\n#         gridwidth= 2,\n#     ),\n    yaxis=dict(\n        title= 'Feature Importance',\n        ticklen= 5,\n        gridwidth= 2\n    ),\n    showlegend= False\n)\nfig = go.Figure(data=data, layout=layout)\npy.iplot(fig,filename='scatter2010')\n\n# Scatter plot \ntrace = go.Scatter(\n    y = feature_dataframe['AdaBoost feature importances'].values,\n    x = feature_dataframe['features'].values,\n    mode='markers',\n    marker=dict(\n        sizemode = 'diameter',\n        sizeref = 1,\n        size = 25,\n#       size= feature_dataframe['AdaBoost feature importances'].values,\n        #color = np.random.randn(500), #set color equal to a variable\n        color = feature_dataframe['AdaBoost feature importances'].values,\n        colorscale='Portland',\n        showscale=True\n    ),\n    text = feature_dataframe['features'].values\n)\ndata = [trace]\n\nlayout= go.Layout(\n    autosize= True,\n    title= 'AdaBoost Feature Importance',\n    hovermode= 'closest',\n#     xaxis= dict(\n#         title= 'Pop',\n#         ticklen= 5,\n#         zeroline= False,\n#         gridwidth= 2,\n#     ),\n    yaxis=dict(\n        title= 'Feature Importance',\n        ticklen= 5,\n        gridwidth= 2\n    ),\n    showlegend= False\n)\nfig = go.Figure(data=data, layout=layout)\npy.iplot(fig,filename='scatter2010')\n\n# Scatter plot \ntrace = go.Scatter(\n    y = feature_dataframe['Gradient Boost feature importances'].values,\n    x = feature_dataframe['features'].values,\n    mode='markers',\n    marker=dict(\n        sizemode = 'diameter',\n        sizeref = 1,\n        size = 25,\n#       size= feature_dataframe['AdaBoost feature importances'].values,\n        #color = np.random.randn(500), #set color equal to a variable\n        color = feature_dataframe['Gradient Boost feature importances'].values,\n        colorscale='Portland',\n        showscale=True\n    ),\n    text = feature_dataframe['features'].values\n)\ndata = [trace]\n\nlayout= go.Layout(\n    autosize= True,\n    title= 'Gradient Boosting Feature Importance',\n    hovermode= 'closest',\n#     xaxis= dict(\n#         title= 'Pop',\n#         ticklen= 5,\n#         zeroline= False,\n#         gridwidth= 2,\n#     ),\n    yaxis=dict(\n        title= 'Feature Importance',\n        ticklen= 5,\n        gridwidth= 2\n    ),\n    showlegend= False\n)\nfig = go.Figure(data=data, layout=layout)\npy.iplot(fig,filename='scatter2010')","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"8a8fe511b229601c962def4f5889a935b4d56bf0"},"cell_type":"markdown","source":"Now let us calculate the mean of all the feature importances and store it as a new column in the feature importance dataframe."},{"metadata":{"trusted":true,"_uuid":"d6918362a9e3ea6670fe621b3057d9f62791c823"},"cell_type":"code","source":"# Create the new column containing the average of values\n\nfeature_dataframe['mean'] = feature_dataframe.mean(axis= 1) # axis = 1 computes the mean row-wise\nfeature_dataframe","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"332dd87ae9d5a9b465b38e99732d8ee2dced4011"},"cell_type":"markdown","source":"**Plotly Barplot of Average Feature Importances**\n\nHaving obtained the mean feature importance across all our classifiers, we can plot them into a Plotly bar plot as follows:"},{"metadata":{"trusted":true,"_uuid":"765487e1e69740fea4c58f89c85a47b7e6a04bc3"},"cell_type":"code","source":"y = feature_dataframe['mean'].values\nx = feature_dataframe['features'].values\ndata = [go.Bar(\n            x= x,\n            y= y,\n            width = 0.5,\n            marker=dict(\n               color = feature_dataframe['mean'].values,\n            colorscale='Portland',\n            showscale=True,\n            reversescale = False\n            ),\n            opacity=0.6\n        )]\n\nlayout= go.Layout(\n    autosize= True,\n    title= 'Barplots of Mean Feature Importance',\n    hovermode= 'closest',\n#     xaxis= dict(\n#         title= 'Pop',\n#         ticklen= 5,\n#         zeroline= False,\n#         gridwidth= 2,\n#     ),\n    yaxis=dict(\n        title= 'Feature Importance',\n        ticklen= 5,\n        gridwidth= 2\n    ),\n    showlegend= False\n)\nfig = go.Figure(data=data, layout=layout)\npy.iplot(fig, filename='bar-direct-labels')","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"6769210bf9644ac72600fa2a2f1b764e047e85c6"},"cell_type":"markdown","source":"# Second-Level Predictions from the First-level Output"},{"metadata":{"_uuid":"af9a757c83cc0ae10a5491ea10b060d3250400a3"},"cell_type":"markdown","source":"**First-level output as new features**\n\nHaving now obtained our first-level predictions, one can think of it as essentially building a new set of features to be used as training data for the next classifier. As per the code below, we are therefore having as our new columns the first-level predictions from our earlier classifiers and we train the next classifier on this."},{"metadata":{"trusted":true,"_uuid":"325ab4423c054eda3d6577a5b97a9087fe546c28"},"cell_type":"code","source":"base_predictions_train = pd.DataFrame( {'RandomForest': rf_oof_train.ravel(),\n     'ExtraTrees': et_oof_train.ravel(),\n     'AdaBoost': ada_oof_train.ravel(),\n     'GradientBoost': gb_oof_train.ravel(),\n     'SVC': svc_oof_train.ravel(),\n     'LogisticRegression': log_oof_train.ravel(),   \n     'KNN': knn_oof_train.ravel()                                   \n    })\nbase_predictions_train.head(10)","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"eca06203750b3d626e8ed4c543485aafa801ccdb"},"cell_type":"markdown","source":"**Correlation Heatmap of the Second Level Training set**"},{"metadata":{"trusted":true,"_uuid":"1a7d779dee33552349a2d6ab571c5eaf08352211","scrolled":false},"cell_type":"code","source":"data = [\n    go.Heatmap(\n        z= base_predictions_train.astype(float).corr().values ,\n        x=base_predictions_train.columns.values,\n        y= base_predictions_train.columns.values,\n          colorscale='Viridis',\n            showscale=True,\n            reversescale = True\n    )\n]\npy.iplot(data, filename='labelled-heatmap')","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"a152c3cf47241091ac85b1fb289e374f40810ccb"},"cell_type":"markdown","source":"There have been quite a few articles and Kaggle competition winner stories about the merits of having trained models that are more uncorrelated with one another producing better scores."},{"metadata":{"trusted":true,"_uuid":"41d7f6a2186abb6884648d5ef124073a7e6f5a6f"},"cell_type":"code","source":"x_train = np.concatenate(( et_oof_train, rf_oof_train, ada_oof_train, gb_oof_train, svc_oof_train, log_oof_train, knn_oof_train), axis=1)\nx_test = np.concatenate(( et_oof_test, rf_oof_test, ada_oof_test, gb_oof_test, svc_oof_test, log_oof_test, knn_oof_test), axis=1)","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"1fe594ec1690491a34b6ce4625dd03b2dae954aa"},"cell_type":"markdown","source":"Having now concatenated and joined both the first-level train and test predictions as x_train and x_test, we can now fit a second-level learning model."},{"metadata":{"_uuid":"6e386b4dd2515241c0de736a944c47b94fa0dd12"},"cell_type":"markdown","source":"### Second level learning model via XGBoost\n\nHere we choose the eXtremely famous library for boosted tree learning model, XGBoost. It was built to optimize large-scale boosted tree algorithms. For further information about the algorithm, check out the [official documentation][1].\n\n  [1]: https://xgboost.readthedocs.io/en/latest/\n\nAnyways, we call an XGBClassifier and fit it to the first-level train and target data and use the learned model to predict the test data as follows:"},{"metadata":{"trusted":true,"_uuid":"7b51d6e6407caf2e972b073f30760cf1a23f8f6f"},"cell_type":"code","source":"#Grid Search\ngbm_param_grid = {\n    'learning_rate': [0.05, 0.08, 0.1, 0.2],\n    'n_estimators': [100, 200, 500],\n    'max_depth': [3, 4, 5],\n    'gamma': [0.8, 0.9] \n}\ngbm = xgb.XGBClassifier()\ngrid_gbm = GridSearchCV(estimator=gbm,\n                       param_grid=gbm_param_grid,\n                       scoring='accuracy',\n                       cv=3,\n                       verbose=1,\n                       n_jobs=4)","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"886f5f56feaaf08249cfb08015efd18ee4d3f80c"},"cell_type":"code","source":"%%time\ngrid_gbm.fit(x_train, y_train)\nprint(\"Best parameters found: \",grid_gbm.best_params_)","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"0af22b3be2e0d9438a198a805042778cce0819aa"},"cell_type":"code","source":"gbm = xgb.XGBClassifier(\n learning_rate = 0.05,\n n_estimators= 100,\n max_depth= 4,\n gamma=0.9,                        \n n_jobs=4\n ).fit(x_train, y_train)\npredictions = gbm.predict(x_test)","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"14488c60ed1d9a90203f1ca9e1bcc6aaf3871322"},"cell_type":"code","source":"acc_gbm = round(gbm.score(x_train, y_train) * 100, 2)\nacc_gbm","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"f0f19030e4183daf1faa4e1132c340195f44c020"},"cell_type":"markdown","source":"Just a quick run down of the XGBoost parameters used in the model:\n\n**max_depth** : How deep you want to grow your tree. Beware if set to too high a number might run the risk of overfitting.\n\n**gamma** : minimum loss reduction required to make a further partition on a leaf node of the tree. The larger, the more conservative the algorithm will be.\n\n**eta** : step size shrinkage used in each boosting step to prevent overfitting"},{"metadata":{"_uuid":"c73e8ce61ea56428719e08385a5df6993bb32e7a"},"cell_type":"markdown","source":"**Producing the Submission file**\n\nFinally having trained and fit all our first-level and second-level models, we can now output the predictions into the proper format for submission to the Titanic competition as follows:"},{"metadata":{"trusted":true,"_uuid":"ace4898907192e0e36426a6dbd82add34a1832c0"},"cell_type":"code","source":"# Generate Submission File \nStackingSubmission = pd.DataFrame({ 'PassengerId': PassengerId,\n                            'Survived': predictions })\nStackingSubmission.to_csv(\"StackingSubmission.csv\", index=False)\n\n","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"717a1d78f6c0e4d1f480ca23271cb261bb62ca91"},"cell_type":"markdown","source":"**Steps for Further Improvement**\n\nAs a closing remark it must be noted that the steps taken above just show a very simple way of producing an ensemble stacker. You hear of ensembles created at the highest level of Kaggle competitions which involves monstrous combinations of stacked classifiers as well as levels of stacking which go to more than 2 levels. \n\nSome additional steps that may be taken to improve one's score could be:\n\n 1. Implementing a good cross-validation strategy in training the models to find optimal parameter values\n 2. Introduce a greater variety of base models for learning. The more uncorrelated the results, the better the final score."},{"metadata":{"_uuid":"80be61943b3f278c2ec8c2ed048adadf93b9c5fa"},"cell_type":"markdown","source":"## References\n\nThis notebook has been created based on great work done solving the Titanic competition and other sources.\n\n- [introduction-to-ensembling-stacking-in-python](https://www.kaggle.com/arthurtok/introduction-to-ensembling-stacking-in-python)\n- [data-science-framework-to-achieve-99-accuracy](https://www.kaggle.com/ldfreeman3/a-data-science-framework-to-achieve-99-accuracy)\n- [a-statistical-analysis-ml-workflow-of-titanic](https://www.kaggle.com/masumrumi/a-statistical-analysis-ml-workflow-of-titanic)\n- [Stacking — A Super Learning Technique](https://medium.com/@gurucharan_33981/stacking-a-super-learning-technique-dbed06b1156d)"},{"metadata":{"trusted":true,"_uuid":"1de5a718b6db4351cae50ba8df24a2547ba8acbc"},"cell_type":"markdown","source":"## Change Log\n- [20190131] fix a bug on family size calculation"},{"metadata":{"trusted":true,"_uuid":"2e1d3ad854bcd9d918be05926eb751223fcce139"},"cell_type":"code","source":"","execution_count":null,"outputs":[]}],"metadata":{"_is_fork":false,"language_info":{"name":"python","version":"3.6.6","mimetype":"text/x-python","codemirror_mode":{"name":"ipython","version":3},"pygments_lexer":"ipython3","nbconvert_exporter":"python","file_extension":".py"},"_change_revision":0,"kernelspec":{"display_name":"Python 3","language":"python","name":"python3"}},"nbformat":4,"nbformat_minor":1}