{"cells":[{"metadata":{"_uuid":"fed5696c67bf55a553d6d04313a77e8c617cad99","_cell_guid":"ea25cdf7-bdbc-3cf1-0737-bc51675e3374"},"cell_type":"markdown","source":" <a id=\"top\"></a> \n# Titanic Data Science Solution(1) - Jimmy Modified\n\nThe notebook walks us through a typical workflow for solving data science competitions at sites like Kaggle. The objective of this notebook is to follow a step-by-step workflow, explaining each step and rationale for every decision we take during solution development. The work is highly based on many good kernels in Kaggle. You can visit them through notebook reference part and the end of this notebook.\n\n## General Data Science Workflow Stages （七步法）\n\nThe competition solution workflow goes through **seven stages** normally used in any Data Science Solution on Kaggle. \n\n1. [Question or problem definition.](#step1) 问题定义\n2. [Acquire training and testing data.](#step2) 获取数据\n3. [Analyze, identify patterns, and explore the data  (Exploratory Data Analysis).](#step3) 分析数据\n4. [Wrangle, prepare, cleanse the data.](#step4) 处理数据\n5. [Model, predict and solve the problem.](#step5) 建模预测\n6. [Visualize, report, and present the problem solving steps and final solution.](#step6) 汇报结果\n7. [Supply or submit the results and optimize.](#step7) 提交结果\n\nThe workflow indicates general sequence of how each stage may follow the other. However there are use cases with exceptions.\n\n- We may combine mulitple workflow stages. We may analyze by visualizing data.\n- Perform a stage earlier than indicated. We may analyze data before and after wrangling.\n- Perform a stage multiple times in our workflow. Visualize stage may be used multiple times.\n- Drop a stage altogether. We may not need supply stage to productize or service enable our dataset for a competition."},{"metadata":{"_uuid":"2001a06bd95c2bbfadf96baeb6f735274b17296e"},"cell_type":"markdown","source":"## Normal Data Science with Python Workflow\n![](https://github.com/icwangjimmy/machine_learning_and_python_in_finance/raw/master/pic/Data_Science_With_Python_Workflow.png)"},{"metadata":{"_uuid":"fed5696c67bf55a553d6d04313a77e8c617cad99","_cell_guid":"ea25cdf7-bdbc-3cf1-0737-bc51675e3374"},"cell_type":"markdown","source":"<a id=\"step1\"></a>\n# 1. Question and problem definition 问题定义\n\nCompetition sites like Kaggle define the problem to solve or questions to ask while providing the datasets for training your data science model and testing the model results against a test dataset. The question or problem definition for Titanic Survival competition is [described here at Kaggle](https://www.kaggle.com/c/titanic).\n![title](https://github.com/icwangjimmy/machine_learning_and_python_in_finance/raw/master/pic/4r79.gif)\n\n> Knowing from a training set of samples listing passengers who survived or did not survive the Titanic disaster, can our model determine based on a given test dataset not containing the survival information, if these passengers in the test dataset survived or not.\n\nWe may also want to develop some early understanding about the domain of our problem. This is described on the [Kaggle competition description page here](https://www.kaggle.com/c/titanic). Here are the highlights to note.\n\n- **On April 15, 1912, during her maiden voyage, the Titanic sank after colliding with an iceberg, killing 1502 out of 2224 passengers and crew. Translated 32% survival rate.**\n- **One of the reasons that the shipwreck led to such loss of life was that there were not enough lifeboats for the passengers and crew.**\n- **Although there was some element of luck involved in surviving the sinking, some groups of people were more likely to survive than others, such as women, children, and the upper-class.**\n"},{"metadata":{"_uuid":"33bccc304602bc338bc2ed67947f2b1a53c045e1"},"cell_type":"markdown","source":""},{"metadata":{"_uuid":"fed5696c67bf55a553d6d04313a77e8c617cad99","_cell_guid":"ea25cdf7-bdbc-3cf1-0737-bc51675e3374"},"cell_type":"markdown","source":"## Workflow goals (7C) 七种数据处理手段\n\nThe data science solutions workflow solves for **seven major goals**.\n\n**Classifying.** We may want to classify or categorize our samples. We may also want to understand the implications or correlation of different classes with our solution goal.\n\n**Correlating.** One can approach the problem based on available features within the training dataset. Which features within the dataset contribute significantly to our solution goal? Statistically speaking is there a [correlation](https://en.wikiversity.org/wiki/Correlation) among a feature and solution goal? As the feature values change does the solution state change as well, and visa-versa? This can be tested both for numerical and categorical features in the given dataset. We may also want to determine correlation among features other than survival for subsequent goals and workflow stages. Correlating certain features may help in creating, completing, or correcting features.\n\n**Converting.** For modeling stage, one needs to prepare the data. Depending on the choice of model algorithm one may require all features to be converted to numerical equivalent values. So for instance converting text categorical values to numeric values.\n\n**Completing.** Data preparation may also require us to estimate any missing values within a feature. Model algorithms may work best when there are no missing values.\n\n**Correcting.** We may also analyze the given training dataset for errors or possibly inaccurate values within features and try to correct these values or exclude the samples containing the errors. One way to do this is to detect any outliers among our samples or features. We may also completely discard a feature if it is not contribting to the analysis or may significantly skew the results.\n\n**Creating.** Can we create new features based on an existing feature or a set of features, such that the new feature follows the correlation, conversion, completeness goals.\n\n**Charting.** How to select the right visualization plots and charts depending on nature of the data and the solution goals."},{"metadata":{"_uuid":"d520194bd554c9ff3e2c1949dd71849291fd47df"},"cell_type":"markdown","source":"## Prepare Python Packages"},{"metadata":{"_uuid":"847a9b3972a6be2d2f3346ff01fea976d92ecdb6","_cell_guid":"5767a33c-8f18-4034-e52d-bf7a8f7d8ab8","trusted":true},"cell_type":"code","source":"# data analysis and wrangling\nimport pandas as pd\nimport numpy as np\n\n# visualization\nimport seaborn as sns\nimport matplotlib.pyplot as plt\n%matplotlib inline\n\n# machine learning algorithms using scikit-learn\nfrom sklearn.linear_model import LogisticRegression\nfrom sklearn.svm import SVC, LinearSVC\nfrom sklearn.ensemble import RandomForestClassifier\nfrom sklearn.neighbors import KNeighborsClassifier\nfrom sklearn.linear_model import Perceptron\nfrom sklearn.linear_model import SGDClassifier\nfrom sklearn.tree import DecisionTreeClassifier\nfrom sklearn.neural_network import MLPClassifier","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"2d307b99ee3d19da3c1cddf509ed179c21dec94a","_cell_guid":"6b5dc743-15b1-aac6-405e-081def6ecca1"},"cell_type":"markdown","source":"<a id=\"step2\"></a>\n# 2. Acquire training and testing data 获取数据\n\nThe Python Pandas packages helps us work with our datasets. We start by acquiring the training and testing datasets into Pandas DataFrames. We also combine these datasets to run certain operations on both datasets together."},{"metadata":{"_uuid":"13f38775c12ad6f914254a08f0d1ef948a2bd453","_cell_guid":"e7319668-86fe-8adc-438d-0eef3fd0a982","trusted":true},"cell_type":"code","source":"train_df = pd.read_csv('../input/train.csv')\ntest_df = pd.read_csv('../input/test.csv')\ncombine = [train_df, test_df]","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"8fd4f208a6e725bb5c478f358b88c6391a82acef"},"cell_type":"code","source":"train_df.info()","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"fd440648da1551ad70821c13563247ea0377f334","scrolled":true},"cell_type":"code","source":"train_df.head()","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"195d47e0187f7beeb598de5d36f294be1a37db72"},"cell_type":"code","source":"test_df.head()","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"79282222056237a52bbbb1dbd831f057f1c23d69","_cell_guid":"3d6188f3-dc82-8ae6-dabd-83e28fcbf10d"},"cell_type":"markdown","source":"<a id=\"step3\"></a>\n# 3. Analyze, identify patterns, and explore the data (Exploratory Data Analysis) 分析数据\n\n## Analyze by describing data\n\nPandas also helps describe the datasets answering following questions early in our project.\n\n**Which features are available in the dataset?**\n\nNoting the feature names for directly manipulating or analyzing these. These feature names are described on the [Kaggle data page here](https://www.kaggle.com/c/titanic/data)."},{"metadata":{"_uuid":"ef106f38a00e162a80c523778af6dcc778ccc1c2","_cell_guid":"ce473d29-8d19-76b8-24a4-48c217286e42","trusted":true},"cell_type":"code","source":"print(train_df.columns.values)","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"2e381c0300aa691453acdf0ea9f5ec30afb30b0a"},"cell_type":"markdown","source":"**VARIABLE DESCRIPTIONS:**\n\nWe've got a sense of our variables, their class type, and the first few observations of each. We know we're working with 1309 observations of 12 variables. To make things a bit more explicit since a couple of the variable names aren't 100% illuminating, here's what we've got to deal with:\n\n\n**Variable Description**\n\n \n1. **Age** :\n    1. Age is fractional if less than 1. If the age is estimated, is it in the form of xx.5\n\n1. **Sibsp** :\n    1. The dataset defines family relations in this way...\n\n        a. Sibling = brother, sister, stepbrother, stepsister\n\n        b. Spouse = husband, wife (mistresses and fiancés were ignored)\n\n1. **Parch**:\n    1. The dataset defines family relations in this way...\n\n        a. Parent = mother, father\n\n        b. Child = daughter, son, stepdaughter, stepson\n\n        c. Some children travelled only with a nanny, therefore parch=0 for them.\n\n1. **Pclass** :\n    *  A proxy for socio-economic status (SES)\n        * 1st = Upper\n        * 2nd = Middle\n        * 3rd = Lower\n1. **Embarked** :\n     * nominal datatype \n1. **Name**: \n    * nominal datatype . It could be used in feature engineering to derive the gender from title\n1. **Sex**: \n   * nominal datatype \n1. **Ticket**:\n    * that have no impact on the outcome variable. Thus, they will be excluded from analysis\n1. **Cabin**: \n    * is a nominal datatype that can be used in feature engineering\n1.  **Fare**:\n    * Indicating the fare\n1. **PassengerID**:\n    * have no impact on the outcome variable. Thus, it will be excluded from analysis\n1. **Survival**:\n    * **[dependent variable](http://www.dailysmarty.com/posts/difference-between-independent-and-dependent-variables-in-machine-learning)** , 0 or 1"},{"metadata":{"_uuid":"a4ad0eb016d1419b90045742e97c09c0985e9e25"},"cell_type":"markdown","source":"## Visualization\n\n**Data visualization**  is the presentation of data in a pictorial or graphical format. It enables decision makers to see analytics presented visually, so they can grasp difficult concepts or identify new patterns.\n\nWith interactive visualization, you can take the concept a step further by using technology to drill down into charts and graphs for more detail, interactively changing what data you see and how it’s processed.\n\n In this section I show you  **15 plots** with **matplotlib** and **seaborn** that is listed here:\n"},{"metadata":{"_uuid":"e484bf32de654ee8bf67868228f676859baeaec2"},"cell_type":"markdown","source":"### 1 Scatter plot\n\nScatter plot Purpose To identify the type of relationship (if any) between two quantitative variables\n"},{"metadata":{"trusted":true,"_uuid":"46689fe551b66ff6b866ba8ed7a4ab0cd020ad81"},"cell_type":"code","source":"# 按照Pclass分类的，Age和Fare的散点图，颜色按照是否survive区分\ng = sns.FacetGrid(train_df, hue=\"Survived\", col=\"Pclass\", margin_titles=True, palette={1:\"seagreen\", 0:\"gray\"}, height=5)\ng=g.map(plt.scatter, \"Fare\", \"Age\", edgecolor=\"w\").add_legend();","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"343ab2b8dcb58c9539c162631a7c155d991a41dc"},"cell_type":"markdown","source":"### 2 Box\nIn descriptive statistics, a **box plot** or boxplot is a method for graphically depicting groups of numerical data through their quartiles. Box plots may also have lines extending vertically from the boxes (whiskers) indicating variability outside the upper and lower quartiles, hence the terms box-and-whisker plot and box-and-whisker diagram.[wikipedia]"},{"metadata":{"trusted":true,"_uuid":"1f044b7589add0d463ce37de5a0334ac633c735c"},"cell_type":"code","source":"#boxplot一般用来寻找outlier的数据点\n#按照Pclass分的Age分布图\nplt.figure(figsize=(12,6))\nax= sns.boxplot(x=\"Pclass\", y=\"Age\", data=train_df)\nax= sns.stripplot(x=\"Pclass\", y=\"Age\", data=train_df, jitter=True, edgecolor=\"gray\")\nplt.show()","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"03647afbbef448b1052e9084dbebc9c8f6dc270c"},"cell_type":"code","source":"#按照Pclass分的Fare分布图\nplt.figure(figsize=(12,6))\nax= sns.boxplot(x=\"Pclass\", y=\"Fare\", data=train_df)\nax= sns.stripplot(x=\"Pclass\", y=\"Fare\", data=train_df, jitter=True, edgecolor=\"gray\")\nplt.show()","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"fbbc02225549054e751719148eb40f8baa2c8f5c"},"cell_type":"markdown","source":"### 3 Histogram\nWe can also create a **histogram** of each input variable to get an idea of the distribution."},{"metadata":{"trusted":true,"_uuid":"14f51d4257eaef94576c2d8b9c501cda9f0e8011"},"cell_type":"code","source":"# histograms 直方图一般用来单变量的分布情况\ntrain_df.hist(figsize=(18,18))\nplt.figure()","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"90085b5003aaeb7bd7a7ef3f9ac6564a58e963fa"},"cell_type":"markdown","source":"It looks like perhaps two of the input variables have a Gaussian distribution. This is useful to note as we can use algorithms that can exploit this assumption.\n"},{"metadata":{"trusted":true,"_uuid":"120775a136fd49af37fce2ec0cfa418106f57f39"},"cell_type":"code","source":"#Age变量可以再深入研究一下\ntrain_df[\"Age\"].hist();","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"86c825d1055ec78ac7a172cb12c08d6b35c73b48"},"cell_type":"code","source":"#按照是否survive来看看age变量的分布\nf,ax=plt.subplots(1,2,figsize=(16,8))\ntrain_df[train_df['Survived']==0].Age.plot.hist(ax=ax[0],bins=20,edgecolor='black',color='red')\nax[0].set_title('Survived= 0')\nx1=list(range(0,85,5))\nax[0].set_xticks(x1)\n\ntrain_df[train_df['Survived']==1].Age.plot.hist(ax=ax[1],color='green',bins=20,edgecolor='black')\nax[1].set_title('Survived= 1')\nx2=list(range(0,85,5))\nax[1].set_xticks(x2)\nplt.show()","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"168a89daeb7a742d10de5b7b73a4ac8153fc1dc5"},"cell_type":"markdown","source":"### 4 Multivariate Plots\nNow we can look at the interactions between the variables.\n\nFirst, let’s look at scatterplots of all pairs of attributes. This can be helpful to spot structured relationships between input variables."},{"metadata":{"trusted":true,"_uuid":"fa9978b3e0186ff2859f9a4a774fa0f18bc54b46"},"cell_type":"code","source":"# scatter plot matrix 散点图矩阵\npd.plotting.scatter_matrix(train_df,figsize=(15,15))\nplt.figure()","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"a09e0a7ceb9ba95037541963cd86236600b63128"},"cell_type":"markdown","source":"Note the diagonal grouping of some pairs of attributes. This suggests a high correlation and a predictable relationship."},{"metadata":{"_uuid":"403df2443bbefb0ca844eecf7e4309f3575cd504"},"cell_type":"markdown","source":"### 5 Violinplots\nDraw a combination of boxplot and kernel density estimate.\n\nA violin plot plays a similar role as a box and whisker plot. It shows the distribution of quantitative data across several levels of one (or more) categorical variables such that those distributions can be compared. Unlike a box plot, in which all of the plot components correspond to actual datapoints, the violin plot features a kernel density estimation of the underlying distribution."},{"metadata":{"trusted":true,"_uuid":"fc6b7e221853e59c5e9bd05429cc46d49b87f68b"},"cell_type":"code","source":"# violinplots 小提琴图和boxplot功能类似，不过可以展示对分布的概率密度估计\n# 这里是按Sex分开的Age分布\nimport warnings\nwarnings.simplefilter(action='ignore', category=FutureWarning)\nplt.figure(figsize=(12,8))\nsns.violinplot(data=train_df,x=\"Sex\", y=\"Age\")","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"33a4d9e10b7e69a173754cc7a6fcacb7e0af3ffd"},"cell_type":"code","source":"#这里更进一步将是否survive分开表示\nf,ax=plt.subplots(1,2,figsize=(18,8))\nsns.violinplot(\"Pclass\",\"Age\", hue=\"Survived\", data=train_df,split=True,ax=ax[0])\nax[0].set_title('Pclass and Age vs Survived')\nax[0].set_yticks(range(0,110,10))\n\nsns.violinplot(\"Sex\",\"Age\", hue=\"Survived\", data=train_df,split=True,ax=ax[1])\nax[1].set_title('Sex and Age vs Survived')\nax[1].set_yticks(range(0,110,10))\nplt.show()","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"5f4510cb60f0131c7391b8d7bb7ce5b47eeac9ea"},"cell_type":"markdown","source":"### 6 Pairplot\nPlot pairwise relationships in a dataset.\n\nBy default, this function will create a grid of Axes such that each variable in data will by shared in the y-axis across a single row and in the x-axis across a single column. The diagonal Axes are treated differently, drawing a plot to show the univariate distribution of the data for the variable in that column."},{"metadata":{"trusted":true,"_uuid":"0ac01878d7fc8ac9c5501e1e258c022c5c48b17a"},"cell_type":"code","source":"# Using seaborn pairplot to see the bivariate relation between each pair of features\n# seaborn也可以做配对关系散点图，显示更好一些\nimport warnings\nwarnings.simplefilter(action='ignore', category=RuntimeWarning)\nsns.pairplot(train_df, hue=\"Sex\")\n# sns.pairplot(train_df, hue=\"Sex\", diag_kind=\"kde\")","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"e07d6b707a8397cbbf7080b0f1cb3014a7328b17"},"cell_type":"markdown","source":"###  7 Kdeplot\nFit and plot a univariate or bivariate kernel density estimate."},{"metadata":{"trusted":true,"_uuid":"489fe753472f3774f7ba40936532d9a8cc8ffc51"},"cell_type":"code","source":"# seaborn's kdeplot, plots univariate or bivariate density estimates.\n#Size can be changed by tweeking the value of height used\n# 这里显示的是按是否survive分开的Fare分布估计\nsns.FacetGrid(train_df, hue=\"Survived\", height=7).map(sns.kdeplot, \"Fare\").add_legend()\nplt.show()","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"66f8146e6910989d8fde4eb6783a0b868ddf6a6b"},"cell_type":"markdown","source":"### 8 Jointplot\nDraw a plot of two variables with bivariate and univariate graphs."},{"metadata":{"trusted":true,"_uuid":"83578fbe44fa17eb59a8520d9194e9b72335abdd"},"cell_type":"code","source":"#讲单变量分布和双变量关系图同时显示\nsns.jointplot(x='Fare',y='Age',data=train_df)","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"7b2578936cbf2ca01b3d32a46140353e870cf129"},"cell_type":"code","source":"#附加了线性回归作为参考线\nsns.jointplot(x='Fare',y='Age' ,data=train_df, kind='reg')","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"e3e4e7308c4a340df3d93415e8c1df7b0abbae03"},"cell_type":"markdown","source":"###  9 Swarm plot\nDraw a categorical scatterplot with non-overlapping points.\n\nThis function is similar to stripplot(), but the points are adjusted (only along the categorical axis) so that they don’t overlap. This gives a better representation of the distribution of values, but it does not scale well to large numbers of observations. This style of plot is sometimes called a “beeswarm”."},{"metadata":{"trusted":true,"_uuid":"fbab4570ef1bd72ce344b753fc0ca095d4c69ef2"},"cell_type":"code","source":"#蜂群图是stripplot的一种改进\nplt.figure(figsize=(12,6))\nsns.swarmplot(x='Pclass',y='Age',data=train_df)","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"fbf45ebac22a704b966c2babd904a0c16f8b7beb"},"cell_type":"markdown","source":"### 10 Heatmap\nPlot rectangular data as a color-encoded matrix."},{"metadata":{"trusted":true,"_uuid":"9bdf55ba79fc5f75af93027c5f258a010fabc7d3"},"cell_type":"code","source":"#热点图用颜色显示栅格数据\nplt.figure(figsize=(10,10)) \nsns.heatmap(train_df.corr(),annot=True,cmap='cubehelix_r') \nplt.show()","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"02a61d8b3f1b3232742aec5f4c72ec4555fa0073"},"cell_type":"markdown","source":"###  11 Bar Plot\n#### Pandas Barplot\nA bar plot is a plot that presents categorical data with rectangular bars with lengths proportional to the values that they represent.\n"},{"metadata":{"trusted":true,"_uuid":"97dac656b07141d07d800d12d4e175a9edf3c645"},"cell_type":"code","source":"#柱状图也是很常用的\nplt.figure(figsize=(6,6)) \ntrain_df['Pclass'].value_counts().plot(kind=\"bar\");","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"04e0dc7b23124073ad62b69df25f5120962cda2a"},"cell_type":"markdown","source":"#### Seaborn Barplot\nShow point estimates and confidence intervals as rectangular bars.\n\nA bar plot represents an estimate of central tendency for a numeric variable with the height of each rectangle and provides some indication of the uncertainty around that estimate using error bars. Bar plots include 0 in the quantitative axis range, and they are a good choice when 0 is a meaningful value for the quantitative variable, and you want to make comparisons against it."},{"metadata":{"trusted":true,"_uuid":"56d20b0564a4fcf6c792bc3f3847f685e17aa0c1"},"cell_type":"code","source":"#带有Fare点估计的柱状图\nplt.figure(figsize=(6,6)) \nsns.barplot(x=\"Pclass\", y=\"Fare\", data=train_df)","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"f4c998ad1cd1255ebd0860ff1b3342c360c2a557"},"cell_type":"code","source":"#带有Age点估计的柱状图\nplt.figure(figsize=(6,6)) \nsns.barplot(x=\"Pclass\", y=\"Age\", data=train_df)","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"625386c6228acb31da3cef5f3ed5ee1d26a3b438"},"cell_type":"markdown","source":"### 12 Pointplot\nShow point estimates and confidence intervals using scatter plot glyphs.\n\nA point plot represents an estimate of central tendency for a numeric variable by the position of scatter plot points and provides some indication of the uncertainty around that estimate using error bars."},{"metadata":{"trusted":true,"_uuid":"0947ccef5e84932504accf17b0fe959cb58cb621"},"cell_type":"code","source":"#如果对零点并不关心，只关心相对大小，是对柱状图的一种替代\nplt.figure(figsize=(6,6)) \nsns.pointplot('Pclass', 'Survived',hue='Sex', data=train_df)\nplt.show()","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"c37c4c38be0045178b2d1b1fe567645dde9cc47d"},"cell_type":"markdown","source":"\n### 13 Distplot\nFlexibly plot a univariate distribution of observations.\n\nThis function combines the matplotlib hist function (with automatic calculation of a good default bin size) with the seaborn kdeplot() and rugplot() functions."},{"metadata":{"trusted":true,"_uuid":"e50783040e329adb46c683a930ae1c75eac6c6b5"},"cell_type":"code","source":"#分布图是对直方图和kde的一种组合\nf,ax=plt.subplots(1,3,figsize=(20,8))\nsns.distplot(train_df[train_df['Pclass']==1].Fare,ax=ax[0])\nax[0].set_title('Fares in Pclass 1')\nsns.distplot(train_df[train_df['Pclass']==2].Fare,ax=ax[1])\nax[1].set_title('Fares in Pclass 2')\nsns.distplot(train_df[train_df['Pclass']==3].Fare,ax=ax[2])\nax[2].set_title('Fares in Pclass 3')\nplt.show()","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"437f5b033cecc9d6264a38c29f351cab17ee611a"},"cell_type":"markdown","source":"### 14 Pie plot\nGenerate a pie plot.\n\nA pie plot is a proportional representation of the numerical data in a column."},{"metadata":{"trusted":true,"_uuid":"d72e594ef43636969f1ec020636f328190f1121c"},"cell_type":"code","source":"#饼图也是常用图的一种\nf,ax=plt.subplots(1,2,figsize=(18,8)) #one row two columns\ntrain_df['Survived'].value_counts().plot.pie(explode=[0,0.1],autopct='%1.1f%%',ax=ax[0],shadow=True)\nax[0].set_title('Survived')\nax[0].set_ylabel('')\n\nsns.countplot('Survived',data=train_df,ax=ax[1])\nax[1].set_title('Survived')\nplt.show()","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"2eeda0ca7c1926e18206131f56b8a826ba6be16b"},"cell_type":"markdown","source":"### 15 Countplot\nShow the counts of observations in each categorical bin using bars."},{"metadata":{"trusted":true,"_uuid":"fbbb693a1ec89bba51374e8fc9bd91374e5e2c08"},"cell_type":"code","source":"# countplot和value_counts功能类似\nf,ax=plt.subplots(1,2,figsize=(18,8))\ntrain_df[['Sex','Survived']].groupby(['Sex']).mean().plot.bar(ax=ax[0])\nax[0].set_title('Survived vs Sex')\n\nsns.countplot('Sex',hue='Survived',data=train_df,ax=ax[1])\nax[1].set_title('Sex:Survived vs Dead')\nplt.show()","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"f89aa98d658140ad5fa863245a4a3cb19d0cefab"},"cell_type":"markdown","source":"## Analysis (数据分析)"},{"metadata":{"_uuid":"1d7acf42af29a63bc038f14eded24e8b8146f541","_cell_guid":"cd19a6f6-347f-be19-607b-dca950590b37"},"cell_type":"markdown","source":"**Which features are categorical?** 分类的数据\n\nThese values classify the samples into sets of similar samples. Within categorical features are the values nominal, ordinal, ratio, or interval based? Among other things this helps us select the appropriate plots for visualization.\n\n- **Categorical**: Survived, Sex, and Embarked. **Ordinal**: Pclass. \n\n**Which features are numerical?** 数值类的数据\n\nWhich features are numerical? These values change from sample to sample. Within numerical features are the values discrete, continuous, or timeseries based? Among other things this helps us select the appropriate plots for visualization.\n\n- **Continous**: Age, Fare. **Discrete**: SibSp, Parch."},{"metadata":{"_uuid":"e068cd3a0465b65a0930a100cb348b9146d5fd2f","_cell_guid":"8d7ac195-ac1a-30a4-3f3f-80b8cf2c1c0f","trusted":true},"cell_type":"code","source":"# preview the data\ntrain_df.head(10)","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"c34fa51a38336d97d5f6a184908cca37daebd584","_cell_guid":"97f4e6f8-2fea-46c4-e4e8-b69062ee3d46"},"cell_type":"markdown","source":"**Which features are mixed data types?** 混合类型的数据\n\nNumerical, alphanumeric data within same feature. These are candidates for correcting goal.\n\n- **Ticket** is a mix of numeric and alphanumeric data types. **Cabin** is alphanumeric.\n\n**Which features may contain errors or typos?** 有错误的特征\n\nThis is harder to review for a large dataset, however reviewing a few samples from a smaller dataset may just tell us outright, which features may require correcting.\n\n- **Name** feature may contain errors or typos as there are several ways used to describe a name including titles, round brackets, and quotes used for alternative or short names."},{"metadata":{"_uuid":"3488e80f309d29f5b68bbcfaba8d78da84f4fb7d","_cell_guid":"f6e761c2-e2ff-d300-164c-af257083bb46","trusted":true},"cell_type":"code","source":"train_df.tail(10)","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"699c52b7a8d076ccd5ea5bc5d606313c558a6e8e","_cell_guid":"8bfe9610-689a-29b2-26ee-f67cd4719079"},"cell_type":"markdown","source":"**Which features contain blank, null or empty values?** 哪些特征数据不全\n\nThese will require correcting.\n\n- Cabin > Age > Embarked features contain a number of null values in that order for the training dataset.\n- Cabin > Age are incomplete in case of test dataset.\n\n**What are the data types for various features?** 特征的数据类型\n\nHelping us during converting goal.\n\n- Seven features are integer or floats. Six in case of test dataset.\n- Five features are strings (object)."},{"metadata":{"_uuid":"817e1cf0ca1cb96c7a28bb81192d92261a8bf427","_cell_guid":"9b805f69-665a-2b2e-f31d-50d87d52865d","trusted":true},"cell_type":"code","source":"train_df.info()\nprint('_'*40)\ntest_df.info()","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"b4cee58eb86d76d0683bf4b47f151de42df11bb1"},"cell_type":"code","source":"print(train_df.isnull().sum())\nprint('_'*40)\nprint(test_df.isnull().sum())","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"2b7c205bf25979e3242762bfebb0e3eb2fd63010","_cell_guid":"859102e1-10df-d451-2649-2d4571e5f082"},"cell_type":"markdown","source":" **What is the distribution of numerical feature values across the samples?** 数值类型的特征分布是怎么样的？\n\nThis helps us determine, among other early insights, how representative is the training dataset of the actual problem domain.\n\n- Total samples are 891 or 40% of the actual number of passengers on board the Titanic (2,224).\n- Survived is a categorical feature with 0 or 1 values.\n- Around 38% samples survived representative of the actual survival rate at 32%.\n- Most passengers (> 75%) did not travel with parents or children.\n- Nearly 30% of the passengers had siblings and/or spouse aboard.\n- Fares varied significantly with few passengers (<1%) paying as high as $512.\n- Few elderly passengers (<1%) within age range 65-80."},{"metadata":{"_uuid":"380251a1c1e0b89147d321968dc739b6cc0eecf2","_cell_guid":"58e387fe-86e4-e068-8307-70e37fe3f37b","trusted":true},"cell_type":"code","source":"# train_df.describe()\n# Review survived rate using `percentiles=[.61, .62]` knowing our problem description mentions 38% survival rate.\ntrain_df.describe(percentiles=[.61, .62])\n# Review Parch distribution using `percentiles=[.75, .8]`\n#train_df.describe(percentiles=[.75, .8])\n# SibSp distribution `[.68, .69]`\n# train_df.describe(percentiles=[.68, .69])\n# Age and Fare `[.1, .2, .3, .4, .5, .6, .7, .8, .9, .99]`\n# train_df.describe(percentiles=[.1, .2, .3, .4, .5, .6, .7, .8, .9, .99])","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"33bbd1709db622978c0c5879e7c5532d4734ade0","_cell_guid":"5462bc60-258c-76bf-0a73-9adc00a2f493"},"cell_type":"markdown","source":"**What is the distribution of categorical features?** 分类类型的特征分布是怎么样的？\n\n- Names are unique across the dataset (count=unique=891)\n- Sex variable as two possible values with 65% male (top=male, freq=577/count=891).\n- Cabin values have several dupicates across samples. Alternatively several passengers shared a cabin.\n- Embarked takes three possible values. S port used by most passengers (top=S)\n- Ticket feature has high ratio (22%) of duplicate values (unique=681)."},{"metadata":{"_uuid":"daa8663f577f9c1a478496cf14fe363570457191","_cell_guid":"8066b378-1964-92e8-1352-dcac934c6af3","trusted":true},"cell_type":"code","source":"train_df.describe(include=['O'])","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"14f19974fe4c396ce22534e56a2277104891cd39"},"cell_type":"markdown","source":"**A heat map of correlation may give us a understanding of which variables are important**  数值类型的变量之间的相关性"},{"metadata":{"trusted":true,"_uuid":"11b61ef3971ad9378719736ce25eabc6716fb120"},"cell_type":"code","source":"corr = train_df.corr()\n_ , ax = plt.subplots( figsize =( 12 , 10 ) )\ncmap = sns.diverging_palette( 220 , 10 , as_cmap = True )\n_ = sns.heatmap(\n        corr, \n        cmap = cmap,\n        square=True, \n        cbar_kws={ 'shrink' : .9 }, \n        ax=ax, \n        annot = True, \n        annot_kws = { 'fontsize' : 12 }\n    )","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"c1d35ebd89a0cf7d7b409470bbb9ecaffd2a9680","_cell_guid":"2cb22b88-937d-6f14-8b06-ea3361357889"},"cell_type":"markdown","source":"## Assumptions based on data analysis （数据分析的初步结果）\n\nWe arrive at following assumptions based on data analysis done so far. We may validate these assumptions further before taking appropriate actions.\n\n**Correlating.** (分析各个特征和生存率survived变量之间相关性)\n\nWe want to know how well does each feature correlate with Survival. We want to do this early in our project and match these quick correlations with modelled correlations later in the project.\n\n**Completing.** (Age和Embarked有missing value，需要补全)\n1. We may want to complete **Age** feature as it is definitely correlated to survival.\n2. We may want to complete the **Embarked** feature as it may also correlate with survival or another important feature.\n\n**Correcting.** (用处不大的变量drop掉或者尽量提取有用信息）\n\n1. **Ticket** feature may be dropped from our analysis as it contains high ratio of duplicates (22%) and there may not be a correlation between Ticket and survival.\n2. **Cabin** feature may be dropped as it is highly incomplete or contains many null values both in training and test dataset.\n3. **PassengerId** may be dropped from training dataset as it does not contribute to survival.\n4. **Name** feature is relatively non-standard, may not contribute directly to survival, so maybe dropped or processed.\n\n**Creating.** （创造一些新的特征来增加数据信息量）\n\n1. We may want to create a new feature called **Family** based on Parch and SibSp to get total count of family members on board.\n2. We may want to engineer the Name feature to extract **Title** as a new feature.\n3. We may want to create new feature for **Age bands**. This turns a continous numerical feature into an ordinal categorical feature.\n4. We may also want to create a **Fare range** feature if it helps our analysis.\n\n**Classifying.** （验证一些分析假设）\n\nWe may also add to our assumptions based on the problem description noted earlier.\n\n1. Women (Sex=female) were more likely to have survived.\n2. Children (Age<?) were more likely to have survived. \n3. The upper-class passengers (Pclass=1) were more likely to have survived."},{"metadata":{"_uuid":"946ee6ca01a3e4eecfa373ca00f88042b683e2ad","_cell_guid":"6db63a30-1d86-266e-2799-dded03c45816"},"cell_type":"markdown","source":"## Analyze by pivoting features （运用图表来验证我们的分析结果）\n\nTo confirm some of our observations and assumptions, we can quickly analyze our feature correlations by pivoting features against each other. We can only do so at this stage for features which do not have any empty values. It also makes sense doing so only for features which are categorical (Sex), ordinal (Pclass) or discrete (SibSp, Parch) type.\n\n- **Pclass** We observe significant correlation (>0.5) among Pclass=1 and Survived (classifying #3). We decide to include this feature in our model.\n- **Sex** We confirm the observation during problem definition that Sex=female had very high survival rate at 74% (classifying #1).\n- **SibSp and Parch** These features have zero correlation for certain values. It may be best to derive a feature or a set of features from these individual features (creating #1)."},{"metadata":{"_uuid":"97a845528ce9f76e85055a4bb9e97c27091f6aa1","_cell_guid":"0964832a-a4be-2d6f-a89e-63526389cee9","trusted":true},"cell_type":"code","source":"train_df[['Pclass', 'Survived']].groupby(['Pclass'], as_index=False).mean().sort_values(by='Survived', ascending=False)","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"00a2f2bca094c5984e6a232c730c8b232e7e20bb","_cell_guid":"68908ba6-bfe9-5b31-cfde-6987fc0fbe9a","trusted":true},"cell_type":"code","source":"train_df[[\"Sex\", \"Survived\"]].groupby(['Sex'], as_index=False).mean().sort_values(by='Survived', ascending=False)","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"a8f7a16c54417dcd86fc48aeef0c4b240d47d71b","_cell_guid":"01c06927-c5a6-342a-5aa8-2e486ec3fd7c","trusted":true},"cell_type":"code","source":"train_df[[\"SibSp\", \"Survived\"]].groupby(['SibSp'], as_index=False).mean().sort_values(by='Survived', ascending=False)","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"d13320e096294f23ceedc41be708d8c774587073"},"cell_type":"code","source":"train_df[[\"SibSp\", \"Survived\"]].groupby(['SibSp'], as_index=False).count()","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"5d953a6779b00b7f3794757dec8744a03162c8fd","_cell_guid":"e686f98b-a8c9-68f8-36a4-d4598638bbd5","trusted":true},"cell_type":"code","source":"train_df[[\"Parch\", \"Survived\"]].groupby(['Parch'], as_index=False).mean().sort_values(by='Survived', ascending=False)","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"2b02a1f849ca6b9688d084a3723148441024e86c"},"cell_type":"code","source":"# Parch_Survived_df = train_df[[\"Parch\", \"Survived\"]].groupby(['Parch'], as_index=False).mean().sort_values(by='Survived', ascending=False)\n# sns.scatterplot(\"Parch\", \"Survived\", data=Parch_Survived_df)","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"9720e826154360d89691a1b4de1e0471a9a431de"},"cell_type":"code","source":"train_df[[\"Parch\", \"Survived\"]].groupby(['Parch'], as_index=False).count()","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"5c6204d01f5a9040cf0bb7c678686ae48daa201f","_cell_guid":"0d43550e-9eff-3859-3568-8856570eff76"},"cell_type":"markdown","source":"## Analyze by visualizing data （运用图表来验证我们的分析结果）\n\nNow we can continue confirming some of our assumptions using visualizations for analyzing the data.\n\n### Correlating numerical features\n\nLet us start by understanding correlations between numerical features and our solution goal (Survived).\n\nA histogram chart is useful for analyzing continous numerical variables like Age where banding or ranges will help identify useful patterns. The histogram can indicate distribution of samples using automatically defined bins or equally ranged bands. This helps us answer questions relating to specific bands (Did infants have better survival rate?)\n\nNote that x-axis in historgram visualizations represents the count of samples or passengers.\n\n**Observations.** (Age的影响）\n\n- Infants (Age <=4) had high survival rate. (小于4岁的乘客生存率高)\n- Oldest passengers (Age = 80) survived. (80高龄的乘客存活了)\n- Large number of 15-25 year olds did not survive. (15-25的青壮年存活率低)\n- Most passengers are in 15-35 age range. (乘客大都在15-35之间)\n\n**Decisions.**\n\nThis simple analysis confirms our assumptions as decisions for subsequent workflow stages.\n\n- We should consider Age (our assumption classifying #2) in our model training.\n- Complete the Age feature for null values (completing #1).\n- We should band age groups (creating #3)."},{"metadata":{"_uuid":"d3a1fa63e9dd4f8a810086530a6363c94b36d030","_cell_guid":"50294eac-263a-af78-cb7e-3778eb9ad41f","trusted":true},"cell_type":"code","source":"g = sns.FacetGrid(train_df, col='Survived', height=5)\ng.map(plt.hist, 'Age', bins=20)","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"9d4f8851ac73f64861bdda89190a759a79b5eccf"},"cell_type":"code","source":"facet = sns.FacetGrid( train_df, aspect=2 , col = \"Survived\", height=5 )\nfacet.map( sns.kdeplot , 'Age' , shade= True )\nfacet.add_legend()\n","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"892259f68c2ecf64fd258965cff1ecfe77dd73a9","_cell_guid":"87096158-4017-9213-7225-a19aea67a800"},"cell_type":"markdown","source":"### Correlating numerical and ordinal features\n\nWe can combine multiple features for identifying correlations using a single plot. This can be done with numerical and categorical features which have numeric values.\n\n**Observations.** （Pclass的影响）\n\n- Pclass=3 had most passengers, however most did not survive. Confirms our classifying assumption #2. （Pclass 3有最多的乘客，不过大部分没能生还）\n- Infant passengers in Pclass=2 and Pclass=3 mostly survived. Further qualifies our classifying assumption #2.（Pclass2和Pclass3的婴幼儿乘客大多生还）\n- Most passengers in Pclass=1 survived. Confirms our classifying assumption #3. （Pclass1的乘客大部分生还）\n- Pclass varies in terms of Age distribution of passengers.（每个class的年龄分布都不一样）\n\n**Decisions.**\n\n- Consider Pclass for model training."},{"metadata":{"_uuid":"4f5bcfa97c8a72f8b413c786954f3a68e135e05a","_cell_guid":"916fdc6b-0190-9267-1ea9-907a3d87330d","trusted":true},"cell_type":"code","source":"# grid = sns.FacetGrid(train_df, col='Pclass', hue='Survived')\ngrid = sns.FacetGrid(train_df, col='Survived', row='Pclass', height=3, aspect=1.6)\ngrid.map(plt.hist, 'Age', alpha=.5, bins=20)\ngrid.add_legend();","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"892ab7ee88b1b1c5f1ac987884fa31e111bb0507","_cell_guid":"36f5a7c0-c55c-f76f-fdf8-945a32a68cb0"},"cell_type":"markdown","source":"### Correlating categorical features\n\nNow we can correlate categorical features with our solution goal.\n\n**Observations.** （Sex和Embarked的影响）\n\n- Female passengers had much better survival rate than males. Confirms classifying (#1). （女性乘客生存率更高）\n- Exception in Embarked=C where males had higher survival rate. （C处上船的男性生存率似乎更高，但可能是由其他原因导致的，比如Pclass）This could be a correlation between Pclass and Embarked and in turn Pclass and Survived, not necessarily direct correlation between Embarked and Survived.\n- Ports of embarkation have varying survival rates for Pclass=3 and among male passengers. Correlating (#1). （不同港口上船的乘客生存率不一样）\n\n**Decisions.**\n\n- Add Sex feature to model training.\n- Complete and add Embarked feature to model training."},{"metadata":{"_uuid":"c0e1f01b3f58e8f31b938b0e5eb1733132edc8ad","_cell_guid":"db57aabd-0e26-9ff9-9ebd-56d401cdf6e8","trusted":true,"scrolled":false},"cell_type":"code","source":"import warnings\nwarnings.filterwarnings('ignore')\n# grid = sns.FacetGrid(train_df, col='Embarked')\ngrid = sns.FacetGrid(train_df, row='Embarked', size=3, aspect=1.6)\ngrid.map(sns.pointplot, 'Pclass', 'Survived', 'Sex', palette='deep')\ngrid.add_legend()","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"4a6ba8ed2abdc13ae73b409c6cb76dc9cb922dec"},"cell_type":"code","source":"plt.figure(figsize=(6,6)) \nsns.countplot('Embarked', hue='Survived', data=train_df)\nplt.show()","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"fd824f937dcb80edd4117a2927cc0d7f99d934b8","_cell_guid":"6b3f73f4-4600-c1ce-34e0-bd7d9eeb074a"},"cell_type":"markdown","source":"### Correlating categorical and numerical features\n\nWe may also want to correlate categorical features (with non-numeric values) and numeric features. We can consider correlating Embarked (Categorical non-numeric), Sex (Categorical non-numeric), Fare (Numeric continuous), with Survived (Categorical numeric).\n\n**Observations.** （Fare的影响）\n\n- Higher fare paying passengers had better survival. Confirms our assumption for creating (#4) fare ranges.（票价高的乘客生存率偏高）\n- Port of embarkation correlates with survival rates. Confirms correlating (#1) and completing (#2). （上船港口和生存率可能有关）\n\n**Decisions.**\n\n- Consider banding Fare feature."},{"metadata":{"_uuid":"c8fd535ac1bc90127369027c2101dbc939db118e","_cell_guid":"a21f66ac-c30d-f429-cc64-1da5460d16a9","trusted":true},"cell_type":"code","source":"# grid = sns.FacetGrid(train_df, col='Embarked', hue='Survived', palette={0: 'k', 1: 'w'})\ngrid = sns.FacetGrid(train_df, row='Embarked', col='Survived', size=3, aspect=1.6)\ngrid.map(sns.barplot, 'Sex', 'Fare', alpha=.5, ci=None)\ngrid.add_legend()","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"73a9111a8dc2a6b8b6c78ef628b6cae2a63fc33f","_cell_guid":"cfac6291-33cc-506e-e548-6cad9408623d"},"cell_type":"markdown","source":"<a id=\"step4\"></a>\n# 4. Wrangle, prepare, cleanse the data 处理数据\n- Wrangle data (Feature Engineering) 特征工程\n\nWe have collected several assumptions and decisions regarding our datasets and solution requirements. So far we did not have to change a single feature or value to arrive at these. Let us now execute our decisions and assumptions for correcting, creating, and completing goals.\n\n### Correcting by dropping features （删除一些特征）\n\nThis is a good starting goal to execute. By dropping features we are dealing with fewer data points. Speeds up our notebook and eases the analysis.\n\nBased on our assumptions and decisions we want to drop the Cabin (correcting #2) and Ticket (correcting #1) features.\n\nNote that where applicable we perform operations on both training and testing datasets together to stay consistent."},{"metadata":{"_uuid":"e328d9882affedcfc4c167aa5bb1ac132547558c","_cell_guid":"da057efe-88f0-bf49-917b-bb2fec418ed9","trusted":true},"cell_type":"code","source":"print(\"Before\", train_df.shape, test_df.shape, combine[0].shape, combine[1].shape)\n\ntrain_df = train_df.drop(['Ticket', 'Cabin'], axis=1)\ntest_df = test_df.drop(['Ticket', 'Cabin'], axis=1)\ncombine = [train_df, test_df]\n\nprint(\"After\", train_df.shape, test_df.shape, combine[0].shape, combine[1].shape)","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"21d5c47ee69f8fbef967f6f41d736b5d4eb6596f","_cell_guid":"6b3a1216-64b6-7fe2-50bc-e89cc964a41c"},"cell_type":"markdown","source":"### Creating new feature extracting from existing （创造一些新的特征）\n\nWe want to analyze if Name feature can be engineered to extract titles and test correlation between titles and survival, before dropping Name and PassengerId features.\n\nIn the following code we extract Title feature using regular expressions. The RegEx pattern `(\\w+\\.)` matches the first word which ends with a dot character within Name feature. The `expand=False` flag returns a DataFrame.\n\n**Observations.** （用Name来产生Title特征）\n\nWhen we plot Title, Age, and Survived, we note the following observations.\n\n- Most titles band Age groups accurately.\n- Survival among Title Age bands varies slightly.\n- Certain titles mostly survived (Mme, Lady, Sir) or did not (Don, Rev, Jonkheer).\n\n**Decision.**\n\n- We decide to retain the new Title feature for model training."},{"metadata":{"_uuid":"c916644bd151f3dc8fca900f656d415b4c55e2bc","_cell_guid":"df7f0cd4-992c-4a79-fb19-bf6f0c024d4b","trusted":true},"cell_type":"code","source":"for dataset in combine:\n    dataset['Title'] = dataset.Name.str.extract(' ([A-Za-z]+)\\.', expand=False)\n\npd.crosstab(train_df['Title'], train_df['Sex'])","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"f766d512ea5bfe60b5eb7a816f482f2ab688fd2f","_cell_guid":"908c08a6-3395-19a5-0cd7-13341054012a"},"cell_type":"markdown","source":"We can replace many titles with a more common name or classify them as `Rare`."},{"metadata":{"_uuid":"b8cd938fba61fb4e226c77521b012f4bb8aa01d0","_cell_guid":"553f56d7-002a-ee63-21a4-c0efad10cfe9","trusted":true},"cell_type":"code","source":"for dataset in combine:\n    dataset['Title'] = dataset['Title'].replace(['Lady', 'Countess','Capt', 'Col', \n                                                 'Don', 'Dr', 'Major', 'Rev', 'Sir', 'Jonkheer', 'Dona'], 'Rare')\n\n    dataset['Title'] = dataset['Title'].replace('Mlle', 'Miss')\n    dataset['Title'] = dataset['Title'].replace('Ms', 'Miss')\n    dataset['Title'] = dataset['Title'].replace('Mme', 'Mrs')\n    \ntrain_df[['Title', 'Survived']].groupby(['Title'], as_index=False).mean()","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"de245fe76474d46995a5acc31b905b8aaa5893f6","_cell_guid":"6d46be9a-812a-f334-73b9-56ed912c9eca"},"cell_type":"markdown","source":"We can convert the categorical titles to ordinal."},{"metadata":{"_uuid":"e805ad52f0514497b67c3726104ba46d361eb92c","_cell_guid":"67444ebc-4d11-bac1-74a6-059133b6e2e8","trusted":true},"cell_type":"code","source":"title_mapping = {\"Mr\": 1, \"Miss\": 2, \"Mrs\": 3, \"Master\": 4, \"Rare\": 5}\nfor dataset in combine:\n    dataset['Title'] = dataset['Title'].map(title_mapping)\n    dataset['Title'] = dataset['Title'].fillna(0)\n\ntrain_df.head()","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"5fefaa1b37c537dda164c87a757fe705a99815d9","_cell_guid":"f27bb974-a3d7-07a1-f7e4-876f6da87e62"},"cell_type":"markdown","source":"Now we can safely drop the Name feature from training and testing datasets. We also do not need the PassengerId feature in the training dataset."},{"metadata":{"_uuid":"1da299cf2ffd399fd5b37d74fb40665d16ba5347","_cell_guid":"9d61dded-5ff0-5018-7580-aecb4ea17506","trusted":true},"cell_type":"code","source":"train_df = train_df.drop(['Name', 'PassengerId'], axis=1)\ntest_df = test_df.drop(['Name'], axis=1)\ncombine = [train_df, test_df]\ntrain_df.shape, test_df.shape","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"a1ac66c79b279d94860e66996d3d8dba801a6d9a","_cell_guid":"2c8e84bb-196d-bd4a-4df9-f5213561b5d3"},"cell_type":"markdown","source":"### Converting a categorical feature (将字符串特征转换成数字特征）\n\nNow we can convert features which contain strings to numerical values. This is required by most model algorithms. Doing so will also help us in achieving the feature completing goal.\n\nLet us start by converting Sex feature to a new feature called Gender where female=1 and male=0."},{"metadata":{"_uuid":"840498eaee7baaca228499b0a5652da9d4edaf37","_cell_guid":"c20c1df2-157c-e5a0-3e24-15a828095c96","trusted":true},"cell_type":"code","source":"for dataset in combine:\n    dataset['Sex'] = dataset['Sex'].map( {'female': 1, 'male': 0} ).astype(int)\n\ntrain_df.head()","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"6da8bfe6c832f4bd2aa1312bdd6b8b4af48a012e","_cell_guid":"d72cb29e-5034-1597-b459-83a9640d3d3a"},"cell_type":"markdown","source":"### Completing a numerical continuous feature （补全missing value）\n\nNow we should start estimating and completing features with missing or null values. We will first do this for the Age feature.\n\nWe can consider three methods to complete a numerical continuous feature.\n\n1. A simple way is to generate random numbers between mean and [standard deviation](https://en.wikipedia.org/wiki/Standard_deviation).\n\n**2. More accurate way of guessing missing values is to use other correlated features. In our case we note correlation among Age, Gender, and Pclass. Guess Age values using [median](https://en.wikipedia.org/wiki/Median) values for Age across sets of Pclass and Gender feature combinations. So, median Age for Pclass=1 and Gender=0, Pclass=1 and Gender=1, and so on...**\n\n3. Combine methods 1 and 2. So instead of guessing age values based on median, use random numbers between mean and standard deviation, based on sets of Pclass and Gender combinations.\n\nMethod 1 and 3 will introduce random noise into our models. The results from multiple executions might vary. We will prefer method 2."},{"metadata":{"_uuid":"345038c8dd1bac9a9bc5e2cfee13fcc1f833eee0","_cell_guid":"c311c43d-6554-3b52-8ef8-533ca08b2f68","trusted":true},"cell_type":"code","source":"# grid = sns.FacetGrid(train_df, col='Pclass', hue='Gender')\ngrid = sns.FacetGrid(train_df, row='Pclass', col='Sex', size=2.2, aspect=1.6)\ngrid.map(plt.hist, 'Age', alpha=.5, bins=20)\ngrid.add_legend()","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"6b22ac53d95c7979d5f4580bd5fd29d27155c347","_cell_guid":"a4f166f9-f5f9-1819-66c3-d89dd5b0d8ff"},"cell_type":"markdown","source":"Let us start by preparing an empty array to contain guessed Age values based on Pclass x Gender combinations. （按Pclass和Gender分成6组）"},{"metadata":{"_uuid":"24a0971daa4cbc3aa700bae42e68c17ce9f3a6e2","_cell_guid":"9299523c-dcf1-fb00-e52f-e2fb860a3920","trusted":true},"cell_type":"code","source":"guess_ages = np.zeros((2,3))\nguess_ages","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"8acd90569767b544f055d573bbbb8f6012853385","_cell_guid":"ec9fed37-16b1-5518-4fa8-0a7f579dbc82"},"cell_type":"markdown","source":"Now we iterate over Sex (0 or 1) and Pclass (1, 2, 3) to calculate guessed values of Age for the six combinations."},{"metadata":{"_uuid":"31198f0ad0dbbb74290ebe135abffa994b8f58f3","_cell_guid":"a4015dfa-a0ab-65bc-0cbe-efecf1eb2569","trusted":true},"cell_type":"code","source":"for dataset in combine:\n    for i in range(0, 2):\n        for j in range(0, 3):\n            guess_df = dataset[(dataset['Sex'] == i) & (dataset['Pclass'] == j+1)]['Age'].dropna()\n\n            # age_mean = guess_df.mean()\n            # age_std = guess_df.std()\n            # age_guess = rnd.uniform(age_mean - age_std, age_mean + age_std)\n\n            age_guess = guess_df.median()\n\n            # Convert random age float to nearest .5 age\n#             guess_ages[i,j] = int( age_guess/0.5 + 0.5 ) * 0.5\n            guess_ages[i,j] = age_guess\n            \n    for i in range(0, 2):\n        for j in range(0, 3):\n            dataset.loc[(dataset.Age.isnull()) & (dataset.Sex == i) & (dataset.Pclass == j+1),'Age'] = guess_ages[i,j]\n\n    dataset['Age'] = dataset['Age'].astype(int)\n\ntrain_df.head()","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"e7c52b44b703f28e4b6f4ddba67ab65f40274550","_cell_guid":"dbe0a8bf-40bc-c581-e10e-76f07b3b71d4"},"cell_type":"markdown","source":"Let us create Age bands and determine correlations with Survived."},{"metadata":{"_uuid":"5c8b4cbb302f439ef0d6278dcfbdafd952675353","_cell_guid":"725d1c84-6323-9d70-5812-baf9994d3aa1","trusted":true},"cell_type":"code","source":"train_df['AgeBand'] = pd.cut(train_df['Age'], 5)\ntrain_df[['AgeBand', 'Survived']].groupby(['AgeBand'], as_index=False).mean().sort_values(by='AgeBand', ascending=True)","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"856392dd415ac14ab74a885a37d068fc7a58f3a5","_cell_guid":"ba4be3a0-e524-9c57-fbec-c8ecc5cde5c6"},"cell_type":"markdown","source":"Let us replace Age with ordinals based on these bands."},{"metadata":{"_uuid":"ee13831345f389db407c178f66c19cc8331445b0","_cell_guid":"797b986d-2c45-a9ee-e5b5-088de817c8b2","trusted":true},"cell_type":"code","source":"for dataset in combine:    \n    dataset.loc[ dataset['Age'] <= 16, 'Age'] = 0\n    dataset.loc[(dataset['Age'] > 16) & (dataset['Age'] <= 32), 'Age'] = 1\n    dataset.loc[(dataset['Age'] > 32) & (dataset['Age'] <= 48), 'Age'] = 2\n    dataset.loc[(dataset['Age'] > 48) & (dataset['Age'] <= 64), 'Age'] = 3\n    dataset.loc[ dataset['Age'] > 64, 'Age']\ntrain_df.head()","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"8e3fbc95e0fd6600e28347567416d3f0d77a24cc","_cell_guid":"004568b6-dd9a-ff89-43d5-13d4e9370b1d"},"cell_type":"markdown","source":"We can now remove the AgeBand feature."},{"metadata":{"_uuid":"1ea01ccc4a24e8951556d97c990aa0136da19721","_cell_guid":"875e55d4-51b0-5061-b72c-8a23946133a3","trusted":true},"cell_type":"code","source":"train_df = train_df.drop(['AgeBand'], axis=1)\ncombine = [train_df, test_df]\ntrain_df.head()","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"e3d4a2040c053fbd0486c8cfc4fec3224bd3ebb3","_cell_guid":"1c237b76-d7ac-098f-0156-480a838a64a9"},"cell_type":"markdown","source":"### Create new feature combining existing features （用现有特征组成新的特征）\n\nWe can create a new feature for FamilySize which combines Parch and SibSp. This will enable us to drop Parch and SibSp from our datasets. (创建一个FamiliSize特征）"},{"metadata":{"_uuid":"33d1236ce4a8ab888b9fac2d5af1c78d174b32c7","_cell_guid":"7e6c04ed-cfaa-3139-4378-574fd095d6ba","trusted":true},"cell_type":"code","source":"for dataset in combine:\n    dataset['FamilySize'] = dataset['SibSp'] + dataset['Parch'] + 1\n\ntrain_df[['FamilySize', 'Survived']].groupby(['FamilySize'], as_index=False).mean().sort_values(by='Survived', ascending=False)","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"67f8e4474cd1ecf4261c153ce8b40ea23cf659e4","_cell_guid":"842188e6-acf8-2476-ccec-9e3451e4fa86"},"cell_type":"markdown","source":"We can create another feature called IsAlone."},{"metadata":{"_uuid":"3b8db81cc3513b088c6bcd9cd1938156fe77992f","_cell_guid":"5c778c69-a9ae-1b6b-44fe-a0898d07be7a","trusted":true},"cell_type":"code","source":"for dataset in combine:\n    dataset['IsAlone'] = 0\n    dataset.loc[dataset['FamilySize'] == 1, 'IsAlone'] = 1\n\ntrain_df[['IsAlone', 'Survived']].groupby(['IsAlone'], as_index=False).mean()","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"3da4204b2c78faa54a94bbad78a8aa85fbf90c87","_cell_guid":"e6b87c09-e7b2-f098-5b04-4360080d26bc"},"cell_type":"markdown","source":"Let us drop Parch, SibSp, and FamilySize features in favor of IsAlone."},{"metadata":{"_uuid":"1e3479690ef7cd8ee10538d4f39d7117246887f0","_cell_guid":"74ee56a6-7357-f3bc-b605-6c41f8aa6566","trusted":true},"cell_type":"code","source":"# train_df = train_df.drop(['Parch', 'SibSp', 'FamilySize'], axis=1)\n# test_df = test_df.drop(['Parch', 'SibSp', 'FamilySize'], axis=1)\ntrain_df = train_df.drop(['Parch', 'SibSp'], axis=1)\ntest_df = test_df.drop(['Parch', 'SibSp'], axis=1)\ncombine = [train_df, test_df]\n\ntrain_df.head()","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"71b800ed96407eba05220f76a1288366a22ec887","_cell_guid":"f890b730-b1fe-919e-fb07-352fbd7edd44"},"cell_type":"markdown","source":"We can also create an artificial feature combining Pclass and Age. （创建一个Age*Class的特征）"},{"metadata":{"_uuid":"aac2c5340c06210a8b0199e15461e9049fbf2cff","_cell_guid":"305402aa-1ea1-c245-c367-056eef8fe453","trusted":true},"cell_type":"code","source":"for dataset in combine:\n    dataset['Age*Class'] = dataset.Age * dataset.Pclass\n\ntrain_df.loc[:, ['Age*Class', 'Age', 'Pclass']].head(10)","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"8264cc5676db8cd3e0b3e3f078cbaa74fd585a3c","_cell_guid":"13292c1b-020d-d9aa-525c-941331bb996a"},"cell_type":"markdown","source":"### Completing a categorical feature （补全Embarked类型特征）\n\nEmbarked feature takes S, Q, C values based on port of embarkation. Our training dataset has two missing values. We simply fill these with the most common occurance."},{"metadata":{"_uuid":"1e3f8af166f60a1b3125a6b046eff5fff02d63cf","_cell_guid":"bf351113-9b7f-ef56-7211-e8dd00665b18","trusted":true},"cell_type":"code","source":"freq_port = train_df.Embarked.dropna().mode()[0]\nfreq_port","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"d85b5575fb45f25749298641f6a0a38803e1ff22","_cell_guid":"51c21fcc-f066-cd80-18c8-3d140be6cbae","trusted":true},"cell_type":"code","source":"for dataset in combine:\n    dataset['Embarked'] = dataset['Embarked'].fillna(freq_port)\n    \ntrain_df[['Embarked', 'Survived']].groupby(['Embarked'], as_index=False).mean().sort_values(by='Survived', ascending=False)","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"d8830e997995145314328b6218b5606df04499b0","_cell_guid":"f6acf7b2-0db3-e583-de50-7e14b495de34"},"cell_type":"markdown","source":"### Converting categorical feature to numeric （将类型特征数值化）\n\nWe can now convert the EmbarkedFill feature by creating a new numeric Port feature."},{"metadata":{"_uuid":"e480a1ef145de0b023821134896391d568a6f4f9","_cell_guid":"89a91d76-2cc0-9bbb-c5c5-3c9ecae33c66","trusted":true},"cell_type":"code","source":"for dataset in combine:\n    dataset['Embarked'] = dataset['Embarked'].map( {'S': 0, 'C': 1, 'Q': 2} ).astype(int)\n\ntrain_df.head()","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"d79834ebc4ab9d48ed404584711475dbf8611b91","_cell_guid":"e3dfc817-e1c1-a274-a111-62c1c814cecf"},"cell_type":"markdown","source":"### Quick completing and converting a numeric feature （补全Fare特征）\n\nWe can now complete the Fare feature for single missing value in test dataset using mode to get the value that occurs most frequently for this feature. We do this in a single line of code.\n\nNote that we are not creating an intermediate new feature or doing any further analysis for correlation to guess missing feature as we are replacing only a single value. The completion goal achieves desired requirement for model algorithm to operate on non-null values.\n\nWe may also want round off the fare to two decimals as it represents currency."},{"metadata":{"_uuid":"aacb62f3526072a84795a178bd59222378bab180","_cell_guid":"3600cb86-cf5f-d87b-1b33-638dc8db1564","trusted":true},"cell_type":"code","source":"# test_df['Fare'].fillna(test_df['Fare'].dropna().median(), inplace=True)\ntest_df['Fare'].fillna(test_df['Fare'].dropna().mode()[0], inplace=True)\ntest_df.head()","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"3466d98e83899d8b38a36ede794c68c5656f48e6","_cell_guid":"4b816bc7-d1fb-c02b-ed1d-ee34b819497d"},"cell_type":"markdown","source":"We can now create FareBand."},{"metadata":{"_uuid":"45a97ad9b33268360c98477697e5a2ad0db6d086"},"cell_type":"markdown","source":"## See difference between qcut and cut (qcut和cut的区别）"},{"metadata":{"trusted":true,"_uuid":"19656333161b614c52ae6b9b8c76e8ec7c74a299"},"cell_type":"code","source":"pd.cut(train_df['Fare'], 4).value_counts()","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"0c7ffb5de8a9f10dc3e9f876fe80a79d75d5c1d3"},"cell_type":"code","source":"pd.qcut(train_df['Fare'], 4).value_counts()","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"b9a78f6b4c72520d4ad99d2c89c84c591216098d","_cell_guid":"0e9018b1-ced5-9999-8ce1-258a0952cbf2","trusted":true},"cell_type":"code","source":"train_df['FareBand'] = pd.qcut(train_df['Fare'], 4)\ntrain_df[['FareBand', 'Survived']].groupby(['FareBand'], as_index=False).mean().sort_values(by='FareBand', ascending=True)","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"89400fba71af02d09ff07adf399fb36ac4913db6","_cell_guid":"d65901a5-3684-6869-e904-5f1a7cce8a6d"},"cell_type":"markdown","source":"Convert the Fare feature to ordinal values based on the FareBand. （创建FareBand特征）"},{"metadata":{"_uuid":"640f305061ec4221a45ba250f8d54bb391035a57","_cell_guid":"385f217a-4e00-76dc-1570-1de4eec0c29c","trusted":true},"cell_type":"code","source":"for dataset in combine:\n    dataset.loc[ dataset['Fare'] <= 7.91, 'Fare'] = 0\n    dataset.loc[(dataset['Fare'] > 7.91) & (dataset['Fare'] <= 14.454), 'Fare'] = 1\n    dataset.loc[(dataset['Fare'] > 14.454) & (dataset['Fare'] <= 31), 'Fare']   = 2\n    dataset.loc[ dataset['Fare'] > 31, 'Fare'] = 3\n    dataset['Fare'] = dataset['Fare'].astype(int)\n\ntrain_df = train_df.drop(['FareBand'], axis=1)\ncombine = [train_df, test_df]\n    \ntrain_df.head(10)","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"531994ed95a3002d1759ceb74d9396db706a41e2","_cell_guid":"27272bb9-3c64-4f9a-4a3b-54f02e1c8289"},"cell_type":"markdown","source":"And the test dataset."},{"metadata":{"_uuid":"8453cecad81fcc44de3f4e4e4c3ce6afa977740d","_cell_guid":"d2334d33-4fe5-964d-beac-6aa620066e15","trusted":true},"cell_type":"code","source":"test_df.head(10)","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"13501e742fba8caf023818fb0a21620bc9a765f4"},"cell_type":"markdown","source":"# We are ready to model!!!"},{"metadata":{"_uuid":"a55f20dd6654610ff2d66c1bf3e4c6c73dcef9e5","_cell_guid":"69783c08-c8cc-a6ca-2a9a-5e75581c6d31"},"cell_type":"markdown","source":"<a id=\"step5\"></a>\n# 5. Model, predict and solve the problem 建模预测\n\nNow we are ready to train a model and predict the required solution. There are 60+ predictive modelling algorithms to choose from. We must understand the type of problem and solution requirement to narrow down to a select few models which we can evaluate. Our problem is a classification and regression problem. We want to identify relationship between output (Survived or not) with other variables or features (Gender, Age, Port...). We are also perfoming a category of machine learning which is called supervised learning as we are training our model with a given dataset. With these two criteria - Supervised Learning plus Classification and Regression, we can narrow down our choice of models to a few. These include:\n\n- Logistic Regression\n- Support Vector Machines\n- Linear SVC\n- KNN or k-Nearest Neighbors\n- Perceptron\n- Stochastic Gradient Descent\n- Decision Tree\n- Random Forrest\n- Artificial neural network"},{"metadata":{"_uuid":"04d2235855f40cffd81f76b977a500fceaae87ad","_cell_guid":"0acf54f9-6cf5-24b5-72d9-29b30052823a","trusted":true},"cell_type":"code","source":"X_train = train_df.drop(\"Survived\", axis=1)\nY_train = train_df[\"Survived\"]\nX_test  = test_df.drop(\"PassengerId\", axis=1).copy()\nX_train.shape, Y_train.shape, X_test.shape","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"782903c09ec9ee4b6f3e03f7c8b5a62c00461deb","_cell_guid":"579bc004-926a-bcfe-e9bb-c8df83356876"},"cell_type":"markdown","source":"Logistic Regression is a useful model to run early in the workflow. Logistic regression measures the relationship between the categorical dependent variable and one or more independent variables (features) by estimating probabilities using a logistic function, which is the cumulative logistic distribution. Reference [Wikipedia](https://en.wikipedia.org/wiki/Logistic_regression).\n![](https://github.com/icwangjimmy/machine_learning_and_python_in_finance/raw/master/pic/LogReg_1.png)\nNote the confidence score generated by the model based on our training dataset."},{"metadata":{"_uuid":"a649b9c53f4c7b40694f60f5c8dc14ec5ef519ec","_cell_guid":"0edd9322-db0b-9c37-172d-a3a4f8dec229","trusted":true},"cell_type":"code","source":"# Logistic Regression\nlogreg = LogisticRegression()\nlogreg.fit(X_train, Y_train)\nY_pred_1 = logreg.predict(X_test)\nacc_log = round(logreg.score(X_train, Y_train) * 100, 2)\nacc_log","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"180e27c96c821656a84889f73986c6ddfff51ed3","_cell_guid":"3af439ae-1f04-9236-cdc2-ec8170a0d4ee"},"cell_type":"markdown","source":"We can use Logistic Regression to validate our assumptions and decisions for feature creating and completing goals. This can be done by calculating the coefficient of the features in the decision function.\n\nPositive coefficients increase the log-odds of the response (and thus increase the probability), and negative coefficients decrease the log-odds of the response (and thus decrease the probability).\n\n- Sex is highest positivie coefficient, implying as the Sex value increases (male: 0 to female: 1), the probability of Survived=1 increases the most.\n- Inversely as Pclass increases, probability of Survived=1 decreases the most.\n- This way Age*Class is a good artificial feature to model as it has second highest negative correlation with Survived.\n- So is Title as second highest positive correlation."},{"metadata":{"trusted":true,"_uuid":"dc9800205a4dc691eb756bcdcbbca69045e93f84"},"cell_type":"code","source":"logreg.coef_[0]","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"6e6f58053fae405fc93d312fc999f3904e708dbe","_cell_guid":"e545d5aa-4767-7a41-5799-a4c5e529ce72","trusted":true},"cell_type":"code","source":"coeff_df = pd.DataFrame(train_df.columns.delete(0))\ncoeff_df.columns = ['Feature']\ncoeff_df[\"Correlation\"] = pd.Series(logreg.coef_[0])\n\ncoeff_df.sort_values(by='Correlation', ascending=False)","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"ccba9ac0a9c3c648ef9bc778977ab99066ab3945","_cell_guid":"ac041064-1693-8584-156b-66674117e4d0"},"cell_type":"markdown","source":"Next we model using Support Vector Machines which are supervised learning models with associated learning algorithms that analyze data used for classification and regression analysis. Given a set of training samples, each marked as belonging to one or the other of **two categories**, an SVM training algorithm builds a model that assigns new test samples to one category or the other, making it a non-probabilistic binary linear classifier. Reference [Wikipedia](https://en.wikipedia.org/wiki/Support_vector_machine).\n![](https://github.com/icwangjimmy/machine_learning_and_python_in_finance/raw/master/pic/svm.PNG)\n\nNote that the model generates a confidence score which is higher than Logistics Regression model."},{"metadata":{"_uuid":"60039d5377da49f1aa9ac4a924331328bd69add1","_cell_guid":"7a63bf04-a410-9c81-5310-bdef7963298f","trusted":true},"cell_type":"code","source":"# Support Vector Machines\nsvc = SVC()\nsvc.fit(X_train, Y_train)\nY_pred_2 = svc.predict(X_test)\nacc_svc = round(svc.score(X_train, Y_train) * 100, 2)\nacc_svc","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"d7a876c0cfde370e3f10a2bcfbfabdbb2a7dc68f"},"cell_type":"code","source":"# Linear SVC\n#Linear Support Vector Classification.\n#Similar to SVC with parameter kernel=’linear’, but implemented in terms of liblinear rather than libsvm, so it has more flexibility \n#in the choice of penalties and loss functions and should scale better to large numbers of samples.\nlinear_svc = LinearSVC()\nlinear_svc.fit(X_train, Y_train)\nY_pred_5 = linear_svc.predict(X_test)\nacc_linear_svc = round(linear_svc.score(X_train, Y_train) * 100, 2)\nacc_linear_svc","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"bb3ed027c45664148b61e3aa5e2ca8111aac8793","_cell_guid":"172a6286-d495-5ac4-1a9c-5b77b74ca6d2"},"cell_type":"markdown","source":"In pattern recognition, the k-Nearest Neighbors algorithm (or k-NN for short) is a non-parametric method used for classification and regression. A sample is classified by a majority vote of its neighbors, with the sample being assigned to the class most common among its k nearest neighbors (k is a positive integer, typically small). If k = 1, then the object is simply assigned to the class of that single nearest neighbor. Reference [Wikipedia](https://en.wikipedia.org/wiki/K-nearest_neighbors_algorithm).\n![](https://github.com/icwangjimmy/machine_learning_and_python_in_finance/raw/master/pic/knn.png)\n\nKNN confidence score is better than Logistics Regression but worse than SVM."},{"metadata":{"_uuid":"54d86cd45703d459d452f89572771deaa8877999","_cell_guid":"ca14ae53-f05e-eb73-201c-064d7c3ed610","trusted":true},"cell_type":"code","source":"# KNN\nknn = KNeighborsClassifier(n_neighbors = 3)\nknn.fit(X_train, Y_train)\nY_pred_3 = knn.predict(X_test)\nacc_knn = round(knn.score(X_train, Y_train) * 100, 2)\nacc_knn","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"df148bf93e11c9ec2c97162d5c0c0605b75d9334","_cell_guid":"1e286e19-b714-385a-fcfa-8cf5ec19956a"},"cell_type":"markdown","source":"The perceptron is an algorithm for supervised learning of binary classifiers (functions that can decide whether an input, represented by a vector of numbers, belongs to some specific class or not). It is a type of linear classifier, i.e. a classification algorithm that makes its predictions based on a linear predictor function combining a set of weights with the feature vector. The algorithm allows for online learning, in that it processes elements in the training set one at a time. Reference [Wikipedia](https://en.wikipedia.org/wiki/Perceptron)."},{"metadata":{"_uuid":"c19d08949f9c3a26931e28adedc848b4deaa8ab6","_cell_guid":"ccc22a86-b7cb-c2dd-74bd-53b218d6ed0d","trusted":true},"cell_type":"code","source":"# Perceptron\n# Perceptron is a classification algorithm which shares the same underlying implementation with SGDClassifier. \n# In fact, Perceptron() is equivalent to SGDClassifier(loss=”perceptron”, eta0=1, learning_rate=”constant”, penalty=None).\nperceptron = Perceptron()\nperceptron.fit(X_train, Y_train)\nY_pred_4 = perceptron.predict(X_test)\nacc_perceptron = round(perceptron.score(X_train, Y_train) * 100, 2)\nacc_perceptron","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"3a016c1f24da59c85648204302d61ea15920e740","_cell_guid":"dc98ed72-3aeb-861f-804d-b6e3d178bf4b","trusted":true},"cell_type":"code","source":"# Stochastic Gradient Descent\n# Linear classifiers (SVM, logistic regression, a.o.) with SGD training.\nsgd = SGDClassifier()\nsgd.fit(X_train, Y_train)\nY_pred_6 = sgd.predict(X_test)\nacc_sgd = round(sgd.score(X_train, Y_train) * 100, 2)\nacc_sgd","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"1c70e99920ae34adce03aaef38d61e2b83ff6a9c","_cell_guid":"bae7f8d7-9da0-f4fd-bdb1-d97e719a18d7"},"cell_type":"markdown","source":"This model uses a decision tree as a predictive model which maps features (tree branches) to conclusions about the target value (tree leaves). Tree models where the target variable can take a finite set of values are called classification trees; in these tree structures, leaves represent class labels and branches represent conjunctions of features that lead to those class labels. Decision trees where the target variable can take continuous values (typically real numbers) are called regression trees. Reference [Wikipedia](https://en.wikipedia.org/wiki/Decision_tree_learning).\n\nThe model confidence score is the highest among models evaluated so far."},{"metadata":{"_uuid":"1f94308b23b934123c03067e84027b507b989e52","_cell_guid":"dd85f2b7-ace2-0306-b4ec-79c68cd3fea0","trusted":true},"cell_type":"code","source":"# Decision Tree\ndecision_tree = DecisionTreeClassifier()\n# decision_tree = DecisionTreeClassifier()\ndecision_tree.fit(X_train, Y_train)\n# Y_pred_7 = decision_tree.predict(X_test)\nacc_decision_tree = round(decision_tree.score(X_train, Y_train) * 100, 2)\nacc_decision_tree","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"11c88ac07a1f82e3be6193a5ab5b4b04db556eef","scrolled":false},"cell_type":"code","source":"from sklearn import tree\nimport graphviz \ndot_data = tree.export_graphviz(decision_tree, out_file=None, \n                                feature_names = X_train.columns.tolist(), class_names = True,\n                                filled = True, rounded = True)\ngraph = graphviz.Source(dot_data) \ngraph","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"24f4e46f202a858076be91752170cad52aa9aefa","_cell_guid":"85693668-0cd5-4319-7768-eddb62d2b7d0"},"cell_type":"markdown","source":"The next model Random Forests is one of the most popular. Random forests or random decision forests are an ensemble learning method for classification, regression and other tasks, that operate by constructing a multitude of decision trees (n_estimators=100) at training time and outputting the class that is the mode of the classes (classification) or mean prediction (regression) of the individual trees. Reference [Wikipedia](https://en.wikipedia.org/wiki/Random_forest).\n![](https://github.com/icwangjimmy/machine_learning_and_python_in_finance/raw/master/pic/rf.png)\nThe model confidence score is the highest among models evaluated so far. We decide to use this model's output (Y_pred) for creating our competition submission of results."},{"metadata":{"_uuid":"483c647d2759a2703d20785a44f51b6dee47d0db","_cell_guid":"f0694a8e-b618-8ed9-6f0d-8c6fba2c4567","trusted":true},"cell_type":"code","source":"# Random Forest\nrandom_forest = RandomForestClassifier(n_estimators=100, max_depth=7)\nrandom_forest.fit(X_train, Y_train)\nY_pred_8 = random_forest.predict(X_test)\nrandom_forest.score(X_train, Y_train)\nacc_random_forest = round(random_forest.score(X_train, Y_train) * 100, 2)\nacc_random_forest","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"0a53ed4e45b6167c9cd4937e5cc3ecc2473bd575"},"cell_type":"code","source":"random_forest.get_params","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"b33d4b21b481dbcfe93eec678d3c4cbb0088678f"},"cell_type":"markdown","source":"## Get some feeling of artificial neural network!\n[google tensorflow playground](https://playground.tensorflow.org/)"},{"metadata":{"trusted":true,"_uuid":"526b9784136e497c1c424dfa22ddcc90967133fc"},"cell_type":"code","source":"#Artificial neural network\n# mlp = MLPClassifier(hidden_layer_sizes=(100, 100), max_iter=400, alpha=1e-4,\n#                     solver='sgd', verbose=10, tol=1e-4, random_state=1)\nmlp = MLPClassifier(hidden_layer_sizes=(100,100,100), max_iter=400, alpha=1e-4,\n                    solver='sgd', verbose=False, tol=1e-4, random_state=1,\n                    learning_rate_init=.1)\n\nmlp.fit(X_train, Y_train)\nY_pred_9 = mlp.predict(X_test)\nmlp.score(X_train, Y_train)\nacc_ann = round(mlp.score(X_train, Y_train) * 100, 2)\nacc_ann","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"2c1428d022430ea594af983a433757e11b47c50c","_cell_guid":"f6c9eef8-83dd-581c-2d8e-ce932fe3a44d"},"cell_type":"markdown","source":"<a id=\"step6\"></a>\n# 6. Visualize, report, and present the problem solving steps and final solution 汇报结果\n- Model evaluation\n\nWe can now rank our evaluation of all the models to choose the best one for our problem. While both Decision Tree and Random Forest score the same, we choose to use Random Forest as they correct for decision trees' habit of overfitting to their training set. "},{"metadata":{"_uuid":"06a52babe50e0dd837b553c78fc73872168e1c7d","_cell_guid":"1f3cebe0-31af-70b2-1ce4-0fd406bcdfc6","trusted":true},"cell_type":"code","source":"models = pd.DataFrame({\n    'Model': ['Support Vector Machines', 'KNN', 'Logistic Regression', \n              'Random Forest', 'Perceptron', 'ANN',\n              'Stochastic Gradient Decent', 'Linear SVC', \n              'Decision Tree'],\n    'Score': [acc_svc, acc_knn, acc_log, \n              acc_random_forest, acc_perceptron, acc_ann,\n              acc_sgd, acc_linear_svc, acc_decision_tree]})\nmodels.sort_values(by='Score', ascending=False)","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"265be95b04e7e02d49209a711f3dfcccfce25d73"},"cell_type":"markdown","source":"## Model Tuning （模型参数调整）"},{"metadata":{"trusted":true,"_uuid":"b576f149a6cbd07bf1446520d65f7e69b5375cc2"},"cell_type":"code","source":"from sklearn.metrics import make_scorer, accuracy_score\nfrom sklearn.model_selection import GridSearchCV","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"9c775ad39636c0f6a6b73230ed656db11f8b386d"},"cell_type":"code","source":"%%time\n# Choose random forest to do model tuning. \nrfc = RandomForestClassifier()\n\n# Choose some parameter combinations to try\nparameters = {'n_estimators': [100, 200, 400], \n              'max_features': ['log2', 'sqrt','auto'], \n              'criterion': ['entropy', 'gini'],\n              'max_depth': [3, 5, 10, 20], \n              'min_samples_split': [2, 3, 5],\n              'min_samples_leaf': [1,5,8]\n             }\n\n# Type of scoring used to compare parameter combinations\nacc_scorer = make_scorer(accuracy_score)\n\n# Run the grid search\ngrid_obj = GridSearchCV(rfc, parameters, scoring=acc_scorer)\ngrid_obj = grid_obj.fit(X_train, Y_train)\n\n# Set the clf to the best combination of parameters\nrfc = grid_obj.best_estimator_\n\n# Fit the best algorithm to the data. \nrfc.fit(X_train, Y_train)","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"a505647c6710af54ad1af84d0ddfbeb48ee2507b"},"cell_type":"code","source":"grid_obj.best_estimator_","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"43905ffc2697f9666eae37cdce5d2c5703860eaf"},"cell_type":"code","source":"Y_pred_10 = rfc.predict(X_test)\nrfc.score(X_train, Y_train)\nacc_rfc = round(rfc.score(X_train, Y_train) * 100, 2)\nacc_rfc","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"c9e4193b0c1d0c481f0cc051e2b63374dfd16cc4"},"cell_type":"markdown","source":"<a id=\"step7\"></a>\n# 7. Supply or submit the results and optimize 提交结果"},{"metadata":{"_uuid":"82b31ea933b3026bd038a8370d651efdcdb3e4d7","_cell_guid":"28854d36-051f-3ef0-5535-fa5ba6a9bef7","trusted":true},"cell_type":"code","source":"submission = pd.DataFrame({\n        \"PassengerId\": test_df[\"PassengerId\"],\n        \"Survived\": Y_pred_10\n    })\nsubmission.to_csv('submission.csv', index=False)","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"cdae56d6adbfb15ff9c491c645ae46e2c91d75ce","_cell_guid":"aeec9210-f9d8-cd7c-c4cf-a87376d5f693"},"cell_type":"markdown","source":"## Notebook References\n\nThis notebook has been created based on great work done solving the Titanic competition and other sources.\n\n- [A journey through Titanic](https://www.kaggle.com/omarelgabry/titanic/a-journey-through-titanic)\n- [Getting Started with Pandas: Kaggle's Titanic Competition](https://www.kaggle.com/c/titanic/details/getting-started-with-random-forests)\n- [Titanic Best Working Classifier](https://www.kaggle.com/sinakhorami/titanic/titanic-best-working-classifier)\n- [An Interactive Data Science Tutorial](https://www.kaggle.com/helgejo/an-interactive-data-science-tutorial)\n- [A Data Science Framework: To Achieve 99% Accuracy](https://www.kaggle.com/ldfreeman3/a-data-science-framework-to-achieve-99-accuracy)\n- [A Comprehensive ML Workflow with Python](https://www.kaggle.com/mjbahmani/a-comprehensive-ml-workflow-with-python)"},{"metadata":{"trusted":true,"_uuid":"1cbf790a374b095ad4f5dc7c3cfbf8e18dd0bc5a"},"cell_type":"code","source":"","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"6f80031b7b312880b781926047ea62ea7b6a25d0"},"cell_type":"code","source":"","execution_count":null,"outputs":[]}],"metadata":{"_is_fork":false,"language_info":{"name":"python","version":"3.6.6","mimetype":"text/x-python","codemirror_mode":{"name":"ipython","version":3},"pygments_lexer":"ipython3","nbconvert_exporter":"python","file_extension":".py"},"_change_revision":0,"kernelspec":{"display_name":"Python 3","language":"python","name":"python3"}},"nbformat":4,"nbformat_minor":1}