{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"pygments_lexer":"ipython3","nbconvert_exporter":"python","version":"3.6.4","file_extension":".py","codemirror_mode":{"name":"ipython","version":3},"name":"python","mimetype":"text/x-python"}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"# Titanic Project Example Walk Through \nIn this notebook, I hope to show how a data scientist would go about working through a problem. The goal is to correctly predict if someone survived the Titanic shipwreck. I thought it would be fun to see how well I could do in this competition without deep learning. \n\n*The accompanying video is located here:* https://www.youtube.com/watch?v=I3FBJdiExcg\n\n**Best results : 79.425 % accuracy (Top 12%)**\n\n## Overview \n### 1) Understand the shape of the data (Histograms, box plots, etc.)\n\n### 2) Data Cleaning \n\n### 3) Data Exploration\n\n### 4) Feature Engineering \n\n### 5) Data Preprocessing for Model\n\n### 6) Basic Model Building \n\n### 7) Model Tuning \n\n### 8) Ensemble Modle Building \n\n### 9) Results ","metadata":{}},{"cell_type":"code","source":"# This Python 3 environment comes with many helpful analytics libraries installed\n# It is defined by the kaggle/python Docker image: https://github.com/kaggle/docker-python\n# For example, here's several helpful packages to load\n\nimport numpy as np # linear algebra\nimport pandas as pd # data processing, CSV file I/O (e.g. pd.read_csv)\nimport seaborn as sns \nimport matplotlib.pyplot as plt\n# Input data files are available in the read-only \"../input/\" directory\n# For example, running this (by clicking run or pressing Shift+Enter) will list all files under the input directory\n\nimport os\nfor dirname, _, filenames in os.walk('/kaggle/input'):\n    for filename in filenames:\n        print(os.path.join(dirname, filename))\n        \n        \n# You can write up to 5GB to the current directory (/kaggle/working/) that gets preserved as output when you create a version using \"Save & Run All\" \n# You can also write temporary files to /kaggle/temp/, but they won't be saved outside of the current session","metadata":{"_uuid":"8f2839f25d086af736a60e9eeb907d3b93b6e0e5","_cell_guid":"b1076dfc-b9ad-4769-8c92-a6c4dae69d19","execution":{"iopub.status.busy":"2022-07-23T10:57:30.317125Z","iopub.execute_input":"2022-07-23T10:57:30.317566Z","iopub.status.idle":"2022-07-23T10:57:31.137987Z","shell.execute_reply.started":"2022-07-23T10:57:30.317528Z","shell.execute_reply":"2022-07-23T10:57:31.136871Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Here we import the data. For this analysis, we will be exclusively working with the Training set. We will be validating based on data from the training set as well. For our final submissions, we will make predictions based on the test set. ","metadata":{}},{"cell_type":"code","source":"training = pd.read_csv('/kaggle/input/titanic/train.csv')\ntest = pd.read_csv('/kaggle/input/titanic/test.csv')\n\ntraining['train_test'] = 1\ntest['train_test'] = 0\ntest['Survived'] = np.NaN\nall_data = pd.concat([training,test])\n\n%matplotlib inline\nall_data.columns","metadata":{"execution":{"iopub.status.busy":"2022-07-23T10:57:31.140851Z","iopub.execute_input":"2022-07-23T10:57:31.141317Z","iopub.status.idle":"2022-07-23T10:57:31.198821Z","shell.execute_reply.started":"2022-07-23T10:57:31.141269Z","shell.execute_reply":"2022-07-23T10:57:31.197895Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## Project Planning\nWhen starting any project, I like to outline the steps that I plan to take. Below is the rough outline that I created for this project using commented cells. ","metadata":{}},{"cell_type":"code","source":"# Understand nature of the data .info() .describe()\n# Histograms and boxplots \n# Value counts \n# Missing data \n# Correlation between the metrics \n# Explore interesting themes \n    # Wealthy survive? \n    # By location \n    # Age scatterplot with ticket price \n    # Young and wealthy Variable? \n    # Total spent? \n# Feature engineering \n# preprocess data together or use a transformer? \n    # use label for train and test   \n# Scaling?\n\n# Model Baseline \n# Model comparison with CV ","metadata":{"execution":{"iopub.status.busy":"2022-07-23T10:57:31.203434Z","iopub.execute_input":"2022-07-23T10:57:31.203741Z","iopub.status.idle":"2022-07-23T10:57:31.208214Z","shell.execute_reply.started":"2022-07-23T10:57:31.203711Z","shell.execute_reply":"2022-07-23T10:57:31.207206Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## Light Data Exploration\n### 1) For numeric data \n* Made histograms to understand distributions \n* Corrplot \n* Pivot table comparing survival rate across numeric variables \n\n\n### 2) For Categorical Data \n* Made bar charts to understand balance of classes \n* Made pivot tables to understand relationship with survival ","metadata":{}},{"cell_type":"code","source":"#quick look at our data types & null counts \ntraining.info()","metadata":{"execution":{"iopub.status.busy":"2022-07-23T10:57:31.211657Z","iopub.execute_input":"2022-07-23T10:57:31.212193Z","iopub.status.idle":"2022-07-23T10:57:31.229850Z","shell.execute_reply.started":"2022-07-23T10:57:31.212141Z","shell.execute_reply":"2022-07-23T10:57:31.229136Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# to better understand the numeric data, we want to use the .describe() method. This gives us an understanding of the central tendencies of the data \ntraining.describe()","metadata":{"execution":{"iopub.status.busy":"2022-07-23T10:57:31.232304Z","iopub.execute_input":"2022-07-23T10:57:31.232824Z","iopub.status.idle":"2022-07-23T10:57:31.274740Z","shell.execute_reply.started":"2022-07-23T10:57:31.232793Z","shell.execute_reply":"2022-07-23T10:57:31.273961Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"#quick way to separate numeric columns\ntraining.describe().columns","metadata":{"execution":{"iopub.status.busy":"2022-07-23T10:57:31.275962Z","iopub.execute_input":"2022-07-23T10:57:31.276448Z","iopub.status.idle":"2022-07-23T10:57:31.304369Z","shell.execute_reply.started":"2022-07-23T10:57:31.276416Z","shell.execute_reply":"2022-07-23T10:57:31.303501Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# look at numeric and categorical values separately \ndf_num = training[['Age','SibSp','Parch','Fare']]\ndf_cat = training[['Survived','Pclass','Sex','Ticket','Cabin','Embarked']]","metadata":{"execution":{"iopub.status.busy":"2022-07-23T10:57:31.305761Z","iopub.execute_input":"2022-07-23T10:57:31.306326Z","iopub.status.idle":"2022-07-23T10:57:31.314186Z","shell.execute_reply.started":"2022-07-23T10:57:31.306293Z","shell.execute_reply":"2022-07-23T10:57:31.313126Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"#distributions for all numeric variables \nfor i in df_num.columns:\n    plt.hist(df_num[i])\n    plt.title(i)\n    plt.show()","metadata":{"execution":{"iopub.status.busy":"2022-07-23T10:57:31.316328Z","iopub.execute_input":"2022-07-23T10:57:31.317152Z","iopub.status.idle":"2022-07-23T10:57:32.038359Z","shell.execute_reply.started":"2022-07-23T10:57:31.317100Z","shell.execute_reply":"2022-07-23T10:57:32.037262Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Perhaps we should take the non-normal distributions and consider normalizing them?","metadata":{"trusted":true}},{"cell_type":"markdown","source":"\n","metadata":{}},{"cell_type":"code","source":"print(df_num.corr())\nsns.heatmap(df_num.corr())","metadata":{"execution":{"iopub.status.busy":"2022-07-23T10:57:32.039959Z","iopub.execute_input":"2022-07-23T10:57:32.040244Z","iopub.status.idle":"2022-07-23T10:57:32.269797Z","shell.execute_reply.started":"2022-07-23T10:57:32.040216Z","shell.execute_reply":"2022-07-23T10:57:32.268644Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# compare survival rate across Age, SibSp, Parch, and Fare \npd.pivot_table(training, index = 'Survived', values = ['Age','SibSp','Parch','Fare'])","metadata":{"execution":{"iopub.status.busy":"2022-07-23T10:57:32.271816Z","iopub.execute_input":"2022-07-23T10:57:32.272286Z","iopub.status.idle":"2022-07-23T10:57:32.301766Z","shell.execute_reply.started":"2022-07-23T10:57:32.272241Z","shell.execute_reply":"2022-07-23T10:57:32.300765Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"for i in df_cat.columns:\n    sns.barplot(df_cat[i].value_counts().index,df_cat[i].value_counts()).set_title(i)\n    plt.show()\\\n    ","metadata":{"execution":{"iopub.status.busy":"2022-07-23T10:57:32.303849Z","iopub.execute_input":"2022-07-23T10:57:32.304388Z","iopub.status.idle":"2022-07-23T10:57:42.643343Z","shell.execute_reply.started":"2022-07-23T10:57:32.304329Z","shell.execute_reply":"2022-07-23T10:57:42.642302Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Cabin and ticket graphs are very messy. This is an area where we may want to do some feature engineering! ","metadata":{}},{"cell_type":"code","source":"# Comparing survival and each of these categorical variables \nprint(pd.pivot_table(training, index = 'Survived', columns = 'Pclass', values = 'Ticket' ,aggfunc ='count'))\nprint()\nprint(pd.pivot_table(training, index = 'Survived', columns = 'Sex', values = 'Ticket' ,aggfunc ='count'))\nprint()\nprint(pd.pivot_table(training, index = 'Survived', columns = 'Embarked', values = 'Ticket' ,aggfunc ='count'))","metadata":{"execution":{"iopub.status.busy":"2022-07-23T10:57:42.644984Z","iopub.execute_input":"2022-07-23T10:57:42.645584Z","iopub.status.idle":"2022-07-23T10:57:42.702909Z","shell.execute_reply.started":"2022-07-23T10:57:42.645538Z","shell.execute_reply":"2022-07-23T10:57:42.701931Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## Feature Engineering \n### 1) Cabin - Simplify cabins (evaluated if cabin letter (cabin_adv) or the purchase of tickets across multiple cabins (cabin_multiple) impacted survival)\n\n### 2) Tickets - Do different ticket types impact survival rates?\n\n### 3) Does a person's title relate to survival rates? ","metadata":{}},{"cell_type":"code","source":"df_cat.Cabin\ntraining['cabin_multiple'] = training.Cabin.apply(lambda x: 0 if pd.isna(x) else len(x.split(' ')))\n# after looking at this, we may want to look at cabin by letter or by number. Let's create some categories for this \n# letters \n# multiple letters \ntraining['cabin_multiple'].value_counts()","metadata":{"execution":{"iopub.status.busy":"2022-07-23T10:57:42.704480Z","iopub.execute_input":"2022-07-23T10:57:42.704973Z","iopub.status.idle":"2022-07-23T10:57:42.718840Z","shell.execute_reply.started":"2022-07-23T10:57:42.704924Z","shell.execute_reply":"2022-07-23T10:57:42.717966Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"pd.pivot_table(training, index = 'Survived', columns = 'cabin_multiple', values = 'Ticket' ,aggfunc ='count')","metadata":{"execution":{"iopub.status.busy":"2022-07-23T10:57:42.720476Z","iopub.execute_input":"2022-07-23T10:57:42.720817Z","iopub.status.idle":"2022-07-23T10:57:42.751678Z","shell.execute_reply.started":"2022-07-23T10:57:42.720780Z","shell.execute_reply":"2022-07-23T10:57:42.750447Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"#creates categories based on the cabin letter (n stands for null)\n#in this case we will treat null values like it's own category\n\ntraining['cabin_adv'] = training.Cabin.apply(lambda x: str(x)[0])\n","metadata":{"execution":{"iopub.status.busy":"2022-07-23T10:57:42.753302Z","iopub.execute_input":"2022-07-23T10:57:42.753740Z","iopub.status.idle":"2022-07-23T10:57:42.761801Z","shell.execute_reply.started":"2022-07-23T10:57:42.753696Z","shell.execute_reply":"2022-07-23T10:57:42.760607Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"#comparing surivial rate by cabin\nprint(training.cabin_adv.value_counts())\npd.pivot_table(training,index='Survived',columns='cabin_adv', values = 'Name', aggfunc='count')","metadata":{"execution":{"iopub.status.busy":"2022-07-23T10:57:42.763302Z","iopub.execute_input":"2022-07-23T10:57:42.763723Z","iopub.status.idle":"2022-07-23T10:57:42.799748Z","shell.execute_reply.started":"2022-07-23T10:57:42.763680Z","shell.execute_reply":"2022-07-23T10:57:42.798890Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"#understand ticket values better \n#numeric vs non numeric \ntraining['numeric_ticket'] = training.Ticket.apply(lambda x: 1 if x.isnumeric() else 0)\ntraining['ticket_letters'] = training.Ticket.apply(lambda x: ''.join(x.split(' ')[:-1]).replace('.','').replace('/','').lower() if len(x.split(' ')[:-1]) >0 else 0)\n","metadata":{"execution":{"iopub.status.busy":"2022-07-23T10:57:42.801326Z","iopub.execute_input":"2022-07-23T10:57:42.801624Z","iopub.status.idle":"2022-07-23T10:57:42.814920Z","shell.execute_reply.started":"2022-07-23T10:57:42.801595Z","shell.execute_reply":"2022-07-23T10:57:42.813109Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"training['numeric_ticket'].value_counts()","metadata":{"execution":{"iopub.status.busy":"2022-07-23T10:57:42.816320Z","iopub.execute_input":"2022-07-23T10:57:42.816645Z","iopub.status.idle":"2022-07-23T10:57:42.827147Z","shell.execute_reply.started":"2022-07-23T10:57:42.816616Z","shell.execute_reply":"2022-07-23T10:57:42.826218Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"#lets us view all rows in dataframe through scrolling. This is for convenience \npd.set_option(\"max_rows\", None)\ntraining['ticket_letters'].value_counts()\n","metadata":{"execution":{"iopub.status.busy":"2022-07-23T10:57:42.828485Z","iopub.execute_input":"2022-07-23T10:57:42.828950Z","iopub.status.idle":"2022-07-23T10:57:42.839955Z","shell.execute_reply.started":"2022-07-23T10:57:42.828918Z","shell.execute_reply":"2022-07-23T10:57:42.838651Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"#difference in numeric vs non-numeric tickets in survival rate \npd.pivot_table(training,index='Survived',columns='numeric_ticket', values = 'Ticket', aggfunc='count')","metadata":{"execution":{"iopub.status.busy":"2022-07-23T10:57:42.841660Z","iopub.execute_input":"2022-07-23T10:57:42.842013Z","iopub.status.idle":"2022-07-23T10:57:42.865368Z","shell.execute_reply.started":"2022-07-23T10:57:42.841982Z","shell.execute_reply":"2022-07-23T10:57:42.864439Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"#survival rate across different tyicket types \npd.pivot_table(training,index='Survived',columns='ticket_letters', values = 'Ticket', aggfunc='count')","metadata":{"execution":{"iopub.status.busy":"2022-07-23T10:57:42.866607Z","iopub.execute_input":"2022-07-23T10:57:42.867019Z","iopub.status.idle":"2022-07-23T10:57:42.907863Z","shell.execute_reply.started":"2022-07-23T10:57:42.866983Z","shell.execute_reply":"2022-07-23T10:57:42.906983Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"#feature engineering on person's title \ntraining.Name.head(50)\ntraining['name_title'] = training.Name.apply(lambda x: x.split(',')[1].split('.')[0].strip())\n#mr., ms., master. etc","metadata":{"execution":{"iopub.status.busy":"2022-07-23T10:57:42.908964Z","iopub.execute_input":"2022-07-23T10:57:42.909403Z","iopub.status.idle":"2022-07-23T10:57:42.917926Z","shell.execute_reply.started":"2022-07-23T10:57:42.909373Z","shell.execute_reply":"2022-07-23T10:57:42.916828Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"training['name_title'].value_counts()","metadata":{"execution":{"iopub.status.busy":"2022-07-23T10:57:42.919676Z","iopub.execute_input":"2022-07-23T10:57:42.920265Z","iopub.status.idle":"2022-07-23T10:57:42.931460Z","shell.execute_reply.started":"2022-07-23T10:57:42.920228Z","shell.execute_reply":"2022-07-23T10:57:42.930327Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## Data Preprocessing for Model \n### 1) Drop null values from Embarked (only 2) \n\n### 2) Include only relevant variables (Since we have limited data, I wanted to exclude things like name and passanger ID so that we could have a reasonable number of features for our models to deal with) \nVariables:  'Pclass', 'Sex','Age', 'SibSp', 'Parch', 'Fare', 'Embarked', 'cabin_adv', 'cabin_multiple', 'numeric_ticket', 'name_title'\n\n### 3) Do categorical transforms on all data. Usually we would use a transformer, but with this approach we can ensure that our traning and test data have the same colums. We also may be able to infer something about the shape of the test data through this method. I will stress, this is generally not recommend outside of a competition (use onehot encoder). \n\n### 4) Impute data with mean for fare and age (Should also experiment with median) \n\n### 5) Normalized fare using logarithm to give more semblance of a normal distribution \n\n### 6) Scaled data 0-1 with standard scaler \n","metadata":{}},{"cell_type":"code","source":"#create all categorical variables that we did above for both training and test sets \nall_data['cabin_multiple'] = all_data.Cabin.apply(lambda x: 0 if pd.isna(x) else len(x.split(' ')))\nall_data['cabin_adv'] = all_data.Cabin.apply(lambda x: str(x)[0])\nall_data['numeric_ticket'] = all_data.Ticket.apply(lambda x: 1 if x.isnumeric() else 0)\nall_data['ticket_letters'] = all_data.Ticket.apply(lambda x: ''.join(x.split(' ')[:-1]).replace('.','').replace('/','').lower() if len(x.split(' ')[:-1]) >0 else 0)\nall_data['name_title'] = all_data.Name.apply(lambda x: x.split(',')[1].split('.')[0].strip())\n\n#impute nulls for continuous data \n#all_data.Age = all_data.Age.fillna(training.Age.mean())\nall_data.Age = all_data.Age.fillna(training.Age.median())\n#all_data.Fare = all_data.Fare.fillna(training.Fare.mean())\nall_data.Fare = all_data.Fare.fillna(training.Fare.median())\n\n#drop null 'embarked' rows. Only 2 instances of this in training and 0 in test \nall_data.dropna(subset=['Embarked'],inplace = True)\n\n#tried log norm of sibsp (not used)\nall_data['norm_sibsp'] = np.log(all_data.SibSp+1)\nall_data['norm_sibsp'].hist()\n\n# log norm of fare (used)\nall_data['norm_fare'] = np.log(all_data.Fare+1)\nall_data['norm_fare'].hist()\n\n# converted fare to category for pd.get_dummies()\nall_data.Pclass = all_data.Pclass.astype(str)\n\n#created dummy variables from categories (also can use OneHotEncoder)\nall_dummies = pd.get_dummies(all_data[['Pclass','Sex','Age','SibSp','Parch','norm_fare','Embarked','cabin_adv','cabin_multiple','numeric_ticket','name_title','train_test']])\n\n#Split to train test again\nX_train = all_dummies[all_dummies.train_test == 1].drop(['train_test'], axis =1)\nX_test = all_dummies[all_dummies.train_test == 0].drop(['train_test'], axis =1)\n\n\ny_train = all_data[all_data.train_test==1].Survived\ny_train.shape","metadata":{"execution":{"iopub.status.busy":"2022-07-23T10:57:42.933190Z","iopub.execute_input":"2022-07-23T10:57:42.933802Z","iopub.status.idle":"2022-07-23T10:57:43.188916Z","shell.execute_reply.started":"2022-07-23T10:57:42.933768Z","shell.execute_reply":"2022-07-23T10:57:43.187924Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Scale data \nfrom sklearn.preprocessing import StandardScaler\nscale = StandardScaler()\nall_dummies_scaled = all_dummies.copy()\nall_dummies_scaled[['Age','SibSp','Parch','norm_fare']]= scale.fit_transform(all_dummies_scaled[['Age','SibSp','Parch','norm_fare']])\nall_dummies_scaled\n\nX_train_scaled = all_dummies_scaled[all_dummies_scaled.train_test == 1].drop(['train_test'], axis =1)\nX_test_scaled = all_dummies_scaled[all_dummies_scaled.train_test == 0].drop(['train_test'], axis =1)\n\ny_train = all_data[all_data.train_test==1].Survived\n","metadata":{"execution":{"iopub.status.busy":"2022-07-23T10:57:43.190409Z","iopub.execute_input":"2022-07-23T10:57:43.190740Z","iopub.status.idle":"2022-07-23T10:57:43.272174Z","shell.execute_reply.started":"2022-07-23T10:57:43.190708Z","shell.execute_reply":"2022-07-23T10:57:43.271097Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## Model Building (Baseline Validation Performance)\nBefore going further, I like to see how various different models perform with default parameters. I tried the following models using 5 fold cross validation to get a baseline. With a validation set basline, we can see how much tuning improves each of the models. Just because a model has a high basline on this validation set doesn't mean that it will actually do better on the eventual test set. \n\n- Naive Bayes (72.6%)\n- Logistic Regression (82.1%)\n- Decision Tree (77.6%)\n- K Nearest Neighbor (80.5%)\n- Random Forest (80.6%)\n- **Support Vector Classifier (83.2%)**\n- Xtreme Gradient Boosting (81.8%)\n- Soft Voting Classifier - All Models (82.8%)","metadata":{}},{"cell_type":"code","source":"from sklearn.model_selection import cross_val_score\nfrom sklearn.naive_bayes import GaussianNB\nfrom sklearn.linear_model import LogisticRegression\nfrom sklearn import tree\nfrom sklearn.neighbors import KNeighborsClassifier\nfrom sklearn.ensemble import RandomForestClassifier\nfrom sklearn.svm import SVC","metadata":{"execution":{"iopub.status.busy":"2022-07-23T10:57:43.273699Z","iopub.execute_input":"2022-07-23T10:57:43.274145Z","iopub.status.idle":"2022-07-23T10:57:43.514029Z","shell.execute_reply.started":"2022-07-23T10:57:43.274100Z","shell.execute_reply":"2022-07-23T10:57:43.513054Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"#I usually use Naive Bayes as a baseline for my classification tasks \ngnb = GaussianNB()\ncv = cross_val_score(gnb,X_train_scaled,y_train,cv=5)\nprint(cv)\nprint(cv.mean())","metadata":{"execution":{"iopub.status.busy":"2022-07-23T10:57:43.515466Z","iopub.execute_input":"2022-07-23T10:57:43.515799Z","iopub.status.idle":"2022-07-23T10:57:43.566714Z","shell.execute_reply.started":"2022-07-23T10:57:43.515766Z","shell.execute_reply":"2022-07-23T10:57:43.565499Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"lr = LogisticRegression(max_iter = 2000)\ncv = cross_val_score(lr,X_train,y_train,cv=5)\nprint(cv)\nprint(cv.mean())","metadata":{"execution":{"iopub.status.busy":"2022-07-23T10:57:43.568762Z","iopub.execute_input":"2022-07-23T10:57:43.569331Z","iopub.status.idle":"2022-07-23T10:57:44.174383Z","shell.execute_reply.started":"2022-07-23T10:57:43.569282Z","shell.execute_reply":"2022-07-23T10:57:44.173214Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"lr = LogisticRegression(max_iter = 2000)\ncv = cross_val_score(lr,X_train_scaled,y_train,cv=5)\nprint(cv)\nprint(cv.mean())","metadata":{"execution":{"iopub.status.busy":"2022-07-23T10:57:44.176185Z","iopub.execute_input":"2022-07-23T10:57:44.176693Z","iopub.status.idle":"2022-07-23T10:57:44.331601Z","shell.execute_reply.started":"2022-07-23T10:57:44.176647Z","shell.execute_reply":"2022-07-23T10:57:44.330509Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"dt = tree.DecisionTreeClassifier(random_state = 1)\ncv = cross_val_score(dt,X_train,y_train,cv=5)\nprint(cv)\nprint(cv.mean())","metadata":{"execution":{"iopub.status.busy":"2022-07-23T10:57:44.333119Z","iopub.execute_input":"2022-07-23T10:57:44.333730Z","iopub.status.idle":"2022-07-23T10:57:44.396240Z","shell.execute_reply.started":"2022-07-23T10:57:44.333682Z","shell.execute_reply":"2022-07-23T10:57:44.395450Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"dt = tree.DecisionTreeClassifier(random_state = 1)\ncv = cross_val_score(dt,X_train_scaled,y_train,cv=5)\nprint(cv)\nprint(cv.mean())","metadata":{"execution":{"iopub.status.busy":"2022-07-23T10:57:44.401623Z","iopub.execute_input":"2022-07-23T10:57:44.402236Z","iopub.status.idle":"2022-07-23T10:57:44.458731Z","shell.execute_reply.started":"2022-07-23T10:57:44.402197Z","shell.execute_reply":"2022-07-23T10:57:44.457960Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"knn = KNeighborsClassifier()\ncv = cross_val_score(knn,X_train,y_train,cv=5)\nprint(cv)\nprint(cv.mean())","metadata":{"execution":{"iopub.status.busy":"2022-07-23T10:57:44.460142Z","iopub.execute_input":"2022-07-23T10:57:44.460656Z","iopub.status.idle":"2022-07-23T10:57:44.567157Z","shell.execute_reply.started":"2022-07-23T10:57:44.460623Z","shell.execute_reply":"2022-07-23T10:57:44.566241Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"knn = KNeighborsClassifier()\ncv = cross_val_score(knn,X_train_scaled,y_train,cv=5)\nprint(cv)\nprint(cv.mean())","metadata":{"execution":{"iopub.status.busy":"2022-07-23T10:57:44.568540Z","iopub.execute_input":"2022-07-23T10:57:44.568845Z","iopub.status.idle":"2022-07-23T10:57:44.689581Z","shell.execute_reply.started":"2022-07-23T10:57:44.568815Z","shell.execute_reply":"2022-07-23T10:57:44.688577Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"rf = RandomForestClassifier(random_state = 1)\ncv = cross_val_score(rf,X_train,y_train,cv=5)\nprint(cv)\nprint(cv.mean())","metadata":{"execution":{"iopub.status.busy":"2022-07-23T10:57:44.691035Z","iopub.execute_input":"2022-07-23T10:57:44.691359Z","iopub.status.idle":"2022-07-23T10:57:45.990683Z","shell.execute_reply.started":"2022-07-23T10:57:44.691328Z","shell.execute_reply":"2022-07-23T10:57:45.989838Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"rf = RandomForestClassifier(random_state = 1)\ncv = cross_val_score(rf,X_train_scaled,y_train,cv=5)\nprint(cv)\nprint(cv.mean())","metadata":{"execution":{"iopub.status.busy":"2022-07-23T10:57:45.992182Z","iopub.execute_input":"2022-07-23T10:57:45.992477Z","iopub.status.idle":"2022-07-23T10:57:47.293411Z","shell.execute_reply.started":"2022-07-23T10:57:45.992448Z","shell.execute_reply":"2022-07-23T10:57:47.292465Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"svc = SVC(probability = True)\ncv = cross_val_score(svc,X_train_scaled,y_train,cv=5)\nprint(cv)\nprint(cv.mean())","metadata":{"execution":{"iopub.status.busy":"2022-07-23T10:57:47.294752Z","iopub.execute_input":"2022-07-23T10:57:47.295101Z","iopub.status.idle":"2022-07-23T10:57:48.083574Z","shell.execute_reply.started":"2022-07-23T10:57:47.295069Z","shell.execute_reply":"2022-07-23T10:57:48.082696Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"from xgboost import XGBClassifier\nxgb = XGBClassifier(random_state =1)\ncv = cross_val_score(xgb,X_train_scaled,y_train,cv=5)\nprint(cv)\nprint(cv.mean())","metadata":{"execution":{"iopub.status.busy":"2022-07-23T10:57:48.084885Z","iopub.execute_input":"2022-07-23T10:57:48.085168Z","iopub.status.idle":"2022-07-23T10:57:49.565975Z","shell.execute_reply.started":"2022-07-23T10:57:48.085141Z","shell.execute_reply":"2022-07-23T10:57:49.565011Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"#Voting classifier takes all of the inputs and averages the results. For a \"hard\" voting classifier each classifier gets 1 vote \"yes\" or \"no\" and the result is just a popular vote. For this, you generally want odd numbers\n#A \"soft\" classifier averages the confidence of each of the models. If a the average confidence is > 50% that it is a 1 it will be counted as such\nfrom sklearn.ensemble import VotingClassifier\nvoting_clf = VotingClassifier(estimators = [('lr',lr),('knn',knn),('rf',rf),('gnb',gnb),('svc',svc),('xgb',xgb)], voting = 'soft') ","metadata":{"execution":{"iopub.status.busy":"2022-07-23T10:57:49.572410Z","iopub.execute_input":"2022-07-23T10:57:49.575841Z","iopub.status.idle":"2022-07-23T10:57:49.600825Z","shell.execute_reply.started":"2022-07-23T10:57:49.575789Z","shell.execute_reply":"2022-07-23T10:57:49.599641Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"cv = cross_val_score(voting_clf,X_train_scaled,y_train,cv=5)\nprint(cv)\nprint(cv.mean())","metadata":{"execution":{"iopub.status.busy":"2022-07-23T10:57:49.608076Z","iopub.execute_input":"2022-07-23T10:57:49.611697Z","iopub.status.idle":"2022-07-23T10:57:53.049475Z","shell.execute_reply.started":"2022-07-23T10:57:49.611641Z","shell.execute_reply":"2022-07-23T10:57:53.048405Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"voting_clf.fit(X_train_scaled,y_train)\ny_hat_base_vc = voting_clf.predict(X_test_scaled).astype(int)\nbasic_submission = {'PassengerId': test.PassengerId, 'Survived': y_hat_base_vc}\nbase_submission = pd.DataFrame(data=basic_submission)\nbase_submission.to_csv('base_submission.csv', index=False)","metadata":{"execution":{"iopub.status.busy":"2022-07-23T10:57:53.050937Z","iopub.execute_input":"2022-07-23T10:57:53.051548Z","iopub.status.idle":"2022-07-23T10:57:53.813537Z","shell.execute_reply.started":"2022-07-23T10:57:53.051502Z","shell.execute_reply":"2022-07-23T10:57:53.812562Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## Model Tuned Performance \nAfter getting the baselines, let's see if we can improve on the indivdual model results!I mainly used grid search to tune the models. I also used Randomized Search for the Random Forest and XG boosted model to simplify testing time. \n\n|Model|Baseline|Tuned Performance|\n|-----|--------|-----------------|\n|Naive Bayes| 72.6%| NA|\n|Logistic Regression| 82.1%| 82.6%|\n|Decision Tree| 77.6%| NA|\n|K Nearest Neighbor| 80.5%|83.0%|\n|Random Forest| 80.6%| 83.6|\n|Support Vector Classifier| 83.2%| 83.2%|\n|Xtreme Gradient Boosting| 81.8%| 85.3%|\n","metadata":{}},{"cell_type":"code","source":"from sklearn.model_selection import GridSearchCV \nfrom sklearn.model_selection import RandomizedSearchCV ","metadata":{"execution":{"iopub.status.busy":"2022-07-23T10:57:53.819036Z","iopub.execute_input":"2022-07-23T10:57:53.822083Z","iopub.status.idle":"2022-07-23T10:57:53.829395Z","shell.execute_reply.started":"2022-07-23T10:57:53.822034Z","shell.execute_reply":"2022-07-23T10:57:53.828348Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"#simple performance reporting function\ndef clf_performance(classifier, model_name):\n    print(model_name)\n    print('Best Score: ' + str(classifier.best_score_))\n    print('Best Parameters: ' + str(classifier.best_params_))","metadata":{"execution":{"iopub.status.busy":"2022-07-23T10:57:53.834725Z","iopub.execute_input":"2022-07-23T10:57:53.838354Z","iopub.status.idle":"2022-07-23T10:57:53.847042Z","shell.execute_reply.started":"2022-07-23T10:57:53.838300Z","shell.execute_reply":"2022-07-23T10:57:53.846008Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"lr = LogisticRegression()\nparam_grid = {'max_iter' : [2000],\n              'penalty' : ['l1', 'l2'],\n              'C' : np.logspace(-4, 4, 20),\n              'solver' : ['liblinear']}\n\nclf_lr = GridSearchCV(lr, param_grid = param_grid, cv = 5, verbose = True, n_jobs = -1)\nbest_clf_lr = clf_lr.fit(X_train_scaled,y_train)\nclf_performance(best_clf_lr,'Logistic Regression')","metadata":{"execution":{"iopub.status.busy":"2022-07-23T10:57:53.851459Z","iopub.execute_input":"2022-07-23T10:57:53.854320Z","iopub.status.idle":"2022-07-23T10:57:57.492078Z","shell.execute_reply.started":"2022-07-23T10:57:53.854271Z","shell.execute_reply":"2022-07-23T10:57:57.491123Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"knn = KNeighborsClassifier()\nparam_grid = {'n_neighbors' : [3,5,7,9],\n              'weights' : ['uniform', 'distance'],\n              'algorithm' : ['auto', 'ball_tree','kd_tree'],\n              'p' : [1,2]}\nclf_knn = GridSearchCV(knn, param_grid = param_grid, cv = 5, verbose = True, n_jobs = -1)\nbest_clf_knn = clf_knn.fit(X_train_scaled,y_train)\nclf_performance(best_clf_knn,'KNN')","metadata":{"execution":{"iopub.status.busy":"2022-07-23T10:57:57.493668Z","iopub.execute_input":"2022-07-23T10:57:57.494233Z","iopub.status.idle":"2022-07-23T10:57:59.660441Z","shell.execute_reply.started":"2022-07-23T10:57:57.494194Z","shell.execute_reply":"2022-07-23T10:57:59.659005Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"svc = SVC(probability = True)\nparam_grid = tuned_parameters = [{'kernel': ['rbf'], 'gamma': [.1,.5,1,2,5,10],\n                                  'C': [.1, 1, 10, 100, 1000]},\n                                 {'kernel': ['linear'], 'C': [.1, 1, 10, 100, 1000]},\n                                 {'kernel': ['poly'], 'degree' : [2,3,4,5], 'C': [.1, 1, 10, 100, 1000]}]\nclf_svc = GridSearchCV(svc, param_grid = param_grid, cv = 5, verbose = True, n_jobs = -1)\nbest_clf_svc = clf_svc.fit(X_train_scaled,y_train)\nclf_performance(best_clf_svc,'SVC')","metadata":{"execution":{"iopub.status.busy":"2022-07-23T10:57:59.662437Z","iopub.execute_input":"2022-07-23T10:57:59.663115Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"#Because the total feature space is so large, I used a randomized search to narrow down the paramters for the model. I took the best model from this and did a more granular search \n\"\"\"\nrf = RandomForestClassifier(random_state = 1)\nparam_grid =  {'n_estimators': [100,500,1000], \n                                  'bootstrap': [True,False],\n                                  'max_depth': [3,5,10,20,50,75,100,None],\n                                  'max_features': ['auto','sqrt'],\n                                  'min_samples_leaf': [1,2,4,10],\n                                  'min_samples_split': [2,5,10]}\n                                  \nclf_rf_rnd = RandomizedSearchCV(rf, param_distributions = param_grid, n_iter = 100, cv = 5, verbose = True, n_jobs = -1)\nbest_clf_rf_rnd = clf_rf_rnd.fit(X_train_scaled,y_train)\nclf_performance(best_clf_rf_rnd,'Random Forest')\"\"\"","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"rf = RandomForestClassifier(random_state = 1)\nparam_grid =  {'n_estimators': [400,450,500,550],\n               'criterion':['gini','entropy'],\n                                  'bootstrap': [True],\n                                  'max_depth': [15, 20, 25],\n                                  'max_features': ['auto','sqrt', 10],\n                                  'min_samples_leaf': [2,3],\n                                  'min_samples_split': [2,3]}\n                                  \nclf_rf = GridSearchCV(rf, param_grid = param_grid, cv = 5, verbose = True, n_jobs = -1)\nbest_clf_rf = clf_rf.fit(X_train_scaled,y_train)\nclf_performance(best_clf_rf,'Random Forest')","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"best_rf = best_clf_rf.best_estimator_.fit(X_train_scaled,y_train)\nfeat_importances = pd.Series(best_rf.feature_importances_, index=X_train_scaled.columns)\nfeat_importances.nlargest(20).plot(kind='barh')","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"\"\"\"xgb = XGBClassifier(random_state = 1)\n\nparam_grid = {\n    'n_estimators': [20, 50, 100, 250, 500,1000],\n    'colsample_bytree': [0.2, 0.5, 0.7, 0.8, 1],\n    'max_depth': [2, 5, 10, 15, 20, 25, None],\n    'reg_alpha': [0, 0.5, 1],\n    'reg_lambda': [1, 1.5, 2],\n    'subsample': [0.5,0.6,0.7, 0.8, 0.9],\n    'learning_rate':[.01,0.1,0.2,0.3,0.5, 0.7, 0.9],\n    'gamma':[0,.01,.1,1,10,100],\n    'min_child_weight':[0,.01,0.1,1,10,100],\n    'sampling_method': ['uniform', 'gradient_based']\n}\n\n#clf_xgb = GridSearchCV(xgb, param_grid = param_grid, cv = 5, verbose = True, n_jobs = -1)\n#best_clf_xgb = clf_xgb.fit(X_train_scaled,y_train)\n#clf_performance(best_clf_xgb,'XGB')\nclf_xgb_rnd = RandomizedSearchCV(xgb, param_distributions = param_grid, n_iter = 1000, cv = 5, verbose = True, n_jobs = -1)\nbest_clf_xgb_rnd = clf_xgb_rnd.fit(X_train_scaled,y_train)\nclf_performance(best_clf_xgb_rnd,'XGB')\"\"\"","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"xgb = XGBClassifier(random_state = 1)\n\nparam_grid = {\n    'n_estimators': [450,500,550],\n    'colsample_bytree': [0.75,0.8,0.85],\n    'max_depth': [None],\n    'reg_alpha': [1],\n    'reg_lambda': [2, 5, 10],\n    'subsample': [0.55, 0.6, .65],\n    'learning_rate':[0.5],\n    'gamma':[.5,1,2],\n    'min_child_weight':[0.01],\n    'sampling_method': ['uniform']\n}\n\nclf_xgb = GridSearchCV(xgb, param_grid = param_grid, cv = 5, verbose = True, n_jobs = -1)\nbest_clf_xgb = clf_xgb.fit(X_train_scaled,y_train)\nclf_performance(best_clf_xgb,'XGB')\n","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"y_hat_xgb = best_clf_xgb.best_estimator_.predict(X_test_scaled).astype(int)\nxgb_submission = {'PassengerId': test.PassengerId, 'Survived': y_hat_xgb}\nsubmission_xgb = pd.DataFrame(data=xgb_submission)\nsubmission_xgb.to_csv('xgb_submission3.csv', index=False)","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## Model Additional Ensemble Approaches \n1) Experimented with a hard voting classifier of three estimators (KNN, SVM, RF) (81.6%)\n\n2) **Experimented with a soft voting classifier of three estimators (KNN, SVM, RF) (82.3%) (Best Performance)**\n\n3) Experimented with soft voting on all estimators performing better than 80% except xgb (KNN, RF, LR, SVC) (82.9%)\n\n4) Experimented with soft voting on all estimators including XGB (KNN, SVM, RF, LR, XGB) (83.5%)","metadata":{}},{"cell_type":"code","source":"best_lr = best_clf_lr.best_estimator_\nbest_knn = best_clf_knn.best_estimator_\nbest_svc = best_clf_svc.best_estimator_\nbest_rf = best_clf_rf.best_estimator_\nbest_xgb = best_clf_xgb.best_estimator_\n\nvoting_clf_hard = VotingClassifier(estimators = [('knn',best_knn),('rf',best_rf),('svc',best_svc)], voting = 'hard') \nvoting_clf_soft = VotingClassifier(estimators = [('knn',best_knn),('rf',best_rf),('svc',best_svc)], voting = 'soft') \nvoting_clf_all = VotingClassifier(estimators = [('knn',best_knn),('rf',best_rf),('svc',best_svc), ('lr', best_lr)], voting = 'soft') \nvoting_clf_xgb = VotingClassifier(estimators = [('knn',best_knn),('rf',best_rf),('svc',best_svc), ('xgb', best_xgb),('lr', best_lr)], voting = 'soft')\n\nprint('voting_clf_hard :',cross_val_score(voting_clf_hard,X_train,y_train,cv=5))\nprint('voting_clf_hard mean :',cross_val_score(voting_clf_hard,X_train,y_train,cv=5).mean())\n\nprint('voting_clf_soft :',cross_val_score(voting_clf_soft,X_train,y_train,cv=5))\nprint('voting_clf_soft mean :',cross_val_score(voting_clf_soft,X_train,y_train,cv=5).mean())\n\nprint('voting_clf_all :',cross_val_score(voting_clf_all,X_train,y_train,cv=5))\nprint('voting_clf_all mean :',cross_val_score(voting_clf_all,X_train,y_train,cv=5).mean())\n\nprint('voting_clf_xgb :',cross_val_score(voting_clf_xgb,X_train,y_train,cv=5))\nprint('voting_clf_xgb mean :',cross_val_score(voting_clf_xgb,X_train,y_train,cv=5).mean())\n","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"#in a soft voting classifier you can weight some models more than others. I used a grid search to explore different weightings\n#no new results here\nparams = {'weights' : [[1,1,1],[1,2,1],[1,1,2],[2,1,1],[2,2,1],[1,2,2],[2,1,2]]}\n\nvote_weight = GridSearchCV(voting_clf_soft, param_grid = params, cv = 5, verbose = True, n_jobs = -1)\nbest_clf_weight = vote_weight.fit(X_train_scaled,y_train)\nclf_performance(best_clf_weight,'VC Weights')\nvoting_clf_sub = best_clf_weight.best_estimator_.predict(X_test_scaled)\n","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"#Make Predictions \nvoting_clf_hard.fit(X_train_scaled, y_train)\nvoting_clf_soft.fit(X_train_scaled, y_train)\nvoting_clf_all.fit(X_train_scaled, y_train)\nvoting_clf_xgb.fit(X_train_scaled, y_train)\n\nbest_rf.fit(X_train_scaled, y_train)\ny_hat_vc_hard = voting_clf_hard.predict(X_test_scaled).astype(int)\ny_hat_rf = best_rf.predict(X_test_scaled).astype(int)\ny_hat_vc_soft =  voting_clf_soft.predict(X_test_scaled).astype(int)\ny_hat_vc_all = voting_clf_all.predict(X_test_scaled).astype(int)\ny_hat_vc_xgb = voting_clf_xgb.predict(X_test_scaled).astype(int)","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"#convert output to dataframe \nfinal_data = {'PassengerId': test.PassengerId, 'Survived': y_hat_rf}\nsubmission = pd.DataFrame(data=final_data)\n\nfinal_data_2 = {'PassengerId': test.PassengerId, 'Survived': y_hat_vc_hard}\nsubmission_2 = pd.DataFrame(data=final_data_2)\n\nfinal_data_3 = {'PassengerId': test.PassengerId, 'Survived': y_hat_vc_soft}\nsubmission_3 = pd.DataFrame(data=final_data_3)\n\nfinal_data_4 = {'PassengerId': test.PassengerId, 'Survived': y_hat_vc_all}\nsubmission_4 = pd.DataFrame(data=final_data_4)\n\nfinal_data_5 = {'PassengerId': test.PassengerId, 'Survived': y_hat_vc_xgb}\nsubmission_5 = pd.DataFrame(data=final_data_5)\n\nfinal_data_comp = {'PassengerId': test.PassengerId, 'Survived_vc_hard': y_hat_vc_hard, 'Survived_rf': y_hat_rf, 'Survived_vc_soft' : y_hat_vc_soft, 'Survived_vc_all' : y_hat_vc_all,  'Survived_vc_xgb' : y_hat_vc_xgb}\ncomparison = pd.DataFrame(data=final_data_comp)","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"#track differences between outputs \ncomparison['difference_rf_vc_hard'] = comparison.apply(lambda x: 1 if x.Survived_vc_hard != x.Survived_rf else 0, axis =1)\ncomparison['difference_soft_hard'] = comparison.apply(lambda x: 1 if x.Survived_vc_hard != x.Survived_vc_soft else 0, axis =1)\ncomparison['difference_hard_all'] = comparison.apply(lambda x: 1 if x.Survived_vc_all != x.Survived_vc_hard else 0, axis =1)\n","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"comparison.difference_hard_all.value_counts()","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"#prepare submission files \nsubmission.to_csv('submission_rf.csv', index =False)\nsubmission_2.to_csv('submission_vc_hard.csv',index=False)\nsubmission_3.to_csv('submission_vc_soft.csv', index=False)\nsubmission_4.to_csv('submission_vc_all.csv', index=False)\nsubmission_5.to_csv('submission_vc_xgb2.csv', index=False)","metadata":{"trusted":true},"execution_count":null,"outputs":[]}]}