{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"pygments_lexer":"ipython3","nbconvert_exporter":"python","version":"3.6.4","file_extension":".py","codemirror_mode":{"name":"ipython","version":3},"name":"python","mimetype":"text/x-python"}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"# My First Kaggle Competition - Titanic Dataset Predictions\n\nThe titanic dataset is a well known machine learning dataset used for starting machine learning studies. In this notebook, I want to apply my learning using:\n* data visualization\n* feature engineering\n* training machine learning models\n\nI will try to do my best to get the maximum accuracy. But, if I won't, this will be a great opportunity to learn from other notebooks and improve my accuracy over time. This notebook is organised in 5 chapters:\n\n1. Reading Data\n2. Exploring Data\n3. Data Visualization\n4. Processing\n5. Modeling\n\nHope you enjoy it!\n\n## 1. Reading Data\n\nThere are two files in this competiton: train and test files in csv.","metadata":{}},{"cell_type":"code","source":"import numpy as np\nimport pandas as pd\nimport matplotlib.pyplot as plt\nimport seaborn as sns\n\n%matplotlib inline\n\nsns.set_style('darkgrid')\n\ntrain_data = pd.read_csv(\"/kaggle/input/titanic/train.csv\")\ntest_data = pd.read_csv(\"/kaggle/input/titanic/test.csv\")\n\ntrain_data.head()","metadata":{"execution":{"iopub.status.busy":"2022-07-19T19:43:30.692309Z","iopub.execute_input":"2022-07-19T19:43:30.692919Z","iopub.status.idle":"2022-07-19T19:43:30.920703Z","shell.execute_reply.started":"2022-07-19T19:43:30.692880Z","shell.execute_reply":"2022-07-19T19:43:30.919869Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## 2. Exploring Data\n\nBefore we get further, it is a good practice to make some questions about data, so we can explore details that can help us modeling better. First, look at the output of train_data.head() above. There are 12 columns in this dataset:\n* PassengerId is the unique identifier of a passanger\n* Survived - value 0 means not survived and value 1 means it did survived\n* Pclass - class in the ship\n* Name - name of a passenger\n* Age\n* SibSp - it is a number that means how many siblins a person have on titanic\n* Parch -\n* Ticket - the ticket\n* Fare - how much someone paid\n* Cabin - the cabin\n* Embarked - where did one embarked\n\nSome questions I want to arise for this notebook:\n* How many survived?\n* How many Pclass do we have?\n* How many male and female?\n* how many survived between man and female?\n* how many survived between pclass?\n* Is there a correlation between fare and survived?\n\nSee info about train dataframe. There are 889 entries in train data.","metadata":{}},{"cell_type":"code","source":"train_data.info()","metadata":{"execution":{"iopub.status.busy":"2022-07-19T19:43:30.922444Z","iopub.execute_input":"2022-07-19T19:43:30.922755Z","iopub.status.idle":"2022-07-19T19:43:30.949978Z","shell.execute_reply.started":"2022-07-19T19:43:30.922714Z","shell.execute_reply":"2022-07-19T19:43:30.949189Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Let us explore variables and the survived class, which is the goal for predictions. Pclass vs Survived. In average, Pclass==1 has more chance of survival than other classes.","metadata":{}},{"cell_type":"code","source":"train_data[['Survived', 'Pclass']].groupby(by='Pclass', as_index=False).describe(include='all')","metadata":{"execution":{"iopub.status.busy":"2022-07-19T19:43:30.951164Z","iopub.execute_input":"2022-07-19T19:43:30.951492Z","iopub.status.idle":"2022-07-19T19:43:31.002662Z","shell.execute_reply.started":"2022-07-19T19:43:30.951458Z","shell.execute_reply":"2022-07-19T19:43:31.001800Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"In average, women have far more chance of survival than men.","metadata":{"execution":{"iopub.status.busy":"2022-05-19T19:55:13.876314Z","iopub.execute_input":"2022-05-19T19:55:13.876581Z","iopub.status.idle":"2022-05-19T19:55:13.882823Z","shell.execute_reply.started":"2022-05-19T19:55:13.876553Z","shell.execute_reply":"2022-05-19T19:55:13.8817Z"}}},{"cell_type":"code","source":"train_data[['Survived', 'Sex']].groupby(by='Sex', as_index=False).describe(include='all')","metadata":{"execution":{"iopub.status.busy":"2022-07-19T19:43:31.004589Z","iopub.execute_input":"2022-07-19T19:43:31.005057Z","iopub.status.idle":"2022-07-19T19:43:31.050974Z","shell.execute_reply.started":"2022-07-19T19:43:31.004977Z","shell.execute_reply":"2022-07-19T19:43:31.050104Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Passengers embarked on 'C' has more chance of survival too.","metadata":{}},{"cell_type":"code","source":"train_data[['Survived', 'Embarked']].groupby(by=['Embarked'], as_index=False).describe()","metadata":{"execution":{"iopub.status.busy":"2022-07-19T19:43:31.069347Z","iopub.execute_input":"2022-07-19T19:43:31.069801Z","iopub.status.idle":"2022-07-19T19:43:31.107085Z","shell.execute_reply.started":"2022-07-19T19:43:31.069764Z","shell.execute_reply":"2022-07-19T19:43:31.106222Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Looking at SibSp, which means how many relatives a passenger has on Titanic, we can see that there may be correlations.","metadata":{}},{"cell_type":"markdown","source":" ","metadata":{}},{"cell_type":"code","source":"train_data[['Survived', 'SibSp']].groupby(by=['SibSp'], as_index=False).mean()","metadata":{"execution":{"iopub.status.busy":"2022-07-19T19:43:31.157238Z","iopub.execute_input":"2022-07-19T19:43:31.157519Z","iopub.status.idle":"2022-07-19T19:43:31.172917Z","shell.execute_reply.started":"2022-07-19T19:43:31.157488Z","shell.execute_reply":"2022-07-19T19:43:31.172122Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Parch is similar to SibSp, we can look further to explore this feature.","metadata":{}},{"cell_type":"code","source":"train_data[['Survived', 'Parch']].groupby(by=['Parch'], as_index=False).mean()","metadata":{"execution":{"iopub.status.busy":"2022-07-19T19:43:31.244670Z","iopub.execute_input":"2022-07-19T19:43:31.244964Z","iopub.status.idle":"2022-07-19T19:43:31.260634Z","shell.execute_reply.started":"2022-07-19T19:43:31.244931Z","shell.execute_reply":"2022-07-19T19:43:31.259818Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## 3. Data Visualization\n\nWe will continue to explore data, but now using data visualization to get a broader idea.\n\nAs we can see in the next figure:\n* Higher Fare increases the survival chance;\n* Pclass 1 has higher Fare and more survival chance;\n* Woman who has higher Fare has more survival chance;\n* Embarked on C who has higher Fare has more survival chance.","metadata":{}},{"cell_type":"code","source":"fig, axes = plt.subplots(1, 3, figsize=(20, 6), sharey=True)\nfig.suptitle('Survived by Fare')\n\nax = sns.boxplot(ax=axes[0],data=train_data, x='Survived', y='Fare', hue='Pclass').set_title('Pclass')\n\nax = sns.boxplot(ax=axes[1], data=train_data, x='Survived', y='Fare', hue='Sex').set_title('Sex')\n\nax = sns.boxplot(ax=axes[2], data=train_data, x='Survived', y='Fare', hue='Embarked').set_title('Embarked')","metadata":{"execution":{"iopub.status.busy":"2022-07-19T19:43:31.328531Z","iopub.execute_input":"2022-07-19T19:43:31.328948Z","iopub.status.idle":"2022-07-19T19:43:32.040693Z","shell.execute_reply.started":"2022-07-19T19:43:31.328908Z","shell.execute_reply":"2022-07-19T19:43:32.039806Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Age looks like a good parameter, as we can see. Children and women have higher survival chances.","metadata":{}},{"cell_type":"code","source":"g = sns.FacetGrid(data=train_data, col='Survived', row='Sex',  height=4, aspect=1.3)\ng.map(plt.hist, 'Age')","metadata":{"execution":{"iopub.status.busy":"2022-07-19T19:43:32.042285Z","iopub.execute_input":"2022-07-19T19:43:32.042533Z","iopub.status.idle":"2022-07-19T19:43:33.111589Z","shell.execute_reply.started":"2022-07-19T19:43:32.042504Z","shell.execute_reply":"2022-07-19T19:43:33.111024Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## 4. Processing Data\n\nWe going to do these tasks:\n* 4.1 Handling with missing values;\n* 4.2 Creating new features;\n* 4.2 Selecting columns and encoding;\n\n### 4.1 Handling with missing values\n\nFirst, look at the dataset size. It has 891 entries and some columns are not filled. Columns with missing data: Age, Cabin and Fare (test dataset).","metadata":{}},{"cell_type":"code","source":"train_data.info()\nprint('-----'*10)\ntest_data.info()","metadata":{"execution":{"iopub.status.busy":"2022-07-19T19:43:33.112829Z","iopub.execute_input":"2022-07-19T19:43:33.113183Z","iopub.status.idle":"2022-07-19T19:43:33.132698Z","shell.execute_reply.started":"2022-07-19T19:43:33.113139Z","shell.execute_reply":"2022-07-19T19:43:33.132061Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"We will calculate the average age regard to Pclass and Sex. Pclass has 3 possible values and Sex has 2 values, so there will be 6 average ages to calculate (Pclass=1 and sex=male, Pclass=1 and sex=female, and so on).\n\n\n* Filling missing values of Age \n\nFor this, we will fill missing values of age with average, based on Pclass and Sex.","metadata":{}},{"cell_type":"code","source":"mean_age = np.zeros((2,3))\n\nsex = {0: 'male', 1: 'female'}\n\nfor i in range(0,3):\n    for j in range(0,2):\n        mean_age[j, i] = int(train_data[(train_data['Sex'] == sex[j]) & (train_data['Pclass'] == i+1)]['Age'].mean())\n        \n        train_data.loc[(train_data['Age'].isna()) & (train_data['Pclass'] == i+1) & (train_data['Sex'] == sex[j]),'Age'] = mean_age[j,i]\n        test_data.loc[(test_data['Age'].isna()) & (test_data['Pclass'] == i+1) & (test_data['Sex'] == sex[j]),'Age'] = mean_age[j,i]","metadata":{"execution":{"iopub.status.busy":"2022-07-19T19:43:33.134429Z","iopub.execute_input":"2022-07-19T19:43:33.134809Z","iopub.status.idle":"2022-07-19T19:43:33.163255Z","shell.execute_reply.started":"2022-07-19T19:43:33.134770Z","shell.execute_reply":"2022-07-19T19:43:33.162469Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"* Filling Missing Embarked values\n\nFor this columns, we will fill with the most common embarked.","metadata":{}},{"cell_type":"code","source":"print(train_data['Embarked'].value_counts())\n\ncommon_embarked = 'S'\ntrain_data.loc[train_data['Embarked'].isnull(), 'Embarked'] = common_embarked","metadata":{"execution":{"iopub.status.busy":"2022-07-19T19:43:33.164298Z","iopub.execute_input":"2022-07-19T19:43:33.164649Z","iopub.status.idle":"2022-07-19T19:43:33.171965Z","shell.execute_reply.started":"2022-07-19T19:43:33.164622Z","shell.execute_reply":"2022-07-19T19:43:33.171166Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"* Filling Missing Fare values\n\nAs we did with Age, we will fill Fare with mean values based on Sex and Pclass columns.","metadata":{}},{"cell_type":"code","source":"mean_fare = np.zeros((2,3))\nsex = {0: 'male', 1: 'female'}\n\nfor i in range(3):\n    for j in range(2):\n        mean_fare[j,i] = train_data.loc[(train_data['Sex'] == sex[j]) & (train_data['Pclass'] == i+1), 'Fare'].mean()\n        \n        train_data.loc[(train_data['Fare'].isnull()) &(train_data['Sex'] == sex[j]) & (train_data['Pclass'] == i+1), 'Fare'] = mean_fare[j,i]\n        test_data.loc[(test_data['Fare'].isnull()) &(test_data['Sex'] == sex[j]) & (test_data['Pclass'] == i+1), 'Fare'] = mean_fare[j,i]","metadata":{"execution":{"iopub.status.busy":"2022-07-19T19:43:33.173594Z","iopub.execute_input":"2022-07-19T19:43:33.174065Z","iopub.status.idle":"2022-07-19T19:43:33.201953Z","shell.execute_reply.started":"2022-07-19T19:43:33.174009Z","shell.execute_reply":"2022-07-19T19:43:33.201187Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"* In order to improve (or try at least) accuracy, we will divide Age and Fare into Range of Values.","metadata":{}},{"cell_type":"code","source":"train_data['AgeRange'] = pd.cut(train_data['Age'], 4)\n\ntrain_data.loc[train_data['Age'] <= 20, 'Age'] = 0\ntrain_data.loc[(train_data['Age'] > 20) & (train_data['Age'] <= 40), 'Age'] = 1\ntrain_data.loc[(train_data['Age'] > 40) & (train_data['Age'] <= 60), 'Age'] = 2\ntrain_data.loc[(train_data['Age'] > 60), 'Age'] = 3\n\ntest_data.loc[test_data['Age'] <= 20, 'Age'] = 0\ntest_data.loc[(test_data['Age'] > 20) & (test_data['Age'] <= 40), 'Age'] = 1\ntest_data.loc[(test_data['Age'] > 40) & (test_data['Age'] <= 60), 'Age'] = 2\ntest_data.loc[(test_data['Age'] > 60), 'Age'] = 3","metadata":{"execution":{"iopub.status.busy":"2022-07-19T19:43:33.203461Z","iopub.execute_input":"2022-07-19T19:43:33.203949Z","iopub.status.idle":"2022-07-19T19:43:33.230295Z","shell.execute_reply.started":"2022-07-19T19:43:33.203894Z","shell.execute_reply":"2022-07-19T19:43:33.229590Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"train_data['FareRange'] = pd.cut(train_data['Fare'], 5)\n\ntrain_data.loc[train_data['Fare'] <= 128.082, 'Fare'] = 0\ntrain_data.loc[(train_data['Fare'] > 128.082) & (train_data['Fare'] <= 256.165), 'Fare'] = 1\ntrain_data.loc[(train_data['Fare'] > 256.165) & (train_data['Fare'] <= 384.247), 'Fare'] = 2\ntrain_data.loc[(train_data['Fare'] > 384.247), 'Fare'] = 3\n\ntest_data.loc[test_data['Fare'] <= 128.082, 'Fare'] = 0\ntest_data.loc[(test_data['Fare'] > 128.082) & (test_data['Fare'] <= 256.165), 'Fare'] = 1\ntest_data.loc[(test_data['Fare'] > 256.165) & (test_data['Fare'] <= 384.247), 'Fare'] = 2\ntest_data.loc[(test_data['Fare'] > 384.247), 'Fare'] = 3","metadata":{"execution":{"iopub.status.busy":"2022-07-19T19:43:33.231372Z","iopub.execute_input":"2022-07-19T19:43:33.231709Z","iopub.status.idle":"2022-07-19T19:43:33.249296Z","shell.execute_reply.started":"2022-07-19T19:43:33.231681Z","shell.execute_reply":"2022-07-19T19:43:33.248598Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### 4.2 Creating new features\n\nWe have columns SibSp and Parch and both are not very informative. So, we will combine them to another columns \"HasFamily\" so 0 means a passenger is Alone and 1 means otherwise. Code below shows that higher survival chances are related to having family onboard.","metadata":{}},{"cell_type":"code","source":"train_data['HasFamily'] = train_data['SibSp'] + train_data['Parch']\ntrain_data['HasFamily'] = train_data['HasFamily'].map(lambda x: 0 if x==0 else 1)\n\ntest_data['HasFamily'] = test_data['SibSp'] + test_data['Parch']\ntest_data['HasFamily'] = test_data['HasFamily'].map(lambda x: 0 if x==0 else 1)\n\nprint(train_data['HasFamily'].value_counts())\ntrain_data[['Survived', 'HasFamily']].groupby('HasFamily', as_index=False).mean()","metadata":{"execution":{"iopub.status.busy":"2022-07-19T19:43:33.250395Z","iopub.execute_input":"2022-07-19T19:43:33.250742Z","iopub.status.idle":"2022-07-19T19:43:33.270772Z","shell.execute_reply.started":"2022-07-19T19:43:33.250714Z","shell.execute_reply":"2022-07-19T19:43:33.269964Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### 4.3 Selecting columns and enconding\n\nWe will choose the columns and encode categorical columns.","metadata":{}},{"cell_type":"code","source":"#selecting columns\nfeatures = ['Pclass', 'HasFamily', 'Fare', 'Age', 'Sex', 'is_C', 'is_Q']\n\n# # we cannot use sex column with text, so we are going to create is_male columns, 1 means male, 0 means female\ntrain_data['Sex'] = train_data['Sex'].map( {'male': 0, 'female': 1} )\ntest_data['Sex'] = test_data['Sex'].map( {'male': 0, 'female': 1} )\n\n# # Embarked has three values: C, S and Q, so we are going to create two columns about this values\ntrain_data['is_C'] = train_data['Embarked'].map(lambda x: 1 if x=='C' else 0)\ntest_data['is_C'] = test_data['Embarked'].map(lambda x: 1 if x=='C' else 0)\n\ntrain_data['is_Q'] = train_data['Embarked'].map(lambda x: 1 if x=='Q' else 0)\ntest_data['is_Q'] = test_data['Embarked'].map(lambda x: 1 if x=='Q' else 0)","metadata":{"execution":{"iopub.status.busy":"2022-07-19T19:43:33.273224Z","iopub.execute_input":"2022-07-19T19:43:33.273430Z","iopub.status.idle":"2022-07-19T19:43:33.286681Z","shell.execute_reply.started":"2022-07-19T19:43:33.273404Z","shell.execute_reply":"2022-07-19T19:43:33.286030Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## 5. Modeling\n\nWe want to test some models and see how each one of them perform on the train data. So, we want to split data into train and validation.\nSecond, we want to test 3 different machine learning algorithms with standard parameters and see their performance.","metadata":{}},{"cell_type":"code","source":"from sklearn.linear_model import LogisticRegression\nfrom sklearn.ensemble import RandomForestClassifier\nfrom sklearn.preprocessing import StandardScaler\nfrom sklearn.model_selection import RepeatedKFold\n\nX = train_data[features]\ny = train_data['Survived']\n\n# Test final data and predict with it\nX_test = test_data[features]\n\n# #Scale data for linear models\n# scaler = StandardScaler()\n# scaler.fit(X)\n# X_transformed = scaler.transform(X)","metadata":{"execution":{"iopub.status.busy":"2022-07-19T19:43:33.287840Z","iopub.execute_input":"2022-07-19T19:43:33.288386Z","iopub.status.idle":"2022-07-19T19:43:33.305564Z","shell.execute_reply.started":"2022-07-19T19:43:33.288348Z","shell.execute_reply":"2022-07-19T19:43:33.304891Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"#### Logistic regression results","metadata":{}},{"cell_type":"code","source":"lr_scores = []\nkf = RepeatedKFold(n_splits=2, n_repeats=5, random_state=0)\n\nfor train_, val_ in kf.split(X):\n\n    X_train, X_val = X.iloc[train_], X.iloc[val_]\n    y_train, y_val = y.iloc[train_], y.iloc[val_]\n\n    modelo = LogisticRegression(max_iter=1000, random_state=0)\n    \n    modelo.fit(X_train, y_train)\n\n    predictions = modelo.predict(X_val)\n\n    accuracy = np.mean(y_val == predictions)\n    lr_scores.append(accuracy)\n    print(\"Accuracy:\", accuracy)\n    \nmean_scores = np.mean(lr_scores)\nprint(f'Logistic Regression mean score {mean_scores}')","metadata":{"execution":{"iopub.status.busy":"2022-07-19T19:43:33.306933Z","iopub.execute_input":"2022-07-19T19:43:33.307365Z","iopub.status.idle":"2022-07-19T19:43:33.430858Z","shell.execute_reply.started":"2022-07-19T19:43:33.307334Z","shell.execute_reply":"2022-07-19T19:43:33.430006Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"#### Random Forest","metadata":{}},{"cell_type":"code","source":"rf_scores = []\nkf = RepeatedKFold(n_splits=2, n_repeats=5, random_state=0)\n\nfor train_, val_ in kf.split(X):\n\n    X_train, X_val = X.iloc[train_], X.iloc[val_]\n    y_train, y_val = y.iloc[train_], y.iloc[val_]\n\n    modelo = RandomForestClassifier(random_state=0)\n    \n    modelo.fit(X_train, y_train)\n\n    predictions = modelo.predict(X_val)\n\n    accuracy = np.mean(y_val == predictions)\n    rf_scores.append(accuracy)\n    print(\"Accuracy:\", accuracy)\n    \nmean_scores = np.mean(rf_scores)\nprint(f'Random Forest mean score {mean_scores}')","metadata":{"execution":{"iopub.status.busy":"2022-07-19T19:43:33.431927Z","iopub.execute_input":"2022-07-19T19:43:33.432160Z","iopub.status.idle":"2022-07-19T19:43:35.483693Z","shell.execute_reply.started":"2022-07-19T19:43:33.432114Z","shell.execute_reply":"2022-07-19T19:43:35.482860Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"X.index.size","metadata":{"execution":{"iopub.status.busy":"2022-07-19T19:43:35.485074Z","iopub.execute_input":"2022-07-19T19:43:35.485453Z","iopub.status.idle":"2022-07-19T19:43:35.492341Z","shell.execute_reply.started":"2022-07-19T19:43:35.485410Z","shell.execute_reply":"2022-07-19T19:43:35.491379Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"from sklearn.ensemble import RandomForestClassifier, GradientBoostingClassifier\nfrom sklearn.linear_model import SGDClassifier\n\ndef train_random_forest(params):\n    max_depth = params[0]\n    min_samples_split = params[1]\n    min_samples_leaf = params[2]\n    max_leaf_nodes = params[3]\n    \n    model = RandomForestClassifier(max_depth=max_depth, min_samples_split=min_samples_split,\n                                 min_samples_leaf=min_samples_leaf, max_leaf_nodes=max_leaf_nodes,\n                                 random_state=0, n_jobs=-1)\n    \n    size = X.index.size\n    X_train, X_val = X[:round(.7*size)], X[round(.7*size):]\n    y_train, y_val = y[:round(.7*size)], y[round(.7*size):]\n    \n    model.fit(X_train, y_train)\n    \n    score = model.score(X_val, y_val)\n    \n    \n    return score\n\ndef train_xboost(params):\n    max_depth = params[0]\n    learning_rate = params[1]\n    colsample_bytree = params[2]\n    subsample = params[3]\n    \n    model = GradientBoostingClassifier(max_depth=max_depth, learning_rate=learning_rate, \n                         subsample=subsample, n_estimators=100, random_state=0)\n    \n    size = X.index.size\n    X_train, X_val = X[:round(.7*size)], X[round(.7*size):]\n    y_train, y_val = y[:round(.7*size)], y[round(.7*size):]\n    \n    model.fit(X_train, y_train)\n    \n    score = model.score(X_val, y_val)\n    \n    \n    return score\n\ndef train_elastic_net(params):\n    alpha = params[0]\n    l1_ratio = params[1]\n    \n    model = SGDClassifier(penalty=\"elasticnet\", l1_ratio=l1_ratio)\n#     model = ElasticNet(alpha=alpha, l1_ratio=l1_ratio, max_iter=10000, random_state=0)\n    size = X.index.size\n    X_train, X_val = X[:round(.7*size)], X[round(.7*size):]\n    y_train, y_val = y[:round(.7*size)], y[round(.7*size):]\n    \n    model.fit(X_train, y_train)\n    \n    score = model.score(X_val, y_val)\n    \n    \n    return score","metadata":{"execution":{"iopub.status.busy":"2022-07-19T19:57:14.905427Z","iopub.execute_input":"2022-07-19T19:57:14.905713Z","iopub.status.idle":"2022-07-19T19:57:14.921721Z","shell.execute_reply.started":"2022-07-19T19:57:14.905683Z","shell.execute_reply":"2022-07-19T19:57:14.920861Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"from skopt import gp_minimize\nimport warnings\nwarnings.filterwarnings('ignore')\n\n# Random Forest params to explore\nspace_rf = [(2, 64, 'log-uniform'), #max_depth\n         (2, 32), # min_sample_split\n         (1, 32), # min_samples\n         (2, 100)] # max_leaf_nodes\n\n# XGBRegressor params to explore\nspace_xboost = [(2, 64), # max depth\n         (1e-5, 1e-1), #learning rate\n         (0.5, 1), #colsample\n         (0.5, 1) #subsample\n        ]\n\n# Elastic net params to explore\nspace_elastic_net = [(0.5, 5.5), # alpha\n         (0.1, 0.8) # l1_ratio\n        ]\n\n# Train models\n\nrf_best_params = gp_minimize(train_random_forest, space_rf, random_state=1, verbose=0, n_calls=30, n_random_starts=10)\n\nxgboost_best_params = gp_minimize(train_xboost, space_xboost, random_state=1, verbose=0, n_calls=30, n_random_starts=10)\n\nelastic_net_best_params = gp_minimize(train_elastic_net, space_elastic_net, random_state=1, verbose=0, n_calls=30, n_random_starts=10)\n\n# Print best params found in each model\nprint(f'Random Forest best params: max_depth={rf_best_params.x[0]}, min_sample_split={rf_best_params.x[1]}, min_samples={rf_best_params.x[2]},\\\n max_leaf_nodes={rf_best_params.x[3]}')\n\nprint(f'XGBoost best params: max_depth={xgboost_best_params.x[0]}, learning_rate={xgboost_best_params.x[1]:.2f}, colsample={xgboost_best_params.x[2]:.2f}, \\\nsubsample={xgboost_best_params.x[3]:.2f}')\n\nprint(f'Elastic net best params: l1_ratio={elastic_net_best_params.x[1]:.2f}')","metadata":{"execution":{"iopub.status.busy":"2022-07-19T19:57:15.558807Z","iopub.execute_input":"2022-07-19T19:57:15.559131Z","iopub.status.idle":"2022-07-19T19:58:14.970785Z","shell.execute_reply.started":"2022-07-19T19:57:15.559096Z","shell.execute_reply":"2022-07-19T19:58:14.969727Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"from sklearn.ensemble import VotingClassifier\n#Creating Voting Classifier with this 3 models\n\nrf = RandomForestClassifier(max_depth=2, min_samples_split=19, min_samples_leaf=1, max_leaf_nodes=2)\n\ngb = GradientBoostingClassifier(max_depth=25, learning_rate=0.0007, subsample=0.51)\n\nel = SGDClassifier(loss=\"hinge\", penalty=\"elasticnet\", l1_ratio=0.44)\n\n# el = ElasticNet(alpha=5.48, l1_ratio=0.75)\n\nvoting_clf = VotingClassifier(\n                estimators=[('rf', rf), ('gb', gb), ('el', el)],\n                voting='hard')\nvoting_clf.fit(X, y)","metadata":{"execution":{"iopub.status.busy":"2022-07-19T20:00:28.251986Z","iopub.execute_input":"2022-07-19T20:00:28.252303Z","iopub.status.idle":"2022-07-19T20:00:28.840898Z","shell.execute_reply.started":"2022-07-19T20:00:28.252264Z","shell.execute_reply":"2022-07-19T20:00:28.840092Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### Train main model and submit","metadata":{}},{"cell_type":"code","source":"# model = RandomForestClassifier(random_state=0, n_jobs=-1)\n# model.fit(X, y)\npredictions = voting_clf.predict(X_test)\n\noutput = pd.DataFrame({'PassengerId': test_data.PassengerId, 'Survived': predictions})\noutput.to_csv('submission.csv', index=False)\nprint(\"Your submission was successfully saved!\")","metadata":{"execution":{"iopub.status.busy":"2022-07-19T20:00:31.853529Z","iopub.execute_input":"2022-07-19T20:00:31.854303Z","iopub.status.idle":"2022-07-19T20:00:31.890295Z","shell.execute_reply.started":"2022-07-19T20:00:31.854260Z","shell.execute_reply":"2022-07-19T20:00:31.889590Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"pd.read_csv('/kaggle/working/submission.csv')","metadata":{"execution":{"iopub.status.busy":"2022-07-19T20:00:34.152041Z","iopub.execute_input":"2022-07-19T20:00:34.152350Z","iopub.status.idle":"2022-07-19T20:00:34.165664Z","shell.execute_reply.started":"2022-07-19T20:00:34.152313Z","shell.execute_reply":"2022-07-19T20:00:34.164838Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"","metadata":{},"execution_count":null,"outputs":[]}]}