{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"pygments_lexer":"ipython3","nbconvert_exporter":"python","version":"3.6.4","file_extension":".py","codemirror_mode":{"name":"ipython","version":3},"name":"python","mimetype":"text/x-python"}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"# Background\n\nIn the notebook [TitanicSurvCom](https://www.kaggle.com/code/jairjiramuoz/titanicsurvcom/notebook) I developed a classifier boosted tree model based on the [XGB](https://xgboost.readthedocs.io/en/stable/tutorials/model.html) implementation for the [Titanic competition](https://www.kaggle.com/competitions/titanic).\nThis model was able to reach a maximum score of 0.77751 accuracy once the hyperparameters were properly tuned. I wondered whether or not a neural network could do better than a boosted tree. Even when I never reached the same accuracy score (about 0.75 on average), I tried another approach: a combined model which takes the predicted probabilities from both classifiers and returns a linear combination of both to obtain a new probability and in turn a new class. In  this notebook I test the performance of every model in order to determine if one is best than another.   \n\n# Data loading\n","metadata":{}},{"cell_type":"code","source":"import numpy as np \nimport pandas as pd \nimport matplotlib.pyplot as plt\nimport seaborn as sns\n\n\ndf_train=pd.read_csv('../input/titanic/train.csv')\ndf_test=pd.read_csv('../input/titanic/test.csv')","metadata":{"execution":{"iopub.status.busy":"2022-07-25T00:41:20.393609Z","iopub.execute_input":"2022-07-25T00:41:20.393964Z","iopub.status.idle":"2022-07-25T00:41:21.903978Z","shell.execute_reply.started":"2022-07-25T00:41:20.393871Z","shell.execute_reply":"2022-07-25T00:41:21.903002Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"#  Imputation methods\n\nIt can be seen that a great deal of the Cabin column is missing. Note that a cabin is in writen in the\nform (A,B,C,D,E,F,G,T) (number) space (A,B,C,D,E,F,G,T) (number) ... \n\nThis might be explained because the letters A to G and T are kind of related to the cabin class and also to the cabin location. Having more than one letter might indicate the cabin is not a single one\n(more expensive cabins might also be more spacious) \n\nThere are lots of imputation methods for numerical features but just a few for categorical data ( unless you get creative with neuron network based models like LSTM). In this case I chose to give up imputation and I tried  to just focus on predicting the letter(s) contained in the cabin number format and the number of times it appeared on. This seemed to me the most useful and informative contribution of the Cabin feature. This will be reffered as the cabin sub model.   \n\nThe age feature also have some missing entries. In this case I deemed best to replace with the mean the missing data. Probably if I had more data it would have been better to infer the age by using the other features (like Name). Besides, in this case the missing data ratio is way lower, so its effect is not as important as the Cabin feature. \n\n","metadata":{}},{"cell_type":"code","source":"#many models to chose from\nfrom sklearn.linear_model import LinearRegression\nfrom sklearn.linear_model import Ridge\nfrom sklearn.linear_model import Lasso\nfrom sklearn.linear_model import MultiTaskLassoCV\nfrom sklearn.linear_model import MultiTaskElasticNet\n\nfrom xgboost import XGBRegressor\n# the implementation of boosted trees by xgboost doesn't allow multioutput so \n# we enable this by using scikit multioutput \nfrom sklearn.multioutput import RegressorChain\n#extract a copy from the data. We do not want to tamper with it by working on an (sub) alternate model\ndf_lab_train=df_train.copy()\ndf_lab_test=df_test.copy()\ndef cabin_transform(df):\n    #delete cabin, will replace this feature with just the informative parts of it\n    df.dropna(subset=['Cabin'],inplace=True)\n    #you could also use other values\n    df_mean=df['Age'].mean()\n    df['Age'].fillna(df_mean,inplace=True)\n    #simple imputation method, the fact that data from this feature is missing might in itself be informative\n    df['Embarked'].fillna('None',inplace=True)\n    #let's reduce the number of features we are working with to the most meaninful ones\n    df=df.loc[:,[ 'Pclass', 'Age', 'SibSp','Parch', 'Fare', 'Cabin']]\n    df['Cabin_class_A']=df['Cabin'].apply(lambda x: x.count('A'))\n    df['Cabin_class_B']=df['Cabin'].apply(lambda x: x.count('B'))\n    df['Cabin_class_C']=df['Cabin'].apply(lambda x: x.count('C'))\n    df['Cabin_class_D']=df['Cabin'].apply(lambda x: x.count('D'))\n    df['Cabin_class_E']=df['Cabin'].apply(lambda x: x.count('E'))\n    df['Cabin_class_F']=df['Cabin'].apply(lambda x: x.count('F'))\n    df['Cabin_class_G']=df['Cabin'].apply(lambda x: x.count('G'))\n    df['Cabin_class_T']=df['Cabin'].apply(lambda x: x.count('T'))\n    df.drop(['Cabin'],axis=1,inplace=True)\n    return df\n\n#in this case the prediction process requires some extra effort\ndef cabin_imp_predict(data,model,data_columns=['Cabin_class_A','Cabin_class_B','Cabin_class_C','Cabin_class_D','Cabin_class_E','Cabin_class_F','Cabin_class_G','Cabin_class_T']):\n    output=pd.DataFrame(model.predict(data),columns=data_columns)\n    #if the output is not a natural number, make it (particularly linear regressor based models have this behaviour )\n    output=output.apply(lambda x: np.floor(abs(x)))\n    #this are frequencies so the must be integers\n    output=output.astype('int32')\n    return output\n\ndf_lab_train=cabin_transform(df_lab_train)\ndf_lab_test=cabin_transform(df_lab_test)\nX_alt_train=df_lab_train.drop(['Cabin_class_A','Cabin_class_B','Cabin_class_C','Cabin_class_D','Cabin_class_E','Cabin_class_F','Cabin_class_G','Cabin_class_T'],axis=1)\nX_alt_test=df_lab_test.drop(['Cabin_class_A','Cabin_class_B','Cabin_class_C','Cabin_class_D','Cabin_class_E','Cabin_class_F','Cabin_class_G','Cabin_class_T'],axis=1)\ny_cabin_train=df_lab_train[['Cabin_class_A','Cabin_class_B','Cabin_class_C','Cabin_class_D','Cabin_class_E','Cabin_class_F','Cabin_class_G','Cabin_class_T']]\ny_cabin_test=df_lab_test[['Cabin_class_A','Cabin_class_B','Cabin_class_C','Cabin_class_D','Cabin_class_E','Cabin_class_F','Cabin_class_G','Cabin_class_T']]\n\n#trying some models\n#cabin_model =  LinearRegression()\n#cabin_model =  Ridge()\n#cabin_model =  MultiTaskLassoCV(cv=5)\n#cabin_model = MultiTaskElasticNet(0.01)\n#BEST, you can tune this parameters by using an optuna study (specified bellow)\ncabin_param={'verbosity':0,'max_depth': 4, 'learning_rate': 0.09953343713958455, 'n_estimators': 5570, 'min_child_weight': 1, 'colsample_bytree': 0.3587655598296461, 'subsample': 0.20853173928072777, 'reg_alpha': 0.0001300620787488413, 'reg_lambda': 1.354795119410994}\n\n#if you use the tuned model set this to True and the next to False\nif False:\n    try:\n        #if cabin_param is missing this won't work\n        cabin_model = RegressorChain(base_estimator=XGBRegressor(**cabin_param))\n    except:\n        cabin_model = RegressorChain(base_estimator=XGBRegressor())\nif True:\n    cabin_model = RegressorChain(base_estimator=XGBRegressor())\n#training the sub model and saving the predictions\ncabin_model.fit(X_alt_train,y_cabin_train)\ny_cabin_train_fit=cabin_imp_predict(X_alt_train,cabin_model)\ny_cabin_test_fit=cabin_imp_predict(X_alt_test,cabin_model)","metadata":{"execution":{"iopub.status.busy":"2022-07-25T00:41:21.905918Z","iopub.execute_input":"2022-07-25T00:41:21.906178Z","iopub.status.idle":"2022-07-25T00:41:26.657358Z","shell.execute_reply.started":"2022-07-25T00:41:21.906145Z","shell.execute_reply":"2022-07-25T00:41:26.656814Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## Verifying the cabin submodel perfomance\n\nIn this section some metrics are computed to measure the submodel performance. You can delete this cell if you are pleased with the current submodel perfomance already or if you chose not use the cabin submodel.","metadata":{}},{"cell_type":"code","source":"from sklearn.metrics import mean_squared_log_error\nfrom sklearn.metrics import mean_absolute_error\nfrom sklearn.metrics import max_error\nprint('train set \\n MSLE {} \\n MAE {}'.format(mean_squared_log_error(y_cabin_train_fit,y_cabin_train),mean_absolute_error(y_cabin_train_fit,y_cabin_train)))\nprint('feature \\t max_error')\nfor col in y_cabin_test.columns:\n    print('{} \\t {}'.format(col,max_error(y_cabin_train_fit[col],y_cabin_train[col])/max(1,y_cabin_train_fit[col].sum()) ))\n\nprint('test set \\n MSLE {} \\n MAE {}'.format(mean_squared_log_error(y_cabin_test_fit,y_cabin_test),mean_absolute_error(y_cabin_test_fit,y_cabin_test)))\nprint('feature \\t max_error')    \nfor col in y_cabin_test.columns:\n    print('{} \\t {}'.format(col,max_error(y_cabin_test_fit[col],y_cabin_test[col])/max(1,y_cabin_test_fit[col].sum())  )) \n","metadata":{"execution":{"iopub.status.busy":"2022-07-25T00:41:26.658619Z","iopub.execute_input":"2022-07-25T00:41:26.658942Z","iopub.status.idle":"2022-07-25T00:41:26.693600Z","shell.execute_reply.started":"2022-07-25T00:41:26.658913Z","shell.execute_reply":"2022-07-25T00:41:26.691894Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"##  Age imputation(optional)\nI tried a simple and naive imputation method for *Age* which did not improve the overall performance in any case based on running a \nsimulation based on the approximation that *Age* is normally distributed (which might not be true) and on the sample parameter distribution\nto fill in the gaps.","metadata":{}},{"cell_type":"code","source":"def ages_imp_sim(df):\n    df_mean=df['Age'].mean()\n    df_std=df['Age'].std()\n    ages_train=list(np.random.normal(loc=df_mean,scale=df_std,size=sum(df['Age'].isna())))\n    ages_train = [x  if 0<x else df_mean for x in ages_train]\n    df.loc[df['Age'].isna(),['Age']] = ages_train\n    return df\ndef ages_imp_sim_alt(df):\n    df_mean=df['Age'].mean()\n    df_std=df['Age'].std()\n    ages_train=list(np.random.normal(loc=df_mean,scale=df_std,size=sum(df['Age'].isna())))\n    ages_train = [x  if 0<x else df_mean for x in ages_train]\n    return ages_train\n\nmissing_age_train = df_train['Age'].isna()\nmissing_age_train = missing_age_train.replace({False:0,True:1})\nmissing_age_test = df_test['Age'].isna()\nmissing_age_test = missing_age_test.replace({False:0,True:1})","metadata":{"execution":{"iopub.status.busy":"2022-07-25T00:41:26.695856Z","iopub.execute_input":"2022-07-25T00:41:26.696347Z","iopub.status.idle":"2022-07-25T00:41:26.712549Z","shell.execute_reply.started":"2022-07-25T00:41:26.696286Z","shell.execute_reply":"2022-07-25T00:41:26.710953Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Data processing\n\nHere other missing values are replaced with meaningful values as well as transformations (like catsting) needed to handle the data more effectively.","metadata":{}},{"cell_type":"code","source":"def process_data(df_train,df_test=None):\n    #df_train['Pclass']=df_train['Pclass'].apply(lambda x: '1' if x==1 else ('2' if x==2 else '3'))\n    df_train['Sex']=df_train['Sex'].apply(lambda x: '1' if x=='female' else '0')\n    #Naive imputation, the case of age is worse so that it changes the distribution completely\n    df_train['Cabin']=df_train['Cabin'].fillna('None')\n    average_age=df_train['Age'].mean()\n    mode_fare=df_train['Fare'].mean()\n    df_train['Age']=df_train['Age'].fillna(average_age)\n    #Failed\n    #df_train.loc[df_train['Age'].isna(),['Age']]=ages_imp_sim_alt(df_train) \n    df_train['Fare']=df_train['Fare'].fillna(mode_fare)\n    df_train['Embarked']=df_train['Embarked'].fillna('S')\n    for name in ['Pclass','Sex','Embarked']:\n            df_train[name] = df_train[name].astype(\"category\")\n    if df_test is not None:\n        #df_test['Pclass']=df_test['Pclass'].apply(lambda x: '1' if x==1 else ('2' if x==2 else '3'))\n        df_test['Sex']=df_test['Sex'].apply(lambda x: '1' if x=='female' else '0')\n        #Naive imputation, the case of age is worse so that it changes the distribution completely\n        df_test['Cabin']=df_test['Cabin'].fillna('None')\n        average_age=df_test['Age'].mean()\n        mode_fare=df_test['Fare'].mean()\n        df_test['Age']=df_test['Age'].fillna(average_age)\n        #Failed\n        #df_test.loc[df_test['Age'].isna(),['Age']]=ages_imp_sim_alt(df_test) \n        df_test['Fare']=df_test['Fare'].fillna(mode_fare)\n        df_test['Embarked']=df_test['Embarked'].fillna('S')\n        for name in ['Pclass','Sex','Embarked']:\n        \n            df_test[name] = df_test[name].astype(\"category\")\n\n    return df_train,df_test    \n\ndf_train,df_test= process_data(df_train,df_test)\nprint(df_train.dtypes)\nprint(df_test.dtypes)","metadata":{"execution":{"iopub.status.busy":"2022-07-25T00:41:26.714072Z","iopub.execute_input":"2022-07-25T00:41:26.716537Z","iopub.status.idle":"2022-07-25T00:41:26.743683Z","shell.execute_reply.started":"2022-07-25T00:41:26.716483Z","shell.execute_reply":"2022-07-25T00:41:26.742889Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Besides the imputation methods from above here is deployed the cabin sub model. This predictions ca be used in the feature creation process (down bellow)","metadata":{}},{"cell_type":"code","source":"#by using the processed data, generate the prediction of the submodel\ndf_train_aux=cabin_imp_predict(df_train[['Pclass', 'Age', 'SibSp', 'Parch', 'Fare']],cabin_model)\ndf_test_aux=cabin_imp_predict(df_test[['Pclass', 'Age', 'SibSp', 'Parch', 'Fare']],cabin_model)\n","metadata":{"execution":{"iopub.status.busy":"2022-07-25T00:41:26.744964Z","iopub.execute_input":"2022-07-25T00:41:26.745937Z","iopub.status.idle":"2022-07-25T00:41:26.831629Z","shell.execute_reply.started":"2022-07-25T00:41:26.745894Z","shell.execute_reply":"2022-07-25T00:41:26.831066Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Encoding the data\n\nIn this case for the features Ticket and Cabin I broke the rule of the train test indepence since some values appear in both and they have not a quite natural way to encode even when intuitively some clusterings might arise (like cabins which have a C o D in them or even two C's). So basically I only matched all the different values that appear on both sets to different integers. If there is any other relation between them and other features the model should be able to spot it. \nThe submodel might also have provided a meaningful coding but I kept the original Feature as well.\n","metadata":{}},{"cell_type":"code","source":"#ENCODING FUNCTIONS\n#saving data sizes, cabin and ticket levels found in the dataset\nm_train=df_train.shape[0]\nm_test=df_test.shape[0]\ncabin_levels_train=df_train['Cabin']\ncabin_levels_test=df_test['Cabin']\nticket_levels_train=df_train['Ticket']\nticket_levels_test=df_test['Ticket']\n#this function only encodes Cabin and Ticket,in that they need some special care\ndef encode_spe_cat(df_train,df_test):\n    train_levels=pd.concat([cabin_levels_train,ticket_levels_train],names=['Cabin','Tickets'],axis=1)\n    #define whether the data belongs to the training set or the test set\n    train_levels['Set']=np.full(m_train,fill_value=0)\n    test_levels=pd.concat([cabin_levels_test,ticket_levels_test],names=['Cabin','Tickets'],axis=1)\n    test_levels['Set']=np.full(m_test,fill_value=1)\n    #merge all the values in a dataframe, we create a comprehensive encoding that involves all the data\n    df_levels = pd.concat([train_levels,test_levels],axis=0,ignore_index=True)\n    df_levels.sort_values(['Cabin','Ticket'],inplace=True)\n    \n    #defining types\n    df_levels['Cabin']=df_levels['Cabin'].astype(\"category\")\n    df_levels['Ticket']=df_levels['Ticket'].astype(\"category\")\n    df_levels['Cabin_codes']=df_levels['Cabin'].cat.codes\n    df_levels['Ticket_codes']=df_levels['Ticket'].cat.codes\n\n    #once, the  codes have been computed for the whole data, separate them again \n    train_codes=df_levels.loc[df_levels['Set']==0,['Cabin','Ticket','Cabin_codes','Ticket_codes']]\n    test_codes=df_levels.loc[df_levels['Set']==1,['Cabin','Ticket','Cabin_codes','Ticket_codes']]\n\n    ticket_train_codes=train_codes[['Ticket','Ticket_codes']]\n    ticket_test_codes=test_codes[['Ticket','Ticket_codes']]\n    cabin_train_codes=train_codes[['Cabin','Cabin_codes']]\n    cabin_test_codes=test_codes[['Cabin','Cabin_codes']]\n    \n    #finally replace the strings with their codes\n    df_train['Ticket']=df_train['Ticket'].replace(dict(ticket_train_codes.to_records(index=False)))\n    df_train['Cabin']=df_train['Cabin'].replace(dict(cabin_train_codes.to_records(index=False)))\n    df_test['Ticket']=df_test['Ticket'].replace(dict(ticket_test_codes.to_records(index=False)))\n    df_test['Cabin']=df_test['Cabin'].replace(dict(cabin_test_codes.to_records(index=False)))\n    \n    return df_train,df_test\n\n#this methods does the whole encoding\ndef encode_cat(df_train,df_test):\n    for colname in ['Pclass','Name','Sex','Embarked']:\n        df_train[colname]=df_train[colname].astype(\"category\")\n        df_train[colname] = df_train[colname].cat.codes\n    for colname in ['Pclass','Name','Sex','Embarked']:\n    \n        df_test[colname]=df_test[colname].astype(\"category\")\n        df_test[colname] = df_test[colname].cat.codes\n    df_train,df_test = encode_spe_cat(df_train,df_test)\n    \n    return df_train,df_test","metadata":{"execution":{"iopub.status.busy":"2022-07-25T00:41:26.832909Z","iopub.execute_input":"2022-07-25T00:41:26.836016Z","iopub.status.idle":"2022-07-25T00:41:26.858286Z","shell.execute_reply.started":"2022-07-25T00:41:26.835950Z","shell.execute_reply":"2022-07-25T00:41:26.857633Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Resampling\n\nResampling becomes critical for imbalanced sets in which there are way more records of one class than from others (in this case less survivors than survivors) when we use neural networks. Boosted trees aren't so much affected, in that case resampling is not needed. In a neural network, not doing this might cause a curious and annoying effect: the neural network always predicts they larger class which  results in a accuracy equal to the ratio of elements of that class in the current dataset. The **ROC curve** is almost a straight line (random classifier) in this case and the **confusion matrix** shows that metrics like recall or precision are not as good as expected.","metadata":{}},{"cell_type":"code","source":"#Let's see the survivors/non rurvivors ratio on the train dataset\nprint('Number of survivors and non survivors in the training set')\nprint(df_train['Survived'].value_counts())\n\ndef resam_custom_naive(df_train,verbose=False):\n    zeros_train={'length':df_train['Survived'].value_counts()[0]}\n    ones_train={'length':df_train['Survived'].value_counts()[1]}\n    if verbose:\n        print('0\\'s: {}\\t1\\'s: {} '.format(zeros_train['length'],ones_train['length']))\n    zeros_train['content']=df_train.loc[df_train['Survived']==0,:]\n    ones_train['content']=df_train.loc[df_train['Survived']==1,:]\n    df_train_resam=pd.concat([ones_train['content'].sample(zeros_train['length']-ones_train['length']),df_train])\n    return df_train_resam\n\n#some data points might be entirely missing, so use only with large datasets\ndef resam_custom_batch(df_train,sample_size=100,num_sample=4,sample_idx=False,verbose=False,include_sample_idx=False):\n    zeros_train={'length':df_train['Survived'].value_counts()[0]}\n    ones_train={'length':df_train['Survived'].value_counts()[1]}\n    if verbose:\n        print('0\\'s: {}\\t1\\'s: {} '.format(zeros_train['length'],ones_train['length']))\n    zeros_train['content']=df_train.loc[df_train['Survived']==0,:]\n    sample_size=sample_size//2\n    ones_train['content']=df_train.loc[df_train['Survived']==1,:]\n    for id_sample in range(1,1+num_sample):\n        df_sample=pd.concat([ones_train['content'].sample(sample_size),zeros_train['content'].sample(sample_size)])\n        if include_sample_idx:\n            df_sample['num_sample']=np.full(sample_size*2,fill_value=id_sample)\n        if id_sample==1:\n            df_train_resam=df_sample\n        else:    \n            df_train_resam=pd.concat([df_train_resam,df_sample])\n    return df_train_resam\n\ndef resam_custom(df_train,sample_size=100,num_sample=4,sample_idx=False,verbose=False):\n    df_train_resam = resam_custom_naive(df_train,verbose)\n    df_extra = resam_custom_batch(df_train,sample_size,num_sample,sample_idx,verbose)\n    df_train_resam =pd.concat([df_train_resam,df_extra])\n    return df_train_resam\n\nresample=True\nif resample:\n    df_train_resam = resam_custom(df_train,sample_size=200,num_sample=6)\nelse:\n    df_train_resam = df_train","metadata":{"execution":{"iopub.status.busy":"2022-07-25T00:41:26.859619Z","iopub.execute_input":"2022-07-25T00:41:26.860011Z","iopub.status.idle":"2022-07-25T00:41:26.935112Z","shell.execute_reply.started":"2022-07-25T00:41:26.859970Z","shell.execute_reply":"2022-07-25T00:41:26.933867Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Feature engineering\n## Feature selection\n\n### Mutual information\nThis function is inspired in that one found at the **Feature engineering course** the main difference is that uses **mutual_info_classif** instead of **mutual_info_regression**, since this is a classification problem.\n","metadata":{}},{"cell_type":"code","source":"from sklearn.feature_selection import mutual_info_regression\nfrom sklearn.feature_selection import mutual_info_classif\n\ncolumns_relation = pd.DataFrame(data=[('PassengerId',False),('Pclass',True),('Name',False),('Sex',True),('Age',False),\n                             ('SibSp',False),('Parch',False),('Ticket',True),('Fare',False),('Cabin',False),('Embarked',True)],columns=['column','is_discrete'])\ncolumns_relation=columns_relation.set_index('column')\n\n#functions to measure perfomance by using cross-val method\ndef make_mi_scores_classif(X, y,discrete_features=columns_relation['is_discrete']):\n    X = X.copy()\n    for colname in X.select_dtypes([\"object\", \"category\"]):\n        X[colname], _ = X[colname].factorize()\n    # All discrete features should now have integer dtypes\n\n    mi_scores =mutual_info_classif(X, y, discrete_features=discrete_features,random_state=0)\n    #mi_scores = mutual_info_regression(X, y,random_state=0)\n    mi_scores = pd.Series(mi_scores, name=\"MI Scores\", index=X.columns)\n    mi_scores = mi_scores.sort_values(ascending=False)\n    return mi_scores\n\n\ndef make_mi_scores(X, y,discrete_features=columns_relation['is_discrete']):\n    X = X.copy()\n    for colname in X.select_dtypes([\"object\", \"category\"]):\n        X[colname], _ = X[colname].factorize()\n    # All discrete features should now have integer dtypes\n\n    mi_scores = mutual_info_regression(X, y, discrete_features=discrete_features,random_state=0)\n    #mi_scores = mutual_info_regression(X, y,random_state=0)\n    mi_scores = pd.Series(mi_scores, name=\"MI Scores\", index=X.columns)\n    mi_scores = mi_scores.sort_values(ascending=False)\n    return mi_scores\n\nX_train = df_train_resam.drop(columns=['Survived'])\ny_train = df_train_resam['Survived']\nX_test = df_test.copy()\n\n\nX_train,X_test= encode_cat(X_train,X_test)\nmake_mi_scores(X_train, y_train)\n","metadata":{"execution":{"iopub.status.busy":"2022-07-25T00:41:26.936149Z","iopub.execute_input":"2022-07-25T00:41:26.936403Z","iopub.status.idle":"2022-07-25T00:41:27.984815Z","shell.execute_reply.started":"2022-07-25T00:41:26.936369Z","shell.execute_reply":"2022-07-25T00:41:27.983898Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"\n## Models and score functions to use\n\nImportant metrics used are:\n\n* accuracy\n* balanced accuracy (which is more informative and accurate provided the 0's an 1's ratio [see sickit learn documentation on this ](https://scikit-learn.org/stable/modules/generated/sklearn.metrics.balanced_accuracy_score.html#sklearn.metrics.balanced_accuracy_score))\n* precision\n* recall\n\nAlso log_loss is used since it changes smoothly with samples taken from the train set. Besides this, we provide the confusion matrix on the cross validation set (not defined yet) to have a graphical reference to this metrics.\n","metadata":{}},{"cell_type":"code","source":"\nfrom sklearn.metrics import log_loss\nfrom sklearn.metrics import accuracy_score\nfrom sklearn.metrics import balanced_accuracy_score\nfrom sklearn.metrics import precision_score\nfrom sklearn.metrics import recall_score\nfrom sklearn.model_selection import train_test_split\nclass Score:\n    def __init__(self,cv_scores,train_scores):\n        self.cv_scores = cv_scores\n        self.train_scores = train_scores\n    def mean_cv(self):\n        return np.mean(self.cv_scores)\n    def mean_train(self):\n        return np.mean(self.train_scores)     ","metadata":{"execution":{"iopub.status.busy":"2022-07-25T00:41:27.987840Z","iopub.execute_input":"2022-07-25T00:41:27.988069Z","iopub.status.idle":"2022-07-25T00:41:27.996064Z","shell.execute_reply.started":"2022-07-25T00:41:27.988040Z","shell.execute_reply":"2022-07-25T00:41:27.994818Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## Feature creation\n\n\n\nHere some functions are defined in order to create new features from the original ones. First, a function to remove all the numbers and spaces of a string is developed, if this process returns an empty strings it defaults to None. This function is applied to *Cabin* values to return values like C,A,CC... which are regarded as meaningful: A,B and C have some relation to the 1st class while other letters are related to lower classes. In the case of Tickets the results are a bit harder to interprete yet they have proved to be useful.\n\nThe functions <code>add_columns</code> and <code>drop_columns</code> hve been added to change the feature set dinamically whithout having to manually repeat all these procedures; nevertheless they need to pass an array <code>columns_relation</code> which is used by the <code>make_mi_scores_classif</code> function to tell apart discrete features.\n\n\n* **ReCabin** is obtained by using <code>pr-string</code> on *Cabin*\n* **ReTicket** is obtained by using pr-string on *Ticket*\n* **lnFare** is the natural logarithm of *Fare* (0 if the input is 0)\n* **Outlier** is 1 is Fare is 0 which is not an expected value of Fare; and 0 otherwise.\n* **Cabin_class_A,Cabin_class_B,**... count how many *A*'s *B*´s ... there are in *Cabin* or if there are None in case of NA. This can be derived\n  either from the *Cabin* feature itself or from the cabin submodel.\n* *Age_missing* expands the imputation used on *Age* by indicating  missing values.\n","metadata":{}},{"cell_type":"code","source":"#categorical clustering\nimport re\ndef pr_string(input_string):\n    all_matches = re.findall(r'[^\\s0-9]*',input_string)\n    output = ''.join(all_matches)\n    if output=='':\n        return 'None'\n    else:\n        return output\n\n\ndef add_columns(X,data,columns,columns_relation=columns_relation):\n    X=pd.concat([X,data],axis=1)\n    columns_relation_add=pd.DataFrame(data=list(columns.items()),columns=['column','is_discrete'])\n    columns_relation_add=columns_relation_add.set_index('column')\n    \n    columns_relation= pd.concat([columns_relation,columns_relation_add],ignore_index=False)\n    return X,columns_relation\ndef drop_columns(X,columns,columns_relation=columns_relation):\n    X=X.drop(columns=columns)\n    columns_relation=columns_relation.drop(labels=columns)\n    return X,columns_relation\n    \ndef feature_creation(X_train,X_test=None,columns_relation=columns_relation):\n    #Case whole model:Don't drop any column\n    #to_drop_columns=[]\n    #Case standard model\n    #to_drop_columns=['PassengerId','Name','Embarked','lnFare','Outlier','SibSp','Cabin_class_T','Cabin_class_G','Cabin_class_E']\n    #baseline\n    #to_drop_columns=[ 'lnFare', 'ReTicket', 'ReCabin','Outlier', 'Cabin_class_A', 'Cabin_class_B', 'Cabin_class_C','Cabin_class_D', 'Cabin_class_E', 'Cabin_class_F', 'Cabin_class_G','Cabin_class_T', 'Cabin_class_None', 'Age_missing']\n    to_drop_columns=['PassengerId','Name','Cabin_class_A', 'Cabin_class_B', 'Cabin_class_C','Cabin_class_D', 'Cabin_class_E', 'Cabin_class_F', 'Cabin_class_G','Cabin_class_T']\n \n    to_add_columns={'lnFare':False,'ReTicket':True,'ReCabin':True,'Cabin_class_A':False,'Cabin_class_B':False,'Cabin_class_C':False,'Cabin_class_D':False,'Cabin_class_E':False,'Cabin_class_F':False,'Cabin_class_G':False,'Cabin_class_T':False,'Cabin_class_None':False,'Age_missing':False,'Outlier':False}\n    \n    old_columns_relation=columns_relation.copy()\n    X_add_train = pd.DataFrame()\n    X_add_train['lnFare']= X_train['Fare'].apply(lambda x:0 if x ==0 else np.log(x))\n    \n    all_re_ticket_levels=['None','A.', 'A..', 'A./.', 'A/', 'A/.', 'A/S', 'AQ/', 'AQ/.', 'C', 'C.A.', 'C.A./SOTON', 'CA', 'CA.', 'F.C.', 'F.C.C.', 'Fa', 'LINE', 'LP' , 'P/PP', 'PC', 'PP', 'S.C./A..', 'S.C./PARIS', 'S.O./P.P.', 'S.O.C.', 'S.O.P.', 'S.P.', 'S.W./PP', 'SC', 'SC/A', 'SC/A.', 'SC/AH', 'SC/AHBasle', 'SC/PARIS', 'SC/Paris', 'SCO/W', 'SO/C', 'SOTON/O', 'SOTON/O.Q.', 'SOTON/OQ', 'STON/O.', 'STON/OQ.', 'SW/PP', 'W./C.', 'W.E.P.', 'W/C', 'WE/P']\n    all_re_cabin_levels=['None','A', 'B', 'BB', 'BBB', 'BBBB', 'C', 'CC', 'CCC', 'D', 'DD', 'E', 'EE', 'F', 'FE', 'FG', 'G', 'T']\n\n    all_re_ticket_levels_encoding = dict(zip(all_re_ticket_levels,range(48)))\n    all_re_cabin_levels_encoding = dict(zip(all_re_cabin_levels,range(18)))\n\n    X_add_train['ReTicket']=ticket_levels_train.apply(pr_string)\n    X_add_train['ReCabin']=cabin_levels_train.apply(pr_string)\n    \n    X_add_train['ReTicket']=X_add_train['ReTicket'].replace(all_re_ticket_levels_encoding)\n    X_add_train['ReCabin']=X_add_train['ReCabin'].replace(all_re_cabin_levels_encoding)\n    \n    X_add_train['Outlier']=X_train['Fare'].apply(lambda x:1 if x==0 else 0 )\n    \n    #if set to True this will use the resul of the Cabin pseudo imputation method (cabin submodel)\n    cabin_class_imp=False\n    if not cabin_class_imp:\n        X_add_train['Cabin_class_A']= cabin_levels_train.apply(lambda x:x.count('A'))\n        X_add_train['Cabin_class_B']= cabin_levels_train.apply(lambda x:x.count('B'))   \n        X_add_train['Cabin_class_C']= cabin_levels_train.apply(lambda x:x.count('C'))   \n        X_add_train['Cabin_class_D']= cabin_levels_train.apply(lambda x:x.count('D'))\n        X_add_train['Cabin_class_E']= cabin_levels_train.apply(lambda x:x.count('E'))\n        X_add_train['Cabin_class_F']= cabin_levels_train.apply(lambda x:x.count('F'))\n        X_add_train['Cabin_class_G']= cabin_levels_train.apply(lambda x:x.count('G'))\n        X_add_train['Cabin_class_T']= cabin_levels_train.apply(lambda x:x.count('T'))\n    if cabin_class_imp:\n        X_add_train['Cabin_class_A']= df_train_aux['Cabin_class_A']\n        X_add_train['Cabin_class_B']= df_train_aux['Cabin_class_B']   \n        X_add_train['Cabin_class_C']= df_train_aux['Cabin_class_C']   \n        X_add_train['Cabin_class_D']= df_train_aux['Cabin_class_D']\n        X_add_train['Cabin_class_E']= df_train_aux['Cabin_class_E']\n        X_add_train['Cabin_class_F']= df_train_aux['Cabin_class_F']\n        X_add_train['Cabin_class_G']= df_train_aux['Cabin_class_G']\n        X_add_train['Cabin_class_T']= df_train_aux['Cabin_class_T']\n    \n    #The none level is added last\n    X_add_train['Cabin_class_None']= cabin_levels_train.apply(lambda x:1 if x=='None' else 0)\n    X_add_train['Age_missing']=missing_age_train\n\n\n    X_train,columns_relation = add_columns(X_train,X_add_train,to_add_columns,columns_relation=columns_relation)\n    if not len(to_drop_columns)==0:\n        X_train,columns_relation= drop_columns(X_train, to_drop_columns,columns_relation=columns_relation)\n    \n    if X_test is not None:\n        X_add_test = pd.DataFrame()\n        X_add_test['lnFare']= X_test['Fare'].apply(lambda x:0 if x ==0 else np.log(x))\n   \n        X_add_test['ReTicket']=ticket_levels_test.apply(pr_string)\n        X_add_test['ReCabin']=cabin_levels_test.apply(pr_string)\n        \n        X_add_test['ReTicket']=X_add_test['ReTicket'].replace(all_re_ticket_levels_encoding)\n        X_add_test['ReCabin']=X_add_test['ReCabin'].replace(all_re_cabin_levels_encoding)\n        \n        X_add_test['Outlier']=X_test['Fare'].apply(lambda x:1 if x==0 else 0 )\n        \n        if not cabin_class_imp:\n            X_add_test['Cabin_class_A']= cabin_levels_test.apply(lambda x:x.count('A'))\n            X_add_test['Cabin_class_B']= cabin_levels_test.apply(lambda x:x.count('B'))   \n            X_add_test['Cabin_class_C']= cabin_levels_test.apply(lambda x:x.count('C'))   \n            X_add_test['Cabin_class_D']= cabin_levels_test.apply(lambda x:x.count('D'))\n            X_add_test['Cabin_class_E']= cabin_levels_test.apply(lambda x:x.count('E'))\n            X_add_test['Cabin_class_F']= cabin_levels_test.apply(lambda x:x.count('F'))\n            X_add_test['Cabin_class_G']= cabin_levels_test.apply(lambda x:x.count('G'))\n            X_add_test['Cabin_class_T']= cabin_levels_test.apply(lambda x:x.count('T'))\n    \n        if cabin_class_imp: \n            X_add_test['Cabin_class_A']= df_test_aux['Cabin_class_A']\n            X_add_test['Cabin_class_B']= df_test_aux['Cabin_class_B']   \n            X_add_test['Cabin_class_C']= df_test_aux['Cabin_class_C']   \n            X_add_test['Cabin_class_D']= df_test_aux['Cabin_class_D']\n            X_add_test['Cabin_class_E']= df_test_aux['Cabin_class_E']\n            X_add_test['Cabin_class_F']= df_test_aux['Cabin_class_F']\n            X_add_test['Cabin_class_G']= df_test_aux['Cabin_class_G']\n            X_add_test['Cabin_class_T']= df_test_aux['Cabin_class_T']\n    \n        X_add_test['Cabin_class_None']= cabin_levels_test.apply(lambda x:1 if x=='None' else 0)\n        X_add_test['Age_missing']=missing_age_test\n\n\n        X_test,old_columns_relation = add_columns(X_test,X_add_test,to_add_columns,columns_relation=old_columns_relation)\n        if not len(to_drop_columns)==0:\n            X_test,_= drop_columns(X_test, to_drop_columns,columns_relation=old_columns_relation)\n        \n    \n    return X_train,X_test,columns_relation","metadata":{"execution":{"iopub.status.busy":"2022-07-25T00:41:27.997949Z","iopub.execute_input":"2022-07-25T00:41:27.998725Z","iopub.status.idle":"2022-07-25T00:41:28.040667Z","shell.execute_reply.started":"2022-07-25T00:41:27.998677Z","shell.execute_reply":"2022-07-25T00:41:28.039412Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Just as a reference we generate an histogram of the *Cabin* one hot encoding features (in the case of the cabin submodel this has a different meaning)","metadata":{}},{"cell_type":"code","source":"X_train,X_test,columns_relation = feature_creation(X_train,X_test,columns_relation=columns_relation)\ntry:\n    print(X_train.loc[:,[col for col in  X_train.columns if col.startswith('Cabin_class_') ]].sum())\nexcept:\n    print('NA, not all the features specified are found in the current dataset')\n    ","metadata":{"execution":{"iopub.status.busy":"2022-07-25T00:41:28.042385Z","iopub.execute_input":"2022-07-25T00:41:28.042908Z","iopub.status.idle":"2022-07-25T00:41:28.163623Z","shell.execute_reply.started":"2022-07-25T00:41:28.042865Z","shell.execute_reply":"2022-07-25T00:41:28.162524Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## Feature selection \n\nThe feature sets that were tested are these:\n\n<div>\n<table>\n<tr><td style=\"background-color:#30db77; font-weight:bold;\">Adding the most representative Cabin features(according to mutual information):</td><td style=\"word-spacing:10px\">\nPclass\nSex\nAge\nParch\nTicket\nFare\nCabin\nReTicket\nReCabin\nCabin_class_A\nCabin_class_B\nCabin_class_C\nCabin_class_D\nCabin_class_F\nCabin_class_None\nAge_missing\n</td></tr>\n<tr><td style=\"background-color:#30db77; font-weight:bold;\">Adding more Cabin and Fare related features:</td><td style=\"word-spacing:10px\">\nPclass\nSex\nAge\nSibSp\nParch\nTicket\nFare\nCabin\nEmbarked\nlnFare\nReTicket\nReCabin\nOutlier\nCabin_class_None\nAge_missing\n</td></tr>\n<tr><td style=\"background-color:#30db77; font-weight:bold;\">Baseline:</td><td style=\"word-spacing:10px\">\nPassengerId\nPclass\nName\nSex\nAge\nSibSp\nParch\nTicket\nFare\nCabin\nEmbarked\n</td></tr>\n<tr><td style=\"background-color:#30db77; font-weight:bold;\">Model with all created features:</td><td style=\"word-spacing:10px\">\nPassengerId\nPclass\nName\nSex\nAge\nSibSp\nParch\nTicket\nFare\nCabin\nEmbarked\nlnFare\nReTicket\nReCabin\nOutlier\nCabin_class_A\nCabin_class_B\nCabin_class_C\nCabin_class_D\nCabin_class_E\nCabin_class_F\nCabin_class_G\nCabin_class_T\nCabin_class_None\nAge_missing\n</td></tr>\n</table>\n</div>\n\nIn the case of Cabin features the result might vary depending on whether the cabin submodel was used or not.\nThe mutual information was computed in all cases. \nWe can use different statistics along side the mutual information, for instance the **chi square statistic** (variance) in order to allow us to select the best features to use. In this case features with large values of variances associated with small **p values** are preferred. ","metadata":{}},{"cell_type":"code","source":"from sklearn.feature_selection import chi2\n\nchi_stat_val,p_values=chi2(X_train,y_train)\nchi_analysis = pd.DataFrame(zip(chi_stat_val,p_values),index=X_train.columns,columns=['chi_stat','p_val']) \n\nprint(chi_analysis.sort_values(by=['p_val']))\n\n\nprint(make_mi_scores_classif(X_train, y_train,discrete_features=columns_relation['is_discrete']))\n#alternatively you can use the boosted tree score functions\n#scores_data=customized_score_cross_val(X_train,y_train,folds=20,threshold=0.7)\n","metadata":{"execution":{"iopub.status.busy":"2022-07-25T00:41:28.164695Z","iopub.execute_input":"2022-07-25T00:41:28.164887Z","iopub.status.idle":"2022-07-25T00:41:28.346063Z","shell.execute_reply.started":"2022-07-25T00:41:28.164862Z","shell.execute_reply":"2022-07-25T00:41:28.344919Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Model selection\n\nIn this section one can choose to use either a **neural network model**, a **boosted tree model** (xgb) or a **combined model**.\nThe model_inclusive option uses both models, if this option is not set to true only one model will be trained and analyzed and the predictions on the test set will only be based on this model. This option also affects the plots.","metadata":{}},{"cell_type":"code","source":"#You can choose a neural network model  or a xgb model  or even both\nneural_network=False\nmodel_inclusive=False\n#if model_inclusive is set to True the result is a mixed model with the probabilities computed as a weighted combination\n#of the two models\nif not model_inclusive:\n    boosted_tree_model= not neural_network\n\n#if set to true both model are used therefire set to true, possibly overriding previous values   \nif  model_inclusive:\n    neural_network=True\n    boosted_tree_model=True\n    \nif (not resample) and neural_network:\n    print('Warning:The model might underperform because the dataset is imbalanced and neural network models are heavily affected by this fact!')\n","metadata":{"execution":{"iopub.status.busy":"2022-07-25T00:41:28.347683Z","iopub.execute_input":"2022-07-25T00:41:28.347912Z","iopub.status.idle":"2022-07-25T00:41:28.354278Z","shell.execute_reply.started":"2022-07-25T00:41:28.347883Z","shell.execute_reply":"2022-07-25T00:41:28.353211Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## Neural Network\n\nFor the neural network model one has to rescale the data to optimize the training process. Besides, before scaling we split the data into the cross validation set and the training set (even for the xgb model)","metadata":{}},{"cell_type":"code","source":"from sklearn.preprocessing import StandardScaler\nX_train, X_cv, y_train, y_cv = train_test_split(X_train, y_train, test_size=0.2)\nif neural_network:\n    scaler_dataset = StandardScaler()\n    scaler_dataset.fit(X_train)\n    X_train_scaled=pd.DataFrame(scaler_dataset.transform(X_train),columns=X_train.columns) \n    X_cv_scaled=pd.DataFrame(scaler_dataset.transform(X_cv),columns=X_cv.columns) ","metadata":{"execution":{"iopub.status.busy":"2022-07-25T00:41:28.355785Z","iopub.execute_input":"2022-07-25T00:41:28.356018Z","iopub.status.idle":"2022-07-25T00:41:28.370882Z","shell.execute_reply.started":"2022-07-25T00:41:28.355988Z","shell.execute_reply":"2022-07-25T00:41:28.369540Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Here we instantiate the neural network model by using the implementation found in tensorflow. Besides we can also include metrics like (binary) accuracy\nin our model,this metrics aren't optimized by the training process (at least not actively but rather as a side effect). We can chose other [metrics](https://keras.io/api/metrics/accuracy_metrics/) as seen in the \nkeras documentation. \nWe use a [adam optimizer](https://www.tensorflow.org/api_docs/python/tf/keras/optimizers/Adam) for finding the values for which\nthe loss function ([cross entropy](https://machinelearningmastery.com/cross-entropy-for-machine-learning/)) is minimum. More information about the  [compile option method](https://www.tensorflow.org/api_docs/python/tf/keras/Model) can be found on the documentation","metadata":{}},{"cell_type":"code","source":"from tensorflow import keras\nfrom keras.models import Sequential\nfrom tensorflow.keras import layers\nfrom tensorflow.keras.optimizers import Adam as op_ad\n\n#for reproducibility (of the neural network model) you can uncomment this lines\n#from numpy.random import seed\n#seed(1)\n#from tensorflow.random import set_seed\n#set_seed(2)\n\n#I've tried wider neural networks but they seem to overfit\n#Even deeper neural networks might have the same issue but the result is not bigger than with wider neural\n#networks. I also used dropout and normalization layers inbetween to optimize training and prevent overfitting\nif neural_network:\n    num_columns=len(X_train_scaled.columns)\n    model = Sequential([\n        layers.Dense(128,activation='relu',input_shape=[num_columns]),\n        layers.Dropout(rate=0.3),\n        layers.BatchNormalization(),\n        #layers.Dense(128,activation='relu'),\n        #layers.Dropout(rate=0.3),\n        #layers.BatchNormalization(),\n        #layers.Dense(128,activation='relu'),\n        #layers.Dropout(rate=0.3),\n        #layers.BatchNormalization(),\n        #I tried selu activation function just for fun\n        #layers.Dense(32,activation='selu'),\n        layers.Dense(1,activation='sigmoid'),\n    ])\n    #here we define the goals of the training process. By adding metrics (like accuracy) we can compute this values\n    #and return them.  \n    model.compile(\n        optimizer=op_ad(learning_rate=0.0005),\n        loss='binary_crossentropy',\n        metrics=['binary_accuracy'],\n        steps_per_execution=2\n    )\n    #additional measure to prevent overfitting\n    early_stopping= keras.callbacks.EarlyStopping(\n        patience=50,\n        min_delta=0.001,\n        restore_best_weights=True\n    )\n","metadata":{"execution":{"iopub.status.busy":"2022-07-25T00:41:28.372226Z","iopub.execute_input":"2022-07-25T00:41:28.372950Z","iopub.status.idle":"2022-07-25T00:41:36.945431Z","shell.execute_reply.started":"2022-07-25T00:41:28.372911Z","shell.execute_reply":"2022-07-25T00:41:36.944403Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"The neural network training starts here. The output is an [History](https://machinelearningmastery.com/display-deep-learning-model-training-history-in-keras/) object which contains a record of the values of the loss and metrics functions on the cross validation and training set required by the training settings and found during this process. This allow us to draw learning curves.","metadata":{}},{"cell_type":"code","source":"if neural_network:\n    history = model.fit(\n        X_train_scaled,y_train,\n        validation_data=(X_cv_scaled,y_cv),\n        batch_size=100,\n        epochs=400,\n        callbacks=[early_stopping],\n        verbose=0\n    )\n","metadata":{"execution":{"iopub.status.busy":"2022-07-25T00:41:36.946755Z","iopub.execute_input":"2022-07-25T00:41:36.947324Z","iopub.status.idle":"2022-07-25T00:41:36.953535Z","shell.execute_reply.started":"2022-07-25T00:41:36.947291Z","shell.execute_reply":"2022-07-25T00:41:36.952109Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### Neural network: Learning curve plots\n\nTo see the training process and detect overfitting we draw the learning curves by using the loss function and metric(s) we have chosen.\nThe information is held in the **history** object.","metadata":{}},{"cell_type":"code","source":"\nif neural_network:\n    history_df = pd.DataFrame(history.history)\n    #this metrics are stored in history by the training process settings\n    history_df.loc[:, ['loss', 'val_loss']].plot();\n    history_df.loc[:, ['binary_accuracy', 'val_binary_accuracy']].plot()\n\n    print(\"Minimum validation loss: {} \\nMaximum vaidation accuracy: {} \".format(history_df['val_loss'].min(),history_df['val_binary_accuracy'].max()))\n    ","metadata":{"execution":{"iopub.status.busy":"2022-07-25T00:41:36.955888Z","iopub.execute_input":"2022-07-25T00:41:36.956789Z","iopub.status.idle":"2022-07-25T00:41:36.977032Z","shell.execute_reply.started":"2022-07-25T00:41:36.956707Z","shell.execute_reply":"2022-07-25T00:41:36.976324Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Making class predictions from probabilities returned by the neural network model we use \nthe lambda expression: \n<code>lambda x: 1 if 0.5 &#60; x else 0 </code>\nto get the class assuming a threshold of 0.5. Note that this also might be a\nmodel hyperparameter but in this notebook we use the fixed value 0.5.","metadata":{}},{"cell_type":"code","source":"from sklearn.metrics import accuracy_score\n\nif neural_network:\n    #here the X_train_scaled is redefined so that we do not change X_train\n    X_train_scaled=pd.DataFrame(scaler_dataset.transform(X_train),columns=X_train.columns) \n    #note that the neural network model returns a probability, since the last layer has a sigmoid activation function \n    prob_train_nn=model.predict(X_train_scaled)\n    #lambda expression to change probabilities into classes, note that we could have used a different threshold \n    prob_to_class= lambda x: 1 if 0.5<x else 0\n    #converting probabilities into classes\n    y_train_pred_nn=[prob_to_class(x) for x in prob_train_nn]\n    \n    #I used both metrics, on the case of the cross validation set they can be very dofferent, particularly\n    #when we do not use resampling\n    print('accuracy score on the train set:{}'.format(accuracy_score(y_train, y_train_pred_nn)))\n    print('balanced accuracy score on the train set:{}'.format(balanced_accuracy_score(y_train, y_train_pred_nn)))","metadata":{"execution":{"iopub.status.busy":"2022-07-25T00:41:36.978040Z","iopub.execute_input":"2022-07-25T00:41:36.979072Z","iopub.status.idle":"2022-07-25T00:41:36.997881Z","shell.execute_reply.started":"2022-07-25T00:41:36.979009Z","shell.execute_reply":"2022-07-25T00:41:36.997304Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"We also make predictions on the cross validation set in order to observe the actual performance of the model. ","metadata":{}},{"cell_type":"code","source":"if neural_network:\n    \n    prob_cv_nn=model.predict(X_cv_scaled)\n    y_cv_pred_nn=[prob_to_class(x) for x in prob_cv_nn]\n    \n    print('accuracy score on  the cross validation set:{}'.format(accuracy_score(y_cv, y_cv_pred_nn))  )\n    print('balanced accuracy score on  the cross validation set:{}'.format(balanced_accuracy_score(y_cv, y_cv_pred_nn)))","metadata":{"scrolled":true,"execution":{"iopub.status.busy":"2022-07-25T00:41:36.998789Z","iopub.execute_input":"2022-07-25T00:41:36.999152Z","iopub.status.idle":"2022-07-25T00:41:37.013605Z","shell.execute_reply.started":"2022-07-25T00:41:36.999113Z","shell.execute_reply":"2022-07-25T00:41:37.012647Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## Boosted random forests (xgb) model\n\nAlternatively, one can use a **boosted tree model**. In this case resampling is not important but the tuning of hyperparameters is \ncritical fot the model performance.\n\nThe parameters are those by running several times the optuna study to tune hyperparameters. The parameters 'use_label_encoder' and 'verbosity' are simply set to avoid warnings and messages while running this algorithms.\n\nThis model is explored a bit further in another notebook [TitanicSurvCom](https://www.kaggle.com/code/jairjiramuoz/titanicsurvcom/notebook)","metadata":{}},{"cell_type":"code","source":"\nfrom xgboost import XGBClassifier\n\n#functions to measure the performace of the xgb model (only)\ndef customized_score_cross_val(X,y,folds=5,threshold=0.5,**kwargs):\n    cv_scores = np.eye(folds)\n    train_scores = np.eye(folds)\n    for step in range(folds):\n        X_train, X_cv, y_train, y_cv = train_test_split(X, y, test_size=0.2)\n        model =  XGBClassifier(**kwargs)\n        model.fit(X_train,y_train)\n        y_fit = model.predict(X_train)\n        y_pred = model.predict(X_cv)\n        train_scores[step]=accuracy_score(y_train, y_fit)\n        cv_scores[step]=accuracy_score(y_cv, y_pred)\n    return  Score(cv_scores,train_scores)   \n\n\ndef smooth_score_cross_val(X,y,folds=5,**kwargs):\n    cv_scores = np.eye(folds)\n    train_scores = np.eye(folds)\n    for step in range(folds):\n        X_train, X_cv, y_train, y_cv = train_test_split(X, y, test_size=0.2)\n        model =  XGBClassifier(**kwargs)\n        model.fit(X_train,y_train)\n        y_fit = model.predict(X_train)\n        y_pred = model.predict(X_cv)\n        train_scores[step]=log_loss(y_train, y_fit)\n        cv_scores[step]=log_loss(y_cv, y_pred)\n    return  Score(cv_scores,train_scores)   \n\n\ndef metrics_cross_val(X,y,folds=5,**kwargs):\n    cv_metrics = np.zeros((folds,5))\n    precision_score\n    accuracy_score\n    balanced_accuracy_score\n    recall_score\n    for step in range(folds):\n        X_train, X_cv, y_train, y_cv = train_test_split(X, y, test_size=0.2)\n        model =  XGBClassifier(**kwargs)\n        model.fit(X_train,y_train)\n        y_fit = model.predict(X_train)\n        y_pred = model.predict(X_cv)\n        cv_metrics[step,0]=log_loss(y_cv, y_pred)\n        cv_metrics[step,1]=precision_score(y_cv, y_pred)\n        cv_metrics[step,2]=accuracy_score(y_cv, y_pred)\n        cv_metrics[step,3]=balanced_accuracy_score(y_cv, y_pred)\n        cv_metrics[step,4]=recall_score(y_cv, y_pred)\n    return  pd.DataFrame(cv_metrics,columns=['log_loss','precision','acuracy','balanced_accuracy','recall'])\n\n\nif boosted_tree_model:\n   \n    #base parameters(same for all training processes):this parameters allow a less verbose training process\n    xgb_params={'use_label_encoder':False,'verbosity':0}\n    #dynamic parameters:this are the real model parameters which we can tune using an optuna study (see bellow)\n    extra_xgb_params = {'max_depth': 5, 'learning_rate': 0.05749106357588373, 'n_estimators': 5682, 'min_child_weight': 5, 'colsample_bytree': 0.8034047905616921, 'subsample': 0.7230299010267518, 'reg_alpha': 0.18436505849416285, 'reg_lambda': 0.7145550969871013}\n    #merging the base and dynamic parameters\n    xgb_params = {**xgb_params,**extra_xgb_params}\n\n    model = XGBClassifier(**xgb_params)\n    model.fit(X_train,y_train)\n    y_train_pred_tree = model.predict(X_train)\n    #the predict_proba method allows to return the classes probabilities. \n    #In this case its important to do this since one can easily combine models by using probabilities instead of classes\n    prob_train_tree=model.predict_proba(X_train)\n    #reshaping this array\n    prob_train_tree=prob_train_tree[:,1]\n    \n    prob_cv_tree=model.predict_proba(X_cv)\n    y_cv_pred_tree=model.predict(X_cv)\n    prob_cv_tree=prob_cv_tree[:,1]\n    \n    #if you want to test xgb model performance alone run this lines\n    #scores_data=metrics_cross_val(X_train,y_train,folds=20,**xgb_params)\n    #print(scores_data.mean())","metadata":{"execution":{"iopub.status.busy":"2022-07-25T00:41:37.015019Z","iopub.execute_input":"2022-07-25T00:41:37.015427Z","iopub.status.idle":"2022-07-25T00:41:53.788265Z","shell.execute_reply.started":"2022-07-25T00:41:37.015393Z","shell.execute_reply":"2022-07-25T00:41:53.787740Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Performance charts\nHere we use the cross validation set to see the performance of our models and some informative charts used for classifier models. For all cases a chart is generated, even when a more complex one (comparative) could have been provided in the case of a combined model. \n## Confusion matrix\n\nIn order to measure the actual model performance one could look for other metrics besides accuracy. This is particularly important because this is an imbalaced set and accuracy alone is not a reliable metric (by itself).","metadata":{}},{"cell_type":"code","source":"from sklearn.metrics import confusion_matrix\nif not model_inclusive:\n    if neural_network:\n        sns.clustermap(confusion_matrix(y_cv,y_cv_pred_nn,normalize='all'), cmap=\"mako\",annot=True)\n    if boosted_tree_model:\n        sns.clustermap(confusion_matrix(y_train,y_train_pred_tree,normalize='all'), cmap=\"mako\",annot=True)\nif model_inclusive:\n    #here we compute the  predictions of the combined model by using a weighted combination of their\n    #respective probabilities. There are other possibilities, however it is important that the output is also a number\n    #between 0 and 1. I used the approach of using weights whose sum is 1. \n    #I have tried many combinations and tested their performance. \n    \n    nn_rate=0.1\n    xgb_rate=0.9\n    \n    #this two numpy arrays come in different shapes so we have to reshape them in order to form a linear\n    #combination and not a concatenation\n    prob_train_comb= nn_rate*prob_train_nn.reshape((1,prob_train_nn.shape[0]))+ xgb_rate*prob_train_tree\n    prob_train_comb=prob_train_comb[0,:]\n    y_train_pred_comb=[prob_to_class(x) for x in prob_train_comb]\n    \n    prob_cv_comb= nn_rate*prob_cv_nn.reshape((1,prob_cv_nn.shape[0]))+ xgb_rate*prob_cv_tree\n    prob_cv_comb=prob_cv_comb[0,:]\n    y_cv_pred_comb=[prob_to_class(x) for x in prob_cv_comb]\n    sns.clustermap(confusion_matrix(y_cv,y_cv_pred_comb,normalize='all'), cmap=\"mako\",annot=True)\n            \n            ","metadata":{"execution":{"iopub.status.busy":"2022-07-25T00:41:53.789449Z","iopub.execute_input":"2022-07-25T00:41:53.789808Z","iopub.status.idle":"2022-07-25T00:41:54.216833Z","shell.execute_reply.started":"2022-07-25T00:41:53.789775Z","shell.execute_reply":"2022-07-25T00:41:54.215780Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## ROC curve \nThe **Receiver operating  characteristic  (ROC)** curve is another informative chart which can be used to evaluate the model perfomance.\nIt compares the  False positives and true positives This can be used to compare different models. A model for which the area under the curve is maximum is then a good model. There are other alternatves to the ROC curve see ***Business Intelligence: data mining and decision making by Carlo Vercellis*** .\nIf one gets a straight line, this means the classifier is a random classifier, which is the same as generating a simulation of the classes based on assumptions about their probability distribution and using this as predictions. One might get something like this if one ommits resampling when using a neural network model: the model learns to return always the most frequent class ( [keras tips on this issue](https://www.tensorflow.org/tutorials/structured_data/imbalanced_data) ).","metadata":{}},{"cell_type":"code","source":"from sklearn.metrics import roc_curve\n\nif not model_inclusive:\n    if  neural_network:\n        pred_prob=model.predict(X_cv)\n        fpr, tpr, thresholds = roc_curve(y_cv, y_cv_pred_nn)\n        plt.figure(figsize=[10.0,5.0])\n        plt.xlabel('False positive rate')\n        plt.ylabel('True positive rate')\n        plt.title('ROC curve')\n        plt.plot(fpr, tpr)\n    if boosted_tree_model:\n        fpr, tpr, thresholds = roc_curve(y_train, prob_train_tree)\n        plt.figure(figsize=[10.0,5.0])\n        plt.xlabel('False positive rate')\n        plt.ylabel('True positive rate')\n        plt.title('ROC curve')\n        plt.plot(fpr, tpr)\nif model_inclusive:\n        fpr, tpr, thresholds = roc_curve(y_cv, prob_cv_comb)\n        plt.figure(figsize=[10.0,5.0])\n        plt.xlabel('False positive rate')\n        plt.ylabel('True positive rate')\n        plt.title('ROC curve')\n        plt.plot(fpr, tpr)\n\n        ","metadata":{"execution":{"iopub.status.busy":"2022-07-25T00:41:54.218194Z","iopub.execute_input":"2022-07-25T00:41:54.218393Z","iopub.status.idle":"2022-07-25T00:41:54.410918Z","shell.execute_reply.started":"2022-07-25T00:41:54.218365Z","shell.execute_reply":"2022-07-25T00:41:54.409174Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Submit file\n\nHere we make predictions for the test set. In the case of a single model (either nn or xgb) this is straight forward, though the combined model  requires\nto compute a linear combination of the predicted probabilities and then obtain the respective classes. ","metadata":{}},{"cell_type":"code","source":"if neural_network:\n    X_scaled_test = scaler_dataset.transform(X_test)\n    prob_test_nn=model.predict(X_scaled_test)\n    y_test_pred_nn=[prob_to_class(x) for x in prob_test_nn]\nif boosted_tree_model:\n    y_test_pred_tree = model.predict(X_test)\n    prob_test_tree=model.predict_proba(X_test)\n    prob_test_tree=prob_test_tree[:,1]\n\n\nif not model_inclusive:    \n    if neural_network:\n        y_pred=y_test_pred_nn\n    if boosted_tree_model:\n        y_pred= y_test_pred_tree\nif model_inclusive:\n    prob_test_comb= nn_rate*prob_test_nn.reshape((1,prob_test_nn.shape[0]))+ xgb_rate*prob_test_tree\n    prob_test_comb=prob_test_comb[0,:]\n    y_pred=[prob_to_class(x) for x in prob_test_comb]\n\ndf_submission=pd.DataFrame(columns=['PassengerId','Survived'])\ndf_submission['PassengerId']=df_test['PassengerId']\ndf_submission['Survived']=y_pred\ndf_submission.to_csv('submission.csv', index=False)\nprint(\"Your submission was successfully saved!\")","metadata":{"execution":{"iopub.status.busy":"2022-07-25T00:41:54.412864Z","iopub.execute_input":"2022-07-25T00:41:54.413220Z","iopub.status.idle":"2022-07-25T00:41:54.541964Z","shell.execute_reply.started":"2022-07-25T00:41:54.413174Z","shell.execute_reply":"2022-07-25T00:41:54.541070Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Comparison of model predictions for the mixed model\n\nIn this case I considered interesting to see how the prediction of every component of the combined model and the real labels (if available) differ on every set.","metadata":{}},{"cell_type":"code","source":"if  model_inclusive: \n    print('Mismatch of predictions on the cross validation set\\n')\n    comparison_results=pd.DataFrame({'neural_network':y_cv_pred_nn,'boosted_tree':y_cv_pred_tree,'mixed_model':y_cv_pred_comb,'actual':y_cv})\n    print(comparison_results.value_counts())\nelse:\n    print('mixed model is not available for the current setting')","metadata":{"execution":{"iopub.status.busy":"2022-07-25T00:41:54.543101Z","iopub.execute_input":"2022-07-25T00:41:54.543290Z","iopub.status.idle":"2022-07-25T00:41:54.548124Z","shell.execute_reply.started":"2022-07-25T00:41:54.543264Z","shell.execute_reply":"2022-07-25T00:41:54.547277Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"if  model_inclusive: \n    print('Mismatch of predictions on the training set\\n')\n    comparison_results=pd.DataFrame({'neural_network':y_train_pred_nn,'boosted_tree':y_train_pred_tree,'mixed_model':y_train_pred_comb,'actual':y_train})\n    print(comparison_results.value_counts())\nelse:\n    print('mixed model is not available for the current setting')","metadata":{"execution":{"iopub.status.busy":"2022-07-25T00:41:54.549185Z","iopub.execute_input":"2022-07-25T00:41:54.549521Z","iopub.status.idle":"2022-07-25T00:41:54.568924Z","shell.execute_reply.started":"2022-07-25T00:41:54.549495Z","shell.execute_reply":"2022-07-25T00:41:54.567609Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"if  model_inclusive: \n    print('Mismatch of predictions on the test set (the actual labels aren\\'t available)\\n')\n    comparison_results=pd.DataFrame({'neural_network':y_test_pred_nn,'boosted_tree':y_test_pred_tree,'mixed_model':y_pred})\n    print(comparison_results.value_counts())\nelse:\n    print('mixed model is not available for the current setting')","metadata":{"execution":{"iopub.status.busy":"2022-07-25T00:41:54.571553Z","iopub.execute_input":"2022-07-25T00:41:54.572057Z","iopub.status.idle":"2022-07-25T00:41:54.580887Z","shell.execute_reply.started":"2022-07-25T00:41:54.572028Z","shell.execute_reply":"2022-07-25T00:41:54.579792Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Tunning hyperparameters\n\nHere we use the implementation of a hyperparameters optimizer by Optuna called [optuna.study](https://optuna.readthedocs.io/en/stable/reference/generated/optuna.study.Study.html#optuna.study.Study) these need a way to computing certain metric or loss functions on the model, the most sensible way is by using a cross validation scheme function for this. Then the algorithm uses a dictionary of hyperparameters with a range of possible values to find the best combination of these.   ","metadata":{}},{"cell_type":"code","source":"import optuna\ndef objective(trial):\n    xgb_params = dict(\n        use_label_encoder=False,\n        verbosity=0,\n        max_depth=trial.suggest_int(\"max_depth\", 2, 10),\n        learning_rate=trial.suggest_float(\"learning_rate\", 1e-4, 1e-1, log=True),\n        n_estimators=trial.suggest_int(\"n_estimators\", 1000, 8000),\n        min_child_weight=trial.suggest_int(\"min_child_weight\", 1, 10),\n        colsample_bytree=trial.suggest_float(\"colsample_bytree\", 0.2, 1.0),\n        subsample=trial.suggest_float(\"subsample\", 0.2, 1.0),\n        reg_alpha=trial.suggest_float(\"reg_alpha\", 1e-4, 1e2, log=True),\n        reg_lambda=trial.suggest_float(\"reg_lambda\", 1e-4, 1e2, log=True),\n        #set seed to reprudocibility\n        #random_state=0\n    )\n    if False:\n        #we intend to optimize accuracy \n        scores_data=customized_score_cross_val(X_train,y_train,folds=15,use_model='other',**xgb_params)\n        return scores_data.mean_cv()\n    else:\n        #we intend to optimize cross entropy\n        scores_data=smooth_score_cross_val(X_train,y_train,folds=15,use_model='other',**xgb_params)\n        return scores_data.mean_cv()\nif False:\n    #if we choose accuracy we set direction to maximize, either way minimize\n    study = optuna.create_study(direction=\"minimize\")\n    study.optimize(objective, n_trials=30)\n    xgb_params = study.best_params\n    print(xgb_params)\n\n\n","metadata":{"execution":{"iopub.status.busy":"2022-07-25T00:41:54.583728Z","iopub.execute_input":"2022-07-25T00:41:54.583939Z","iopub.status.idle":"2022-07-25T00:41:55.704456Z","shell.execute_reply.started":"2022-07-25T00:41:54.583913Z","shell.execute_reply":"2022-07-25T00:41:55.703633Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## The tunning of the (sub) model cabin  hyperparameters\nThis can be ommited, since there is little improvement in the submodel perfomance.","metadata":{}},{"cell_type":"code","source":"def cabin_score_cross_val(X,y,folds=5,**kwargs):\n    cv_scores = np.eye(folds)\n    train_scores = np.eye(folds)\n    for step in range(folds):\n        X_train, X_cv, y_train, y_cv = train_test_split(X, y, test_size=0.2)\n        model =  XGBRegressor(**kwargs)\n        model.fit(X_train,y_train)\n        y_fit = model.predict(X_train)\n        y_pred = model.predict(X_cv)\n        y_fit= cabin_imp_predict(X_train,model)\n        y_pred = cabin_imp_predict(X_cv,model)\n        train_scores[step]=mean_squared_log_error(y_train, y_fit)\n        cv_scores[step]=mean_squared_log_error(y_cv, y_pred)\n    return  Score(cv_scores,train_scores)\ndef cabin_hyp_tun(trial):\n    xgb_params = dict(\n        #use_label_encoder=False,\n        #verbosity=0,\n        max_depth=trial.suggest_int(\"max_depth\", 2, 10),\n        learning_rate=trial.suggest_float(\"learning_rate\", 1e-4, 1e-1, log=True),\n        n_estimators=trial.suggest_int(\"n_estimators\", 1000, 8000),\n        min_child_weight=trial.suggest_int(\"min_child_weight\", 1, 10),\n        colsample_bytree=trial.suggest_float(\"colsample_bytree\", 0.2, 1.0),\n        subsample=trial.suggest_float(\"subsample\", 0.2, 1.0),\n        reg_alpha=trial.suggest_float(\"reg_alpha\", 1e-4, 1e2, log=True),\n        reg_lambda=trial.suggest_float(\"reg_lambda\", 1e-4, 1e2, log=True),\n    )\n    scores_data=cabin_score_cross_val(X_alt_train,y_cabin_train,folds=10,**xgb_params)\n    scores_data.mean_cv()\n    return scores_data.mean_cv()\nif False:\n    study = optuna.create_study(direction=\"minimize\")\n    study.optimize(cabin_hyp_tun, n_trials=20)\n    xgb_params = study.best_params\n    print(xgb_params)","metadata":{"execution":{"iopub.status.busy":"2022-07-25T00:41:55.705722Z","iopub.execute_input":"2022-07-25T00:41:55.706001Z","iopub.status.idle":"2022-07-25T00:41:55.720554Z","shell.execute_reply.started":"2022-07-25T00:41:55.705961Z","shell.execute_reply":"2022-07-25T00:41:55.719743Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## Updating hyperparameters from various optuna studies","metadata":{}},{"cell_type":"markdown","source":"This lines of code intend to define fucntions that average the hyperparameters for the xgb model: Once some values are obtained by an **optuna study**  one can run again another study and obtain another values which can in turn be combined with previous ones to get new ones.","metadata":{}},{"cell_type":"code","source":"def bag_param_dicts(param_one,param_two):\n    param_output={}\n    if param_one.keys()==param_two.keys():\n        for param in param_one.keys():\n            param_output[param]=(param_one[param]+param_two[param])/2\n            if  param in {'max_depth','n_estimators','min_child_weight'}:\n                param_output[param]= int(param_output[param])\n        return param_output\n    else:\n        print('not equal parameter dictionaries')\n\ndef  bag_multiple_param_dicts(param_dict_list):\n    param_output={}\n    if len(param_dict_list)==0:\n        return {}\n    if len(param_dict_list)==1:\n        return param_dict_list[0]\n    current_param_dict=param_dict_list[0]\n    for dict_id in range(1,len(param_dict_list)):\n        param_output=bag_param_dicts(current_param_dict,param_dict_list[dict_id])\n        current_param_dict=param_dict_list[dict_id]\n    return param_output\n\n#here use the xgb_params found in the latest optuna study, you can set this or use the output of the current optuna study (if available)\n#xgb_params={'max_depth': 5, 'learning_rate': 0.04256324369543573, 'n_estimators': 7837, 'min_child_weight': 1, 'colsample_bytree': 0.36036725390453006, 'subsample': 0.28020014927658576, 'reg_alpha': 0.0008411237345127937, 'reg_lambda': 0.001545649282090005}\ntry:\n    #remember that extra_xgb_params are the part of the current hyperparameters been used by the xgb model\n    print(bag_param_dicts(extra_xgb_params, xgb_params))\nexcept:\n    print('NA: \\'extra_xgb_params\\' or \\'xgb_params\\' do not exist or have a different in the current scope')","metadata":{"execution":{"iopub.status.busy":"2022-07-25T00:41:55.721953Z","iopub.execute_input":"2022-07-25T00:41:55.722449Z","iopub.status.idle":"2022-07-25T00:41:55.738218Z","shell.execute_reply.started":"2022-07-25T00:41:55.722408Z","shell.execute_reply":"2022-07-25T00:41:55.737305Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Comparison of models (Anova)\nWe define a model hyperparameter object based on this template\n\n* cabin_submodel: bool whether or not the cabin submodel was used\n* resample\n  * sample_size: the size of the block of repeated rows \n  * num_sample: number of blocks of repeated rows\n* features\n  * PassengerId: original feature\n  * Pclass: original feature\n  * Name: original feature\n  * Sex: original feature\n  * Age: original feature\n  * SibSp: original feature \n  * Parch: original feature\n  * Ticket: original feature\n  * Fare: original feature \n  * Cabin: original feature \n  * Embarked: original feature\n  * lnFare: the natural logarithm of Fare if Fare is not 0, 0 otherwise\n  * ReTicket: The codes assigned from the union of all Ticket levels\n  * ReCabin: The codes assigned from the union of all Cabin levels\n  * Outlier: marks rows where Fare is 0\n  * Cabin_class_A: counts the number of times the letter A appears on Cabin if the cabin_submodel is True this is predicted by the submodel\n  * Cabin_class_B: counts the number of times the letter B appears on Cabin if the cabin_submodel is True this is predicted by the submodel\n  * Cabin_class_C: counts the number of times the letter C appears on Cabin if the cabin_submodel is True this is predicted by the submodel\n  * Cabin_class_D: counts the number of times the letter D appears on Cabin if the cabin_submodel is True this is predicted by the submodel \n  * Cabin_class_E: counts the number of times the letter E appears on Cabin if the cabin_submodel is True this is predicted by the submodel\n  * Cabin_class_F: counts the number of times the letter F appears on Cabin if the cabin_submodel is True this is predicted by the submodel\n  * Cabin_class_G: counts the number of times the letter G appears on Cabin if the cabin_submodel is True this is predicted by the submodel\n  * Cabin_class_T: counts the number of times the letter T appears on Cabin if the cabin_submodel is True this is predicted by the submodel\n  * Cabin_class_None: used to represent a missed Cabin value  \n  * Age_missing: used to represent a missed Age value\n* neural_network\n  * contribution: How much of this model is used in the combined model. If is equal to one a pure neural network model is used instead \n  * layers: array of numbers indicating the number of neurons in each layer if the model has less layers all the remaning values are set to NaN  (maximum 3)\n* boosted_tree \n  * contribution: How much of this model is used in the combined model. If is equal to one a pure boosted model is used instead   \n  * params\n    * max_depth: The depth of the tree (as a graph)\n    * learning_rate: gradient descent parameter\n    * n_estimators: the number of trees used\n    * min_child_weight: minimum of instance weight needed in a child\n    * colsample_bytree: subsample ratio of columns for each tree\n    * subsample: subsample ratio of the training instance\n    * reg_alpha: regularization term\n    * reg_lambda: regularization term\n* test_accuracy: The score obtained on the test set\n\nA set of feature would be chosen and repeated a number of times (since both boosted trees and neural network models have no deterministic training processes). The result was noted .  The sublevels neural_network->params and boosted_tree->params could have been set to an empty object {} to save memory but I kept them even when a pure neural network or tree model were used in order to simplify implementation. \n","metadata":{}},{"cell_type":"code","source":"import sys\nimport json\nwhole_model_features=['PassengerId','Pclass', 'Name', 'Sex', 'Age', 'SibSp', 'Parch','Ticket', 'Fare', 'Cabin', 'Embarked', 'lnFare', 'ReTicket', 'ReCabin','Outlier','Cabin_class_A','Cabin_class_B','Cabin_class_C','Cabin_class_D', 'Cabin_class_E','Cabin_class_F','Cabin_class_G','Cabin_class_T','Cabin_class_None','Age_missing']\n#used to automate the process of noting the features used in the current model\ncurrent_model_features=[True if (col in X_train.columns)  else False for col in whole_model_features ]\n\n#this was used in order to update the data even when the actual accuracy wasn't known by then\ncurrent_model_info={'cabin_submodel':False,\n 'resample':{'sample_size':200,'num_sample':6},\n 'features': dict(zip(whole_model_features,current_model_features)),\n'neural_network':{'contribution':0.1,'layers':[128]},\n 'boosted_tree':{'contribution':0.9,'params':{'max_depth': 5, 'learning_rate': 0.05749106357588373, 'n_estimators': 5682, 'min_child_weight': 5, 'colsample_bytree': 0.8034047905616921, 'subsample': 0.7230299010267518, 'reg_alpha': 0.18436505849416285, 'reg_lambda': 0.7145550969871013}},\n 'test_accuracy':0.73\n}\n\n#this is the result of the tests according to the template being used\nperfomance_data=[{'cabin_submodel': False, 'resample': {'sample_size': 200, 'num_sample': 6}, 'features': {'PassengerId': True, 'Pclass': True, 'Name': True, 'Sex': True, 'Age': True, 'SibSp': True, 'Parch': True, 'Ticket': True, 'Fare': True, 'Cabin': True, 'Embarked': True, 'lnFare': True, 'ReTicket': True, 'ReCabin': True, 'Outlier': True, 'Cabin_class_A': True, 'Cabin_class_B': True, 'Cabin_class_C': True, 'Cabin_class_D': True, 'Cabin_class_E': True, 'Cabin_class_F': True, 'Cabin_class_G': True, 'Cabin_class_T': True, 'Cabin_class_None': True, 'Age_missing': True}, 'neural_network': {'contribution': 0.1, 'layers': [128, 128, 128]}, 'boosted_tree': {'contribution': 0.9, 'params': {'max_depth': 5, 'learning_rate': 0.05749106357588373, 'n_estimators': 5682, 'min_child_weight': 5, 'colsample_bytree': 0.8034047905616921, 'subsample': 0.7230299010267518, 'reg_alpha': 0.18436505849416285, 'reg_lambda': 0.7145550969871013}}, 'test_accuracy': 0.74401},\n{'cabin_submodel': False, 'resample': {'sample_size': 200, 'num_sample': 6}, 'features': {'PassengerId': False, 'Pclass': True, 'Name': False, 'Sex': True, 'Age': True, 'SibSp': False, 'Parch': True, 'Ticket': True, 'Fare': True, 'Cabin': True, 'Embarked': False, 'lnFare': False, 'ReTicket': True, 'ReCabin': True, 'Outlier': False, 'Cabin_class_A': True, 'Cabin_class_B': True, 'Cabin_class_C': True, 'Cabin_class_D': True, 'Cabin_class_E': False, 'Cabin_class_F': True, 'Cabin_class_G': False, 'Cabin_class_T': False, 'Cabin_class_None': True, 'Age_missing': True}, 'neural_network': {'contribution': 0.1, 'layers': [128, 128, 128]}, 'boosted_tree': {'contribution': 0.9, 'params': {'max_depth': 5, 'learning_rate': 0.05749106357588373, 'n_estimators': 5682, 'min_child_weight': 5, 'colsample_bytree': 0.8034047905616921, 'subsample': 0.7230299010267518, 'reg_alpha': 0.18436505849416285, 'reg_lambda': 0.7145550969871013}}, 'test_accuracy': 0.75598},\n{'cabin_submodel': False, 'resample': {'sample_size': 200, 'num_sample': 6}, 'features': {'PassengerId': False, 'Pclass': True, 'Name': False, 'Sex': True, 'Age': True, 'SibSp': False, 'Parch': True, 'Ticket': True, 'Fare': True, 'Cabin': True, 'Embarked': False, 'lnFare': False, 'ReTicket': True, 'ReCabin': True, 'Outlier': False, 'Cabin_class_A': True, 'Cabin_class_B': True, 'Cabin_class_C': True, 'Cabin_class_D': True, 'Cabin_class_E': False, 'Cabin_class_F': True, 'Cabin_class_G': False, 'Cabin_class_T': False, 'Cabin_class_None': True, 'Age_missing': True}, 'neural_network': {'contribution': 1.0, 'layers': [128, 128, 128]}, 'boosted_tree': {'contribution': 0.0, 'params': {'max_depth': 5, 'learning_rate': 0.05749106357588373, 'n_estimators': 5682, 'min_child_weight': 5, 'colsample_bytree': 0.8034047905616921, 'subsample': 0.7230299010267518, 'reg_alpha': 0.18436505849416285, 'reg_lambda': 0.7145550969871013}}, 'test_accuracy': 0.74641},\n{'cabin_submodel': False, 'resample': {'sample_size': 200, 'num_sample': 6}, 'features': {'PassengerId': False, 'Pclass': True, 'Name': False, 'Sex': True, 'Age': True, 'SibSp': False, 'Parch': True, 'Ticket': True, 'Fare': True, 'Cabin': True, 'Embarked': False, 'lnFare': False, 'ReTicket': True, 'ReCabin': True, 'Outlier': False, 'Cabin_class_A': True, 'Cabin_class_B': True, 'Cabin_class_C': True, 'Cabin_class_D': True, 'Cabin_class_E': False, 'Cabin_class_F': True, 'Cabin_class_G': False, 'Cabin_class_T': False, 'Cabin_class_None': True, 'Age_missing': True}, 'neural_network': {'contribution': 1.0, 'layers': [128, 128, 128]}, 'boosted_tree': {'contribution': 0.0, 'params': {'max_depth': 5, 'learning_rate': 0.05749106357588373, 'n_estimators': 5682, 'min_child_weight': 5, 'colsample_bytree': 0.8034047905616921, 'subsample': 0.7230299010267518, 'reg_alpha': 0.18436505849416285, 'reg_lambda': 0.7145550969871013}}, 'test_accuracy': 0.74401},\n{'cabin_submodel': False, 'resample': {'sample_size': 200, 'num_sample': 6}, 'features': {'PassengerId': False, 'Pclass': True, 'Name': False, 'Sex': True, 'Age': True, 'SibSp': False, 'Parch': True, 'Ticket': True, 'Fare': True, 'Cabin': True, 'Embarked': False, 'lnFare': False, 'ReTicket': True, 'ReCabin': True, 'Outlier': False, 'Cabin_class_A': True, 'Cabin_class_B': True, 'Cabin_class_C': True, 'Cabin_class_D': True, 'Cabin_class_E': False, 'Cabin_class_F': True, 'Cabin_class_G': False, 'Cabin_class_T': False, 'Cabin_class_None': True, 'Age_missing': True}, 'neural_network': {'contribution': 0.0, 'layers': [128, 128, 128]}, 'boosted_tree': {'contribution': 1.0, 'params': {'max_depth': 5, 'learning_rate': 0.05749106357588373, 'n_estimators': 5682, 'min_child_weight': 5, 'colsample_bytree': 0.8034047905616921, 'subsample': 0.7230299010267518, 'reg_alpha': 0.18436505849416285, 'reg_lambda': 0.7145550969871013}}, 'test_accuracy': 0.75837},\n{'cabin_submodel': False, 'resample': {'sample_size': 200, 'num_sample': 6}, 'features': {'PassengerId': False, 'Pclass': True, 'Name': False, 'Sex': True, 'Age': True, 'SibSp': False, 'Parch': True, 'Ticket': True, 'Fare': True, 'Cabin': True, 'Embarked': False, 'lnFare': False, 'ReTicket': True, 'ReCabin': True, 'Outlier': False, 'Cabin_class_A': True, 'Cabin_class_B': True, 'Cabin_class_C': True, 'Cabin_class_D': True, 'Cabin_class_E': False, 'Cabin_class_F': True, 'Cabin_class_G': False, 'Cabin_class_T': False, 'Cabin_class_None': True, 'Age_missing': True}, 'neural_network': {'contribution': 0.0, 'layers': [128, 128, 128]}, 'boosted_tree': {'contribution': 1.0, 'params': {'max_depth': 5, 'learning_rate': 0.05749106357588373, 'n_estimators': 5682, 'min_child_weight': 5, 'colsample_bytree': 0.8034047905616921, 'subsample': 0.7230299010267518, 'reg_alpha': 0.18436505849416285, 'reg_lambda': 0.7145550969871013}}, 'test_accuracy': 0.75598},\n{'cabin_submodel': True, 'resample': {'sample_size': 200, 'num_sample': 6}, 'features': {'PassengerId': False, 'Pclass': True, 'Name': False, 'Sex': True, 'Age': True, 'SibSp': False, 'Parch': True, 'Ticket': True, 'Fare': True, 'Cabin': True, 'Embarked': False, 'lnFare': False, 'ReTicket': True, 'ReCabin': True, 'Outlier': False, 'Cabin_class_A': True, 'Cabin_class_B': True, 'Cabin_class_C': True, 'Cabin_class_D': True, 'Cabin_class_E': False, 'Cabin_class_F': True, 'Cabin_class_G': False, 'Cabin_class_T': False, 'Cabin_class_None': True, 'Age_missing': True}, 'neural_network': {'contribution': 0.1, 'layers': [128, 128, 128]}, 'boosted_tree': {'contribution': 0.9,'params': {'max_depth': 5, 'learning_rate': 0.05749106357588373, 'n_estimators': 5682, 'min_child_weight': 5, 'colsample_bytree': 0.8034047905616921, 'subsample': 0.7230299010267518, 'reg_alpha': 0.18436505849416285, 'reg_lambda': 0.7145550969871013}}, 'test_accuracy': 0.73923},\n{'cabin_submodel': True, 'resample': {'sample_size': 200, 'num_sample': 6}, 'features': {'PassengerId': False, 'Pclass': True, 'Name': False, 'Sex': True, 'Age': True, 'SibSp': False, 'Parch': True, 'Ticket': True, 'Fare': True, 'Cabin': True, 'Embarked': False, 'lnFare': False, 'ReTicket': True, 'ReCabin': True, 'Outlier': False, 'Cabin_class_A': True, 'Cabin_class_B': True, 'Cabin_class_C': True, 'Cabin_class_D': True, 'Cabin_class_E': False, 'Cabin_class_F': True, 'Cabin_class_G': False, 'Cabin_class_T': False, 'Cabin_class_None': True, 'Age_missing': True}, 'neural_network': {'contribution': 0.1, 'layers': [128, 128, 128]}, 'boosted_tree': {'contribution': 0.9, 'params': {'max_depth': 5, 'learning_rate': 0.05749106357588373, 'n_estimators': 5682, 'min_child_weight': 5, 'colsample_bytree': 0.8034047905616921, 'subsample': 0.7230299010267518, 'reg_alpha': 0.18436505849416285, 'reg_lambda': 0.7145550969871013}}, 'test_accuracy': 0.75119},\n{'cabin_submodel': True, 'resample': {'sample_size': 200, 'num_sample': 6}, 'features': {'PassengerId': False, 'Pclass': True, 'Name': False, 'Sex': True, 'Age': True, 'SibSp': False, 'Parch': True, 'Ticket': True, 'Fare': True, 'Cabin': True, 'Embarked': False, 'lnFare': False, 'ReTicket': True, 'ReCabin': True, 'Outlier': False, 'Cabin_class_A': True, 'Cabin_class_B': True, 'Cabin_class_C': True, 'Cabin_class_D': True, 'Cabin_class_E': False, 'Cabin_class_F': True, 'Cabin_class_G': False, 'Cabin_class_T': False, 'Cabin_class_None': True, 'Age_missing': True}, 'neural_network': {'contribution': 1.0, 'layers': [128, 128, 128]}, 'boosted_tree': {'contribution': 0.0, 'params': {'max_depth': 5, 'learning_rate': 0.05749106357588373, 'n_estimators': 5682, 'min_child_weight': 5, 'colsample_bytree': 0.8034047905616921, 'subsample': 0.7230299010267518, 'reg_alpha': 0.18436505849416285, 'reg_lambda': 0.7145550969871013}}, 'test_accuracy': 0.72248},\n{'cabin_submodel': True, 'resample': {'sample_size': 200, 'num_sample': 6}, 'features': {'PassengerId': False, 'Pclass': True, 'Name': False, 'Sex': True, 'Age': True, 'SibSp': False, 'Parch': True, 'Ticket': True, 'Fare': True, 'Cabin': True, 'Embarked': False, 'lnFare': False, 'ReTicket': True, 'ReCabin': True, 'Outlier': False, 'Cabin_class_A': True, 'Cabin_class_B': True, 'Cabin_class_C': True, 'Cabin_class_D': True, 'Cabin_class_E': False, 'Cabin_class_F': True, 'Cabin_class_G': False, 'Cabin_class_T': False, 'Cabin_class_None': True, 'Age_missing': True}, 'neural_network': {'contribution': 1.0, 'layers': [128, 128, 128]}, 'boosted_tree': {'contribution': 0.0, 'params': {'max_depth': 5, 'learning_rate': 0.05749106357588373, 'n_estimators': 5682, 'min_child_weight': 5, 'colsample_bytree': 0.8034047905616921, 'subsample': 0.7230299010267518, 'reg_alpha': 0.18436505849416285, 'reg_lambda': 0.7145550969871013}}, 'test_accuracy': 0.7177},\n{'cabin_submodel': True, 'resample': {'sample_size': 200, 'num_sample': 6}, 'features': {'PassengerId': False, 'Pclass': True, 'Name': False, 'Sex': True, 'Age': True, 'SibSp': False, 'Parch': True, 'Ticket': True, 'Fare': True, 'Cabin': True, 'Embarked': False, 'lnFare': False, 'ReTicket': True, 'ReCabin': True, 'Outlier': False, 'Cabin_class_A': True, 'Cabin_class_B': True, 'Cabin_class_C': True, 'Cabin_class_D': True, 'Cabin_class_E': False, 'Cabin_class_F': True, 'Cabin_class_G': False, 'Cabin_class_T': False, 'Cabin_class_None': True, 'Age_missing': True}, 'neural_network': {'contribution': 0.0, 'layers': [128, 128, 128]}, 'boosted_tree': {'contribution': 1.0, 'params': {'max_depth': 5, 'learning_rate': 0.05749106357588373, 'n_estimators': 5682, 'min_child_weight': 5, 'colsample_bytree': 0.8034047905616921, 'subsample': 0.7230299010267518, 'reg_alpha': 0.18436505849416285, 'reg_lambda': 0.7145550969871013}}, 'test_accuracy': 0.7488},\n{'cabin_submodel': True, 'resample': {'sample_size': 200, 'num_sample': 6}, 'features': {'PassengerId': False, 'Pclass': True, 'Name': False, 'Sex': True, 'Age': True, 'SibSp': False, 'Parch': True, 'Ticket': True, 'Fare': True, 'Cabin': True, 'Embarked': False, 'lnFare': False, 'ReTicket': True, 'ReCabin': True, 'Outlier': False, 'Cabin_class_A': True, 'Cabin_class_B': True, 'Cabin_class_C': True, 'Cabin_class_D': True, 'Cabin_class_E': False, 'Cabin_class_F': True, 'Cabin_class_G': False, 'Cabin_class_T': False, 'Cabin_class_None': True, 'Age_missing': True}, 'neural_network': {'contribution': 0.0, 'layers': [128, 128, 128]}, 'boosted_tree': {'contribution': 1.0, 'params': {'max_depth': 5, 'learning_rate': 0.05749106357588373, 'n_estimators': 5682, 'min_child_weight': 5, 'colsample_bytree': 0.8034047905616921, 'subsample': 0.7230299010267518, 'reg_alpha': 0.18436505849416285, 'reg_lambda': 0.7145550969871013}}, 'test_accuracy': 0.7488},\n{'cabin_submodel':True,'resample':{'sample_size':200,'num_sample':6},'features': {'PassengerId': False, 'Pclass': True, 'Name': False, 'Sex': True, 'Age': True, 'SibSp': False, 'Parch': True, 'Ticket': True, 'Fare': True, 'Cabin': True, 'Embarked': False, 'lnFare': False, 'ReTicket': True, 'ReCabin': True, 'Outlier': False, 'Cabin_class_A': True, 'Cabin_class_B': True, 'Cabin_class_C': True, 'Cabin_class_D': True, 'Cabin_class_E': False, 'Cabin_class_F': True, 'Cabin_class_G': False, 'Cabin_class_T': False, 'Cabin_class_None': True, 'Age_missing': True},'neural_network':{'contribution':1.0,'layers':[128,128]},'boosted_tree':{'contribution':0.0,'params':{'max_depth': 5, 'learning_rate': 0.05749106357588373, 'n_estimators': 5682, 'min_child_weight': 5, 'colsample_bytree': 0.8034047905616921, 'subsample': 0.7230299010267518, 'reg_alpha': 0.18436505849416285, 'reg_lambda': 0.7145550969871013}},'test_accuracy':0.72248},\n{'cabin_submodel':False,'resample':{'sample_size':200,'num_sample':6},'features': {'PassengerId': False, 'Pclass': True, 'Name': False, 'Sex': True, 'Age': True, 'SibSp': False, 'Parch': True, 'Ticket': True, 'Fare': True, 'Cabin': True, 'Embarked': False, 'lnFare': False, 'ReTicket': True, 'ReCabin': True, 'Outlier': False, 'Cabin_class_A': True, 'Cabin_class_B': True, 'Cabin_class_C': True, 'Cabin_class_D': True, 'Cabin_class_E': False, 'Cabin_class_F': True, 'Cabin_class_G': False, 'Cabin_class_T': False, 'Cabin_class_None': True, 'Age_missing': True},'neural_network':{'contribution':1.0,'layers':[128,128]},'boosted_tree':{'contribution':0.0,'params':{'max_depth': 5, 'learning_rate': 0.05749106357588373, 'n_estimators': 5682, 'min_child_weight': 5, 'colsample_bytree': 0.8034047905616921, 'subsample': 0.7230299010267518, 'reg_alpha': 0.18436505849416285, 'reg_lambda': 0.7145550969871013}},'test_accuracy':0.72966},\n{'cabin_submodel':False,'resample':{'sample_size':200,'num_sample':6},'features': {'PassengerId': False, 'Pclass': True, 'Name': False, 'Sex': True, 'Age': True, 'SibSp': False, 'Parch': True, 'Ticket': True, 'Fare': True, 'Cabin': True, 'Embarked': False, 'lnFare': False, 'ReTicket': True, 'ReCabin': True, 'Outlier': False, 'Cabin_class_A': True, 'Cabin_class_B': True, 'Cabin_class_C': True, 'Cabin_class_D': True, 'Cabin_class_E': False, 'Cabin_class_F': True, 'Cabin_class_G': False, 'Cabin_class_T': False, 'Cabin_class_None': True, 'Age_missing': True},'neural_network':{'contribution':1.0,'layers':[128,128]},'boosted_tree':{'contribution':0.0,'params':{'max_depth': 5, 'learning_rate': 0.05749106357588373, 'n_estimators': 5682, 'min_child_weight': 5, 'colsample_bytree': 0.8034047905616921, 'subsample': 0.7230299010267518, 'reg_alpha': 0.18436505849416285, 'reg_lambda': 0.7145550969871013}},'test_accuracy':0.73684},\n{'cabin_submodel':False,'resample':{'sample_size':200,'num_sample':6},'features': {'PassengerId': False, 'Pclass': True, 'Name': False, 'Sex': True, 'Age': True, 'SibSp': False, 'Parch': True, 'Ticket': True, 'Fare': True, 'Cabin': True, 'Embarked': False, 'lnFare': False, 'ReTicket': True, 'ReCabin': True, 'Outlier': False, 'Cabin_class_A': True, 'Cabin_class_B': True, 'Cabin_class_C': True, 'Cabin_class_D': True, 'Cabin_class_E': False, 'Cabin_class_F': True, 'Cabin_class_G': False, 'Cabin_class_T': False, 'Cabin_class_None': True, 'Age_missing': True},'neural_network':{'contribution':1.0,'layers':[128,256]},'boosted_tree':{'contribution':0.0,'params':{'max_depth': 5, 'learning_rate': 0.05749106357588373, 'n_estimators': 5682, 'min_child_weight': 5, 'colsample_bytree': 0.8034047905616921, 'subsample': 0.7230299010267518, 'reg_alpha': 0.18436505849416285, 'reg_lambda': 0.7145550969871013}},'test_accuracy':0.73444},\n{'cabin_submodel':False,'resample':{'sample_size':200,'num_sample':6},'features': {'PassengerId': False, 'Pclass': True, 'Name': False, 'Sex': True, 'Age': True, 'SibSp': False, 'Parch': True, 'Ticket': True, 'Fare': True, 'Cabin': True, 'Embarked': False, 'lnFare': False, 'ReTicket': True, 'ReCabin': True, 'Outlier': False, 'Cabin_class_A': True, 'Cabin_class_B': True, 'Cabin_class_C': True, 'Cabin_class_D': True, 'Cabin_class_E': False, 'Cabin_class_F': True, 'Cabin_class_G': False, 'Cabin_class_T': False, 'Cabin_class_None': True, 'Age_missing': True},'neural_network':{'contribution':1.0,'layers':[128,256]},'boosted_tree':{'contribution':0.0,'params':{'max_depth': 5, 'learning_rate': 0.05749106357588373, 'n_estimators': 5682, 'min_child_weight': 5, 'colsample_bytree': 0.8034047905616921, 'subsample': 0.7230299010267518, 'reg_alpha': 0.18436505849416285, 'reg_lambda': 0.7145550969871013}},'test_accuracy':0.75358},\n{'cabin_submodel':False,'resample':{'sample_size':200,'num_sample':6},'features': {'PassengerId': False, 'Pclass': True, 'Name': False, 'Sex': True, 'Age': True, 'SibSp': False, 'Parch': True, 'Ticket': True, 'Fare': True, 'Cabin': True, 'Embarked': False, 'lnFare': False, 'ReTicket': True, 'ReCabin': True, 'Outlier': False, 'Cabin_class_A': True, 'Cabin_class_B': True, 'Cabin_class_C': True, 'Cabin_class_D': True, 'Cabin_class_E': False, 'Cabin_class_F': True, 'Cabin_class_G': False, 'Cabin_class_T': False, 'Cabin_class_None': True, 'Age_missing': True},'neural_network':{'contribution':1.0,'layers':[128,256]},'boosted_tree':{'contribution':0.0,'params':{'max_depth': 5, 'learning_rate': 0.05749106357588373, 'n_estimators': 5682, 'min_child_weight': 5, 'colsample_bytree': 0.8034047905616921, 'subsample': 0.7230299010267518, 'reg_alpha': 0.18436505849416285, 'reg_lambda': 0.7145550969871013}},'test_accuracy':0.7488},\n{'cabin_submodel':False,'resample':{'sample_size':200,'num_sample':6},'features': {'PassengerId': False, 'Pclass': True, 'Name': False, 'Sex': True, 'Age': True, 'SibSp': False, 'Parch': True, 'Ticket': True, 'Fare': True, 'Cabin': True, 'Embarked': False, 'lnFare': False, 'ReTicket': True, 'ReCabin': True, 'Outlier': False, 'Cabin_class_A': True, 'Cabin_class_B': True, 'Cabin_class_C': True, 'Cabin_class_D': True, 'Cabin_class_E': False, 'Cabin_class_F': True, 'Cabin_class_G': False, 'Cabin_class_T': False, 'Cabin_class_None': True, 'Age_missing': True},'neural_network':{'contribution':1.0,'layers':[128]},'boosted_tree':{'contribution':0.0,'params':{'max_depth': 5, 'learning_rate': 0.05749106357588373, 'n_estimators': 5682, 'min_child_weight': 5, 'colsample_bytree': 0.8034047905616921, 'subsample': 0.7230299010267518, 'reg_alpha': 0.18436505849416285, 'reg_lambda': 0.7145550969871013}},'test_accuracy':0.75837},\n{'cabin_submodel':False,'resample':{'sample_size':200,'num_sample':6},'features': {'PassengerId': False, 'Pclass': True, 'Name': False, 'Sex': True, 'Age': True, 'SibSp': False, 'Parch': True, 'Ticket': True, 'Fare': True, 'Cabin': True, 'Embarked': False, 'lnFare': False, 'ReTicket': True, 'ReCabin': True, 'Outlier': False, 'Cabin_class_A': True, 'Cabin_class_B': True, 'Cabin_class_C': True, 'Cabin_class_D': True, 'Cabin_class_E': False, 'Cabin_class_F': True, 'Cabin_class_G': False, 'Cabin_class_T': False, 'Cabin_class_None': True, 'Age_missing': True},'neural_network':{'contribution':1.0,'layers':[128]},'boosted_tree':{'contribution':0.0,'params':{'max_depth': 5, 'learning_rate': 0.05749106357588373, 'n_estimators': 5682, 'min_child_weight': 5, 'colsample_bytree': 0.8034047905616921, 'subsample': 0.7230299010267518, 'reg_alpha': 0.18436505849416285, 'reg_lambda': 0.7145550969871013}},'test_accuracy':0.74641},\n{'cabin_submodel':False,'resample':{'sample_size':200,'num_sample':6},'features': {'PassengerId': False, 'Pclass': True, 'Name': False, 'Sex': True, 'Age': True, 'SibSp': False, 'Parch': True, 'Ticket': True, 'Fare': True, 'Cabin': True, 'Embarked': False, 'lnFare': False, 'ReTicket': True, 'ReCabin': True, 'Outlier': False, 'Cabin_class_A': True, 'Cabin_class_B': True, 'Cabin_class_C': True, 'Cabin_class_D': True, 'Cabin_class_E': False, 'Cabin_class_F': True, 'Cabin_class_G': False, 'Cabin_class_T': False, 'Cabin_class_None': True, 'Age_missing': True},'neural_network':{'contribution':1.0,'layers':[128]},'boosted_tree':{'contribution':0.0,'params':{'max_depth': 5, 'learning_rate': 0.05749106357588373, 'n_estimators': 5682, 'min_child_weight': 5, 'colsample_bytree': 0.8034047905616921, 'subsample': 0.7230299010267518, 'reg_alpha': 0.18436505849416285, 'reg_lambda': 0.7145550969871013}},'test_accuracy':0.75598},\n{'cabin_submodel':False,'resample':{'sample_size':200,'num_sample':6},'features': {'PassengerId': False, 'Pclass': True, 'Name': False, 'Sex': True, 'Age': True, 'SibSp': False, 'Parch': True, 'Ticket': True, 'Fare': True, 'Cabin': True, 'Embarked': False, 'lnFare': False, 'ReTicket': True, 'ReCabin': True, 'Outlier': False, 'Cabin_class_A': True, 'Cabin_class_B': True, 'Cabin_class_C': True, 'Cabin_class_D': True, 'Cabin_class_E': False, 'Cabin_class_F': True, 'Cabin_class_G': False, 'Cabin_class_T': False, 'Cabin_class_None': True, 'Age_missing': True},'neural_network':{'contribution':0.1,'layers':[128]},'boosted_tree':{'contribution':0.9,'params':{'max_depth': 5, 'learning_rate': 0.05749106357588373, 'n_estimators': 5682, 'min_child_weight': 5, 'colsample_bytree': 0.8034047905616921, 'subsample': 0.7230299010267518, 'reg_alpha': 0.18436505849416285, 'reg_lambda': 0.7145550969871013}},'test_accuracy':0.76076},\n{'cabin_submodel':False,'resample':{'sample_size':200,'num_sample':6},'features': {'PassengerId': False, 'Pclass': True, 'Name': False, 'Sex': True, 'Age': True, 'SibSp': False, 'Parch': True, 'Ticket': True, 'Fare': True, 'Cabin': True, 'Embarked': False, 'lnFare': False, 'ReTicket': True, 'ReCabin': True, 'Outlier': False, 'Cabin_class_A': True, 'Cabin_class_B': True, 'Cabin_class_C': True, 'Cabin_class_D': True, 'Cabin_class_E': False, 'Cabin_class_F': True, 'Cabin_class_G': False, 'Cabin_class_T': False, 'Cabin_class_None': True, 'Age_missing': True},'neural_network':{'contribution':0.1,'layers':[128]},'boosted_tree':{'contribution':0.9,'params':{'max_depth': 5, 'learning_rate': 0.05749106357588373, 'n_estimators': 5682, 'min_child_weight': 5, 'colsample_bytree': 0.8034047905616921, 'subsample': 0.7230299010267518, 'reg_alpha': 0.18436505849416285, 'reg_lambda': 0.7145550969871013}},'test_accuracy':0.75119},\n{'cabin_submodel':False,'resample':{'sample_size':200,'num_sample':6},'features': {'PassengerId': False, 'Pclass': True, 'Name': False, 'Sex': True, 'Age': True, 'SibSp': False, 'Parch': True, 'Ticket': True, 'Fare': True, 'Cabin': True, 'Embarked': False, 'lnFare': False, 'ReTicket': True, 'ReCabin': True, 'Outlier': False, 'Cabin_class_A': True, 'Cabin_class_B': True, 'Cabin_class_C': True, 'Cabin_class_D': True, 'Cabin_class_E': False, 'Cabin_class_F': True, 'Cabin_class_G': False, 'Cabin_class_T': False, 'Cabin_class_None': True, 'Age_missing': True},'neural_network':{'contribution':0.2,'layers':[128]},'boosted_tree':{'contribution':0.8,'params':{'max_depth': 5, 'learning_rate': 0.05749106357588373, 'n_estimators': 5682, 'min_child_weight': 5, 'colsample_bytree': 0.8034047905616921, 'subsample': 0.7230299010267518, 'reg_alpha': 0.18436505849416285, 'reg_lambda': 0.7145550969871013}},'test_accuracy':0.72727},\n{'cabin_submodel':False,'resample':{'sample_size':200,'num_sample':6},'features': {'PassengerId': False, 'Pclass': True, 'Name': False, 'Sex': True, 'Age': True, 'SibSp': False, 'Parch': True, 'Ticket': True, 'Fare': True, 'Cabin': True, 'Embarked': False, 'lnFare': False, 'ReTicket': True, 'ReCabin': True, 'Outlier': False, 'Cabin_class_A': True, 'Cabin_class_B': True, 'Cabin_class_C': True, 'Cabin_class_D': True, 'Cabin_class_E': False, 'Cabin_class_F': True, 'Cabin_class_G': False, 'Cabin_class_T': False, 'Cabin_class_None': True, 'Age_missing': True},'neural_network':{'contribution':0.2,'layers':[128]},'boosted_tree':{'contribution':0.8,'params':{'max_depth': 5, 'learning_rate': 0.05749106357588373, 'n_estimators': 5682, 'min_child_weight': 5, 'colsample_bytree': 0.8034047905616921, 'subsample': 0.7230299010267518, 'reg_alpha': 0.18436505849416285, 'reg_lambda': 0.7145550969871013}},'test_accuracy':0.73923},\n{'cabin_submodel':False,'resample':{'sample_size':200,'num_sample':6},'features': {'PassengerId': False, 'Pclass': True, 'Name': False, 'Sex': True, 'Age': True, 'SibSp': False, 'Parch': True, 'Ticket': True, 'Fare': True, 'Cabin': True, 'Embarked': False, 'lnFare': False, 'ReTicket': True, 'ReCabin': True, 'Outlier': False, 'Cabin_class_A': True, 'Cabin_class_B': True, 'Cabin_class_C': True, 'Cabin_class_D': True, 'Cabin_class_E': False, 'Cabin_class_F': True, 'Cabin_class_G': False, 'Cabin_class_T': False, 'Cabin_class_None': True, 'Age_missing': True},'neural_network':{'contribution':0.3,'layers':[128]},'boosted_tree':{'contribution':0.7,'params':{'max_depth': 5, 'learning_rate': 0.05749106357588373, 'n_estimators': 5682, 'min_child_weight': 5, 'colsample_bytree': 0.8034047905616921, 'subsample': 0.7230299010267518, 'reg_alpha': 0.18436505849416285, 'reg_lambda': 0.7145550969871013}},'test_accuracy':0.7177},\n{'cabin_submodel':False,'resample':{'sample_size':200,'num_sample':6},'features': {'PassengerId': False, 'Pclass': True, 'Name': False, 'Sex': True, 'Age': True, 'SibSp': False, 'Parch': True, 'Ticket': True, 'Fare': True, 'Cabin': True, 'Embarked': False, 'lnFare': False, 'ReTicket': True, 'ReCabin': True, 'Outlier': False, 'Cabin_class_A': True, 'Cabin_class_B': True, 'Cabin_class_C': True, 'Cabin_class_D': True, 'Cabin_class_E': False, 'Cabin_class_F': True, 'Cabin_class_G': False, 'Cabin_class_T': False, 'Cabin_class_None': True, 'Age_missing': True},'neural_network':{'contribution':0.3,'layers':[128]},'boosted_tree':{'contribution':0.7,'params':{'max_depth': 5, 'learning_rate': 0.05749106357588373, 'n_estimators': 5682, 'min_child_weight': 5, 'colsample_bytree': 0.8034047905616921, 'subsample': 0.7230299010267518, 'reg_alpha': 0.18436505849416285, 'reg_lambda': 0.7145550969871013}},'test_accuracy':0.73205},\n{'cabin_submodel': False, 'resample': {'sample_size': 200, 'num_sample': 6}, 'features': {'PassengerId': True, 'Pclass': True, 'Name': True, 'Sex': True, 'Age': True, 'SibSp': True, 'Parch': True, 'Ticket': True, 'Fare': True, 'Cabin': True, 'Embarked': True, 'lnFare': False, 'ReTicket': False, 'ReCabin': False, 'Outlier': False, 'Cabin_class_A': False, 'Cabin_class_B': False, 'Cabin_class_C': False, 'Cabin_class_D': False, 'Cabin_class_E': False, 'Cabin_class_F': False, 'Cabin_class_G': False, 'Cabin_class_T': False, 'Cabin_class_None': False, 'Age_missing': False}, 'neural_network': {'contribution': 1.0, 'layers': [128]}, 'boosted_tree': {'contribution': 0.0, 'params': {'max_depth': 5, 'learning_rate': 0.05749106357588373, 'n_estimators': 5682, 'min_child_weight': 5, 'colsample_bytree': 0.8034047905616921, 'subsample': 0.7230299010267518, 'reg_alpha': 0.18436505849416285, 'reg_lambda': 0.7145550969871013}}, 'test_accuracy': 0.7488},\n{'cabin_submodel': False, 'resample': {'sample_size': 200, 'num_sample': 6}, 'features': {'PassengerId': True, 'Pclass': True, 'Name': True, 'Sex': True, 'Age': True, 'SibSp': True, 'Parch': True, 'Ticket': True, 'Fare': True, 'Cabin': True, 'Embarked': True, 'lnFare': False, 'ReTicket': False, 'ReCabin': False, 'Outlier': False, 'Cabin_class_A': False, 'Cabin_class_B': False, 'Cabin_class_C': False, 'Cabin_class_D': False, 'Cabin_class_E': False, 'Cabin_class_F': False, 'Cabin_class_G': False, 'Cabin_class_T': False, 'Cabin_class_None': False, 'Age_missing': False}, 'neural_network': {'contribution': 1.0, 'layers': [128]}, 'boosted_tree': {'contribution': 0.0, 'params': {'max_depth': 5, 'learning_rate': 0.05749106357588373, 'n_estimators': 5682, 'min_child_weight': 5, 'colsample_bytree': 0.8034047905616921, 'subsample': 0.7230299010267518, 'reg_alpha': 0.18436505849416285, 'reg_lambda': 0.7145550969871013}}, 'test_accuracy': 0.746162},\n{'cabin_submodel': False, 'resample': {'sample_size': 200, 'num_sample': 6}, 'features': {'PassengerId': True, 'Pclass': True, 'Name': True, 'Sex': True, 'Age': True, 'SibSp': True, 'Parch': True, 'Ticket': True, 'Fare': True, 'Cabin': True, 'Embarked': True, 'lnFare': False, 'ReTicket': False, 'ReCabin': False, 'Outlier': False, 'Cabin_class_A': False, 'Cabin_class_B': False, 'Cabin_class_C': False, 'Cabin_class_D': False, 'Cabin_class_E': False, 'Cabin_class_F': False, 'Cabin_class_G': False, 'Cabin_class_T': False, 'Cabin_class_None': False, 'Age_missing': False}, 'neural_network': {'contribution': 1.0, 'layers': [128]}, 'boosted_tree': {'contribution': 0.0, 'params': {'max_depth': 5, 'learning_rate': 0.05749106357588373, 'n_estimators': 5682, 'min_child_weight': 5, 'colsample_bytree': 0.8034047905616921, 'subsample': 0.7230299010267518, 'reg_alpha': 0.18436505849416285, 'reg_lambda': 0.7145550969871013}}, 'test_accuracy': 0.75119},\n{'cabin_submodel': False, 'resample': {'sample_size': 200, 'num_sample': 6}, 'features': {'PassengerId': True, 'Pclass': True, 'Name': True, 'Sex': True, 'Age': True, 'SibSp': True, 'Parch': True, 'Ticket': True, 'Fare': True, 'Cabin': True, 'Embarked': True, 'lnFare': False, 'ReTicket': False, 'ReCabin': False, 'Outlier': False, 'Cabin_class_A': False, 'Cabin_class_B': False, 'Cabin_class_C': False, 'Cabin_class_D': False, 'Cabin_class_E': False, 'Cabin_class_F': False, 'Cabin_class_G': False, 'Cabin_class_T': False, 'Cabin_class_None': False, 'Age_missing': False}, 'neural_network': {'contribution': 0.0, 'layers': [128]}, 'boosted_tree': {'contribution': 1.0, 'params': {'max_depth': 5, 'learning_rate': 0.05749106357588373, 'n_estimators': 5682, 'min_child_weight': 5, 'colsample_bytree': 0.8034047905616921, 'subsample': 0.7230299010267518, 'reg_alpha': 0.18436505849416285, 'reg_lambda': 0.7145550969871013}}, 'test_accuracy': 0.77511},\n{'cabin_submodel': False, 'resample': {'sample_size': 200, 'num_sample': 6}, 'features': {'PassengerId': True, 'Pclass': True, 'Name': True, 'Sex': True, 'Age': True, 'SibSp': True, 'Parch': True, 'Ticket': True, 'Fare': True, 'Cabin': True, 'Embarked': True, 'lnFare': False, 'ReTicket': False, 'ReCabin': False, 'Outlier': False, 'Cabin_class_A': False, 'Cabin_class_B': False, 'Cabin_class_C': False, 'Cabin_class_D': False, 'Cabin_class_E': False, 'Cabin_class_F': False, 'Cabin_class_G': False, 'Cabin_class_T': False, 'Cabin_class_None': False, 'Age_missing': False}, 'neural_network': {'contribution': 0.0, 'layers': [128]}, 'boosted_tree': {'contribution': 1.0, 'params': {'max_depth': 5, 'learning_rate': 0.05749106357588373, 'n_estimators': 5682, 'min_child_weight': 5, 'colsample_bytree': 0.8034047905616921, 'subsample': 0.7230299010267518, 'reg_alpha': 0.18436505849416285, 'reg_lambda': 0.7145550969871013}}, 'test_accuracy': 0.77511},\n{'cabin_submodel': False, 'resample': {'sample_size': 200, 'num_sample': 6}, 'features': {'PassengerId': True, 'Pclass': True, 'Name': True, 'Sex': True, 'Age': True, 'SibSp': True, 'Parch': True, 'Ticket': True, 'Fare': True, 'Cabin': True, 'Embarked': True, 'lnFare': False, 'ReTicket': False, 'ReCabin': False, 'Outlier': False, 'Cabin_class_A': False, 'Cabin_class_B': False, 'Cabin_class_C': False, 'Cabin_class_D': False, 'Cabin_class_E': False, 'Cabin_class_F': False, 'Cabin_class_G': False, 'Cabin_class_T': False, 'Cabin_class_None': False, 'Age_missing': False}, 'neural_network': {'contribution': 0.1, 'layers': [128]}, 'boosted_tree': {'contribution': 0.9, 'params': {'max_depth': 5, 'learning_rate': 0.05749106357588373, 'n_estimators': 5682, 'min_child_weight': 5, 'colsample_bytree': 0.8034047905616921, 'subsample': 0.7230299010267518, 'reg_alpha': 0.18436505849416285, 'reg_lambda': 0.7145550969871013}}, 'test_accuracy': 0.76794},\n{'cabin_submodel': False, 'resample': {'sample_size': 200, 'num_sample': 6}, 'features': {'PassengerId': True, 'Pclass': True, 'Name': True, 'Sex': True, 'Age': True, 'SibSp': True, 'Parch': True, 'Ticket': True, 'Fare': True, 'Cabin': True, 'Embarked': True, 'lnFare': False, 'ReTicket': False, 'ReCabin': False, 'Outlier': False, 'Cabin_class_A': False, 'Cabin_class_B': False, 'Cabin_class_C': False, 'Cabin_class_D': False, 'Cabin_class_E': False, 'Cabin_class_F': False, 'Cabin_class_G': False, 'Cabin_class_T': False, 'Cabin_class_None': False, 'Age_missing': False}, 'neural_network': {'contribution': 0.1, 'layers': [128]}, 'boosted_tree': {'contribution': 0.9, 'params': {'max_depth': 5, 'learning_rate': 0.05749106357588373, 'n_estimators': 5682, 'min_child_weight': 5, 'colsample_bytree': 0.8034047905616921, 'subsample': 0.7230299010267518, 'reg_alpha': 0.18436505849416285, 'reg_lambda': 0.7145550969871013}}, 'test_accuracy': 0.76794},\n{'cabin_submodel': False, 'resample': {'sample_size': 200, 'num_sample': 6}, 'features': {'PassengerId': True, 'Pclass': True, 'Name': True, 'Sex': True, 'Age': True, 'SibSp': True, 'Parch': True, 'Ticket': True, 'Fare': True, 'Cabin': True, 'Embarked': True, 'lnFare': False, 'ReTicket': False, 'ReCabin': False, 'Outlier': False, 'Cabin_class_A': False, 'Cabin_class_B': False, 'Cabin_class_C': False, 'Cabin_class_D': False, 'Cabin_class_E': False, 'Cabin_class_F': False, 'Cabin_class_G': False, 'Cabin_class_T': False, 'Cabin_class_None': False, 'Age_missing': False}, 'neural_network': {'contribution': 0.1, 'layers': [128]}, 'boosted_tree': {'contribution': 0.9, 'params': {'max_depth': 5, 'learning_rate': 0.05749106357588373, 'n_estimators': 5682, 'min_child_weight': 5, 'colsample_bytree': 0.8034047905616921, 'subsample': 0.7230299010267518, 'reg_alpha': 0.18436505849416285, 'reg_lambda': 0.7145550969871013}}, 'test_accuracy': 0.76555},\n{'cabin_submodel': False, 'resample': {'sample_size': 200, 'num_sample': 6}, 'features': {'PassengerId': True, 'Pclass': True, 'Name': True, 'Sex': True, 'Age': True, 'SibSp': True, 'Parch': True, 'Ticket': True, 'Fare': True, 'Cabin': True, 'Embarked': True, 'lnFare': True, 'ReTicket': True, 'ReCabin': True, 'Outlier': True, 'Cabin_class_A': True, 'Cabin_class_B': True, 'Cabin_class_C': True, 'Cabin_class_D': True, 'Cabin_class_E': True, 'Cabin_class_F': True, 'Cabin_class_G': True, 'Cabin_class_T': True, 'Cabin_class_None': True, 'Age_missing': True}, 'neural_network': {'contribution': 1.0, 'layers': [128]}, 'boosted_tree': {'contribution': 0.0, 'params': {'max_depth': 5, 'learning_rate': 0.05749106357588373, 'n_estimators': 5682, 'min_child_weight': 5, 'colsample_bytree': 0.8034047905616921, 'subsample': 0.7230299010267518, 'reg_alpha': 0.18436505849416285, 'reg_lambda': 0.7145550969871013}}, 'test_accuracy': 0.74401},\n{'cabin_submodel': False, 'resample': {'sample_size': 200, 'num_sample': 6}, 'features': {'PassengerId': True, 'Pclass': True, 'Name': True, 'Sex': True, 'Age': True, 'SibSp': True, 'Parch': True, 'Ticket': True, 'Fare': True, 'Cabin': True, 'Embarked': True, 'lnFare': True, 'ReTicket': True, 'ReCabin': True, 'Outlier': True, 'Cabin_class_A': True, 'Cabin_class_B': True, 'Cabin_class_C': True, 'Cabin_class_D': True, 'Cabin_class_E': True, 'Cabin_class_F': True, 'Cabin_class_G': True, 'Cabin_class_T': True, 'Cabin_class_None': True, 'Age_missing': True}, 'neural_network': {'contribution': 1.0, 'layers': [128]}, 'boosted_tree': {'contribution': 0.0, 'params': {'max_depth': 5, 'learning_rate': 0.05749106357588373, 'n_estimators': 5682, 'min_child_weight': 5, 'colsample_bytree': 0.8034047905616921, 'subsample': 0.7230299010267518, 'reg_alpha': 0.18436505849416285, 'reg_lambda': 0.7145550969871013}}, 'test_accuracy': 0.74401},\n{'cabin_submodel': False, 'resample': {'sample_size': 200, 'num_sample': 6}, 'features': {'PassengerId': True, 'Pclass': True, 'Name': True, 'Sex': True, 'Age': True, 'SibSp': True, 'Parch': True, 'Ticket': True, 'Fare': True, 'Cabin': True, 'Embarked': True, 'lnFare': True, 'ReTicket': True, 'ReCabin': True, 'Outlier': True, 'Cabin_class_A': True, 'Cabin_class_B': True, 'Cabin_class_C': True, 'Cabin_class_D': True, 'Cabin_class_E': True, 'Cabin_class_F': True, 'Cabin_class_G': True, 'Cabin_class_T': True, 'Cabin_class_None': True, 'Age_missing': True}, 'neural_network': {'contribution': 0.0, 'layers': [128]}, 'boosted_tree': {'contribution': 1.0, 'params': {'max_depth': 5, 'learning_rate': 0.05749106357588373, 'n_estimators': 5682, 'min_child_weight': 5, 'colsample_bytree': 0.8034047905616921, 'subsample': 0.7230299010267518, 'reg_alpha': 0.18436505849416285, 'reg_lambda': 0.7145550969871013}}, 'test_accuracy': 0.77751},\n{'cabin_submodel': False, 'resample': {'sample_size': 200, 'num_sample': 6}, 'features': {'PassengerId': True, 'Pclass': True, 'Name': True, 'Sex': True, 'Age': True, 'SibSp': True, 'Parch': True, 'Ticket': True, 'Fare': True, 'Cabin': True, 'Embarked': True, 'lnFare': True, 'ReTicket': True, 'ReCabin': True, 'Outlier': True, 'Cabin_class_A': True, 'Cabin_class_B': True, 'Cabin_class_C': True, 'Cabin_class_D': True, 'Cabin_class_E': True, 'Cabin_class_F': True, 'Cabin_class_G': True, 'Cabin_class_T': True, 'Cabin_class_None': True, 'Age_missing': True}, 'neural_network': {'contribution': 0.0, 'layers': [128]}, 'boosted_tree': {'contribution': 1.0, 'params': {'max_depth': 5, 'learning_rate': 0.05749106357588373, 'n_estimators': 5682, 'min_child_weight': 5, 'colsample_bytree': 0.8034047905616921, 'subsample': 0.7230299010267518, 'reg_alpha': 0.18436505849416285, 'reg_lambda': 0.7145550969871013}}, 'test_accuracy': 0.75358},\n{'cabin_submodel': False, 'resample': {'sample_size': 200, 'num_sample': 6}, 'features': {'PassengerId': True, 'Pclass': True, 'Name': True, 'Sex': True, 'Age': True, 'SibSp': True, 'Parch': True, 'Ticket': True, 'Fare': True, 'Cabin': True, 'Embarked': True, 'lnFare': True, 'ReTicket': True, 'ReCabin': True, 'Outlier': True, 'Cabin_class_A': True, 'Cabin_class_B': True, 'Cabin_class_C': True, 'Cabin_class_D': True, 'Cabin_class_E': True, 'Cabin_class_F': True, 'Cabin_class_G': True, 'Cabin_class_T': True, 'Cabin_class_None': True, 'Age_missing': True}, 'neural_network': {'contribution': 0.0, 'layers': [128]}, 'boosted_tree': {'contribution': 1.0, 'params': {'max_depth': 5, 'learning_rate': 0.05749106357588373, 'n_estimators': 5682, 'min_child_weight': 5, 'colsample_bytree': 0.8034047905616921, 'subsample': 0.7230299010267518, 'reg_alpha': 0.18436505849416285, 'reg_lambda': 0.7145550969871013}}, 'test_accuracy': 0.75837},\n{'cabin_submodel': False, 'resample': {'sample_size': 200, 'num_sample': 6}, 'features': {'PassengerId': True, 'Pclass': True, 'Name': True, 'Sex': True, 'Age': True, 'SibSp': True, 'Parch': True, 'Ticket': True, 'Fare': True, 'Cabin': True, 'Embarked': True, 'lnFare': True, 'ReTicket': True, 'ReCabin': True, 'Outlier': True, 'Cabin_class_A': True, 'Cabin_class_B': True, 'Cabin_class_C': True, 'Cabin_class_D': True, 'Cabin_class_E': True, 'Cabin_class_F': True, 'Cabin_class_G': True, 'Cabin_class_T': True, 'Cabin_class_None': True, 'Age_missing': True}, 'neural_network': {'contribution': 0.1, 'layers': [128]}, 'boosted_tree': {'contribution': 0.9, 'params': {'max_depth': 5, 'learning_rate': 0.05749106357588373, 'n_estimators': 5682, 'min_child_weight': 5, 'colsample_bytree': 0.8034047905616921, 'subsample': 0.7230299010267518, 'reg_alpha': 0.18436505849416285, 'reg_lambda': 0.7145550969871013}}, 'test_accuracy': 0.75837},\n{'cabin_submodel': False, 'resample': {'sample_size': 200, 'num_sample': 6}, 'features': {'PassengerId': True, 'Pclass': True, 'Name': True, 'Sex': True, 'Age': True, 'SibSp': True, 'Parch': True, 'Ticket': True, 'Fare': True, 'Cabin': True, 'Embarked': True, 'lnFare': True, 'ReTicket': True, 'ReCabin': True, 'Outlier': True, 'Cabin_class_A': True, 'Cabin_class_B': True, 'Cabin_class_C': True, 'Cabin_class_D': True, 'Cabin_class_E': True, 'Cabin_class_F': True, 'Cabin_class_G': True, 'Cabin_class_T': True, 'Cabin_class_None': True, 'Age_missing': True}, 'neural_network': {'contribution': 0.1, 'layers': [128]}, 'boosted_tree': {'contribution': 0.9, 'params': {'max_depth': 5, 'learning_rate': 0.05749106357588373, 'n_estimators': 5682, 'min_child_weight': 5, 'colsample_bytree': 0.8034047905616921, 'subsample': 0.7230299010267518, 'reg_alpha': 0.18436505849416285, 'reg_lambda': 0.7145550969871013}}, 'test_accuracy': 0.76315},\n{'cabin_submodel': False, 'resample': {'sample_size': 200, 'num_sample': 6}, 'features': {'PassengerId': False, 'Pclass': True, 'Name': False, 'Sex': True, 'Age': True, 'SibSp': True, 'Parch': True, 'Ticket': True, 'Fare': True, 'Cabin': True, 'Embarked': True, 'lnFare': True, 'ReTicket': True, 'ReCabin': True, 'Outlier': True, 'Cabin_class_A': False, 'Cabin_class_B': False, 'Cabin_class_C': False, 'Cabin_class_D': False, 'Cabin_class_E': False, 'Cabin_class_F': False, 'Cabin_class_G': False, 'Cabin_class_T': False, 'Cabin_class_None': True, 'Age_missing': True}, 'neural_network': {'contribution': 1.0, 'layers': [128]}, 'boosted_tree': {'contribution': 0.0, 'params': {'max_depth': 5, 'learning_rate': 0.05749106357588373, 'n_estimators': 5682, 'min_child_weight': 5, 'colsample_bytree': 0.8034047905616921, 'subsample': 0.7230299010267518, 'reg_alpha': 0.18436505849416285, 'reg_lambda': 0.7145550969871013}}, 'test_accuracy': 0.74162},\n{'cabin_submodel': False, 'resample': {'sample_size': 200, 'num_sample': 6}, 'features': {'PassengerId': False, 'Pclass': True, 'Name': False, 'Sex': True, 'Age': True, 'SibSp': True, 'Parch': True, 'Ticket': True, 'Fare': True, 'Cabin': True, 'Embarked': True, 'lnFare': True, 'ReTicket': True, 'ReCabin': True, 'Outlier': True, 'Cabin_class_A': False, 'Cabin_class_B': False, 'Cabin_class_C': False, 'Cabin_class_D': False, 'Cabin_class_E': False, 'Cabin_class_F': False, 'Cabin_class_G': False, 'Cabin_class_T': False, 'Cabin_class_None': True, 'Age_missing': True}, 'neural_network': {'contribution': 1.0, 'layers': [128]}, 'boosted_tree': {'contribution': 0.0, 'params': {'max_depth': 5, 'learning_rate': 0.05749106357588373, 'n_estimators': 5682, 'min_child_weight': 5, 'colsample_bytree': 0.8034047905616921, 'subsample': 0.7230299010267518, 'reg_alpha': 0.18436505849416285, 'reg_lambda': 0.7145550969871013}}, 'test_accuracy': 0.74641},\n{'cabin_submodel': False, 'resample': {'sample_size': 200, 'num_sample': 6}, 'features': {'PassengerId': False, 'Pclass': True, 'Name': False, 'Sex': True, 'Age': True, 'SibSp': True, 'Parch': True, 'Ticket': True, 'Fare': True, 'Cabin': True, 'Embarked': True, 'lnFare': True, 'ReTicket': True, 'ReCabin': True, 'Outlier': True, 'Cabin_class_A': False, 'Cabin_class_B': False, 'Cabin_class_C': False, 'Cabin_class_D': False, 'Cabin_class_E': False, 'Cabin_class_F': False, 'Cabin_class_G': False, 'Cabin_class_T': False, 'Cabin_class_None': True, 'Age_missing': True}, 'neural_network': {'contribution': 1.0, 'layers': [128]}, 'boosted_tree': {'contribution': 0.0, 'params': {'max_depth': 5, 'learning_rate': 0.05749106357588373, 'n_estimators': 5682, 'min_child_weight': 5, 'colsample_bytree': 0.8034047905616921, 'subsample': 0.7230299010267518, 'reg_alpha': 0.18436505849416285, 'reg_lambda': 0.7145550969871013}}, 'test_accuracy': 0.7488},\n{'cabin_submodel': False, 'resample': {'sample_size': 200, 'num_sample': 6}, 'features': {'PassengerId': False, 'Pclass': True, 'Name': False, 'Sex': True, 'Age': True, 'SibSp': True, 'Parch': True, 'Ticket': True, 'Fare': True, 'Cabin': True, 'Embarked': True, 'lnFare': True, 'ReTicket': True, 'ReCabin': True, 'Outlier': True, 'Cabin_class_A': False, 'Cabin_class_B': False, 'Cabin_class_C': False, 'Cabin_class_D': False, 'Cabin_class_E': False, 'Cabin_class_F': False, 'Cabin_class_G': False, 'Cabin_class_T': False, 'Cabin_class_None': True, 'Age_missing': True}, 'neural_network': {'contribution': 0.0, 'layers': [128]}, 'boosted_tree': {'contribution': 1.0, 'params': {'max_depth': 5, 'learning_rate': 0.05749106357588373, 'n_estimators': 5682, 'min_child_weight': 5, 'colsample_bytree': 0.8034047905616921, 'subsample': 0.7230299010267518, 'reg_alpha': 0.18436505849416285, 'reg_lambda': 0.7145550969871013}}, 'test_accuracy': 0.76076},\n{'cabin_submodel': False, 'resample': {'sample_size': 200, 'num_sample': 6}, 'features': {'PassengerId': False, 'Pclass': True, 'Name': False, 'Sex': True, 'Age': True, 'SibSp': True, 'Parch': True, 'Ticket': True, 'Fare': True, 'Cabin': True, 'Embarked': True, 'lnFare': True, 'ReTicket': True, 'ReCabin': True, 'Outlier': True, 'Cabin_class_A': False, 'Cabin_class_B': False, 'Cabin_class_C': False, 'Cabin_class_D': False, 'Cabin_class_E': False, 'Cabin_class_F': False, 'Cabin_class_G': False, 'Cabin_class_T': False, 'Cabin_class_None': True, 'Age_missing': True}, 'neural_network': {'contribution': 0.0, 'layers': [128]}, 'boosted_tree': {'contribution': 1.0, 'params': {'max_depth': 5, 'learning_rate': 0.05749106357588373, 'n_estimators': 5682, 'min_child_weight': 5, 'colsample_bytree': 0.8034047905616921, 'subsample': 0.7230299010267518, 'reg_alpha': 0.18436505849416285, 'reg_lambda': 0.7145550969871013}}, 'test_accuracy': 0.75837},\n{'cabin_submodel': False, 'resample': {'sample_size': 200, 'num_sample': 6}, 'features': {'PassengerId': False, 'Pclass': True, 'Name': False, 'Sex': True, 'Age': True, 'SibSp': True, 'Parch': True, 'Ticket': True, 'Fare': True, 'Cabin': True, 'Embarked': True, 'lnFare': True, 'ReTicket': True, 'ReCabin': True, 'Outlier': True, 'Cabin_class_A': False, 'Cabin_class_B': False, 'Cabin_class_C': False, 'Cabin_class_D': False, 'Cabin_class_E': False, 'Cabin_class_F': False, 'Cabin_class_G': False, 'Cabin_class_T': False, 'Cabin_class_None': True, 'Age_missing': True}, 'neural_network': {'contribution': 0.0, 'layers': [128]}, 'boosted_tree': {'contribution': 1.0, 'params': {'max_depth': 5, 'learning_rate': 0.05749106357588373, 'n_estimators': 5682, 'min_child_weight': 5, 'colsample_bytree': 0.8034047905616921, 'subsample': 0.7230299010267518, 'reg_alpha': 0.18436505849416285, 'reg_lambda': 0.7145550969871013}}, 'test_accuracy': 0.75837},\n{'cabin_submodel': False, 'resample': {'sample_size': 200, 'num_sample': 6}, 'features': {'PassengerId': False, 'Pclass': True, 'Name': False, 'Sex': True, 'Age': True, 'SibSp': True, 'Parch': True, 'Ticket': True, 'Fare': True, 'Cabin': True, 'Embarked': True, 'lnFare': True, 'ReTicket': True, 'ReCabin': True, 'Outlier': True, 'Cabin_class_A': False, 'Cabin_class_B': False, 'Cabin_class_C': False, 'Cabin_class_D': False, 'Cabin_class_E': False, 'Cabin_class_F': False, 'Cabin_class_G': False, 'Cabin_class_T': False, 'Cabin_class_None': True, 'Age_missing': True}, 'neural_network': {'contribution': 0.1, 'layers': [128]}, 'boosted_tree': {'contribution': 0.9, 'params': {'max_depth': 5, 'learning_rate': 0.05749106357588373, 'n_estimators': 5682, 'min_child_weight': 5, 'colsample_bytree': 0.8034047905616921, 'subsample': 0.7230299010267518, 'reg_alpha': 0.18436505849416285, 'reg_lambda': 0.7145550969871013}}, 'test_accuracy': 0.7488},\n{'cabin_submodel': False, 'resample': {'sample_size': 200, 'num_sample': 6}, 'features': {'PassengerId': False, 'Pclass': True, 'Name': False, 'Sex': True, 'Age': True, 'SibSp': True, 'Parch': True, 'Ticket': True, 'Fare': True, 'Cabin': True, 'Embarked': True, 'lnFare': True, 'ReTicket': True, 'ReCabin': True, 'Outlier': True, 'Cabin_class_A': False, 'Cabin_class_B': False, 'Cabin_class_C': False, 'Cabin_class_D': False, 'Cabin_class_E': False, 'Cabin_class_F': False, 'Cabin_class_G': False, 'Cabin_class_T': False, 'Cabin_class_None': True, 'Age_missing': True}, 'neural_network': {'contribution': 0.1, 'layers': [128]}, 'boosted_tree': {'contribution': 0.9, 'params': {'max_depth': 5, 'learning_rate': 0.05749106357588373, 'n_estimators': 5682, 'min_child_weight': 5, 'colsample_bytree': 0.8034047905616921, 'subsample': 0.7230299010267518, 'reg_alpha': 0.18436505849416285, 'reg_lambda': 0.7145550969871013}}, 'test_accuracy': 0.73923},\n{'cabin_submodel': False, 'resample': {'sample_size': 200, 'num_sample': 6}, 'features': {'PassengerId': False, 'Pclass': True, 'Name': False, 'Sex': True, 'Age': True, 'SibSp': True, 'Parch': True, 'Ticket': True, 'Fare': True, 'Cabin': True, 'Embarked': True, 'lnFare': True, 'ReTicket': True, 'ReCabin': True, 'Outlier': True, 'Cabin_class_A': False, 'Cabin_class_B': False, 'Cabin_class_C': False, 'Cabin_class_D': False, 'Cabin_class_E': False, 'Cabin_class_F': False, 'Cabin_class_G': False, 'Cabin_class_T': False, 'Cabin_class_None': True, 'Age_missing': True}, 'neural_network': {'contribution': 0.1, 'layers': [128]}, 'boosted_tree': {'contribution': 0.9, 'params': {'max_depth': 5, 'learning_rate': 0.05749106357588373, 'n_estimators': 5682, 'min_child_weight': 5, 'colsample_bytree': 0.8034047905616921, 'subsample': 0.7230299010267518, 'reg_alpha': 0.18436505849416285, 'reg_lambda': 0.7145550969871013}}, 'test_accuracy': 0.75358}]\n\n#just to print some records, we might use it to update the accuracy of the last test\nfor record_index in range(len(perfomance_data)-4,len(perfomance_data)):\n    print(perfomance_data[record_index])\ntry:\n    #the results are also stored in a json file\n    perfomance_data_file= open('perfomanceData.json','x')\n    perfomance_data_file.write(json.dumps(perfomance_data)) \n    perfomance_data_file.close()\n    print('\\x1b[42mSUCCESSFULLY WROTE TO \\x1b[5;35;42mperfomanceData.json\\x1b[m\\x1b[m')\n    \nexcept:\n    print('\\x1b[41mCouldn\\'t write to \\x1b[5;37;41mperfomanceData.json\\x1b[m\\x1b[m')\n        ","metadata":{"execution":{"iopub.status.busy":"2022-07-25T00:41:55.739905Z","iopub.execute_input":"2022-07-25T00:41:55.740157Z","iopub.status.idle":"2022-07-25T00:41:55.928199Z","shell.execute_reply.started":"2022-07-25T00:41:55.740124Z","shell.execute_reply":"2022-07-25T00:41:55.927447Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"In order to decide whether or not the models show a significant difference the test accuracy was computed in each of the next cases:\n1. Is there any difference using the **cabin sub model** or not?\n2. Is any of the models **boosted tree**, **neural network** and **combined model**, any better?  \n3. Neural model architecture: Is a deeper or wider neural network better?\n4. Feature choice: By adding new features to the former ones, is there any perfomance improvement?\nEven when the assumptions for the Anova aren't properly checked I added the standard deviation and skewness to verify that the data don't deviate much from this assumptions.  There could also have been a test for variations on the resample parameters but this was not made because this were kept constant along the tests.","metadata":{}},{"cell_type":"code","source":"from scipy.stats import f_oneway,skew,kurtosis\n\n\nsummary_performance=[]\n\n#these lambda functions are used to get the elements(number of cells) of the layers array, it has to be done in this way since not all layers lists have the\n#same length in which case NaN is returned\nlayer_one=lambda x:x['neural_network']['layers'][0] if 0<x['neural_network']['contribution'] else np.nan    \nlayer_two=lambda x: x['neural_network']['layers'][1] if (0<x['neural_network']['contribution'] and 2<=len(x['neural_network']['layers'])) else np.nan   \nlayer_three=lambda x: x['neural_network']['layers'][2] if (0<x['neural_network']['contribution'] and 3<=len(x['neural_network']['layers'])) else np.nan   \n\n#extracting a summary dataframe to perform the anova tests \nfor record in perfomance_data:\n    summary_performance.append([val for val in record['features'].values()]+[record['cabin_submodel'],record['neural_network']['contribution'],record['boosted_tree']['contribution'],layer_one(record),layer_two(record),layer_three(record),record['test_accuracy']])\nsummary_performance=pd.DataFrame(summary_performance,columns=whole_model_features+['cabin_submodel','nn_contribution','tree_contribution','layer_one','layer_two','layer_three','test_accuracy'])\n\n\nprint(\"\\x1b[44mCabin submodel  comparison\\x1b[m\")\ntemp_df=summary_performance[['cabin_submodel','test_accuracy']].groupby(by=['cabin_submodel'])\ndisplay(temp_df.aggregate(['mean','count','std',skew,kurtosis]))\n\nprint(\"\\n\\x1b[42mAnova for cabin submodel  comparison\\x1b[m\\n\")\nanova_test=f_oneway(*[temp_df.get_group(x)['test_accuracy'] for x in temp_df.groups])\nprint('Statistic:{} p-value:{}'.format(anova_test[0],anova_test[1]))   \n\nprint(\"\\n\\x1b[44mModel  comparison\\x1b[m\\n\")\ntemp_df=summary_performance[['nn_contribution','tree_contribution','test_accuracy']].groupby(by=['nn_contribution','tree_contribution'])\ndisplay(temp_df.aggregate(['mean','count','std',skew,kurtosis]))\nprint(\"\\n\\x1b[42mAnova for model  comparison\\x1b[m\\n\")\nanova_test=f_oneway(*[temp_df.get_group(x)['test_accuracy'] for x in temp_df.groups])\nprint('Statistic:{} p-value:{}'.format(anova_test[0],anova_test[1]))\nprint('\\n\\x1b[44mNeural network architecture comparison\\x1b[m\\n')\ntemp_df= summary_performance[['layer_one','layer_two','layer_three','test_accuracy']]\n#Changed the NaN's to 0.0.Needed to run an anova test or rather for the inbetween step just before it (inside list comprehension). \ntemp_df=temp_df.fillna(0)\n#if you ommit last step in other to get the NaN rows you used the argument dropna=False\ntemp_df=temp_df.groupby(by=['layer_one','layer_two','layer_three'])\ndisplay(temp_df.aggregate(['mean','count','std',skew,kurtosis]))\nprint(\"\\n\\x1b[42mAnova for neural network  comparison\\x1b[m\\n\")\nanova_test=f_oneway(*[temp_df.get_group(x)['test_accuracy'] for x in temp_df.groups])\nprint('Statistic:{} p-value:{}'.format(anova_test[0],anova_test[1]))  \n\nprint('\\n\\x1b[44mFeature choice comparison\\x1b[m\\n')\ntemp_df=summary_performance[whole_model_features+['test_accuracy']].groupby(by=whole_model_features)\ndisplay(temp_df.aggregate(['mean','count','std',skew,kurtosis]))\nprint(\"\\n\\x1b[42mAnova for feature  comparison\\x1b[m\\n\")\nanova_test=f_oneway(*[temp_df.get_group(x)['test_accuracy'] for x in temp_df.groups])\nprint('Statistic:{} p-value:{}'.format(anova_test[0],anova_test[1]))  \ndel summary_performance,layer_one,layer_two,layer_three,temp_df,anova_test","metadata":{"execution":{"iopub.status.busy":"2022-07-25T00:41:55.929529Z","iopub.execute_input":"2022-07-25T00:41:55.930316Z","iopub.status.idle":"2022-07-25T00:41:56.082035Z","shell.execute_reply.started":"2022-07-25T00:41:55.930284Z","shell.execute_reply":"2022-07-25T00:41:56.080712Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Besides here is the statistics from the whole test_accuracy distribution,the mean in particularhas a value of of **0749**. Aside from this is noted that the model consisting of zeros returns an accuracy of **0.62200** while the random classifier that predicts ***0*** with **0.622** probability has an average accuracy of **0.54**. Thus, the gain from these models are **20.42%** and **38.74%** respectively which are  modest gains. ","metadata":{}},{"cell_type":"code","source":"all_test_accuraccy=[record['test_accuracy'] for record in perfomance_data]\nprint('mean:{}\\nstandard variation:{}\\nskew:{}\\nkurtosis:{}\\nmin value:{}\\nmax value:{}'.format(np.mean(all_test_accuraccy),np.std(all_test_accuraccy),skew(all_test_accuraccy),kurtosis(all_test_accuraccy),min(all_test_accuraccy),max(all_test_accuraccy)))\nsns.kdeplot(all_test_accuraccy)\n","metadata":{"execution":{"iopub.status.busy":"2022-07-25T00:41:56.083581Z","iopub.execute_input":"2022-07-25T00:41:56.083809Z","iopub.status.idle":"2022-07-25T00:41:56.265931Z","shell.execute_reply.started":"2022-07-25T00:41:56.083782Z","shell.execute_reply":"2022-07-25T00:41:56.265039Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Conclusion\n\nFor the tests and cases tested there was no real difference detected for the cases used in this work. Even when there is some models  that seem to do better than others, on average the difference is not that great so as to be signicant. The boosted tree model tree only has the advantage of not requiring so much data preprocessing as a neural network model but the downside of this model is that its hyperparameters have to be tuned. Besides there could be an important differece if we hadn't used resample while using the neural network model, as the model tends to predict all clases as 0's .It is possible that the features created aren't meaninful or that there are some errors affecting the performance of the models. There might be some other models or features or additional data that really show a significant improvement, nevertherless, is worth noting that the dataset is rather small and there is a lot of data missing, mainly from Cabin, which seems to be a critical feature. \n","metadata":{}}]}