{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"pygments_lexer":"ipython3","nbconvert_exporter":"python","version":"3.6.4","file_extension":".py","codemirror_mode":{"name":"ipython","version":3},"name":"python","mimetype":"text/x-python"}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"I want to thank :\n- @odins0n for his inspiring notebook https://www.kaggle.com/code/odins0n/spaceship-titanic-eda-27-different-models\n- @alexteboul for having remind us few features engineering tricks https://www.kaggle.com/competitions/spaceship-titanic/discussion/309693#1863503\n- @kalelpark for the discussion about encodings, and the advise of summing expenses https://www.kaggle.com/competitions/spaceship-titanic/discussion/309803#1863873\n- @rishirajacharya for having made me discover phiK correlation https://www.kaggle.com/competitions/spaceship-titanic/discussion/309513#1863874","metadata":{}},{"cell_type":"markdown","source":"# <u><b>0) Prerequisite for this code</b></u>\n# ","metadata":{}},{"cell_type":"markdown","source":"#### <u><b>Definition of activation booleans and file pathes</b></u> ","metadata":{}},{"cell_type":"code","source":"# Booléen d'activation (ou non) de l'installation des modules/librairies nécessaires au bon fonctionnement de ce Notebook (cf partie 1.1) du Notebook).\ninstallation = False\n\n### Tout ce qui concerne la base de donnée\npath_b2d = \"/kaggle/input/spaceship-titanic/\"\n","metadata":{},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"###### \n#### <u><b>Occasional installation of missing python libraries</b></u> ","metadata":{}},{"cell_type":"code","source":"if installation :\n    !pip install missingno\n    !pip install phik","metadata":{},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"###### \n#### <u><b>Importation of needed libraries/modules/functions</b></u> ","metadata":{}},{"cell_type":"code","source":"import numpy as np\nimport pandas as pd\nimport missingno as mgno\nimport matplotlib.pyplot as plt\nimport seaborn as sns\nfrom re import split\nfrom phik import phik_matrix","metadata":{},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"import warnings\n#warnings.filterwarnings(message=\"The max_iter was reached which means \", action='ignore')\n#warnings.filterwarnings(UserWarning, action='ignore')\nwarnings.filterwarnings(action='ignore')\nfrom sklearn.preprocessing import (\n    LabelEncoder, \n    OneHotEncoder, \n    StandardScaler, \n    QuantileTransformer\n)\nfrom sklearn.linear_model import LogisticRegression\nfrom sklearn.model_selection import GridSearchCV\nfrom sklearn.ensemble import AdaBoostClassifier\nfrom sklearn.neural_network import MLPClassifier\nfrom sklearn.impute import KNNImputer","metadata":{},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"###### \n#### <u><b>Definition of our own personnal constants</b></u> ","metadata":{}},{"cell_type":"code","source":"### Freezing randomness of numpy and sklearn with an integer\nrgn = 420\nnp.random.seed(rgn)","metadata":{},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# \n# <u><b>I) Exploratory Data Analysis</b></u>\n# ","metadata":{}},{"cell_type":"markdown","source":"## <u><b>I.A) Discovering datasets</b></u> \n###### \n","metadata":{}},{"cell_type":"markdown","source":"### <u><b>I.A.1) Loading data</b></u>\n","metadata":{}},{"cell_type":"markdown","source":"Datasets have been uploaded from Kaggle to my own personnal computer. The path to access them is given as the local variable (string) <i>path_fig</i>, defined in the <b>0</b> section.","metadata":{}},{"cell_type":"code","source":"# Respectively : train set, test set, sample_submission\ndf_train, df_test, samsub = (\n    pd.read_csv(path_b2d + 'train.csv'),\n    pd.read_csv(path_b2d + 'test.csv'),\n    pd.read_csv(path_b2d + 'sample_submission.csv')    \n)","metadata":{},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Let's take a quick look at the datasets, loaded as pandas.Datframe. First, the train set :","metadata":{}},{"cell_type":"code","source":"# Train set\ndf_train","metadata":{},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"We can see there is a mix of categorical, numerical, and boolean features as columns of this Dataframe, including - at last postion as \"Transported\" - the target vector.\n\nThen, the test set :","metadata":{}},{"cell_type":"code","source":"# Test set\ndf_test","metadata":{},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Test set gets about two times less elements than train set. We can see that the last column \"Transported\" is missing, since WE HAVE to predict its value.","metadata":{}},{"cell_type":"code","source":"# Sample submission\nsamsub","metadata":{},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"samsub['Transported'].value_counts()","metadata":{},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"As we can see, this is just the format under which our predictions from test set are expected to be put at the end of this notebook. The fact that the unique value is \"False\" is just a - pessimistic -  intitialisation.\n\nWe will no longer use this variable until our job is done.","metadata":{}},{"cell_type":"markdown","source":"###### \n### <u><b>I.A.2) Nature of datasets features</b></u> \n","metadata":{}},{"cell_type":"markdown","source":"We can now go further into analysis. Let's take a look first at the neature of all parameters of this dataest (features AND target). Please note that we don't use the easiest way - Dataframe.dtypes - since that solution would consider features <i>CryoSleep</i> and <i>VIP</i> as \"object\" and not \"bool\", thus making no difference with string features, categorized as \"object\" as well. ","metadata":{}},{"cell_type":"code","source":"print('Column\\t\\tParameter type\\n-------------------------------')\nfor val,col in zip(df_train.loc[0], df_train.columns) :\n    print(f'{col : <15}', type(val)) \ndel val,col","metadata":{"tags":[]},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"We have :\n- 6 quantitative features ;\n- 7 categorical features (5 are strings, the 2 remaining are booleans).\n\nThis is important to note, because we will have to transform those categorical features into numerical features when algorithm time will come. If the transformation is obvious for booleans (True = 1, False = 0 or -1), string objects might be heavier to deal with, regarding the number of classes per features, and the possible relationships between those classes. As a matter of fact, no ordinal relationship would lead us to use <i>one hot encoding</i>, whereas ordinal or hierarchical relationships would make us prefer using <i>label encoding</i>. In the cas of using <i>one hot encoding</i>, the number of features would increase drastically, which brings back memories of curse of dimensionality, we then should be very carefull, and would have to consider if all features are worth being taken as inputs for the algorithm.","metadata":{}},{"cell_type":"markdown","source":"###### \n### <u><b>I.A.3) Missing values</b></u> \n","metadata":{}},{"cell_type":"markdown","source":"Using graphical functions frm Missingno library, we check how many missing values are hidden within datasets and where. ","metadata":{}},{"cell_type":"code","source":"### Ploting bar and matrix diagrams of both train set (lefthand side) and test set (righthand side)\nplt.figure('Missing values', figsize=(40,20)), plt.clf()\nfor i, (df, ds) in enumerate(\n    zip([df_train, df_test], ['train','test'])\n):\n    axe_up, axe_down = plt.subplot(2,2,i+1), plt.subplot(2,2,i+3)\n    axe_up.set_title('Dataset = '+ds+' set', weight='bold', fontsize=20)\n    mgno.bar(df=df, ax=axe_down), mgno.matrix(df=df, ax=axe_up)\ndel axe_down, axe_up, i, df, ds","metadata":{},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"For both datasets, missing values represent never more than few percents of each parameters, with <i>PassengerId</i> and <i>Transported</i> being completely spared bay missing values.\n\nLet's evaluate what would be the impact of cleaning datasets from lines with any missing values, by calculating the percentage of complete lines :","metadata":{}},{"cell_type":"code","source":"for df, ds in zip([df_train, df_test], ['train','test']):\n    df_without_nan = df.dropna(axis=0, how='any')\n    print(\n        f'For {ds} set, {len(df_without_nan)/len(df)*100 : .0f}% of rows don`t have any NaN'\n    )\ndel df, df_without_nan, ds","metadata":{},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"We cannot tolerate to lose 23-24% of our rows, considering the initial size of datasets that aren't that big (in such data classification, datasets can easily have hundreds of thousand of inputs).\n\nWe conclude that we will HAVE to find a way to fill these values.","metadata":{}},{"cell_type":"markdown","source":"######\n## <u><b>I.B) Feature engineering (1)</b></u> \n###### \n","metadata":{}},{"cell_type":"markdown","source":"### <u><b>I.B.1) Concatenating datasets</b></u>\n","metadata":{}},{"cell_type":"markdown","source":"To make the upcomming univariate analysis easier (and then multivariate analysis), we concatenate both train and test dataframes. Furthermore, we create a new column containing the origin of each row (from train ou from test set).","metadata":{}},{"cell_type":"code","source":"# Concatenation of both datasets, along indexes\ndf_all = pd.concat([df_train, df_test], ignore_index=True)\n# Sorting indexes along PassengerId values\ndf_all.sort_values(by='PassengerId', inplace=True)\n# Adding a column, to know from which dataset the curent line comes\ndf_all.insert(\n    df_train.shape[1],\n    'TrainOrTest', \n    ['Test' if np.isnan(val) else 'Train' for val in df_all['Transported']]\n             )\n# Lets take a look at this global dataset\ndf_all","metadata":{},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"We can now erase original datasets to free a bit of space within kernel memory. If necessary, we can still re-upload them quite easily.","metadata":{}},{"cell_type":"code","source":"del df_train, df_test","metadata":{},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"###### \n### <u><b>I.B.2) Categorical features cardinality</b></u>\n","metadata":{}},{"cell_type":"markdown","source":"For the reasons written few sections above, we have to check the cardinality of each numerical features, to know what to do of each feature.","metadata":{}},{"cell_type":"code","source":"print('Column\\t\\tNumber of unique values\\n-------------------------')\nfor col in df_all.columns[:-2] :\n    if df_all.dtypes[col] == 'object' :\n        print(f'{col : <10}\\t{df_all[col].value_counts().shape[0]}') \ndel col","metadata":{},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"As we can read it :\n- <i>PassengerId</i> is made of unique values ONLY, it's an identifiant, therfore it sould not been taken as a feature for our future algorithm\n- <i>HomePlanet</i>, <i>CryoSleep</i>, <i>Destination</i> and <i>VIP</i> contain an handable number of categories. \n\n#### <b>We should investigate at least two things</b> :\n#### - Why does <i>Name</i>'s cardinality isn't equal to 12970, like <i>PassengerId</i> ? Due to <i>Name</i> meaning, we could expect the feature to be an identifiant as well...\n#### - Can we transform <i>Cabin</i> into new features ?","metadata":{}},{"cell_type":"markdown","source":"###### \n### <u><b>I.B.3) <i>Name</i> non unique values</b></u>","metadata":{}},{"cell_type":"markdown","source":"We create a list of all the names that aren't unique.","metadata":{}},{"cell_type":"code","source":"names_multiple = list(\n    df_all['Name'].value_counts()[df_all['Name'].value_counts()>1].index\n)","metadata":{},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"And we use it to focus on corresponding rows :","metadata":{}},{"cell_type":"code","source":"df_all[\n    df_all['Name'].apply(lambda x: x in names_multiple)\n].sort_values(by='Name')","metadata":{},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"It doesn't seem that rows with the same <i>Name</i> value are identical. In fact, those passengers frequently don't even come from the same planet ! \n\nLet's ensure us that there isn't duplicated rows :","metadata":{}},{"cell_type":"code","source":"df_all[\n    df_all['Name'].apply(lambda x: x in names_multiple)\n].sort_values(by='Name').drop_duplicates()","metadata":{},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"The dataset remains unchanged.\n#### <b>CONCLUSIONS :</b> We should no longer bother with <i>Name</i> non unique values. We can almost consider <i>Name</i> as an identifioant, and sort this columns out features list. ","metadata":{}},{"cell_type":"markdown","source":"###### \n### <u><b>I.B.4) Transforming <i>Name</i> paramater into two new features.</b></u>\n","metadata":{}},{"cell_type":"markdown","source":"Nevertheless, we might be in need to split <i>Name</i> into :\n- <i>FirstName</i> ;\n- <i>LastName</i>.\nIndeed, we had seen that, beyond duplicated names, it seemed to have groups of people with same last name, but different forst names, living in the same cabins, suggesting families. Therefore it might help for filling some missing values in the future, who knows ? \n\nLet's do it. We will use python library <i>re</i> and all regular expression related functions to do so.","metadata":{}},{"cell_type":"code","source":"# Transforming a string into a list of 2 strings (or a nan into list of 2 nans)\ndf_all['Name'] = df_all['Name'].apply(\n    lambda x: split(string=x, pattern=' ') if pd.notna(x) else [np.nan, np.nan]\n)","metadata":{},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"### Insertion of news columns, one for each aimed feature\n# Letter corresponding to the Deck\ndf_all.insert(13, 'FirstName', df_all['Name'].apply(lambda x: x[0]))\n# Letter corresponding to the Side\ndf_all.insert(14, 'LastName', df_all['Name'].apply(lambda x: x[1]))","metadata":{},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Droping column corresponding to the old Name feature\ndf_all.drop(axis=1, columns=['Name'], inplace=True)","metadata":{},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"###### \n### <u><b>I.B.5) Transforming <i>Cabin</i> paramater into several new features.</b></u>","metadata":{}},{"cell_type":"markdown","source":"As we saw above, <i>Cabin</i> gets far too many unique values (more than 9000) to be useful in this shape. Indeed, we cannot do one hot encoding on that, this would create uge sparse matrix of numerical features, and be a waste of both computing ressources and time.\n\nThe might be a way to figure it out, considering the fact that any value of <i>Cabin</i> take the shape of a <i>deck</i>/<i>num</i>/<i>side</i> (string/integer/string). \n\nConsidering that \n- the position on the 1912 original Titanic was possibly correlated to both social inequalities among passengers, \n- the proximity with lifeboats (specially the position on which deck), \n\nit may be intersting to split <i>Cabin</i> into 3 features <i>Deck</i>, <i>Num</i> and <i>Side</i> (<i>Deck</i> and <i>Side</i> being categorical with handable numbers of classes), and investigate if those parameters are relevant to be taken account as features for our algorithm.\n\nLet's do it. We will use python library <i>re</i> and all regular expression related functions to do so.","metadata":{}},{"cell_type":"code","source":"# Transforming a string into a list of 3 strings (or a nan into list of 3 nans)\ndf_all['Cabin'] = df_all['Cabin'].apply(\n    lambda x: split(string=x, pattern='/') if pd.notna(x) else [\n        np.nan, np.nan, np.nan\n    ]\n)","metadata":{},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"### Insertion of news columns, one for each aimed feature\n# Letter corresponding to the Deck\ndf_all.insert(4, 'Deck', df_all['Cabin'].apply(lambda x: x[0]))\n# Integer corresponding to the number of Cabin\ndf_all.insert(\n    5, \n    'NumCabin',\n    df_all['Cabin'].apply(\n        lambda x: int(x[1]) if pd.notna(x[1]) else x[1]\n    )\n)\n# Letter corresponding to the Side\ndf_all.insert(6, 'Side', df_all['Cabin'].apply(lambda x: x[2]))","metadata":{},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Droping column corresponding to the old Cabin feature\ndf_all.drop(axis=1, columns=['Cabin'], inplace=True)","metadata":{},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"###### \n### <u><b>I.B.6) Transforming <i>PassengerId</i> paramater into two new features.</b></u>\n","metadata":{}},{"cell_type":"markdown","source":"We have read on Kaggle discussion that some people encourage to proceed similarly but to divide <i>PassengerId</i> into two numbers, splitting each value along the undescore.\n\nLet's do it. We will use python library <i>re</i> and all regular expression related functions to do so. We split <i>PassengerId</i> into :\n- <i>PId1</i>, first part of <i>PassengerId</i>, left size of the underscore ;\n- <i>PId2</i>, second part of <i>PassengerId</i>, right size of the underscore.","metadata":{}},{"cell_type":"code","source":"### Insertion of news columns, one for each aimed feature\n# String corresponding to the first \ndf_all.insert(\n    1, \n    'PId1', \n    df_all['PassengerId'].apply(\n        lambda x : split(string=x, pattern='_')\n    ).apply(lambda x : x[0])\n)\n# String corresponding to the 2nd part of PassengerId\ndf_all.insert(\n    2, \n    'PId2', \n    df_all['PassengerId'].apply(\n        lambda x : split(string=x, pattern='_')\n    ).apply(lambda x : x[1])\n)","metadata":{},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"As we did previously, we will drop the initial <i>PassengerId</i> column, but this time, we explicitly conserve this information, copying <i>PassengerId</i> into a new mono-column pandas.Dataframe object before droping it from <i>df_all</i>.","metadata":{}},{"cell_type":"code","source":"df_pid = df_all['PassengerId'].copy(deep=True)\ndf_all.drop(axis=1, columns=['PassengerId'], inplace=True)","metadata":{},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Please notice that for now, we leave the values taken by these 2 new columns as strings, but that we keep in mind the possibility to turn then to integers later in this notebook, if we think it can be relevant.","metadata":{}},{"cell_type":"code","source":"print('Column\\t\\tNumber of unique values\\n--------------------')\nfor col in df_all.columns[:2] :\n    print(f'{col: <10}\\t{df_all[col].value_counts().shape[0]}') \ndel col","metadata":{},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Ok, due to those numbers of unique value, we will HAVE to turn <i>PId1</i> into a numerical feature. For <i>PId2</i>, it's far less urgent, but we decide to do so, in order to anticipate a potential label encoding in the future.","metadata":{}},{"cell_type":"code","source":"df_all['PId1'] = df_all['PId1'].apply(lambda x: int(x))\ndf_all['PId2'] = df_all['PId2'].apply(lambda x: int(x))","metadata":{},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"###### \n### <u><b>I.B.7) Summing the five luxury expenses features to create a supplementary column.</b></u>\n","metadata":{}},{"cell_type":"code","source":"df_all.insert(\n    15, \n    'TotalExpenses', \n    df_all[\n        ['RoomService', 'FoodCourt', 'ShoppingMall', 'Spa', 'VRDeck']\n    ].sum(axis=1)\n)","metadata":{},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"#### <u><b>Displaying the dataset after this first step of features engineering</b></u>","metadata":{}},{"cell_type":"code","source":"df_all","metadata":{},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"######\n## <u><b>I.C) Univariate EDA.</b></u> \n###### \n","metadata":{}},{"cell_type":"markdown","source":"### <u><b>I.C.1) Using feature cardinality.</b></u>\n","metadata":{}},{"cell_type":"markdown","source":"Rather than <i>dtype</i>, we will use each features' cardinality to split columns between numerical and categorical features. As a matter of fact, a numerical feature taking very few unique values can be considered as a categorical feature...","metadata":{}},{"cell_type":"code","source":"print('Column\\t\\tNumber of unique values\\n--------------------')\nfor col in df_all.columns :\n    print(f'{col: <10}\\t{df_all[col].value_counts().shape[0]}') \ndel col","metadata":{},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Ok, we made our decision, we can explain how we will proceed :\n- <i>PId1</i> is left aside, and is not considered as a feature for our future algorithm, due to its very high number of unique values (far more than any other numerical parameter).\n- <i>PId2</i>, <i>HomePlanet</i>, <i>CryoSleep</i>, <i>Deck</i>, <i>Side</i>, <i>Destination</i> and <i>VIP</i> are considered as categorical features.\n- <i>NumCabin</i>, <i>RoomService</i>, <i>FoodCourt</i>, <i>ShoppingMall</i>, <i>Spa</i> and <i>VRDeck</i> are considered as ordinal features.\n- <i>FirstName</i>, <i>LastName</i> and <i>TrainOrTest</i> won't be use as features for the algorithm, but will be used as tool of EDA (at least during multi-variate EDA).\n- <i>Transported</i> will be our target, of course.","metadata":{}},{"cell_type":"code","source":"num_col = [\n    'Age', 'NumCabin', 'TotalExpenses', 'RoomService', \n    'FoodCourt', 'ShoppingMall', 'Spa', 'VRDeck'\n]\ncat_col = [\n    'PId2', 'HomePlanet', 'CryoSleep', 'Deck', \n    'Side', 'Destination', 'VIP'\n]","metadata":{},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"######\n### <u><b>I.C.2) Ordinal/numerical features</b></u>\n","metadata":{}},{"cell_type":"code","source":"plt.figure('UV EDA, num features', figsize=(28, 14)), plt.clf()\nfor i, col in enumerate(num_col):\n    axe1 = plt.subplot(2, 4, i+1)\n    sns.histplot(\n        data=df_all, x=col, ax=axe1, log_scale=(False, True), \n        label=str(df_all[col].describe())\n    )\n    axe1.legend(loc='best', fontsize=12)\n    if df_all[col].std()>1.5*df_all[col].mean():\n        axe1.set_xlim([-0.1*df_all[col].std(), 5*df_all[col].std()])\ndel i, col, axe1","metadata":{},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Those features take distribution overwhelmed by one mode, which dwells among low taken values (please notice that the y-axis is log scaled to emphasis high range taken values whose counts are really low). This is especially true the five luxury expenses features (<i>RoomService</i>, <i>FoodCourt</i>, <i>ShoppingMall</i>, <i>Spa</i> and <i>VRDeck</i>), for whom 0 is by far the most frequently taken value :","metadata":{}},{"cell_type":"code","source":"print('Column\\t\\tMost freq.\\tOccurences',\n      '\\n----------------------------------------------------')\nfor col in df_all[num_col]:\n    dfvc = df_all[col].value_counts()\n    print(f'{col: <10}\\t{dfvc.index[0]}\\t\\t{dfvc[0]}') \ndel col, dfvc","metadata":{},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"The important imbalance between std and mean should push us to expect potential non linear correlations betweens thos features and the target. It may be advise in the future to do some feature pre-processing to work with less biaised distributions.\n\nIt would interesting to draw the univariate distribution of the sum of these five last features, to see how the overwhelming importance of 0 is reduced.","metadata":{}},{"cell_type":"markdown","source":"It's better, even if 0 is alway the most frequently taken values. We can see that the median is now a non nul value, and that the ration std/mean has decreased to be now around 2, as regard of higher values of ration for the initial five distribution. \n\nWe should consider in the future taking this sum as a supplementary feature for our algorithm. Further analysis should help us decide if taking this new feature :\n- in addition to \n- inspite of \n\nthe five type of original luxury expenses.","metadata":{}},{"cell_type":"markdown","source":"At last, we must check if the univariate distributions remain approximatively identical bewteen train and test sets. We use groupby and describe functions to do so, focusing on mean, std and median of each feature.","metadata":{}},{"cell_type":"code","source":"df_all.groupby('TrainOrTest')[num_col].describe(percentiles=[]).drop(\n    axis=1, level=1, columns=['count', 'min', 'max', 'mean']\n)","metadata":{},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"That's the case, we can't find any significant difference between the two sets.","metadata":{}},{"cell_type":"markdown","source":"######\n### <u><b>I.C.3 Categorical features</b></u>\n","metadata":{}},{"cell_type":"code","source":"plt.figure('UV EDA, cat features', figsize=(28, 14)), plt.clf()\nfor i, col in enumerate(cat_col):\n    axe1 = plt.subplot(2, 4, i+1)\n    sns.countplot(\n        data=df_all, x=col, ax=axe1, label=str(df_all[col].describe())\n    )\n    axe1.legend(loc='best', fontsize=12)\ndel i, col, axe1","metadata":{},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Contrary to ordinal features' case, we have a lot of different situations for categorical features' distributions :\n- <i>PId2</i> sees its taken values' counts decreasing with the actual taken value, 1 being the first and overwhelming mode.\n- <i>HomePlanet</i> takes the value <i>Earth</i> about one out of two, each <i>Europa</i> and <i>Mars</i> counting for approximatively 25% of the occurences, which means that <i>HomePlanet</i>'s distribution is inballanced, but it remains sufficent in term of well representation of each value.\n- Same situation for the boolean <i>CryoSleep</i>, <i>False</i> counting for ~2/3 of the occurences.\n- <i>Deck</i> has two principal mode (<i>F</i> and <i>G</i>, counting together for 66% of the occurences), and on the other hand, one categorie is almost not represented (<i>T</i>) \n- <i>Side</i> is almost perfectly ballanced between its two taken value.\n- <i>Destination</i> starts being very inballanced, even betwenn the two less frequently taken value.\n- <i>VIP</i> is the most inballanced features of all categorical features, to the point where we could start having doubts about its utility as a feature.","metadata":{}},{"cell_type":"markdown","source":"We can conclude that, in term of handable distributions :\n- <i>HomePlanet</i>, <i>CryoSleep</i> and <i>Side</i> are good features to provide for classification.\n- <i>PId2</i>, <i>Deck</i> and <i>Destination</i> might need further transformations to lessen their inballances (for example : considering all unfreqently taken values as one single more frequently taken value).\n- <i>VIP</i> is a terrible feature, maybe we should just drop it out of the features list.","metadata":{}},{"cell_type":"markdown","source":"At least, those inballances are proportionnaly the same between train and test set :","metadata":{}},{"cell_type":"code","source":"plt.figure('UV EDA, cat features, hue train test', figsize=(28, 14))\nplt.clf()\nfor i, col in enumerate(cat_col):\n    axe1 = plt.subplot(2,4,i+1)\n    sns.countplot(data=df_all, x=col, ax=axe1, hue='TrainOrTest')\n    axe1.legend(loc='best', fontsize=12)\ndel i, col, axe1","metadata":{},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"######\n## <u><b>I.D) Multi-variate EDA.</b></u> \n###### \n","metadata":{}},{"cell_type":"markdown","source":"### <u><b>I.D.1) Phi-K correlation</b></u>\n","metadata":{}},{"cell_type":"code","source":"phik_corr_matrix = df_all[df_all['TrainOrTest'] == 'Train'].drop(\n    axis=1,columns=['TrainOrTest', 'PId1', 'FirstName','LastName']\n).phik_matrix().round(2)","metadata":{},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"plt.figure('PhiK CorrMatrix - train set', figsize=(10,10)), plt.clf()\naxe1 = plt.subplot(111)\naxe1.set_title(\n    r'$\\phi$K correlation matrix - train set', fontsize=15, fontweight='bold'\n)\nsns.heatmap(phik_corr_matrix, cbar=True, square=True, annot=True, ax=axe1)\ndel axe1","metadata":{},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"###### \n### <u><b>I.D.2) Focussing on <i>Transported</i></b></u>\n","metadata":{}},{"cell_type":"code","source":"plt.figure('PhiK correlation coefficients with Transported', figsize=(10,7))\nplt.clf()\naxe1 = plt.subplot(111)\naxe1.set_title(\n    r'$\\phi$K coefficents with Transported - train set', \n    fontsize=15, fontweight='bold'\n)\naxe1.set_ylabel(r'$\\phi$K coefficents', fontsize=12, fontweight='bold')\naxe1.set_xlabel('Other features', fontsize=12, fontweight='bold')\nphik_corr_matrix['Transported'][:-1].sort_values().plot.bar(\n    ax=axe1, rot=45\n)\ndel axe1","metadata":{},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"df = df_all[df_all['TrainOrTest'] == 'Train'].drop(\n    axis=1,columns=['TrainOrTest', 'PId1', 'FirstName','LastName']\n)\n### Graphical display\nplt.figure('Transported as hue', figsize=(5*6, 2*6)), plt.clf()\nfor i, col in enumerate(['NumCabin', 'Age', 'TotalExpenses']):\n    axe1 = plt.subplot(2, 5, i+1)\n    sns.histplot(\n        data=df, x=col, ax=axe1, log_scale=(False, True), \n        hue='Transported', palette='tab10'\n    )\n    if df[col].std() > 1.5*df[col].mean():\n        axe1.set_xlim([-0.1*df_all[col].std(), 5*df_all[col].std()])\nfor i, col in enumerate(cat_col):\n    axe1 = plt.subplot(2, 5, i+4)\n    sns.countplot(\n        data=df, x=col, ax=axe1, hue='Transported', palette='tab10'\n    )\ndel i, col, axe1, df","metadata":{},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"###### \n### <u><b>I.D.2) Ordinal features distributions as a function of categorical features</b></u>\n","metadata":{}},{"cell_type":"code","source":"df = df_all[df_all['TrainOrTest'] == 'Train'].drop(\n    axis=1,columns=['TrainOrTest', 'PId1', 'FirstName','LastName']\n)\n### Graphical display\nplt.figure(\n    'MV EDA, num features - categorical hue - few number of cat', \n    figsize=(5*6, 3*6)\n)\nplt.clf()\nfor i, col in enumerate(['Age', 'NumCabin', 'TotalExpenses']):#, 'RoomService', 'Spa', 'VRDeck']):\n    logscale = (False, False) if i < 2 else (False, True)\n    for j, hue in enumerate(\n        ['HomePlanet', 'CryoSleep', 'Side', 'Destination', 'VIP']\n    ):\n        axe1 = plt.subplot(3, 5, i*5 + j + 1)\n        sns.histplot(data=df, x=col, ax=axe1, log_scale=logscale, \n                     hue=hue, palette='tab10')\n        if i == 2:\n            axe1.set_xlim([-0.1*df_all[col].std(), 5*df_all[col].std()])\ndel i, j, col, hue, axe1, logscale, df","metadata":{},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"df = df_all[df_all['TrainOrTest'] == 'Train'].drop(\n    axis=1,columns=['TrainOrTest', 'PId1', 'FirstName','LastName']\n)\n### Graphical display\nplt.figure(\n    'MV EDA, num features - categorical hue - 8 number of cat', \n    figsize=(3*4, 3*4)\n)\nplt.clf()\nfor i, col in enumerate(['Age', 'NumCabin', 'TotalExpenses']):#, 'RoomService', 'Spa', 'VRDeck']):\n    logscale = (False, False) if i < 2 else (False, True)\n    for j, hue in enumerate(['PId2', 'Deck']):\n        axe1 = plt.subplot(3, 2, i*2 + j + 1)\n        sns.histplot(\n            data=df, x=col, ax=axe1, log_scale=logscale, \n            hue=hue, palette='tab10'\n        )\n        if i == 2:\n            axe1.set_xlim([-0.1*df_all[col].std(), 5*df_all[col].std()])        \ndel i, col, hue, axe1, logscale, df","metadata":{},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"###### \n### <u><b>I.D.3) Categorical features distributions as a function of categorical features</b></u>\n","metadata":{}},{"cell_type":"code","source":"df = df_all[df_all['TrainOrTest'] == 'Train'].drop(\n    axis=1, columns=['TrainOrTest', 'PId1', 'FirstName','LastName']\n)\n### Graphical display\nplt.figure('Cat, cat hue', figsize=(4*len(cat_col), 4*len(cat_col)))\nplt.clf()\nfor i, col in enumerate(cat_col):\n    for j, hue in enumerate(cat_col):\n        axe1 = plt.subplot(\n            len(cat_col), len(cat_col), i*len(cat_col) + j + 1\n        )\n        if hue != col:\n            sns.countplot(data=df, x=col, ax=axe1, hue=hue, palette='tab10')\n        else:\n            axe1.axis('off')\ndel i, j, col, hue, axe1, df","metadata":{},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"###### \n### <u><b>I.D.4) An accurated look taken at famillies of passengers.</b></u>\n","metadata":{}},{"cell_type":"markdown","source":"We noticed that a lot of passenger shared the same <i>LastName</i> value, infact, passengers without namesakes are quite rare...","metadata":{}},{"cell_type":"code","source":"### Variables needed for graphical display\nnames_sup = df_all['LastName'].value_counts()[:50].index\ndf1 = df_all[df_all['LastName'].apply(lambda x: x in names_sup)]\ndf2 = df_all['LastName'].value_counts()\n### Graphical display\nplt.figure('Same lastnames', figsize=(14, 7)), plt.clf()\naxe1, axe2 = plt.subplot(121), plt.subplot(122)\n# 50 most frequent names counting\nsns.countplot(data=df1, x='LastName', ax=axe1)\naxe1.set_xticklabels(\n    labels=axe1.get_xticklabels(), fontdict={'rotation': 90}\n)\naxe1.set_title('Occurences of the 50 most frequent names', weight='bold')\n# Occurences of number of passenger with the same lastname\nsns.countplot(data=df2, x=df2.index)\naxe2.set_title('Occurences of namesakes among passengers', weight='bold')\naxe2.set_xlabel('Number of namesakes')\ndel axe1, axe2, df1, df2, names_sup","metadata":{},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"That is to say : there should be famillies among passengers, and probably a lot of famillies.\n\nAnother proof of that would be the following statistics :","metadata":{}},{"cell_type":"code","source":"df_all.groupby(\n    ['Deck','LastName','Side']\n)['NumCabin'].value_counts().value_counts()","metadata":{},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"This means that more than 80% of passengers with the same <i>LastName</i> on the same <i>Deck</i> and the same <i>Side</i> DO SHARE THE SAME CABIN, hence high probability of theses passengers being members of the same famillies.Beyond this interpretation, this an important relationship between those 4 features, and we should keep it in mind when the time to fill missing values will come.\n\nAnother relationship of the same kind has been enlightened by XXX on the Kaggle discussion XXX. It shows that 1st part of <i>PassengerId</i> (<i>ie</i> our personnal column <i>PId1</i>) is quite correlated with <i>LastName</i>, more than 75% of passengers with the same <i>LastName</i> being also in the same group of <i>PId1</i> :","metadata":{}},{"cell_type":"code","source":"df_all.groupby(['PId1'])['LastName'].value_counts().value_counts()","metadata":{},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"A percentage which grows even bigger when <i>Deck</i> parameter is also taken into account :","metadata":{}},{"cell_type":"code","source":"df_all.groupby(\n    ['Deck','PId1']\n)['LastName'].describe()['unique'].value_counts()","metadata":{},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"And whose inverse is almost equivalent :","metadata":{}},{"cell_type":"code","source":"df_all.groupby(\n    ['LastName','PId1']\n)['Deck'].describe()['unique'].value_counts()","metadata":{},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# \n# <u><b>II) Pre-processing datasets</b></u>\n# ","metadata":{}},{"cell_type":"markdown","source":"## <u><b>II.A) Filling missing value</b></u> \n###### \n","metadata":{}},{"cell_type":"code","source":"df_complete = df_all.copy(deep=True)","metadata":{},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"###### \n### <u><b>II.A.1) Most important feature : <i>CryoSleep</i>.</b></u>\n","metadata":{}},{"cell_type":"code","source":"print(\n    'There was %.0f NaN among CrySleep feature'%(\n        df_complete['CryoSleep'].isna().sum()\n    )\n)","metadata":{},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Thanks to EDA, we saw that <i>CryoSleep</i> was highly absolute correlated with <i>Deck</i> and the luxury expenses features. \n\nEspecialy, <i>CryoSleep</i>=True implies that the depenses are all null, and the inverse is really close to be true all the time :","metadata":{}},{"cell_type":"code","source":"df_all[\n    (\n        df_all['RoomService'] == 0\n    )&(\n        df_all['FoodCourt'] == 0\n    )&(\n        df_all['ShoppingMall'] == 0\n    )&(\n        df_all['Spa'] == 0\n    )&(\n        df_all['VRDeck'] == 0\n    )\n]['CryoSleep'].describe()","metadata":{},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"On the other hand, we saw that <i>Deck</i>=F or D or E highly increase the risk of not being cryoslept.","metadata":{}},{"cell_type":"markdown","source":"Thefore we decided to use this strategy to fill <i>CryoSleep</i> missing value :\n- 1) if all luxury expenses are null, we put <i>CryoSleep</i>=True ; on the contrary, if any value among those five feature is > 0, then the put <i>CryoSleep</i>=False ;\n- 2) if not, we pass the most frequent value of the corresponding kind of <i>Deck</i>.","metadata":{}},{"cell_type":"code","source":"### Passing False if any expenses is != 0 \nindx = df_complete[\n    (\n        df_complete['TotalExpenses'] > 0\n    )&(\n        df_complete['CryoSleep'].isna()\n    )\n].index\ndf_complete.loc[indx, 'CryoSleep'] = False\nprint(\n    'It remains %.0f NaN among CrySleep feature'%(\n        df_complete['CryoSleep'].isna().sum()\n    )\n)\ndel indx","metadata":{},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"### Passing True if all expenses are = 0 \nindx = df_complete[\n    (\n        df_complete['RoomService'] == 0\n    )&(\n        df_complete['FoodCourt'] == 0\n    )&(\n        df_complete['ShoppingMall'] == 0\n    )&(\n        df_complete['Spa'] == 0\n    )&(\n        df_complete['VRDeck'] == 0\n    )&(\n        df_complete['CryoSleep'].isna()\n    )\n].index\ndf_complete.loc[indx, 'CryoSleep'] = True\nprint(\n    'It remains %.0f NaN among CrySleep feature'%(\n        df_complete['CryoSleep'].isna().sum()\n    )\n)\ndel indx","metadata":{},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"indx = df_complete[df_complete['CryoSleep'].isna()].index\ntops = df_complete.groupby('Deck')['CryoSleep'].describe()['top']\nfor i in list(indx):\n    df_complete.loc[i, 'CryoSleep'] = tops[df_complete.loc[i, 'Deck']]\nprint(\n    'After this, it remains %.0f NaN among CrySleep feature'%(\n        df_complete['CryoSleep'].isna().sum()\n    )\n)\ndel indx, tops, i ","metadata":{},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"###### \n### <u><b>II.A.2) <i>Deck</i>, <i>NumCabin</i> and <i>Side</i></b></u>\n","metadata":{}},{"cell_type":"code","source":"print(\n    'There was %.0f/%.0f/%.0f NaN among Deck/Cabin/Side feature'%(\n        df_complete['Deck'].isna().sum(), \n        df_complete['NumCabin'].isna().sum(), \n        df_complete['Side'].isna().sum()\n    )\n)","metadata":{},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"#### <u><b>First choice method :</b></u>\nWe use the fact that passengers with same <i>PId1</i> most of the time are family members sharing the same cabin","metadata":{}},{"cell_type":"code","source":"# Indexes where Deck/Num/Side are NaN\nindx = df_all[pd.isna(df_all['Deck'])].index","metadata":{},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"for i in indx :\n    id1 = df_all.loc[i,'PId1']\n    # If there is at least one other passenger with this PId1\n    if df_all[df_all['PId1'] == id1].shape[0] > 1:\n        # we store corresponding values of deck/cabin/side\n        d_c_s = df_all[\n            df_all['PId1'] == id1\n        ][\n            ['Deck', 'NumCabin', 'Side']\n        ].dropna()\n        # we pass those value for curent passenger\n        if d_c_s.shape[0] > 0:\n            d_c_s = d_c_s.drop_duplicates().values[0]\n            df_complete.loc[i, ['Deck','NumCabin','Side']] = d_c_s\ndel i, id1, d_c_s","metadata":{},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"print(\n    'It remains %.0f/%.0f/%.0f NaN among Deck/Cabin/Side feature'%(\n        df_complete['Deck'].isna().sum(), \n        df_complete['NumCabin'].isna().sum(), \n        df_complete['Side'].isna().sum()\n    )\n)","metadata":{},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"#### <u><b>Second choice method :</b></u>\nThis is more complex for second choice method.\n\nFirst of all, lets etablish a dictionnary. Each key is a Deck. Values corresponding to each key are the vacant cabin per Deck. To infer such values, we use the fact that the maximum NumCabin depends on which deck the cabin are. This way, we can deduce by the list of existing cabins per deck what should be the other possible cabins, that is to say, the cabins that remain vacant.","metadata":{}},{"cell_type":"code","source":"empty_dcs_byd = {}\nfor d in df_all['Deck'].dropna().unique():\n    # initialisation of list hich will contain vacant (deck,cabin,side) \n    list_dcs = []\n    # maximum number of cabins by deck\n    max_num = df_all[df_all['Deck'] == d]['NumCabin'].max() + 1\n    for s in ['S','P']:\n        # deck/cabin/side already taken\n        taken_cabins = np.array(\n            df_complete[\n                df_complete['Deck'] == d\n            ][\n                ['Side','NumCabin']\n            ].value_counts().sort_index()[s].index\n        )\n        if len(taken_cabins) < max_num:\n            free_cabs = np.setxor1d(\n                taken_cabins, np.arange(max_num)\n            )\n            list_dcs.append([(d,cab,s) for cab in free_cabs])\n    if len(list_dcs) > 0:\n        empty_dcs_byd[d] = list(np.concatenate(list_dcs))\n    else :\n        empty_dcs_byd[d] = []\ndel d, list_dcs, s, taken_cabins, free_cabs, max_num","metadata":{},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Then, we regroup all passengers by (<i>Deck, HomepLanet</i>), and calculate the median value of <i>TotalExpenses<i> for each ones of those groups","metadata":{}},{"cell_type":"code","source":"# indexes of NaN for Deck AND not NaN for HomePlanet\nindx = df_complete[\n    (df_complete['Deck'].isna())&(df_complete['HomePlanet'].notna())\n].index\n# dictionnary : all decks possible for a given HomePlanet\ndecks_by_planet = {\n    planet : list(\n        df_complete.groupby('HomePlanet')['Deck'].value_counts()[planet].index\n    ) for planet in df_all['HomePlanet'].unique() if pd.notna(planet)\n}\n# TotalExpenses means per kind of deck\ntotexps_by_deck = df_complete.groupby('Deck')['TotalExpenses'].median()","metadata":{},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"At last, given the taken value of <i>HomePlanet</i>, we knwo that only a few number of Deck values remain possible. \nWe chose the <i>Deck</i> for which the <i>TotalExpenses</i>'median from previous grouping is the clostest to the curent value of <i>TotalExpenses</i> taken by current passenger.","metadata":{}},{"cell_type":"code","source":"for i in list(indx):\n    # We take HomePlanet's and TotalExpenses' values\n    planet, totexp = (\n        df_complete.loc[i, 'HomePlanet'], \n        df_complete.loc[i, 'TotalExpenses']\n    )\n    # Potential decks given the planet\n    potential_decks = decks_by_planet[planet]\n    # Deck whose TotalExpenses'median is the closest to TotalExpenses' curent value\n    sorted_decks = list(totexps_by_deck[potential_decks].index[\n        abs(totexps_by_deck[potential_decks].values - totexp).argsort()\n    ])\n    j = 0\n    while len(empty_dcs_byd[sorted_decks[j]]) == 0:\n        j += 1\n        if j == len(sorted_decks):\n            break\n    if j < len(sorted_decks):\n        df_complete.loc[\n            i, ['Deck','NumCabin','Side']\n        ] = empty_dcs_byd[sorted_decks[j]].pop(-1)\ndel i, planet, totexp, potential_decks, sorted_decks, j ","metadata":{},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"print(\n    'It remains %.0f/%.0f/%.0f NaN among Deck/Cabin/Side feature'%(\n        df_complete['Deck'].isna().sum(), \n        df_complete['NumCabin'].isna().sum(), \n        df_complete['Side'].isna().sum()\n    )\n)","metadata":{},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"del decks_by_planet","metadata":{},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"#### <u><b>Third choice method :</b></u>\nIn case <i>HomePlanet</i> is a NaN too, we simply group passenger by <i>Deck</i>. The rest of processus is unchanged.","metadata":{}},{"cell_type":"code","source":"# indexes of NaN for Deck \nindx = df_complete[df_complete['Deck'].isna()].index\nfor i in list(indx):\n    # We take TotalExpenses' value\n    totexp = df_complete.loc[i, 'TotalExpenses']\n    # Deck whose TotalExpenses'median is the closest to TotalExpenses' curent value\n    sorted_decks = list(totexps_by_deck.index[\n        abs(totexps_by_deck.values-totexp).argsort()\n    ])\n    j = 0\n    while len(empty_dcs_byd[sorted_decks[j]]) == 0:\n        j += 1\n        if j == len(sorted_decks):\n            break\n    if j < len(sorted_decks):\n        df_complete.loc[\n            i, ['Deck','NumCabin','Side']\n        ] = empty_dcs_byd[sorted_decks[j]].pop(-1)\ndel i, totexp, sorted_decks, j","metadata":{},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"print(\n    'It remains %.0f/%.0f/%.0f NaN among Deck/Cabin/Side feature'%(\n        df_complete['Deck'].isna().sum(), \n        df_complete['NumCabin'].isna().sum(), \n        df_complete['Side'].isna().sum()\n    )\n)","metadata":{},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"del totexps_by_deck","metadata":{},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"###### \n### <u><b>II.A.3) Luxury expenses features.</b></u>\n","metadata":{}},{"cell_type":"code","source":"print(\n    'There was %.0f NaN among luxury expenses features'%(\n        df_complete[\n            ['RoomService', 'FoodCourt', 'ShoppingMall', 'Spa', 'VRDeck']\n        ].isna().sum().sum()\n    )\n)","metadata":{},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"For those featues, we already know that <i>CryoSleep</i>=True can be an easy way to fill NaN with 0, since we have shown that - logically - being cryoslept means spending no money. If we take a look at <i>CryoSleep</i> distribution for rows with at least one NaN among luxury expenses :","metadata":{}},{"cell_type":"code","source":"df_complete[\n    (\n        df_complete['RoomService'].isna()\n    )|(\n        df_complete['FoodCourt'].isna()\n    )|(\n        df_complete['ShoppingMall'].isna()\n    )|(\n        df_complete['Spa'].isna()\n    )|(\n        df_complete['VRDeck'].isna()\n    )\n]['CryoSleep'].value_counts()","metadata":{},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Well, there are 521 rows to fill easily with zeros. \n\nFor the rest of rows corresponding to people with <i>CryoSleep</i>=False, we choose to use again the quite good correlation between those expenses and features such as <i>Deck</i> and/or <i>HomePlanet</i>, by passing the closest means per group (<i>Deck</i> , <i>HomePlanet</i>).","metadata":{}},{"cell_type":"code","source":"indx = df_complete[\n    (\n        df_complete['RoomService'].isna()\n    )|(\n        df_complete['FoodCourt'].isna()\n    )|(\n        df_complete['ShoppingMall'].isna()\n    )|(\n        df_complete['Spa'].isna()\n    )|(\n        df_complete['VRDeck'].isna()\n    )\n].index","metadata":{},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"#### <u><b>First choice method :</b></u>","metadata":{}},{"cell_type":"code","source":"indx1 = df_complete.loc[indx][\n    df_complete.loc[indx,'CryoSleep'] == True\n].index\nfor col in [\n    'RoomService', 'FoodCourt', 'ShoppingMall', 'Spa', 'VRDeck'\n]:\n    df_complete.loc[indx1, col] = 0\ndel  col","metadata":{},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"print(\n    'It remains now %.0f NaN among luxury expenses features'%(\n        df_complete[\n            ['RoomService', \n             'FoodCourt', \n             'ShoppingMall', \n             'Spa', \n             'VRDeck']\n        ].isna().sum().sum()\n    )\n)","metadata":{},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"#### <u><b>Second choice method :</b></u>","metadata":{}},{"cell_type":"code","source":"indx2 = df_complete.loc[indx][\n    df_complete.loc[indx,'CryoSleep'] == False\n].index","metadata":{},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# median expenses, by kind of deck\nexpenses_by_deck = df_complete.groupby(['Deck'])[\n    ['RoomService', 'FoodCourt', 'ShoppingMall', 'Spa', 'VRDeck']\n].median()\n# median expenses, by group of (planet, deck)\nexpenses_by_planetdeck = df_complete.groupby(['HomePlanet','Deck'])[\n    ['RoomService', 'FoodCourt', 'ShoppingMall', 'Spa', 'VRDeck']\n].median()#.loc[('Earth','F')]","metadata":{},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"for i in list(indx2):\n    if pd.notna(df_complete.loc[i, 'HomePlanet']):\n        df_complete.loc[\n            i, ['RoomService', 'FoodCourt', 'ShoppingMall', 'Spa', 'VRDeck']\n        ] = expenses_by_planetdeck.loc[\n            (df_complete.loc[i, 'HomePlanet'], df_complete.loc[i, 'Deck'])\n        ]\n    else :\n        df_complete.loc[\n            i, ['RoomService', 'FoodCourt', 'ShoppingMall', 'Spa', 'VRDeck']\n        ] = expenses_by_deck.loc[df_complete.loc[i, 'Deck']]\ndel i, indx2, expenses_by_deck, expenses_by_planetdeck","metadata":{},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"print(\n    'It remains now %.0f NaN among luxury expenses features'%(\n        df_complete[\n            ['RoomService', \n             'FoodCourt', \n             'ShoppingMall', \n             'Spa', \n             'VRDeck']\n        ].isna().sum().sum()\n    )\n)","metadata":{},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"We can hereby re-calculate the sum of all those expenses (<i>ie</i> the parameter <i>TotalExpenses</i>), now that we have juts updated and filled all of the 5 columns used to calculate this parameter.","metadata":{}},{"cell_type":"code","source":"df_complete['TotalExpenses'] = df_complete[\n    ['RoomService', 'FoodCourt', 'ShoppingMall', 'Spa', 'VRDeck']\n].sum(axis=1)","metadata":{},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"###### \n### <u><b>II.A.4) <i>HomePlanet</i> and <i>Destination</i></b></u>\n","metadata":{}},{"cell_type":"markdown","source":"#### <u><b><i>HomePlanet</i></b></u> ","metadata":{}},{"cell_type":"code","source":"print(\n    'There was %.0f NaN among HomePlanet'%(\n        df_complete['HomePlanet'].isna().sum().sum()\n    )\n)","metadata":{},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Even if <i>HomePlanet</i> has a low average phiK correlation with the target <i>Transported</i>, it's an important paramter to fill missing value, due to its good correlations with a lot of other features.\n\nBeside, <i>TotalExepenses</i> and <i>Deck</i>, whose correlations with <i>HomePlanet</i> are huge, are now completely filled (and update in the case of <i>TotalExepenses</i>). Thus, we have now the possibility to fill <i>HomePlanet</i>'s missing values in one single shot.\n\nWe will use a method similar to what we did previously : we will choose the planet corresponding to the group of (<i>Deck</i>, <i>HomePlanet</i>) whose <i>TotalExepenses</i>'s median is the closest of <i>TotalExepenses</i> curent taken value (the potential groups will be determined using <i>Deck</i> curent taken value)","metadata":{}},{"cell_type":"code","source":"# medians of each groups of (deck, planet)\ntotexp_by_deckplanet = df_complete.groupby(\n    ['Deck', 'HomePlanet']\n)['TotalExpenses'].median()\n# indexes\nindx = df_complete[df_complete['HomePlanet'].isna()].index\n### Actual filling\nfor i in list(indx):\n    # We take TotalExpenses' and Deck's curent values\n    deck, totexp = (\n        df_complete.loc[i, 'Deck'], \n        df_complete.loc[i, 'TotalExpenses']\n    )\n    # We create local variable : TotalExpenses' medians by planet, for a given deck\n    totexp_planets = totexp_by_deckplanet.loc[deck]\n    # Attribution of value instead of NaN\n    df_complete.loc[i, 'HomePlanet'] = list(totexp_planets.index)[\n        abs(totexp_planets.values - totexp).argmin()\n    ]\nprint(\n    'It remains %.0f NaN among HomePLanet feature'%(\n        df_complete['HomePlanet'].isna().sum()\n    )\n)\ndel i, totexp, deck, totexp_planets","metadata":{},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"#### <u><b><i>Destination</i></b></u> ","metadata":{}},{"cell_type":"code","source":"print(\n    'There was %.0f NaN among Destination'%(\n        df_complete['Destination'].isna().sum().sum()\n    )\n)","metadata":{},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"For the same reasons than before, we will use the same method again : we will choose the destination corresponding to the group of (<i>Deck</i>, <i>Destination</i>) whose <i>TotalExepenses</i>'s median is the closest of <i>TotalExepenses</i> curent taken value (the potential groups will be determined using <i>Deck</i> curent taken value)","metadata":{}},{"cell_type":"code","source":"# medians of each groups of (deck, planet)\ntotexp_by_deckplanet = df_complete.groupby(\n    ['Deck', 'Destination']\n)['TotalExpenses'].median()\n# indexes\nindx = df_complete[df_complete['Destination'].isna()].index\n### Actual filling\nfor i in list(indx):\n    # We take TotalExpenses' and Deck's curent values\n    deck, totexp = (\n        df_complete.loc[i, 'Deck'], \n        df_complete.loc[i, 'TotalExpenses']\n    )\n    # We create local variable : TotalExpenses' medians by planet, for a given deck\n    totexp_planets = totexp_by_deckplanet.loc[deck]\n    # Attribution of value instead of NaN\n    df_complete.loc[i, 'Destination'] = list(totexp_planets.index)[\n        abs(totexp_planets.values - totexp).argmin()\n    ]\nprint(\n    'It remains %.0f NaN among Destination feature'%(\n        df_complete['Destination'].isna().sum()\n    )\n)\ndel i, totexp, deck, totexp_planets","metadata":{},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"###### \n### <u><b>II.A.5) <i>Age</i></b></u>\n","metadata":{}},{"cell_type":"code","source":"kimp = KNNImputer(n_neighbors=10)","metadata":{},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"df_complete['Age'] = kimp.fit_transform(\n    df_complete[\n        ['Age', 'RoomService', 'FoodCourt', 'ShoppingMall', 'Spa', 'VRDeck']\n    ]\n)[:,0]","metadata":{},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"print(\n    'It remains %.0f NaN among Age feature'%(\n        df_complete['Age'].isna().sum()\n    )\n)","metadata":{},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"###### \n### <u><b>II.A.6) <i>VIP</i></b></u>\n","metadata":{}},{"cell_type":"code","source":"indx = df_complete[df_complete['VIP'].isna()].index\ndf_complete.loc[indx, 'VIP'] = False","metadata":{},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"print(\n    'It remains %.0f NaN among VIP feature'%(\n        df_complete['VIP'].isna().sum()\n    )\n)","metadata":{},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"df_complete","metadata":{},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"######\n## <u><b>II.B) Creating train and test sets.</b></u> \n###### \n","metadata":{}},{"cell_type":"markdown","source":"### <u><b>II.B.1) Choosing our features</b></u>\n","metadata":{}},{"cell_type":"markdown","source":"#### <u><b>Re-calculating phi-k correlations.</b></u>","metadata":{}},{"cell_type":"code","source":"phik_corr_matrix = df_complete[df_complete['TrainOrTest'] == 'Train'].drop(\n    axis=1, columns=['TrainOrTest', 'PId1', 'FirstName','LastName']\n).phik_matrix().round(2)","metadata":{},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"plt.figure(\n    'PhiK correlation coefficients - after filling missing values', \n    figsize=(20,7)\n)\nplt.clf()\naxe1, axe2 = plt.subplot(121), plt.subplot(122)\naxe1.set_title(\n    r'$\\phi$K correlation matrix - train set', \n    fontsize=15, fontweight='bold'\n)\nsns.heatmap(\n    phik_corr_matrix, cbar=True, square=True, annot=True, ax=axe1\n)\naxe2.set_title(\n    r'$\\phi$K coefficents with Transported - train set', \n    fontsize=15, fontweight='bold'\n)\naxe2.set_ylabel(\n    r'$\\phi$K coefficents', fontsize=12, fontweight='bold'\n)\naxe2.set_xlabel(\n    'Other features', fontsize=12, fontweight='bold'\n)\nphik_corr_matrix['Transported'][:-1].sort_values().plot.bar(\n    ax=axe2, rot=45\n)\ndel axe1, axe2","metadata":{},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Filling missing values didn't change that much phi-k correlation coefficients between parameters. Thus, we can now use those values, and especially correlations coefficients with our taregt <i>Transported</i> to select most pertinent features.","metadata":{}},{"cell_type":"markdown","source":"#### <u><b>Actual features selection.</b></u>","metadata":{}},{"cell_type":"markdown","source":"Rather than selecting relevant features, it might be easier to exclude obviously irrelevant features.\n\nFor exemple, we don't take too many risks saying that <i>VIP</i>, taking almost every time the value False, and having a low phi-K correlation coefficient with our target. Therefore, we can legitimately decide to check <i>VIP</i> out of the relevant features list.\n\nDue to the second reason that we have just risen, we can aslo check <i>ShoppingMall</i> out, its phi-K corr. coef. with our target being even smaller.\n\n<i>FoodCourt</i> phi-K corr. coef. with <i>Transported</i> is barely better. Furthermore, <i>FoodCourt</i> is (just as <i>ShoppingMall</i> is) strongly correlated with <i>TotalExpenses</i>, whose phi-K corr. coef. with <i>Transported</i> is higher than the sum of <i>ShoppingMall</i> + <i>FoodCourt</i> coefficents. By saying this, we mean that if <i>ShoppingMall</i> and <i>FoodCourt</i> should have any influence on <i>Transported</i>, there is a solid probability that it expresses itsleft through <i>TotalExpenses</i>. Thus, we can check out <i>FoodCourt</i> without further delay.\n\nThe last parameter we can check out following that kind of logic would be <i>Destination</i>, its phi-K corr. coef. with <i>Transported</i> being barely better - again. \n\nWe complete the list to cheked out parameters with <i>PId1</i>, <i>FirstName</i>, <i>LastName</i> (too much unique values, since they are identifiants rather than features)","metadata":{}},{"cell_type":"code","source":"features = df_complete.columns.copy(deep=True)\nfeatures = list(\n    features.drop(\n        [\n            'PId1', 'FirstName', 'Destination', 'VIP', 'FoodCourt', \n            'ShoppingMall', 'FirstName', 'LastName', 'Transported', \n            'TrainOrTest'\n        ]\n    )\n)","metadata":{},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"######\n### <u><b>II.B.2) Actual creating train and test sets.</b></u>\n","metadata":{}},{"cell_type":"code","source":"# list of indexes from train or test set\nindxs = [df_complete.index[\n    df_complete['TrainOrTest'].apply(lambda x: x == val)\n] for val in ['Train', 'Test']]","metadata":{},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# division of features matrix between train and test sets\nX_train, X_test = [\n    df_complete.loc[indx, features] for indx in indxs\n]","metadata":{},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# division of features matrix between train and test sets\ny_train, y_test = [\n    df_complete.loc[indx, 'Transported'] for indx in indxs\n]","metadata":{},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"X_train","metadata":{},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"######\n## <u><b>II.C) Encoding and/or rescaling features.</b></u> \n###### \n","metadata":{}},{"cell_type":"markdown","source":"### <u><b>II.C.1) Encoding labelled features</b></u>\n","metadata":{}},{"cell_type":"markdown","source":"#### <u><b>Label encoding</b></u>","metadata":{}},{"cell_type":"code","source":"deck_labenc = LabelEncoder()\ndeck_labenc.fit(X_train['Deck'])","metadata":{},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"X_train['Deck'], X_test['Deck'] = [\n    deck_labenc.transform(x) for x in [X_train['Deck'], X_test['Deck']]\n]","metadata":{},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"#### <u><b>One hot encoding</b></u>","metadata":{}},{"cell_type":"code","source":"hp_ohe = OneHotEncoder(sparse=False)\nhp_ohe.fit(X_train[['HomePlanet']])\ntrain_ohe, test_ohe = [\n    hp_ohe.transform(x[['HomePlanet']]) for x in [X_train, X_test]\n]\nfor i, name in enumerate(hp_ohe.get_feature_names()):\n    X_train.insert(1, name, train_ohe[:, i])\n    X_test.insert(1, name, test_ohe[:, i])\ndel train_ohe, test_ohe, i, name\nX_train.drop(columns='HomePlanet', inplace=True)\nX_test.drop(columns='HomePlanet', inplace=True)","metadata":{},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"print(\n    'CryoSleep is now column n°', \n    np.arange(len(X_train.columns))[X_train.columns == 'CryoSleep'][0]\n)","metadata":{},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"cs_ohe = OneHotEncoder(sparse=False)\ncs_ohe.fit(X_train[['CryoSleep']])\ntrain_ohe, test_ohe = [\n    cs_ohe.transform(x[['CryoSleep']]) for x in [X_train, X_test]\n]\nfor i, name in enumerate(cs_ohe.get_feature_names()):\n    X_train.insert(4, name, train_ohe[:, i])\n    X_test.insert(4, name, test_ohe[:, i])\ndel train_ohe, test_ohe, i, name\nX_train.drop(columns='CryoSleep', inplace=True)\nX_test.drop(columns='CryoSleep', inplace=True)","metadata":{},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"print(\n    'Side is now column n°', \n    np.arange(len(X_train.columns))[X_train.columns == 'Side'][0]\n)","metadata":{},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"sp_ohe = OneHotEncoder(sparse=False)\nsp_ohe.fit(X_train[['Side']])\ntrain_ohe, test_ohe = [\n    sp_ohe.transform(x[['Side']]) for x in [X_train, X_test]\n]\nfor i, name in enumerate(sp_ohe.get_feature_names()):\n    X_train.insert(4, name, train_ohe[:, i])\n    X_test.insert(4, name, test_ohe[:, i])\ndel train_ohe, test_ohe, i, name\nX_train.drop(columns='Side', inplace=True)\nX_test.drop(columns='Side', inplace=True)","metadata":{},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"###### \n### <u><b>II.C.2) Rescalling features</b></u>\n","metadata":{}},{"cell_type":"markdown","source":"#### <u><b>Visible impact of encoding on train set</b></u>","metadata":{}},{"cell_type":"code","source":"X_train","metadata":{},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"We have to explicitly convert <i>NumCabin</i>'s type into integer (less than hundred of its taken values were actually strings) :","metadata":{}},{"cell_type":"code","source":"print(\n    'We discovered that NumCain type was :', \n    X_train['NumCabin'].dtype, \n    ', which means some values aren`t considered as float'\n)\nX_train['NumCabin'] = pd.to_numeric(\n    X_train['NumCabin']\n).convert_dtypes(convert_integer=True)\nprint('Now, NumCain type has become :', X_train['NumCabin'].dtype)","metadata":{},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Let's make a statistical description of the train set :","metadata":{}},{"cell_type":"code","source":"X_train.describe(percentiles=[])","metadata":{},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"It shows that there are huge differences between features, regarding if their come initially from categorical or ordinal features (before we did the encondings). \nParticullary, standard deviations and maxima differ drastically, which should warn us that this could lead to huge biaises during training phase, if we decided to give those values to the algorithm as inputs.","metadata":{}},{"cell_type":"markdown","source":"#### <u><b>Keeping curent scaling</b></u>","metadata":{}},{"cell_type":"markdown","source":"Contrary to what we thought previously, we have shown in a previous version of this notebook that the curent scaling leads to better validations scores than 2 kinds of common scaling. \n\n(The following link leads to our Github's branch dedicated to testing different scalings efficiency in terms of training a simple classifier : \n\nhttps://github.com/LukDuth/OC_Student_ML_P8_Titanic_Spaceship/blob/encoding_and_testing_scalings/Duthoit_Luke_1_Code_072022.ipynb). \n\nIn this sake, we reproduce below the hyper-parameters optimisation and cross-validation of a logistic regression with the curent scaling.","metadata":{}},{"cell_type":"code","source":"### Initilisation of models\nlogreg = LogisticRegression(random_state=rgn)\ngrid = GridSearchCV(estimator=logreg, \n                    param_grid={\n                        'penalty':('l1','l2'), \n                        'C':np.logspace(-4, 1, 6), \n                        'solver':('liblinear', 'saga')\n                    }, \n                    cv=5)\n### cross validation and optimisation of hyper-parameters\ngrid.fit(\n    X=X_train, y=y_train.apply(lambda x : str(x))\n)\nprint(grid.best_params_, grid.best_score_)\ndel logreg, grid","metadata":{},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"This score is honorable enough, considering the simplicty of the classifier, to keep the curent scaling, and not spending much time trying different ones.\n\nOnce may notice that L1 penalty is adequate, since we handle features with important different ranges of values, therefore L2 penalty would have probabily enhanced the importance of biaises and outliers within the learning process.","metadata":{}},{"cell_type":"markdown","source":"# \n# <u><b>III) Actual machine learning</b></u>\n# ","metadata":{}},{"cell_type":"markdown","source":"#### <u><b>Presentation of our purpouse</b></u>","metadata":{}},{"cell_type":"markdown","source":"We are going to test only two kind of classifiers. Indeed, we saw on that kaggle kernel https://www.kaggle.com/code/odins0n/spaceship-titanic-eda-27-different-models that a lot of classifiers could led to approximatively the same precision score. So we decided not to spend much time on puting to test a lot of different classifiers. Instead, we rather spent much time on hyper-parameters optimization, and cross-validation.\n\nThe classifiers we have choosen are :\n- AdaBoost (it was one of the best classifiers from https://www.kaggle.com/code/odins0n/spaceship-titanic-eda-27-different-models, and moreover one that was directly accessible with sklearn)\n- MLPClassifier (it seemed that the author of https://www.kaggle.com/code/odins0n/spaceship-titanic-eda-27-different-models used a model called \"perceptron\", but we don't know if it's litteraly a unique perceptron, or layers of perceptrons).","metadata":{}},{"cell_type":"markdown","source":"#### <u><b>Best model</b></u>","metadata":{}},{"cell_type":"markdown","source":"We displayed our full work - hyper-parameters optimisations, cross-validation and final best performances comparisons between classifiers - with the branch \"testing_classifiers\" in the github repository we used to updated our code? You can read it following this link :\n\nhttps://github.com/LukDuth/OC_Student_ML_P8_Titanic_Spaceship/blob/testing_classifiers/Duthoit_Luke_1_Code_072022.ipynb\n\nNevertheless, we show a bit of it below, focussing on our best model so far. This is the second phase of its hyper-parameters optimisation, after most of them were fine tunned enough :","metadata":{}},{"cell_type":"code","source":"### Initialisation\nmlp_1 = MLPClassifier(\n    random_state=rgn, solver='adam', learning_rate='invscaling', \n    activation='logistic', learning_rate_init=0.001, \n    beta_2=np.logspace(-4, np.log10(0.99), 3)[1], max_iter=40, \n    early_stopping=True, n_iter_no_change=10)\ngrid_mlp_1 = GridSearchCV(\n    estimator=mlp_1, \n    param_grid={\n        'hidden_layer_sizes':[(88,), (100,), (112,)], \n        'beta_1':np.logspace(-4.5, -3.5, 5)\n    }, \n    cv=5, \n    n_jobs=-1)\n### cross validation and optimisation of hyper-parameters\ngrid_mlp_1.fit(X=X_train, y=y_train.apply(lambda x: str(x)))\nprint(grid_mlp_1.best_params_, grid_mlp_1.best_score_)","metadata":{},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"If we display its parameters and performances :","metadata":{}},{"cell_type":"code","source":"pd.DataFrame(\n    data={\n        'mean_test_score': grid_mlp_1.cv_results_[\n            'mean_test_score'\n        ][grid_mlp_1.best_index_], \n        'std_test_score': grid_mlp_1.cv_results_[\n            'std_test_score'\n        ][grid_mlp_1.best_index_],     \n        'mean_fit_time [s]': grid_mlp_1.cv_results_[\n            'mean_fit_time'\n        ][grid_mlp_1.best_index_],     \n        'std_fit_time [s]': grid_mlp_1.cv_results_[\n            'std_fit_time'\n        ][grid_mlp_1.best_index_],     \n    }, index=[grid_mlp_1.best_estimator_]\n)","metadata":{},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# \n# <u><b>IV) Submission</b></u>\n# ","metadata":{}},{"cell_type":"code","source":"pred = pd.Series(\n    data=grid_mlp_1.best_estimator_.predict(X=X_test), \n    index=samsub.index,\n    name='Transported'\n)\n### changing string to bool\npred = pred.apply(lambda x: True if x == 'True' else False)\n### Display\npred","metadata":{},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"pred.value_counts()","metadata":{},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"samsub.drop(columns=['Transported'], inplace=True)\n# Adding predicted values\nsamsub.insert(loc=1, column='Transported', value=pred.values)\n# checking it's correct\nsamsub['Transported'].value_counts()","metadata":{},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"samsub","metadata":{},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"samsub.to_csv('sample_submission.csv', index=False)","metadata":{},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"","metadata":{},"execution_count":null,"outputs":[]}]}