{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"pygments_lexer":"ipython3","nbconvert_exporter":"python","version":"3.6.4","file_extension":".py","codemirror_mode":{"name":"ipython","version":3},"name":"python","mimetype":"text/x-python"}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"code","source":"import os\n\nimport numpy as np \nimport pandas as pd \n\nimport matplotlib.pyplot as plt\nimport seaborn as sns\n\nseed = 42\nsns.set_theme(style=\"whitegrid\")\nDIR_NAME = \"/kaggle/input/titanic\"\n\ntrain_df = pd.read_csv(os.path.join(DIR_NAME, \"train.csv\"))\ntest_df = pd.read_csv(os.path.join(DIR_NAME, \"test.csv\"))","metadata":{"_uuid":"8f2839f25d086af736a60e9eeb907d3b93b6e0e5","_cell_guid":"b1076dfc-b9ad-4769-8c92-a6c4dae69d19","execution":{"iopub.status.busy":"2022-07-30T05:59:40.507624Z","iopub.execute_input":"2022-07-30T05:59:40.508271Z","iopub.status.idle":"2022-07-30T05:59:41.607028Z","shell.execute_reply.started":"2022-07-30T05:59:40.508182Z","shell.execute_reply":"2022-07-30T05:59:41.605954Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## Titanic - Machine Learning from Disaster\n\n\n**Task:** <br> \nThe Titanic was a luxury British steamship. In 1909, during the voyage, the Titanic hit an iceberg and sank. When the Titanic sank it killed 1502 out of 2224 passengers and crew. During this chellenge we will use passangers data to predict whether they survived.\n\n\nLet's start by loading Titanic dataset and checking few samples.","metadata":{}},{"cell_type":"code","source":"train_df.sample(5, random_state=seed)","metadata":{"execution":{"iopub.status.busy":"2022-07-30T05:59:41.610308Z","iopub.execute_input":"2022-07-30T05:59:41.610824Z","iopub.status.idle":"2022-07-30T05:59:41.640506Z","shell.execute_reply.started":"2022-07-30T05:59:41.610780Z","shell.execute_reply":"2022-07-30T05:59:41.638815Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Our dataset contains numerical and categorical columns. We will process both type separetly. Moreover there are columns that are irrelevant to our task e.g. PassangerId column. We will drop this columns before futher analysis.","metadata":{}},{"cell_type":"code","source":"print(f\"The number of train samples: {train_df.shape[0]}\")\nprint(f\"The number of test samples: {test_df.shape[0]}\")","metadata":{"execution":{"iopub.status.busy":"2022-07-30T05:59:41.642934Z","iopub.execute_input":"2022-07-30T05:59:41.643781Z","iopub.status.idle":"2022-07-30T05:59:41.649716Z","shell.execute_reply.started":"2022-07-30T05:59:41.643714Z","shell.execute_reply":"2022-07-30T05:59:41.648795Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"train_df.drop([\"PassengerId\", \"Cabin\", \"Ticket\"], axis=1, inplace=True)\ntest_df.drop([\"Cabin\", \"Ticket\"], axis=1, inplace=True)\ntest_ids = test_df.pop(\"PassengerId\")","metadata":{"execution":{"iopub.status.busy":"2022-07-30T05:59:41.652446Z","iopub.execute_input":"2022-07-30T05:59:41.654070Z","iopub.status.idle":"2022-07-30T05:59:41.668264Z","shell.execute_reply.started":"2022-07-30T05:59:41.654025Z","shell.execute_reply":"2022-07-30T05:59:41.667077Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## Data visualization","metadata":{}},{"cell_type":"markdown","source":"We will plot the remaining data and check their value distributions. We can also verify how they are related to Survived column. ","metadata":{}},{"cell_type":"code","source":"print(f\"Min fare price: {train_df.Fare.min()}, max fare price: {train_df.Fare.max()}\\n\")\n\nfig, axis = plt.subplots(1, 2, figsize=(14, 4))\nsns.boxplot(data=train_df, x=\"Fare\", ax=axis[0])\nsns.kdeplot(data=train_df, x=\"Fare\", fill=True, hue=\"Survived\", multiple=\"stack\", ax=axis[1])\nplt.show()","metadata":{"execution":{"iopub.status.busy":"2022-07-30T05:59:41.670187Z","iopub.execute_input":"2022-07-30T05:59:41.670957Z","iopub.status.idle":"2022-07-30T05:59:42.118181Z","shell.execute_reply.started":"2022-07-30T05:59:41.670913Z","shell.execute_reply":"2022-07-30T05:59:42.117461Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"We can see that most of Fare values are smaller than 100. But there are some outliers. We can expect that Fare values are correlated to Pclass values. ","metadata":{}},{"cell_type":"code","source":"print(f\"Min age: {train_df.Age.min()}, max age: {train_df.Age.max()}\\n\")\n\nfig, axis = plt.subplots(1, 2, figsize=(14, 4))\nsns.boxplot(data=train_df, x=\"Age\", ax=axis[0])\nsns.kdeplot(data=train_df, x=\"Age\", fill=True, hue=\"Survived\", multiple=\"stack\", ax=axis[1])\nplt.show()","metadata":{"execution":{"iopub.status.busy":"2022-07-30T05:59:42.119279Z","iopub.execute_input":"2022-07-30T05:59:42.119788Z","iopub.status.idle":"2022-07-30T05:59:42.554068Z","shell.execute_reply.started":"2022-07-30T05:59:42.119757Z","shell.execute_reply":"2022-07-30T05:59:42.552483Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"From Age column distribution we can see that most of passangers were between 20 and 40 years old. Again we have same outliers with maximum passanger age 80 years.","metadata":{}},{"cell_type":"code","source":"fig, axis = plt.subplots(1, 2, figsize=(14, 4))\nax = sns.countplot(data=train_df, x=\"Sex\", hue=\"Survived\", ax=axis[0])\nax = sns.countplot(data=train_df, x=\"Pclass\", hue=\"Survived\", ax=axis[1])\nplt.show()","metadata":{"execution":{"iopub.status.busy":"2022-07-30T05:59:42.555860Z","iopub.execute_input":"2022-07-30T05:59:42.556170Z","iopub.status.idle":"2022-07-30T05:59:42.890954Z","shell.execute_reply.started":"2022-07-30T05:59:42.556143Z","shell.execute_reply":"2022-07-30T05:59:42.889789Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"If we compare count of Pclass and Sex survived we can observe that women were more likely to survive than man. Also there was higher chance for passangers from first class to survive than for third class passangers.","metadata":{}},{"cell_type":"markdown","source":"## Data preprocessing","metadata":{}},{"cell_type":"markdown","source":"The Titanic dataset contains some null values. Mostly for Age and Cabin columns. We also have Name column from which we can extract passanger title.","metadata":{}},{"cell_type":"code","source":"train_df.isnull().sum()","metadata":{"execution":{"iopub.status.busy":"2022-07-30T05:59:42.892245Z","iopub.execute_input":"2022-07-30T05:59:42.893043Z","iopub.status.idle":"2022-07-30T05:59:42.903370Z","shell.execute_reply.started":"2022-07-30T05:59:42.893011Z","shell.execute_reply":"2022-07-30T05:59:42.902143Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"category_columns = [\"Pclass\", \"Sex\", \"Embarked\"]\nnumeric_columns = [\"Age\", \"Fare\", \"SibSp\", \"Parch\"]","metadata":{"execution":{"iopub.status.busy":"2022-07-30T05:59:42.904660Z","iopub.execute_input":"2022-07-30T05:59:42.905155Z","iopub.status.idle":"2022-07-30T05:59:42.914411Z","shell.execute_reply.started":"2022-07-30T05:59:42.905123Z","shell.execute_reply":"2022-07-30T05:59:42.913381Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"There are also less common titles as Jonkheer or Mlle. We will use only top n common titles and rest of them we will replace with 'Rest' keyword. For this purpose we create custom Estimator class. Estimator fit method counts each title in dataset and store the most common ones. The transform method then directly replace Name value with passanger title.","metadata":{}},{"cell_type":"code","source":"from collections import Counter\nfrom sklearn.base import BaseEstimator, TransformerMixin\n\nclass TitleExtractor(BaseEstimator, TransformerMixin):\n    def __init__(self, top_n=8):\n        self.top_n = top_n\n             \n    def fit(self, X, y=None):        \n        title_count = Counter([name.split(\", \")[1].split(\".\")[0] for name in X])\n        \n        common_titles = []\n        for title, count in title_count.most_common(self.top_n):\n            common_titles.append(title)\n           \n        self.title_count = title_count\n        self.common_titles = common_titles \n        return self\n        \n    def transform(self, X):       \n        titles = []\n        for name in X:\n            name = name.split(\", \")[1].split(\".\")[0]\n            if name not in self.common_titles:\n                name = \"Rare\"\n            titles.append(name)\n        return np.reshape(titles, (-1, 1))\n                           \n    def get_feature_names(self):\n         return self.title_count","metadata":{"execution":{"iopub.status.busy":"2022-07-30T05:59:42.919396Z","iopub.execute_input":"2022-07-30T05:59:42.919835Z","iopub.status.idle":"2022-07-30T05:59:43.035281Z","shell.execute_reply.started":"2022-07-30T05:59:42.919803Z","shell.execute_reply":"2022-07-30T05:59:43.033912Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"For convenient processing of numeric and categorical columns we will use Pipelines. The numerical columns contains null values so we impute them using KNNImputer first. This will impute missing values using mean of neighbours values. Then we scale values with RobustScaler which is robust to ouliers. For the categorical columns we first fill missing values using most frequent class and then replace classes with ordinal encoding. Name column will be handled separetely using our custom estimator.","metadata":{}},{"cell_type":"code","source":"from sklearn.pipeline import Pipeline\nfrom sklearn.impute import KNNImputer, SimpleImputer\nfrom sklearn import preprocessing\n\n\nage_origin = train_df[\"Age\"].to_list()\ninputet_age = list(KNNImputer().fit_transform(train_df[numeric_columns])[:, 0])\n\n\nnumeric_pipe = Pipeline([('fillna', KNNImputer(n_neighbors=5)),\n                         ('scaler', preprocessing.RobustScaler())])\n                         \ncategoric_pipe = Pipeline([('imputer', SimpleImputer(strategy='most_frequent')),\n                           ('ordinal_encoder', preprocessing.OrdinalEncoder())])\n                         \nname_pipe = Pipeline([('title_extractor', TitleExtractor()),\n                      ('ordinal_encoder', preprocessing.OrdinalEncoder())])","metadata":{"execution":{"iopub.status.busy":"2022-07-30T05:59:43.037215Z","iopub.execute_input":"2022-07-30T05:59:43.038247Z","iopub.status.idle":"2022-07-30T05:59:43.326397Z","shell.execute_reply.started":"2022-07-30T05:59:43.038201Z","shell.execute_reply":"2022-07-30T05:59:43.325069Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"from sklearn.compose import ColumnTransformer\n\nct = ColumnTransformer(\n    [(\"numeric\", numeric_pipe, numeric_columns),\n     (\"categorical\", categoric_pipe, category_columns),\n     (\"name\", name_pipe, \"Name\")])","metadata":{"execution":{"iopub.status.busy":"2022-07-30T05:59:43.330801Z","iopub.execute_input":"2022-07-30T05:59:43.332211Z","iopub.status.idle":"2022-07-30T05:59:43.347006Z","shell.execute_reply.started":"2022-07-30T05:59:43.332158Z","shell.execute_reply":"2022-07-30T05:59:43.345555Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"train_full = ct.fit_transform(train_df)\ntest_full = ct.transform(test_df)\n\ny_train_full = train_df[\"Survived\"]","metadata":{"execution":{"iopub.status.busy":"2022-07-30T05:59:43.348424Z","iopub.execute_input":"2022-07-30T05:59:43.349828Z","iopub.status.idle":"2022-07-30T05:59:43.498651Z","shell.execute_reply.started":"2022-07-30T05:59:43.349784Z","shell.execute_reply":"2022-07-30T05:59:43.497099Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"We use only eight the most common titles. Other titles are pretty rare in our training dataset.","metadata":{}},{"cell_type":"code","source":"titles_count = ct.transformers_[2][1][0].get_feature_names().items()\npd.DataFrame(titles_count, columns=['Title', 'Count']).sort_values(\"Count\", ascending=False)","metadata":{"execution":{"iopub.status.busy":"2022-07-30T05:59:43.500552Z","iopub.execute_input":"2022-07-30T05:59:43.501050Z","iopub.status.idle":"2022-07-30T05:59:43.539428Z","shell.execute_reply.started":"2022-07-30T05:59:43.501004Z","shell.execute_reply":"2022-07-30T05:59:43.538255Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"We can compare how KNNImputer changed our age distribution in training dataset.","metadata":{}},{"cell_type":"code","source":"age_df = pd.DataFrame({\"Age\": age_origin + inputet_age, \n                       \"target\": [\"Origin\" if i < len(age_origin) else \"Processed\" \n                                  for i in range(len(age_origin)+len(inputet_age))]})\nsns.displot(age_df, kind=\"kde\", x=\"Age\", hue=\"target\", height=4, aspect=2, fill=True)\nplt.show()","metadata":{"execution":{"iopub.status.busy":"2022-07-30T05:59:43.540702Z","iopub.execute_input":"2022-07-30T05:59:43.541137Z","iopub.status.idle":"2022-07-30T05:59:44.102583Z","shell.execute_reply.started":"2022-07-30T05:59:43.541094Z","shell.execute_reply":"2022-07-30T05:59:44.101266Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## Correlation matrix","metadata":{}},{"cell_type":"markdown","source":"Let's check how strong our variables are related to each others using correlation matrix.","metadata":{}},{"cell_type":"code","source":"train_full = pd.DataFrame(train_full, columns=numeric_columns + category_columns + [\"Name\"])\n\nplt.figure(figsize=(12, 8))\nsns.heatmap(pd.concat([train_full, train_df.Survived], axis=1).corr(), annot=True)\nplt.show()","metadata":{"execution":{"iopub.status.busy":"2022-07-30T05:59:44.104136Z","iopub.execute_input":"2022-07-30T05:59:44.104561Z","iopub.status.idle":"2022-07-30T05:59:44.781094Z","shell.execute_reply.started":"2022-07-30T05:59:44.104518Z","shell.execute_reply":"2022-07-30T05:59:44.779831Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"We can see strong negative correlation between columns Survived and Sex. There is also expected correlation between Pclass and Fare.","metadata":{}},{"cell_type":"markdown","source":"## Model selection","metadata":{}},{"cell_type":"markdown","source":"We will test several models on our dataset. For this purpose we will use cross validation and store mean of model scores.","metadata":{}},{"cell_type":"code","source":"pd.DataFrame(train_full, columns=numeric_columns + category_columns + [\"Name\"]).sample(5, random_state=seed)","metadata":{"execution":{"iopub.status.busy":"2022-07-30T05:59:44.782520Z","iopub.execute_input":"2022-07-30T05:59:44.783001Z","iopub.status.idle":"2022-07-30T05:59:44.812269Z","shell.execute_reply.started":"2022-07-30T05:59:44.782956Z","shell.execute_reply":"2022-07-30T05:59:44.811111Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"from sklearn.model_selection import train_test_split\n\nX_train, X_valid, y_train, y_valid = train_test_split(train_full.to_numpy(), train_df[\"Survived\"], \n                                                      test_size=0.2, random_state=seed)","metadata":{"execution":{"iopub.status.busy":"2022-07-30T05:59:44.813868Z","iopub.execute_input":"2022-07-30T05:59:44.814547Z","iopub.status.idle":"2022-07-30T05:59:44.824457Z","shell.execute_reply.started":"2022-07-30T05:59:44.814499Z","shell.execute_reply":"2022-07-30T05:59:44.823065Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"from functools import partial \nfrom sklearn.model_selection import cross_val_score\nfrom sklearn.exceptions import ConvergenceWarning\nfrom sklearn.utils._testing import ignore_warnings\n\n@ignore_warnings(category=ConvergenceWarning)\ndef evaluate_model(estimator,  X_train, y_train, cv=5):\n    score = cross_val_score(estimator,  X_train, y_train, cv=cv)        \n    return round(score.mean(), 4)\n\n\nevaluate = partial(evaluate_model, X_train=train_full, y_train=y_train_full)","metadata":{"execution":{"iopub.status.busy":"2022-07-30T05:59:44.826334Z","iopub.execute_input":"2022-07-30T05:59:44.827808Z","iopub.status.idle":"2022-07-30T05:59:44.984374Z","shell.execute_reply.started":"2022-07-30T05:59:44.827759Z","shell.execute_reply":"2022-07-30T05:59:44.983017Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"from sklearn.svm import SVC\nfrom sklearn.linear_model import LogisticRegression\nfrom sklearn.ensemble import RandomForestClassifier, AdaBoostClassifier\nfrom sklearn.metrics import accuracy_score\nfrom sklearn.neural_network import MLPClassifier\n\n\nscores = {}\nscores[\"LR\"] = evaluate(LogisticRegression(random_state=seed, max_iter=500))\nscores[\"SVC\"] = evaluate(SVC(random_state=seed))\nscores[\"ada_boost\"] = evaluate(AdaBoostClassifier(random_state=seed))\nscores[\"RF\"] = evaluate(RandomForestClassifier(max_depth=8, random_state=seed))\nscores[\"MLP\"] = evaluate(MLPClassifier(random_state=seed))\n\npd.DataFrame(scores.items(), columns=['Model', 'Score']).sort_values(\"Score\", ascending=False)","metadata":{"execution":{"iopub.status.busy":"2022-07-30T05:59:44.986677Z","iopub.execute_input":"2022-07-30T05:59:44.987129Z","iopub.status.idle":"2022-07-30T05:59:53.020578Z","shell.execute_reply.started":"2022-07-30T05:59:44.987087Z","shell.execute_reply":"2022-07-30T05:59:53.019414Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"ax = sns.barplot(x=\"Model\", y=\"Score\", data=pd.DataFrame({\"Model\": scores.keys(), \"Score\": scores.values()}))\nax.set_ylim(bottom=0.75, top=0.85)\nplt.show()","metadata":{"execution":{"iopub.status.busy":"2022-07-30T05:59:53.022597Z","iopub.execute_input":"2022-07-30T05:59:53.023546Z","iopub.status.idle":"2022-07-30T05:59:53.387235Z","shell.execute_reply.started":"2022-07-30T05:59:53.023498Z","shell.execute_reply":"2022-07-30T05:59:53.386153Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"We can see that RandomForestClassifier has reached the highest score. We will use this model for final predictions but first we will apply hyperparameters seach.","metadata":{}},{"cell_type":"markdown","source":"## Hyperparameters search","metadata":{}},{"cell_type":"markdown","source":"We will search for the best number of estimators, max depth of three and others hyperparameters for RandomForestClassifier model. For hyperparameters search we can use GridSearchCV.","metadata":{}},{"cell_type":"code","source":"from sklearn.model_selection import GridSearchCV\n\nparams = {\n    'n_estimators': [20, 50, 100, 120],\n    'max_depth': [8, 12, 18, 24, 30],\n    'max_features': [\"auto\", \"sqrt\"],\n    'min_samples_split': [2, 6, 10],\n    'min_samples_leaf': [1, 3, 4]\n}\n\nclf = GridSearchCV(\n    estimator=RandomForestClassifier(random_state=seed),\n    param_grid=params,\n    cv=5,\n    n_jobs=5,\n    verbose=1\n)\n\nclf.fit(X_train, y_train)","metadata":{"execution":{"iopub.status.busy":"2022-07-30T05:59:53.388827Z","iopub.execute_input":"2022-07-30T05:59:53.389256Z","iopub.status.idle":"2022-07-30T06:01:42.283195Z","shell.execute_reply.started":"2022-07-30T05:59:53.389212Z","shell.execute_reply":"2022-07-30T06:01:42.282021Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"clf_best_params = clf.best_params_\nprint(clf.best_params_)","metadata":{"execution":{"iopub.status.busy":"2022-07-30T06:01:42.285047Z","iopub.execute_input":"2022-07-30T06:01:42.285375Z","iopub.status.idle":"2022-07-30T06:01:42.291033Z","shell.execute_reply.started":"2022-07-30T06:01:42.285345Z","shell.execute_reply":"2022-07-30T06:01:42.290007Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## Evaluation","metadata":{}},{"cell_type":"markdown","source":"We split our full training dataset to training and validation partitions. So we can use training set to fit our model with best params and then evaluate model performance on validation set.","metadata":{}},{"cell_type":"code","source":"clf = RandomForestClassifier(**clf_best_params)\nclf.fit(X_train, y_train)\n\ny_pred = clf.predict(X_valid)","metadata":{"execution":{"iopub.status.busy":"2022-07-30T06:01:42.292549Z","iopub.execute_input":"2022-07-30T06:01:42.292904Z","iopub.status.idle":"2022-07-30T06:01:42.562545Z","shell.execute_reply.started":"2022-07-30T06:01:42.292875Z","shell.execute_reply":"2022-07-30T06:01:42.561161Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"from sklearn.metrics import accuracy_score, classification_report\n\nacc = accuracy_score(y_valid, y_pred)\nprint(classification_report(y_valid, y_pred))\nprint(f\"Validation accuracy: {acc:.4f}\")","metadata":{"execution":{"iopub.status.busy":"2022-07-30T06:01:42.564544Z","iopub.execute_input":"2022-07-30T06:01:42.565264Z","iopub.status.idle":"2022-07-30T06:01:42.577955Z","shell.execute_reply.started":"2022-07-30T06:01:42.565228Z","shell.execute_reply":"2022-07-30T06:01:42.576802Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"from sklearn.metrics import confusion_matrix\n \nconfuction_matrix = confusion_matrix(y_valid, y_pred)\nsns.heatmap(confuction_matrix, annot=True, annot_kws={\"size\": 16})\nplt.show()","metadata":{"execution":{"iopub.status.busy":"2022-07-30T06:01:42.579602Z","iopub.execute_input":"2022-07-30T06:01:42.579992Z","iopub.status.idle":"2022-07-30T06:01:42.819863Z","shell.execute_reply.started":"2022-07-30T06:01:42.579962Z","shell.execute_reply":"2022-07-30T06:01:42.818772Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Finally we can use our trained model on test set.","metadata":{}},{"cell_type":"code","source":"y_test = clf.predict(test_full)\ny_test = pd.DataFrame({\"PassengerId\": test_ids, \"Survived\": y_test})\ny_test.to_csv('submission.csv', index=False)","metadata":{"execution":{"iopub.status.busy":"2022-07-30T06:01:42.820990Z","iopub.execute_input":"2022-07-30T06:01:42.821270Z","iopub.status.idle":"2022-07-30T06:01:42.856036Z","shell.execute_reply.started":"2022-07-30T06:01:42.821244Z","shell.execute_reply":"2022-07-30T06:01:42.855010Z"},"trusted":true},"execution_count":null,"outputs":[]}]}