{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"pygments_lexer":"ipython3","nbconvert_exporter":"python","version":"3.6.4","file_extension":".py","codemirror_mode":{"name":"ipython","version":3},"name":"python","mimetype":"text/x-python"}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"File and Data Field Descriptions\ntrain.csv - Personal records for about two-thirds (~8700) of the passengers, to be used as training data.\n- PassengerId - A unique Id for each passenger. Each Id takes the form gggg_pp where gggg indicates a group the passenger is travelling with and pp is their number within the group. People in a group are often family members, but not always.\n- HomePlanet - The planet the passenger departed from, typically their planet of permanent residence.\n- CryoSleep - Indicates whether the passenger elected to be put into suspended animation for the duration of the voyage. Passengers in cryosleep are confined to their cabins.\n- Cabin - The cabin number where the passenger is staying. Takes the form deck/num/side, where side can be either P for Port or S for Starboard.\n- Destination - The planet the passenger will be debarking to.\n- Age - The age of the passenger.\n- VIP - Whether the passenger has paid for special VIP service during the voyage.\n- RoomService, FoodCourt, ShoppingMall, Spa, VRDeck - Amount the passenger has billed at each of the Spaceship Titanic's many luxury amenities.\n- Name - The first and last names of the passenger.\n- Transported - Whether the passenger was transported to another dimension. This is the target, the column you are trying to predict.\n\ntest.csv - Personal records for the remaining one-third (~4300) of the passengers, to be used as test data. Your task is to predict the value of Transported for the passengers in this set.\n\nsample_submission.csv - A submission file in the correct format.\n- PassengerId - Id for each passenger in the test set.\n- Transported - The target. For each passenger, predict either True or False.","metadata":{}},{"cell_type":"code","source":"!pip install dataprep","metadata":{"_kg_hide-output":true,"_kg_hide-input":false,"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"## imports\n\nimport pandas as pd\nimport numpy as np\n\nimport seaborn as sns\nimport matplotlib.pyplot as plt\nfrom dataprep.eda import *\n\nfrom sklearn.ensemble import RandomForestClassifier\nfrom sklearn.model_selection import RandomizedSearchCV,GridSearchCV\nfrom sklearn.metrics import accuracy_score\nfrom xgboost import XGBClassifier\n\nimport warnings\n\n%matplotlib inline\nsns.set()\nwarnings.filterwarnings('ignore')\n\n","metadata":{"execution":{"iopub.status.busy":"2022-07-19T20:09:44.699760Z","iopub.execute_input":"2022-07-19T20:09:44.700312Z","iopub.status.idle":"2022-07-19T20:09:46.879183Z","shell.execute_reply.started":"2022-07-19T20:09:44.700258Z","shell.execute_reply":"2022-07-19T20:09:46.877994Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# get data\n# local\n#train_df = pd.read_csv('./Data/train.csv')\n#test_df = pd.read_csv('./Data/test.csv')\n# kaggle notebook\ntrain_df = pd.read_csv('../input/spaceship-titanic/train.csv')\ntest_df = pd.read_csv('../input/spaceship-titanic/test.csv')\ncombine = [train_df, test_df]","metadata":{"execution":{"iopub.status.busy":"2022-07-19T20:09:46.880551Z","iopub.execute_input":"2022-07-19T20:09:46.880898Z","iopub.status.idle":"2022-07-19T20:09:46.927949Z","shell.execute_reply.started":"2022-07-19T20:09:46.880864Z","shell.execute_reply":"2022-07-19T20:09:46.926807Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"plot(train_df)","metadata":{"execution":{"iopub.status.busy":"2022-07-19T20:09:46.929963Z","iopub.execute_input":"2022-07-19T20:09:46.930520Z","iopub.status.idle":"2022-07-19T20:09:49.446536Z","shell.execute_reply.started":"2022-07-19T20:09:46.930486Z","shell.execute_reply":"2022-07-19T20:09:49.444574Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"missing values in all features","metadata":{}},{"cell_type":"code","source":"pd.crosstab(train_df['HomePlanet'], train_df['CryoSleep'], normalize='index')","metadata":{"execution":{"iopub.status.busy":"2022-07-19T20:09:49.447993Z","iopub.execute_input":"2022-07-19T20:09:49.448359Z","iopub.status.idle":"2022-07-19T20:09:49.479385Z","shell.execute_reply.started":"2022-07-19T20:09:49.448327Z","shell.execute_reply":"2022-07-19T20:09:49.478113Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"pd.crosstab(train_df['Destination'], train_df['CryoSleep'], normalize='index')","metadata":{"execution":{"iopub.status.busy":"2022-07-19T20:09:49.482413Z","iopub.execute_input":"2022-07-19T20:09:49.482988Z","iopub.status.idle":"2022-07-19T20:09:49.510045Z","shell.execute_reply.started":"2022-07-19T20:09:49.482946Z","shell.execute_reply":"2022-07-19T20:09:49.508791Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"plot_correlation(train_df)","metadata":{"execution":{"iopub.status.busy":"2022-07-19T20:09:49.511619Z","iopub.execute_input":"2022-07-19T20:09:49.512158Z","iopub.status.idle":"2022-07-19T20:09:50.049111Z","shell.execute_reply.started":"2022-07-19T20:09:49.512115Z","shell.execute_reply":"2022-07-19T20:09:50.046676Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"##### HasSpend","metadata":{}},{"cell_type":"markdown","source":"Create Spend feature by adding all credit features, might be good to have, and lets add an HasSpend feature as well","metadata":{}},{"cell_type":"code","source":"for dataset in combine:\n    dataset['Spend'] = (dataset['RoomService']\n                        + dataset['FoodCourt']\n                        + dataset['ShoppingMall']\n                        + dataset['Spa']\n                        + dataset['VRDeck'])\n    \n    dataset['HasSpend'] = 0\n    dataset.loc[(dataset['Spend'] > 0),'HasSpend'] = 1\n\ntrain_df.head()","metadata":{"execution":{"iopub.status.busy":"2022-07-19T20:09:50.050834Z","iopub.execute_input":"2022-07-19T20:09:50.051741Z","iopub.status.idle":"2022-07-19T20:09:50.080751Z","shell.execute_reply.started":"2022-07-19T20:09:50.051699Z","shell.execute_reply":"2022-07-19T20:09:50.079696Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"plot(train_df, 'Spend')","metadata":{"execution":{"iopub.status.busy":"2022-07-19T20:09:50.081955Z","iopub.execute_input":"2022-07-19T20:09:50.082733Z","iopub.status.idle":"2022-07-19T20:09:50.709099Z","shell.execute_reply.started":"2022-07-19T20:09:50.082693Z","shell.execute_reply":"2022-07-19T20:09:50.707328Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"plot(train_df, 'Spend', 'Transported')","metadata":{"execution":{"iopub.status.busy":"2022-07-19T20:09:50.712678Z","iopub.execute_input":"2022-07-19T20:09:50.713048Z","iopub.status.idle":"2022-07-19T20:09:51.207961Z","shell.execute_reply.started":"2022-07-19T20:09:50.713013Z","shell.execute_reply":"2022-07-19T20:09:51.205413Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"for dataset in combine:\n    dataset['Spend'] = np.log(dataset['Spend'])\n\ntrain_df.head()","metadata":{"execution":{"iopub.status.busy":"2022-07-19T20:09:51.209758Z","iopub.execute_input":"2022-07-19T20:09:51.210123Z","iopub.status.idle":"2022-07-19T20:09:51.234641Z","shell.execute_reply.started":"2022-07-19T20:09:51.210090Z","shell.execute_reply":"2022-07-19T20:09:51.233444Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"plt.figure(figsize=(10,5))\nsns.histplot(data=train_df, x='Spend', hue='Transported', bins=30)","metadata":{"execution":{"iopub.status.busy":"2022-07-19T20:09:51.236076Z","iopub.execute_input":"2022-07-19T20:09:51.236419Z","iopub.status.idle":"2022-07-19T20:09:51.861509Z","shell.execute_reply.started":"2022-07-19T20:09:51.236388Z","shell.execute_reply":"2022-07-19T20:09:51.860079Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"for dataset in combine:\n    dataset['Spend'] = dataset['Spend'].fillna(0)\n    dataset['Spend'] = dataset['Spend'].replace([np.inf, -np.inf], 0)","metadata":{"execution":{"iopub.status.busy":"2022-07-19T20:09:51.863532Z","iopub.execute_input":"2022-07-19T20:09:51.864043Z","iopub.status.idle":"2022-07-19T20:09:51.873949Z","shell.execute_reply.started":"2022-07-19T20:09:51.863992Z","shell.execute_reply":"2022-07-19T20:09:51.872539Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Lets check the all credit features as well","metadata":{}},{"cell_type":"code","source":"dataset['Spend'].isna().sum()","metadata":{"execution":{"iopub.status.busy":"2022-07-19T20:09:51.875969Z","iopub.execute_input":"2022-07-19T20:09:51.876523Z","iopub.status.idle":"2022-07-19T20:09:51.890822Z","shell.execute_reply.started":"2022-07-19T20:09:51.876473Z","shell.execute_reply":"2022-07-19T20:09:51.889511Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"CreditFeatures = ['RoomService','FoodCourt','ShoppingMall','Spa','VRDeck']\n\nfor var in CreditFeatures:\n    plt.figure(figsize=(10,5))\n    sns.histplot(data=train_df, x=var, hue='Transported', bins=30)","metadata":{"execution":{"iopub.status.busy":"2022-07-19T20:09:51.892126Z","iopub.execute_input":"2022-07-19T20:09:51.893002Z","iopub.status.idle":"2022-07-19T20:09:53.827474Z","shell.execute_reply.started":"2022-07-19T20:09:51.892955Z","shell.execute_reply":"2022-07-19T20:09:53.826622Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"for dataset in combine:\n    for col in dataset[CreditFeatures]:\n        dataset[col] = np.log(pd.to_numeric(dataset[col], errors='coerce'))\n        dataset[col].replace([-np.inf], 0, inplace=True)\n\ntrain_df.head()","metadata":{"execution":{"iopub.status.busy":"2022-07-19T20:09:53.829126Z","iopub.execute_input":"2022-07-19T20:09:53.829486Z","iopub.status.idle":"2022-07-19T20:09:53.867151Z","shell.execute_reply.started":"2022-07-19T20:09:53.829440Z","shell.execute_reply":"2022-07-19T20:09:53.865848Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"##### PaxGroup","metadata":{}},{"cell_type":"markdown","source":"Get passengers in group","metadata":{}},{"cell_type":"code","source":"# get first four chars for PassegerId to create PaxGroup\nfor dataset in combine:\n    dataset['PaxGroup'] = dataset['PassengerId'].str[:4]\n\ntrain_df.head()    ","metadata":{"execution":{"iopub.status.busy":"2022-07-19T20:09:53.868802Z","iopub.execute_input":"2022-07-19T20:09:53.869277Z","iopub.status.idle":"2022-07-19T20:09:53.901655Z","shell.execute_reply.started":"2022-07-19T20:09:53.869217Z","shell.execute_reply":"2022-07-19T20:09:53.900511Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# create PaxGroupSize\nfor dataset in combine:\n    dataset['PaxGroupSize'] = dataset['PaxGroup'].map(dataset['PaxGroup'].value_counts())\n\ntrain_df.head()","metadata":{"execution":{"iopub.status.busy":"2022-07-19T20:09:53.904421Z","iopub.execute_input":"2022-07-19T20:09:53.905268Z","iopub.status.idle":"2022-07-19T20:09:53.939483Z","shell.execute_reply.started":"2022-07-19T20:09:53.905223Z","shell.execute_reply":"2022-07-19T20:09:53.938550Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"train_df[['PaxGroupSize','Transported']].groupby(['PaxGroupSize'], as_index=False).mean().sort_values(by='Transported', ascending=False)","metadata":{"execution":{"iopub.status.busy":"2022-07-19T20:09:53.940900Z","iopub.execute_input":"2022-07-19T20:09:53.941214Z","iopub.status.idle":"2022-07-19T20:09:53.958489Z","shell.execute_reply.started":"2022-07-19T20:09:53.941186Z","shell.execute_reply":"2022-07-19T20:09:53.957412Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# bin PaxGroupSize\nfor dataset in combine:    \n    dataset.loc[(dataset['PaxGroupSize'] == 1), 'PaxGroupSize'] = 0\n    dataset.loc[(dataset['PaxGroupSize'] == 2) | (dataset['PaxGroupSize'] == 7), 'PaxGroupSize'] = 1\n    dataset.loc[(dataset['PaxGroupSize'] == 3) | (dataset['PaxGroupSize'] == 5) | (dataset['PaxGroupSize'] == 6), 'PaxGroupSize'] = 2\n    dataset.loc[(dataset['PaxGroupSize'] == 4), 'PaxGroupSize'] = 3\n    dataset.loc[(dataset['PaxGroupSize'] == 8), 'PaxGroupSize'] = 4\n\ntrain_df.head()","metadata":{"execution":{"iopub.status.busy":"2022-07-19T20:09:53.959891Z","iopub.execute_input":"2022-07-19T20:09:53.960196Z","iopub.status.idle":"2022-07-19T20:09:53.995273Z","shell.execute_reply.started":"2022-07-19T20:09:53.960169Z","shell.execute_reply":"2022-07-19T20:09:53.994227Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"plot(train_df,'PaxGroupSize')","metadata":{"execution":{"iopub.status.busy":"2022-07-19T20:09:53.996965Z","iopub.execute_input":"2022-07-19T20:09:53.997319Z","iopub.status.idle":"2022-07-19T20:09:54.707023Z","shell.execute_reply.started":"2022-07-19T20:09:53.997289Z","shell.execute_reply":"2022-07-19T20:09:54.704944Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"plot(train_df,'PaxGroupSize','Transported')","metadata":{"execution":{"iopub.status.busy":"2022-07-19T20:09:54.708474Z","iopub.execute_input":"2022-07-19T20:09:54.708855Z","iopub.status.idle":"2022-07-19T20:09:55.018709Z","shell.execute_reply.started":"2022-07-19T20:09:54.708813Z","shell.execute_reply":"2022-07-19T20:09:55.015803Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"##### Name","metadata":{}},{"cell_type":"markdown","source":"might be useful, split into given name and family name","metadata":{}},{"cell_type":"code","source":"for dataset in combine:\n    # given name\n    dataset['GivenName'] = dataset['Name'].str.split(' ', expand=True)[0]\n\n    # family name\n    dataset['FamilyName'] = dataset['Name'].str.split(' ', expand=True)[1]\n\ntrain_df.head()","metadata":{"execution":{"iopub.status.busy":"2022-07-19T20:09:55.020532Z","iopub.execute_input":"2022-07-19T20:09:55.020974Z","iopub.status.idle":"2022-07-19T20:09:55.103222Z","shell.execute_reply.started":"2022-07-19T20:09:55.020934Z","shell.execute_reply":"2022-07-19T20:09:55.102302Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"##### Deck","metadata":{}},{"cell_type":"code","source":"# get first chars from Cabin into a Deck feature\nfor dataset in combine:\n    dataset['Deck'] = dataset['Cabin'].str[:1]\n\ntrain_df.head()","metadata":{"execution":{"iopub.status.busy":"2022-07-19T20:09:55.104481Z","iopub.execute_input":"2022-07-19T20:09:55.105120Z","iopub.status.idle":"2022-07-19T20:09:55.142527Z","shell.execute_reply.started":"2022-07-19T20:09:55.105083Z","shell.execute_reply":"2022-07-19T20:09:55.140252Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"plot(train_df,'Deck')","metadata":{"execution":{"iopub.status.busy":"2022-07-19T20:09:55.144061Z","iopub.execute_input":"2022-07-19T20:09:55.144947Z","iopub.status.idle":"2022-07-19T20:09:55.907560Z","shell.execute_reply.started":"2022-07-19T20:09:55.144906Z","shell.execute_reply":"2022-07-19T20:09:55.905366Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"plot(train_df,'Deck','Transported')","metadata":{"execution":{"iopub.status.busy":"2022-07-19T20:09:55.910005Z","iopub.execute_input":"2022-07-19T20:09:55.910395Z","iopub.status.idle":"2022-07-19T20:09:56.208659Z","shell.execute_reply.started":"2022-07-19T20:09:55.910360Z","shell.execute_reply":"2022-07-19T20:09:56.206348Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"plot_missing(train_df,'Deck','Transported')","metadata":{"execution":{"iopub.status.busy":"2022-07-19T20:09:56.210743Z","iopub.execute_input":"2022-07-19T20:09:56.211923Z","iopub.status.idle":"2022-07-19T20:09:57.151825Z","shell.execute_reply.started":"2022-07-19T20:09:56.211876Z","shell.execute_reply":"2022-07-19T20:09:57.146635Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"train_df[['Deck','Transported']].groupby('Deck').mean().sort_values('Transported', ascending=False)","metadata":{"execution":{"iopub.status.busy":"2022-07-19T20:09:57.161518Z","iopub.execute_input":"2022-07-19T20:09:57.163031Z","iopub.status.idle":"2022-07-19T20:09:57.180486Z","shell.execute_reply.started":"2022-07-19T20:09:57.162982Z","shell.execute_reply":"2022-07-19T20:09:57.179371Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"train_df['Deck'].isna().sum()","metadata":{"execution":{"iopub.status.busy":"2022-07-19T20:09:57.181699Z","iopub.execute_input":"2022-07-19T20:09:57.182761Z","iopub.status.idle":"2022-07-19T20:09:57.191900Z","shell.execute_reply.started":"2022-07-19T20:09:57.182709Z","shell.execute_reply":"2022-07-19T20:09:57.190924Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# simple impute with mode\nfor dataset in combine:\n    dataset['Deck'] = dataset['Deck'].transform(lambda x: x.fillna(x.mode()[0]))","metadata":{"execution":{"iopub.status.busy":"2022-07-19T20:09:57.192938Z","iopub.execute_input":"2022-07-19T20:09:57.193547Z","iopub.status.idle":"2022-07-19T20:09:57.206915Z","shell.execute_reply.started":"2022-07-19T20:09:57.193500Z","shell.execute_reply":"2022-07-19T20:09:57.205618Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"train_df['Deck'].isna().sum()","metadata":{"execution":{"iopub.status.busy":"2022-07-19T20:09:57.209256Z","iopub.execute_input":"2022-07-19T20:09:57.210741Z","iopub.status.idle":"2022-07-19T20:09:57.225172Z","shell.execute_reply.started":"2022-07-19T20:09:57.210690Z","shell.execute_reply":"2022-07-19T20:09:57.224098Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"##### DeckSide","metadata":{}},{"cell_type":"code","source":"# get last chars from Cabin into a Deck feature\nfor dataset in combine:\n    dataset['DeckSide'] = dataset['Cabin'].str[-1:]\n\ntrain_df.head()","metadata":{"execution":{"iopub.status.busy":"2022-07-19T20:09:57.226844Z","iopub.execute_input":"2022-07-19T20:09:57.227619Z","iopub.status.idle":"2022-07-19T20:09:57.273708Z","shell.execute_reply.started":"2022-07-19T20:09:57.227578Z","shell.execute_reply":"2022-07-19T20:09:57.272442Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"plot(train_df,'DeckSide')","metadata":{"execution":{"iopub.status.busy":"2022-07-19T20:09:57.275221Z","iopub.execute_input":"2022-07-19T20:09:57.275644Z","iopub.status.idle":"2022-07-19T20:09:58.112153Z","shell.execute_reply.started":"2022-07-19T20:09:57.275604Z","shell.execute_reply":"2022-07-19T20:09:58.110583Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"plot(train_df,'DeckSide','Transported')","metadata":{"execution":{"iopub.status.busy":"2022-07-19T20:09:58.114233Z","iopub.execute_input":"2022-07-19T20:09:58.114657Z","iopub.status.idle":"2022-07-19T20:09:58.452084Z","shell.execute_reply.started":"2022-07-19T20:09:58.114620Z","shell.execute_reply":"2022-07-19T20:09:58.450835Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"plot_missing(train_df,'DeckSide','Transported')","metadata":{"execution":{"iopub.status.busy":"2022-07-19T20:09:58.453881Z","iopub.execute_input":"2022-07-19T20:09:58.454217Z","iopub.status.idle":"2022-07-19T20:09:59.118207Z","shell.execute_reply.started":"2022-07-19T20:09:58.454187Z","shell.execute_reply":"2022-07-19T20:09:59.117220Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"train_df[['DeckSide','Transported']].groupby('DeckSide').mean().sort_values('Transported', ascending=False)","metadata":{"execution":{"iopub.status.busy":"2022-07-19T20:09:59.120180Z","iopub.execute_input":"2022-07-19T20:09:59.120976Z","iopub.status.idle":"2022-07-19T20:09:59.139597Z","shell.execute_reply.started":"2022-07-19T20:09:59.120926Z","shell.execute_reply":"2022-07-19T20:09:59.138011Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"train_df['DeckSide'].isna().sum()","metadata":{"execution":{"iopub.status.busy":"2022-07-19T20:09:59.141316Z","iopub.execute_input":"2022-07-19T20:09:59.142017Z","iopub.status.idle":"2022-07-19T20:09:59.150370Z","shell.execute_reply.started":"2022-07-19T20:09:59.141975Z","shell.execute_reply":"2022-07-19T20:09:59.148921Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# simple impute with mode\nfor dataset in combine:\n    dataset['DeckSide'] = dataset['DeckSide'].transform(lambda x: x.fillna(x.mode()[0]))","metadata":{"execution":{"iopub.status.busy":"2022-07-19T20:09:59.151784Z","iopub.execute_input":"2022-07-19T20:09:59.152344Z","iopub.status.idle":"2022-07-19T20:09:59.170310Z","shell.execute_reply.started":"2022-07-19T20:09:59.152310Z","shell.execute_reply":"2022-07-19T20:09:59.168622Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"train_df['DeckSide'].isna().sum()","metadata":{"execution":{"iopub.status.busy":"2022-07-19T20:09:59.171641Z","iopub.execute_input":"2022-07-19T20:09:59.172496Z","iopub.status.idle":"2022-07-19T20:09:59.186272Z","shell.execute_reply.started":"2022-07-19T20:09:59.172427Z","shell.execute_reply":"2022-07-19T20:09:59.185334Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"##### CabinNumber","metadata":{}},{"cell_type":"code","source":"for dataset in combine:\n    dataset['CabinNumber'] = dataset['Cabin'].str.split('/', expand=True)[1].astype('float').astype('Int16')\n    \ntrain_df.head()","metadata":{"execution":{"iopub.status.busy":"2022-07-19T20:09:59.187534Z","iopub.execute_input":"2022-07-19T20:09:59.188617Z","iopub.status.idle":"2022-07-19T20:09:59.248259Z","shell.execute_reply.started":"2022-07-19T20:09:59.188577Z","shell.execute_reply":"2022-07-19T20:09:59.247011Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# simple impute with mode\nfor dataset in combine:\n    dataset['CabinNumber'] = dataset['CabinNumber'].transform(lambda x: x.fillna(x.mode()[0]))","metadata":{"execution":{"iopub.status.busy":"2022-07-19T20:09:59.250080Z","iopub.execute_input":"2022-07-19T20:09:59.250422Z","iopub.status.idle":"2022-07-19T20:09:59.262415Z","shell.execute_reply.started":"2022-07-19T20:09:59.250393Z","shell.execute_reply":"2022-07-19T20:09:59.261159Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"plot(train_df,'CabinNumber')","metadata":{"execution":{"iopub.status.busy":"2022-07-19T20:09:59.264290Z","iopub.execute_input":"2022-07-19T20:09:59.265261Z","iopub.status.idle":"2022-07-19T20:09:59.912625Z","shell.execute_reply.started":"2022-07-19T20:09:59.265209Z","shell.execute_reply":"2022-07-19T20:09:59.910003Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"plot(train_df,'CabinNumber','Transported')","metadata":{"execution":{"iopub.status.busy":"2022-07-19T20:09:59.913993Z","iopub.execute_input":"2022-07-19T20:09:59.914331Z","iopub.status.idle":"2022-07-19T20:10:00.399350Z","shell.execute_reply.started":"2022-07-19T20:09:59.914300Z","shell.execute_reply":"2022-07-19T20:10:00.396439Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"CabinNumber might matter, look like survival alters per ~300","metadata":{}},{"cell_type":"code","source":"# bin CabinNumber\nfor dataset in combine:\n     dataset.loc[(dataset['CabinNumber'] <= 300), 'CabinNumber'] = 0\n     dataset.loc[(dataset['CabinNumber'] > 300) & (dataset['CabinNumber'] <= 600), 'CabinNumber'] = 1\n     dataset.loc[(dataset['CabinNumber'] > 600) & (dataset['CabinNumber'] <= 900), 'CabinNumber'] = 2\n     dataset.loc[(dataset['CabinNumber'] > 900) & (dataset['CabinNumber'] <= 1200), 'CabinNumber'] = 3\n     dataset.loc[(dataset['CabinNumber'] > 1200) & (dataset['CabinNumber'] <= 1500), 'CabinNumber'] = 4\n     dataset.loc[(dataset['CabinNumber'] > 1500) & (dataset['CabinNumber'] <= 1800), 'CabinNumber'] = 5\n     dataset.loc[(dataset['CabinNumber'] > 1800) , 'CabinNumber'] = 6\n\ntrain_df.head()","metadata":{"execution":{"iopub.status.busy":"2022-07-19T20:10:00.401125Z","iopub.execute_input":"2022-07-19T20:10:00.401707Z","iopub.status.idle":"2022-07-19T20:10:00.462877Z","shell.execute_reply.started":"2022-07-19T20:10:00.401632Z","shell.execute_reply":"2022-07-19T20:10:00.461535Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"plot(train_df,'CabinNumber','Transported')","metadata":{"execution":{"iopub.status.busy":"2022-07-19T20:10:00.464492Z","iopub.execute_input":"2022-07-19T20:10:00.464977Z","iopub.status.idle":"2022-07-19T20:10:00.813566Z","shell.execute_reply.started":"2022-07-19T20:10:00.464942Z","shell.execute_reply":"2022-07-19T20:10:00.811051Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"train_df[['CabinNumber','Transported']].groupby('CabinNumber').mean().sort_values('Transported', ascending=False)","metadata":{"execution":{"iopub.status.busy":"2022-07-19T20:10:00.815183Z","iopub.execute_input":"2022-07-19T20:10:00.815719Z","iopub.status.idle":"2022-07-19T20:10:00.837103Z","shell.execute_reply.started":"2022-07-19T20:10:00.815675Z","shell.execute_reply":"2022-07-19T20:10:00.836130Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"##### Age","metadata":{}},{"cell_type":"code","source":"plot(train_df,'Age')","metadata":{"execution":{"iopub.status.busy":"2022-07-19T20:10:00.838416Z","iopub.execute_input":"2022-07-19T20:10:00.839565Z","iopub.status.idle":"2022-07-19T20:10:01.447311Z","shell.execute_reply.started":"2022-07-19T20:10:00.839524Z","shell.execute_reply":"2022-07-19T20:10:01.443640Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"plot(train_df,'Age','Transported')","metadata":{"execution":{"iopub.status.busy":"2022-07-19T20:10:01.448598Z","iopub.execute_input":"2022-07-19T20:10:01.449239Z","iopub.status.idle":"2022-07-19T20:10:01.906068Z","shell.execute_reply.started":"2022-07-19T20:10:01.449200Z","shell.execute_reply":"2022-07-19T20:10:01.902389Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"plot_missing(train_df,'Age','Transported')","metadata":{"execution":{"iopub.status.busy":"2022-07-19T20:10:01.908067Z","iopub.execute_input":"2022-07-19T20:10:01.908520Z","iopub.status.idle":"2022-07-19T20:10:02.531330Z","shell.execute_reply.started":"2022-07-19T20:10:01.908479Z","shell.execute_reply":"2022-07-19T20:10:02.529411Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# update age = 0 to NaN\nfor dataset in combine:\n    dataset['Age'].replace(0, np.nan, inplace=True)","metadata":{"execution":{"iopub.status.busy":"2022-07-19T20:10:02.534019Z","iopub.execute_input":"2022-07-19T20:10:02.534384Z","iopub.status.idle":"2022-07-19T20:10:02.542240Z","shell.execute_reply.started":"2022-07-19T20:10:02.534351Z","shell.execute_reply":"2022-07-19T20:10:02.540960Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"train_df['Age'].isna().sum()","metadata":{"execution":{"iopub.status.busy":"2022-07-19T20:10:02.544370Z","iopub.execute_input":"2022-07-19T20:10:02.545025Z","iopub.status.idle":"2022-07-19T20:10:02.560177Z","shell.execute_reply.started":"2022-07-19T20:10:02.544970Z","shell.execute_reply":"2022-07-19T20:10:02.558828Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# impute with median (skew < 0.5)\nfor dataset in combine:\n    dataset['Age'] = dataset['Age'].fillna(dataset.groupby(['HomePlanet'\n                                            ,'Destination'\n                                            ,'VIP'\n                                            ,'HasSpend'\n                                            ,'PaxGroupSize'\n                                            ,'Deck'\n                                            ,'DeckSide'])['Age'].transform('mean'))","metadata":{"execution":{"iopub.status.busy":"2022-07-19T20:10:02.561584Z","iopub.execute_input":"2022-07-19T20:10:02.562038Z","iopub.status.idle":"2022-07-19T20:10:02.592954Z","shell.execute_reply.started":"2022-07-19T20:10:02.561996Z","shell.execute_reply":"2022-07-19T20:10:02.592025Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"train_df['Age'].isna().sum()","metadata":{"execution":{"iopub.status.busy":"2022-07-19T20:10:02.595009Z","iopub.execute_input":"2022-07-19T20:10:02.596217Z","iopub.status.idle":"2022-07-19T20:10:02.605578Z","shell.execute_reply.started":"2022-07-19T20:10:02.596146Z","shell.execute_reply":"2022-07-19T20:10:02.604648Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"train_df.loc[(train_df['Age'].isna())]","metadata":{"execution":{"iopub.status.busy":"2022-07-19T20:10:02.607571Z","iopub.execute_input":"2022-07-19T20:10:02.607905Z","iopub.status.idle":"2022-07-19T20:10:02.655737Z","shell.execute_reply.started":"2022-07-19T20:10:02.607874Z","shell.execute_reply":"2022-07-19T20:10:02.654599Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Impute with median with less groupby features\nfor dataset in combine:\n    dataset['Age'] = dataset['Age'].fillna(dataset.groupby(['HasSpend'\n                                            ,'PaxGroupSize'\n                                            ,'CryoSleep'])['Age'].transform('mean'))","metadata":{"execution":{"iopub.status.busy":"2022-07-19T20:10:02.657689Z","iopub.execute_input":"2022-07-19T20:10:02.658263Z","iopub.status.idle":"2022-07-19T20:10:02.673941Z","shell.execute_reply.started":"2022-07-19T20:10:02.658226Z","shell.execute_reply":"2022-07-19T20:10:02.672357Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"train_df['Age'].isna().sum()","metadata":{"execution":{"iopub.status.busy":"2022-07-19T20:10:02.675715Z","iopub.execute_input":"2022-07-19T20:10:02.676210Z","iopub.status.idle":"2022-07-19T20:10:02.685926Z","shell.execute_reply.started":"2022-07-19T20:10:02.676162Z","shell.execute_reply":"2022-07-19T20:10:02.685009Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"train_df['AgeBand'] = pd.cut(train_df['Age'], 5)\ntrain_df[['AgeBand', 'Transported']].groupby(['AgeBand'], as_index=False).mean().sort_values(by='AgeBand', ascending=True)","metadata":{"execution":{"iopub.status.busy":"2022-07-19T20:10:02.687428Z","iopub.execute_input":"2022-07-19T20:10:02.688175Z","iopub.status.idle":"2022-07-19T20:10:02.715978Z","shell.execute_reply.started":"2022-07-19T20:10:02.688131Z","shell.execute_reply":"2022-07-19T20:10:02.715093Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# bin age\nfor dataset in combine:    \n    dataset.loc[(dataset['Age'] <= 17), 'Age'] = 0\n    dataset.loc[(dataset['Age'] > 17), 'Age'] = 1\n\n    dataset['Age'] = dataset['Age'].astype('int')\n\ntrain_df.head()","metadata":{"execution":{"iopub.status.busy":"2022-07-19T20:10:02.717551Z","iopub.execute_input":"2022-07-19T20:10:02.718059Z","iopub.status.idle":"2022-07-19T20:10:02.750668Z","shell.execute_reply.started":"2022-07-19T20:10:02.718028Z","shell.execute_reply":"2022-07-19T20:10:02.749477Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"plot(train_df,'Age','Transported')","metadata":{"execution":{"iopub.status.busy":"2022-07-19T20:10:02.752418Z","iopub.execute_input":"2022-07-19T20:10:02.753329Z","iopub.status.idle":"2022-07-19T20:10:03.064528Z","shell.execute_reply.started":"2022-07-19T20:10:02.753276Z","shell.execute_reply":"2022-07-19T20:10:03.061681Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"##### CryoSleep","metadata":{}},{"cell_type":"code","source":"plot(train_df,'CryoSleep')","metadata":{"execution":{"iopub.status.busy":"2022-07-19T20:10:03.065863Z","iopub.execute_input":"2022-07-19T20:10:03.066199Z","iopub.status.idle":"2022-07-19T20:10:03.779388Z","shell.execute_reply.started":"2022-07-19T20:10:03.066170Z","shell.execute_reply":"2022-07-19T20:10:03.775265Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"plot(train_df,'CryoSleep','Transported')","metadata":{"execution":{"iopub.status.busy":"2022-07-19T20:10:03.780970Z","iopub.execute_input":"2022-07-19T20:10:03.781431Z","iopub.status.idle":"2022-07-19T20:10:04.082664Z","shell.execute_reply.started":"2022-07-19T20:10:03.781400Z","shell.execute_reply":"2022-07-19T20:10:04.079466Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"plot_missing(train_df,'CryoSleep','Transported')","metadata":{"execution":{"iopub.status.busy":"2022-07-19T20:10:04.084168Z","iopub.execute_input":"2022-07-19T20:10:04.085053Z","iopub.status.idle":"2022-07-19T20:10:04.800191Z","shell.execute_reply.started":"2022-07-19T20:10:04.085012Z","shell.execute_reply":"2022-07-19T20:10:04.798847Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"As espected, passengers in CryoSleep have not spend any credits","metadata":{}},{"cell_type":"markdown","source":"Update rows with CryoSleep = True","metadata":{}},{"cell_type":"code","source":"for dataset in combine:\n    dataset.loc[(dataset['CryoSleep'].isna()) & (dataset['HasSpend'] == 0), 'CryoSleep'] = 1","metadata":{"execution":{"iopub.status.busy":"2022-07-19T20:10:04.801718Z","iopub.execute_input":"2022-07-19T20:10:04.802515Z","iopub.status.idle":"2022-07-19T20:10:04.812766Z","shell.execute_reply.started":"2022-07-19T20:10:04.802444Z","shell.execute_reply":"2022-07-19T20:10:04.811025Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"train_df['CryoSleep'].isna().sum()","metadata":{"execution":{"iopub.status.busy":"2022-07-19T20:10:04.814475Z","iopub.execute_input":"2022-07-19T20:10:04.815698Z","iopub.status.idle":"2022-07-19T20:10:04.830346Z","shell.execute_reply.started":"2022-07-19T20:10:04.815645Z","shell.execute_reply":"2022-07-19T20:10:04.829009Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"False if have spent any credits","metadata":{}},{"cell_type":"code","source":"for dataset in combine:\n    dataset.loc[(dataset['CryoSleep'].isna()) & (dataset['HasSpend'] != 0), 'CryoSleep'] = 0","metadata":{"execution":{"iopub.status.busy":"2022-07-19T20:10:04.831786Z","iopub.execute_input":"2022-07-19T20:10:04.832158Z","iopub.status.idle":"2022-07-19T20:10:04.847993Z","shell.execute_reply.started":"2022-07-19T20:10:04.832126Z","shell.execute_reply":"2022-07-19T20:10:04.846700Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"train_df['CryoSleep'].isna().sum()","metadata":{"execution":{"iopub.status.busy":"2022-07-19T20:10:04.849398Z","iopub.execute_input":"2022-07-19T20:10:04.849838Z","iopub.status.idle":"2022-07-19T20:10:04.870874Z","shell.execute_reply.started":"2022-07-19T20:10:04.849802Z","shell.execute_reply":"2022-07-19T20:10:04.869684Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"train_df['CryoSleep']","metadata":{"execution":{"iopub.status.busy":"2022-07-19T20:10:04.872415Z","iopub.execute_input":"2022-07-19T20:10:04.873278Z","iopub.status.idle":"2022-07-19T20:10:04.889719Z","shell.execute_reply.started":"2022-07-19T20:10:04.873234Z","shell.execute_reply":"2022-07-19T20:10:04.887828Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"for dataset in combine:\n    dataset['CryoSleep'] = dataset['CryoSleep'].replace({'True':True,'False':False}).astype('int')","metadata":{"execution":{"iopub.status.busy":"2022-07-19T20:10:04.891226Z","iopub.execute_input":"2022-07-19T20:10:04.892016Z","iopub.status.idle":"2022-07-19T20:10:04.912776Z","shell.execute_reply.started":"2022-07-19T20:10:04.891967Z","shell.execute_reply":"2022-07-19T20:10:04.911528Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"train_df['CryoSleep']","metadata":{"execution":{"iopub.status.busy":"2022-07-19T20:10:04.914327Z","iopub.execute_input":"2022-07-19T20:10:04.915762Z","iopub.status.idle":"2022-07-19T20:10:04.926993Z","shell.execute_reply.started":"2022-07-19T20:10:04.915709Z","shell.execute_reply":"2022-07-19T20:10:04.925771Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"##### HomePlanet","metadata":{}},{"cell_type":"code","source":"plot(train_df,'HomePlanet')","metadata":{"execution":{"iopub.status.busy":"2022-07-19T20:10:04.928563Z","iopub.execute_input":"2022-07-19T20:10:04.929404Z","iopub.status.idle":"2022-07-19T20:10:05.668801Z","shell.execute_reply.started":"2022-07-19T20:10:04.929350Z","shell.execute_reply":"2022-07-19T20:10:05.665478Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"plot(train_df,'HomePlanet','Transported')","metadata":{"execution":{"iopub.status.busy":"2022-07-19T20:10:05.670925Z","iopub.execute_input":"2022-07-19T20:10:05.671354Z","iopub.status.idle":"2022-07-19T20:10:06.198496Z","shell.execute_reply.started":"2022-07-19T20:10:05.671315Z","shell.execute_reply":"2022-07-19T20:10:06.194965Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"plot_missing(train_df,'HomePlanet','Transported')","metadata":{"execution":{"iopub.status.busy":"2022-07-19T20:10:06.200223Z","iopub.execute_input":"2022-07-19T20:10:06.200620Z","iopub.status.idle":"2022-07-19T20:10:06.867565Z","shell.execute_reply.started":"2022-07-19T20:10:06.200585Z","shell.execute_reply":"2022-07-19T20:10:06.866295Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"train_df[['HomePlanet','Transported']].groupby(['HomePlanet'], as_index=False).mean().sort_values(by='HomePlanet', ascending=False)","metadata":{"execution":{"iopub.status.busy":"2022-07-19T20:10:06.868973Z","iopub.execute_input":"2022-07-19T20:10:06.869309Z","iopub.status.idle":"2022-07-19T20:10:06.885808Z","shell.execute_reply.started":"2022-07-19T20:10:06.869279Z","shell.execute_reply":"2022-07-19T20:10:06.884749Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"pd.pivot_table(train_df\n                , values='PassengerId'\n                , index='Deck'\n                , columns='HomePlanet'\n                , aggfunc='count')","metadata":{"execution":{"iopub.status.busy":"2022-07-19T20:10:06.887203Z","iopub.execute_input":"2022-07-19T20:10:06.887937Z","iopub.status.idle":"2022-07-19T20:10:06.913501Z","shell.execute_reply.started":"2022-07-19T20:10:06.887898Z","shell.execute_reply":"2022-07-19T20:10:06.912671Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"for dataset in combine:\n    # fix missing from deck G\n    dataset.loc[(dataset['Deck'] == 'G') & (dataset['HomePlanet'].isna()), 'HomePlanet'] = 'Earth'\n\n    # fix missing from deck A,B,C,T\n    dataset.loc[((dataset['Deck'] == 'A') | (dataset['Deck'] == 'B') | (dataset['Deck'] == 'C') | (dataset['Deck'] == 'T')) & (dataset['HomePlanet'].isna()), 'HomePlanet'] = 'Europa'","metadata":{"execution":{"iopub.status.busy":"2022-07-19T20:10:06.914815Z","iopub.execute_input":"2022-07-19T20:10:06.915282Z","iopub.status.idle":"2022-07-19T20:10:06.936571Z","shell.execute_reply.started":"2022-07-19T20:10:06.915249Z","shell.execute_reply":"2022-07-19T20:10:06.935321Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"train_df['HomePlanet'].isna().sum()","metadata":{"execution":{"iopub.status.busy":"2022-07-19T20:10:06.938038Z","iopub.execute_input":"2022-07-19T20:10:06.938385Z","iopub.status.idle":"2022-07-19T20:10:06.946244Z","shell.execute_reply.started":"2022-07-19T20:10:06.938352Z","shell.execute_reply":"2022-07-19T20:10:06.945210Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"if in same group are all from same HomePlanet?\n","metadata":{}},{"cell_type":"code","source":"tmp = train_df.loc[(~pd.isna(train_df['HomePlanet'])),['HomePlanet','PaxGroup']]\nresult = tmp.groupby(['PaxGroup','HomePlanet'])['HomePlanet'].count().unstack().isna().sum(axis=1)\n\nresult.describe()","metadata":{"execution":{"iopub.status.busy":"2022-07-19T20:10:06.947645Z","iopub.execute_input":"2022-07-19T20:10:06.948901Z","iopub.status.idle":"2022-07-19T20:10:06.977874Z","shell.execute_reply.started":"2022-07-19T20:10:06.948862Z","shell.execute_reply":"2022-07-19T20:10:06.976826Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"two na columns per row, all i a group is from same HomePlanet","metadata":{}},{"cell_type":"code","source":"labels = {'Earth':1,'Europa':2,'Mars':3}\n\nfor dataset in combine:\n    dataset['HomePlanet'] = dataset['HomePlanet'].map(labels)\n    tmp = dataset[['HomePlanet','PaxGroup']]\n    tmp['HomePlanet'] = tmp.groupby('PaxGroup').transform(lambda x: x.fillna(x.mean()))\n    dataset['HomePlanet'] = tmp['HomePlanet']","metadata":{"execution":{"iopub.status.busy":"2022-07-19T20:10:06.979500Z","iopub.execute_input":"2022-07-19T20:10:06.979863Z","iopub.status.idle":"2022-07-19T20:10:21.808944Z","shell.execute_reply.started":"2022-07-19T20:10:06.979832Z","shell.execute_reply":"2022-07-19T20:10:21.807988Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"train_df['HomePlanet'].isna().sum()","metadata":{"execution":{"iopub.status.busy":"2022-07-19T20:10:21.810375Z","iopub.execute_input":"2022-07-19T20:10:21.810956Z","iopub.status.idle":"2022-07-19T20:10:21.817851Z","shell.execute_reply.started":"2022-07-19T20:10:21.810920Z","shell.execute_reply":"2022-07-19T20:10:21.816835Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# simple impute with mode\nfor dataset in combine:\n    dataset['HomePlanet'] = dataset['HomePlanet'].transform(lambda x: x.fillna(x.mode()[0]))","metadata":{"execution":{"iopub.status.busy":"2022-07-19T20:10:21.819746Z","iopub.execute_input":"2022-07-19T20:10:21.820563Z","iopub.status.idle":"2022-07-19T20:10:21.839496Z","shell.execute_reply.started":"2022-07-19T20:10:21.820523Z","shell.execute_reply":"2022-07-19T20:10:21.838132Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"train_df['HomePlanet'].isna().sum()","metadata":{"execution":{"iopub.status.busy":"2022-07-19T20:10:21.842647Z","iopub.execute_input":"2022-07-19T20:10:21.843242Z","iopub.status.idle":"2022-07-19T20:10:21.853214Z","shell.execute_reply.started":"2022-07-19T20:10:21.843191Z","shell.execute_reply":"2022-07-19T20:10:21.852397Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"##### Destination","metadata":{}},{"cell_type":"code","source":"plot(train_df,'Destination')","metadata":{"execution":{"iopub.status.busy":"2022-07-19T20:10:21.863731Z","iopub.execute_input":"2022-07-19T20:10:21.864346Z","iopub.status.idle":"2022-07-19T20:10:22.620021Z","shell.execute_reply.started":"2022-07-19T20:10:21.864312Z","shell.execute_reply":"2022-07-19T20:10:22.619093Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"plot(train_df,'Destination','Transported')","metadata":{"execution":{"iopub.status.busy":"2022-07-19T20:10:22.621094Z","iopub.execute_input":"2022-07-19T20:10:22.621615Z","iopub.status.idle":"2022-07-19T20:10:22.925644Z","shell.execute_reply.started":"2022-07-19T20:10:22.621580Z","shell.execute_reply":"2022-07-19T20:10:22.922996Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"plot_missing(train_df,'Destination','Transported')","metadata":{"execution":{"iopub.status.busy":"2022-07-19T20:10:22.926871Z","iopub.execute_input":"2022-07-19T20:10:22.927776Z","iopub.status.idle":"2022-07-19T20:10:23.603609Z","shell.execute_reply.started":"2022-07-19T20:10:22.927731Z","shell.execute_reply":"2022-07-19T20:10:23.600436Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"train_df['Destination'].isna().sum()","metadata":{"execution":{"iopub.status.busy":"2022-07-19T20:10:23.605177Z","iopub.execute_input":"2022-07-19T20:10:23.605680Z","iopub.status.idle":"2022-07-19T20:10:23.614969Z","shell.execute_reply.started":"2022-07-19T20:10:23.605641Z","shell.execute_reply":"2022-07-19T20:10:23.613689Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"pd.pivot_table(train_df\n                , values='PassengerId'\n                , index='Deck'\n                , columns='Destination'\n                , aggfunc='count')","metadata":{"execution":{"iopub.status.busy":"2022-07-19T20:10:23.616398Z","iopub.execute_input":"2022-07-19T20:10:23.616812Z","iopub.status.idle":"2022-07-19T20:10:23.647389Z","shell.execute_reply.started":"2022-07-19T20:10:23.616777Z","shell.execute_reply":"2022-07-19T20:10:23.646595Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"pd.pivot_table(train_df\n                , values='PassengerId'\n                , index='HomePlanet'\n                , columns='Destination'\n                , aggfunc='count')","metadata":{"execution":{"iopub.status.busy":"2022-07-19T20:10:23.648440Z","iopub.execute_input":"2022-07-19T20:10:23.649117Z","iopub.status.idle":"2022-07-19T20:10:23.672633Z","shell.execute_reply.started":"2022-07-19T20:10:23.649081Z","shell.execute_reply":"2022-07-19T20:10:23.671377Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# simple impute with mode\nfor dataset in combine:\n    dataset['Destination'] = dataset['Destination'].transform(lambda x: x.fillna(x.mode()[0]))","metadata":{"execution":{"iopub.status.busy":"2022-07-19T20:10:23.674318Z","iopub.execute_input":"2022-07-19T20:10:23.674703Z","iopub.status.idle":"2022-07-19T20:10:23.684520Z","shell.execute_reply.started":"2022-07-19T20:10:23.674668Z","shell.execute_reply":"2022-07-19T20:10:23.683337Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"train_df['Destination'].isna().sum()","metadata":{"execution":{"iopub.status.busy":"2022-07-19T20:10:23.685997Z","iopub.execute_input":"2022-07-19T20:10:23.686384Z","iopub.status.idle":"2022-07-19T20:10:23.697720Z","shell.execute_reply.started":"2022-07-19T20:10:23.686351Z","shell.execute_reply":"2022-07-19T20:10:23.696789Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"##### VIP","metadata":{}},{"cell_type":"code","source":"plot(train_df,'VIP')","metadata":{"execution":{"iopub.status.busy":"2022-07-19T20:10:23.699132Z","iopub.execute_input":"2022-07-19T20:10:23.699546Z","iopub.status.idle":"2022-07-19T20:10:24.401903Z","shell.execute_reply.started":"2022-07-19T20:10:23.699513Z","shell.execute_reply":"2022-07-19T20:10:24.398639Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"plot(train_df,'VIP','Transported')","metadata":{"execution":{"iopub.status.busy":"2022-07-19T20:10:24.403533Z","iopub.execute_input":"2022-07-19T20:10:24.404011Z","iopub.status.idle":"2022-07-19T20:10:24.709812Z","shell.execute_reply.started":"2022-07-19T20:10:24.403974Z","shell.execute_reply":"2022-07-19T20:10:24.706903Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"plot_missing(train_df,'VIP','Transported')","metadata":{"execution":{"iopub.status.busy":"2022-07-19T20:10:24.711574Z","iopub.execute_input":"2022-07-19T20:10:24.711944Z","iopub.status.idle":"2022-07-19T20:10:25.382074Z","shell.execute_reply.started":"2022-07-19T20:10:24.711907Z","shell.execute_reply":"2022-07-19T20:10:25.379169Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# simple impute with mode\nfor dataset in combine:\n    dataset['VIP'] = dataset['VIP'].transform(lambda x: x.fillna(x.mode()[0]))","metadata":{"execution":{"iopub.status.busy":"2022-07-19T20:10:25.383958Z","iopub.execute_input":"2022-07-19T20:10:25.385007Z","iopub.status.idle":"2022-07-19T20:10:25.399549Z","shell.execute_reply.started":"2022-07-19T20:10:25.384957Z","shell.execute_reply":"2022-07-19T20:10:25.398120Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"train_df['VIP'].isna().sum()","metadata":{"execution":{"iopub.status.busy":"2022-07-19T20:10:25.401468Z","iopub.execute_input":"2022-07-19T20:10:25.401858Z","iopub.status.idle":"2022-07-19T20:10:25.416476Z","shell.execute_reply.started":"2022-07-19T20:10:25.401820Z","shell.execute_reply":"2022-07-19T20:10:25.415542Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"for dataset in combine:\n    dataset['VIP'] = dataset['VIP'].replace({'True':1,'False':0}).astype('int')","metadata":{"execution":{"iopub.status.busy":"2022-07-19T20:10:25.417502Z","iopub.execute_input":"2022-07-19T20:10:25.417846Z","iopub.status.idle":"2022-07-19T20:10:25.438414Z","shell.execute_reply.started":"2022-07-19T20:10:25.417816Z","shell.execute_reply":"2022-07-19T20:10:25.437503Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"##### FamiliySize","metadata":{}},{"cell_type":"code","source":"for dataset in combine:\n    dataset['FamilySize'] = dataset.groupby('FamilyName')['FamilyName'].transform('count')\n\ntrain_df.head()","metadata":{"execution":{"iopub.status.busy":"2022-07-19T20:10:25.440187Z","iopub.execute_input":"2022-07-19T20:10:25.440799Z","iopub.status.idle":"2022-07-19T20:10:25.483341Z","shell.execute_reply.started":"2022-07-19T20:10:25.440752Z","shell.execute_reply":"2022-07-19T20:10:25.482443Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"plot(train_df,'FamilySize')","metadata":{"execution":{"iopub.status.busy":"2022-07-19T20:10:25.485296Z","iopub.execute_input":"2022-07-19T20:10:25.485685Z","iopub.status.idle":"2022-07-19T20:10:26.113251Z","shell.execute_reply.started":"2022-07-19T20:10:25.485647Z","shell.execute_reply":"2022-07-19T20:10:26.111748Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"plot(train_df,'FamilySize','Transported')","metadata":{"execution":{"iopub.status.busy":"2022-07-19T20:10:26.114713Z","iopub.execute_input":"2022-07-19T20:10:26.115092Z","iopub.status.idle":"2022-07-19T20:10:26.571445Z","shell.execute_reply.started":"2022-07-19T20:10:26.115057Z","shell.execute_reply":"2022-07-19T20:10:26.570430Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"train_df[['FamilySize','Transported']].groupby(['FamilySize'], as_index=False).mean().sort_values(by='FamilySize', ascending=False)","metadata":{"execution":{"iopub.status.busy":"2022-07-19T20:10:26.572860Z","iopub.execute_input":"2022-07-19T20:10:26.574175Z","iopub.status.idle":"2022-07-19T20:10:26.598374Z","shell.execute_reply.started":"2022-07-19T20:10:26.574119Z","shell.execute_reply":"2022-07-19T20:10:26.597351Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Impute with median\nfor dataset in combine:\n    dataset['FamilySize'] = dataset['FamilySize'].fillna(dataset['FamilySize'].median()).astype('int')","metadata":{"execution":{"iopub.status.busy":"2022-07-19T20:10:26.600227Z","iopub.execute_input":"2022-07-19T20:10:26.602157Z","iopub.status.idle":"2022-07-19T20:10:26.616805Z","shell.execute_reply.started":"2022-07-19T20:10:26.601929Z","shell.execute_reply":"2022-07-19T20:10:26.615599Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"plot(train_df,'FamilySize','Transported')","metadata":{"execution":{"iopub.status.busy":"2022-07-19T20:10:26.618174Z","iopub.execute_input":"2022-07-19T20:10:26.619214Z","iopub.status.idle":"2022-07-19T20:10:27.069728Z","shell.execute_reply.started":"2022-07-19T20:10:26.619174Z","shell.execute_reply":"2022-07-19T20:10:27.067881Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"train_df['FamilySize'].isna().sum()","metadata":{"execution":{"iopub.status.busy":"2022-07-19T20:10:27.071320Z","iopub.execute_input":"2022-07-19T20:10:27.072437Z","iopub.status.idle":"2022-07-19T20:10:27.082520Z","shell.execute_reply.started":"2022-07-19T20:10:27.072387Z","shell.execute_reply":"2022-07-19T20:10:27.081115Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"##### Finishing up","metadata":{}},{"cell_type":"code","source":"train_df.head()","metadata":{"execution":{"iopub.status.busy":"2022-07-19T20:10:27.084644Z","iopub.execute_input":"2022-07-19T20:10:27.085170Z","iopub.status.idle":"2022-07-19T20:10:27.122403Z","shell.execute_reply.started":"2022-07-19T20:10:27.085110Z","shell.execute_reply":"2022-07-19T20:10:27.121537Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# drop features\ntrain_df = train_df.drop(['Cabin','Name',col,'PaxGroup','AgeBand','GivenName','FamilyName','VIP','HasSpend','CabinNumber'], axis=1)\ntest_df = test_df.drop(['Cabin','Name',col,'PaxGroup','GivenName','FamilyName','VIP','HasSpend','CabinNumber'], axis=1)\n\ntrain_df.head()","metadata":{"execution":{"iopub.status.busy":"2022-07-19T20:10:27.123731Z","iopub.execute_input":"2022-07-19T20:10:27.124604Z","iopub.status.idle":"2022-07-19T20:10:27.150062Z","shell.execute_reply.started":"2022-07-19T20:10:27.124564Z","shell.execute_reply":"2022-07-19T20:10:27.148923Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# encode\ntrain_df = pd.get_dummies(train_df, columns=['CryoSleep','HomePlanet','Destination','DeckSide','Deck'], drop_first=True)\ntest_df = pd.get_dummies(test_df, columns=['CryoSleep','HomePlanet','Destination','DeckSide','Deck'], drop_first=True)\n\ntrain_df.head()","metadata":{"execution":{"iopub.status.busy":"2022-07-19T20:10:27.151388Z","iopub.execute_input":"2022-07-19T20:10:27.152505Z","iopub.status.idle":"2022-07-19T20:10:27.196882Z","shell.execute_reply.started":"2022-07-19T20:10:27.152439Z","shell.execute_reply":"2022-07-19T20:10:27.195833Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"#### Model","metadata":{}},{"cell_type":"code","source":"X_train = train_df.drop(['Transported','PassengerId'], axis=1)\ny_train = train_df['Transported']\nX_test  = test_df.drop(['PassengerId'], axis=1).copy()\nX_train.shape, y_train.shape, X_test.shape","metadata":{"execution":{"iopub.status.busy":"2022-07-19T20:10:27.198980Z","iopub.execute_input":"2022-07-19T20:10:27.200972Z","iopub.status.idle":"2022-07-19T20:10:27.214069Z","shell.execute_reply.started":"2022-07-19T20:10:27.200926Z","shell.execute_reply":"2022-07-19T20:10:27.212706Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"##### Random Forest","metadata":{}},{"cell_type":"code","source":"# random search parameters\n\n#n_estimators = [int(x) for x in np.linspace(start = 0, stop = 2000, num = 10)]\n#max_features = ['auto', 'sqrt']\n#max_depth = [int(x) for x in np.linspace(10, 110, num = 11)]\n#max_depth.append(None)\n#min_samples_split = [2, 5, 10]\n#min_samples_leaf = [1, 2, 4]\n#bootstrap = [True, False]\n\n\n#parameters = {'n_estimators': n_estimators,\n#                'max_features': max_features,\n#                'max_depth': max_depth,\n#                'min_samples_split': min_samples_split,\n#                'min_samples_leaf': min_samples_leaf,\n#                'bootstrap': bootstrap}\n\n#random_forest = RandomForestClassifier()\n#random_forest_rand = RandomizedSearchCV(random_forest, param_distributions  = parameters, cv = 5, scoring = 'accuracy', n_jobs= -1, )\n#random_forest_rand.fit(X_train, y_train)\n\n#acc_random_forest = round(random_forest_rand.best_score_ * 100, 2)\n#print(acc_random_forest)\n#print()\n#print(random_forest_rand.best_params_)","metadata":{"execution":{"iopub.status.busy":"2022-07-19T20:10:27.216091Z","iopub.execute_input":"2022-07-19T20:10:27.216528Z","iopub.status.idle":"2022-07-19T20:10:27.222074Z","shell.execute_reply.started":"2022-07-19T20:10:27.216488Z","shell.execute_reply":"2022-07-19T20:10:27.220884Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# # grid search parameters\n# parameters = {\n#     'n_estimators': [2000] \n#     ,'max_depth': [65,70,75]\n#     ,'min_samples_leaf' : [3,4,5]\n#     ,'min_samples_split' : [8,10,12]\n#     ,'bootstrap' : ['True']\n#     ,'max_features' : ['sqrt']}\n\n# random_forest = RandomForestClassifier()\n# Random_forest_grid = GridSearchCV(random_forest, param_grid=parameters, cv=5, scoring='accuracy', n_jobs=-1)\n# Random_forest_grid.fit(X_train, y_train)\n\n# acc_random_forest = round(Random_forest_grid.best_score_ * 100, 2)\n# print(acc_random_forest)\n# print()\n# print(Random_forest_grid.best_params_)","metadata":{"execution":{"iopub.status.busy":"2022-07-19T20:10:27.223145Z","iopub.execute_input":"2022-07-19T20:10:27.224047Z","iopub.status.idle":"2022-07-19T20:10:27.238339Z","shell.execute_reply.started":"2022-07-19T20:10:27.224011Z","shell.execute_reply":"2022-07-19T20:10:27.236588Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# # test\n# random_forest = RandomForestClassifier(max_depth=70, min_samples_leaf=3, min_samples_split=12, n_estimators=2000, bootstrap=True, max_features='sqrt')\n# random_forest.fit(X_train, y_train)\n# y_pred = random_forest.predict(X_test)","metadata":{"execution":{"iopub.status.busy":"2022-07-19T20:10:27.240187Z","iopub.execute_input":"2022-07-19T20:10:27.240605Z","iopub.status.idle":"2022-07-19T20:10:27.257778Z","shell.execute_reply.started":"2022-07-19T20:10:27.240571Z","shell.execute_reply":"2022-07-19T20:10:27.256711Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# importances = random_forest.feature_importances_\n# sorted_indices = np.argsort(importances)[::-1]\n\n# plt.figure(figsize=(10,8))\n# plt.bar(range(X_train.shape[1]), importances[sorted_indices], align='center')\n# plt.xticks(range(X_train.shape[1]), X_train.columns[sorted_indices], rotation=90)\n# plt.show()","metadata":{"execution":{"iopub.status.busy":"2022-07-19T20:10:27.259311Z","iopub.execute_input":"2022-07-19T20:10:27.259999Z","iopub.status.idle":"2022-07-19T20:10:27.276118Z","shell.execute_reply.started":"2022-07-19T20:10:27.259954Z","shell.execute_reply":"2022-07-19T20:10:27.273892Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"##### XGBoost","metadata":{}},{"cell_type":"code","source":"# random search parameters\n#parameters = {'gamma': [0,1,5]\n#                ,'random_state':[42]\n#                ,'eval_metric':['auc','error]\n#                ,'objective':['binary:logistic']\n#                ,'min_child_weight': [1]\n#                ,'subsample': np.linspace(0,1,11)\n#                ,'colsample_bytree': np.linspace(0,1,11)\n#                ,'max_depth': range(3, 11)\n#                ,'n_estimators': range(50,1000,50)\n#                ,'learning_rate': np.linspace(0,1,11)}\n\n#xgb = XGBClassifier()\n#xgb_rand = RandomizedSearchCV(xgb, param_distributions=parameters, cv=5, scoring='accuracy', n_jobs=-1)\n#xgb_rand.fit(X_train, y_train)\n\n#acc_random_forest = round(xgb_rand.best_score_ * 100, 2)\n#print(acc_random_forest)\n#print()\n#print(xgb_rand.best_params_)","metadata":{"execution":{"iopub.status.busy":"2022-07-19T20:10:27.278710Z","iopub.execute_input":"2022-07-19T20:10:27.279771Z","iopub.status.idle":"2022-07-19T20:10:27.293188Z","shell.execute_reply.started":"2022-07-19T20:10:27.279711Z","shell.execute_reply":"2022-07-19T20:10:27.292165Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# grid search parameters\nparameters = {'gamma': [5]\n                 ,'random_state':[42]\n                 ,'eval_metric':['auc']\n                 ,'objective':['binary:logistic']\n                 ,'min_child_weight': [1]\n                 ,'subsample': [1]\n                 ,'colsample_bytree': [0.7]\n                 ,'max_depth': [8]\n                 ,'n_estimators': [1000]\n                 ,'learning_rate': [0.1]}\n\nxgb = XGBClassifier()\nxgb_cv = GridSearchCV(estimator=xgb, cv=5, param_grid=parameters, n_jobs=-1)\nxgb_cv.fit(X_train,y_train)\n\nprint(xgb_cv.best_score_)\nprint()\nprint(xgb_cv.best_params_)","metadata":{"execution":{"iopub.status.busy":"2022-07-19T20:10:27.294746Z","iopub.execute_input":"2022-07-19T20:10:27.295394Z","iopub.status.idle":"2022-07-19T20:11:12.626988Z","shell.execute_reply.started":"2022-07-19T20:10:27.295318Z","shell.execute_reply":"2022-07-19T20:11:12.625505Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# test\ny_pred = xgb_cv.predict(X_test).astype('bool')","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"importances = xgb_cv.best_estimator_.feature_importances_\nsorted_indices = np.argsort(importances)[::-1]\n\nplt.figure(figsize=(10,8))\nplt.bar(range(X_train.shape[1]), importances[sorted_indices], align='center')\nplt.xticks(range(X_train.shape[1]), X_train.columns[sorted_indices], rotation=90)\nplt.show()","metadata":{"execution":{"iopub.status.busy":"2022-07-19T20:11:37.883643Z","iopub.execute_input":"2022-07-19T20:11:37.884133Z","iopub.status.idle":"2022-07-19T20:11:38.464778Z","shell.execute_reply.started":"2022-07-19T20:11:37.884095Z","shell.execute_reply":"2022-07-19T20:11:38.463925Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"##### Submission","metadata":{}},{"cell_type":"code","source":"submission = pd.DataFrame({\n        'PassengerId': test_df['PassengerId'],\n        'Transported': y_pred\n    })","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# local\nsubmission.to_csv('./data/submission.csv', index=False)\n# kaggle notebook\n#submission.to_csv('/kaggle/working/submission.csv', index=False)","metadata":{"trusted":true},"execution_count":null,"outputs":[]}]}