{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"pygments_lexer":"ipython3","nbconvert_exporter":"python","version":"3.6.4","file_extension":".py","codemirror_mode":{"name":"ipython","version":3},"name":"python","mimetype":"text/x-python"}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"<center><img src=\"https://d.newsweek.com/en/full/1859515/screengrab-titanic-movie.jpg\" width=500></center>\n\n<h1><center>🛥️ Titanic - Machine Learning from Disaster 🛥️</center></h1>","metadata":{"_cell_guid":"b1076dfc-b9ad-4769-8c92-a6c4dae69d19","_uuid":"8f2839f25d086af736a60e9eeb907d3b93b6e0e5","papermill":{"duration":0.011613,"end_time":"2022-07-11T03:11:31.529325","exception":false,"start_time":"2022-07-11T03:11:31.517712","status":"completed"},"tags":[]}},{"cell_type":"markdown","source":"# 1. Introduction\n\nThe sinking of the Titanic is one of the most infamous shipwrecks in history. 🧊🛥️\n\nOn April 15, 1912, during her maiden voyage, the widely considered “unsinkable” RMS Titanic sank after colliding with an iceberg. Unfortunately, there weren’t enough lifeboats for everyone onboard, resulting in the death of 1502 out of 2224 passengers and crew.\n\nWhile there was some element of luck involved in surviving, certain groups of people were more likely to survive than others, and while some passenger's hearts would go on... not everyone's did, sorry jack!  💔\n\n### Libraries 📚⬇","metadata":{"papermill":{"duration":0.009939,"end_time":"2022-07-11T03:11:31.549649","exception":false,"start_time":"2022-07-11T03:11:31.53971","status":"completed"},"tags":[]}},{"cell_type":"code","source":"# Data Manipulation\nimport numpy as np\nimport pandas as pd\n\n# Visualization\n%matplotlib inline\nimport matplotlib.pyplot as plt\nimport missingno\nimport seaborn as sns\nplt.style.use('fivethirtyeight')\n\n# Preprocessing\nfrom sklearn.preprocessing import OneHotEncoder, LabelEncoder, label_binarize\n\n# Machine learning\nimport xgboost as xgb\nfrom sklearn.preprocessing import StandardScaler\nfrom sklearn.model_selection import train_test_split\nfrom sklearn import model_selection, tree, preprocessing, metrics, linear_model\nfrom sklearn.svm import LinearSVC\nfrom sklearn.neighbors import KNeighborsClassifier\nfrom sklearn.naive_bayes import GaussianNB\nfrom sklearn.linear_model import LinearRegression, LogisticRegression, SGDClassifier\nfrom sklearn.tree import DecisionTreeClassifier\nfrom catboost import CatBoostClassifier, Pool, cv\nfrom sklearn.ensemble import RandomForestClassifier, AdaBoostClassifier, GradientBoostingClassifier, ExtraTreesClassifier\n\n# Ignore warnings\nimport warnings\nwarnings.filterwarnings('ignore')","metadata":{"papermill":{"duration":1.796884,"end_time":"2022-07-11T03:11:33.356716","exception":false,"start_time":"2022-07-11T03:11:31.559832","status":"completed"},"tags":[],"execution":{"iopub.status.busy":"2022-07-11T03:24:15.825189Z","iopub.execute_input":"2022-07-11T03:24:15.825586Z","iopub.status.idle":"2022-07-11T03:24:15.839743Z","shell.execute_reply.started":"2022-07-11T03:24:15.825554Z","shell.execute_reply":"2022-07-11T03:24:15.838895Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# 2. The .csv files 📁\n\n> 📌**Note**:\n* `train.csv` contains the outcome for each passenger. It contains datapoints in  unique columns.\n* `test.csv` does not contain the outcome for each passenger. This will be used to see how well our model performs on unseen data.\n* `gender_submission.csv` servers as an example of what the submission file should look like. \n\n<div class=\"alert alert-block alert-info\">\n<b>Note:</b> I had previously run Pandas Profiling <code>df.profile_report()</code> to create a basic report on the input <code>train.csv</code> DataFrame. This is similar to using <code>df.describe()</code> but provides some more information fairly quickly. By running this I was able to see that. The Age column has <b>177 (19.9%)</b> missing values and Cabin has <b>687 (77.1%)</b> missing values. For this notebook I've left Pandas Profiling out, but provided a link below for anyone who wants to learn more. \n</div>\n \n### Links 🔗\n* [Pandas Profiling](https://pandas-profiling.ydata.ai/docs/master/index.html)","metadata":{"papermill":{"duration":0.011424,"end_time":"2022-07-11T03:11:33.378987","exception":false,"start_time":"2022-07-11T03:11:33.367563","status":"completed"},"tags":[]}},{"cell_type":"code","source":"# Import train and test data\ntrain = pd.read_csv(\"../input/titanic/train.csv\")\ntest = pd.read_csv(\"../input/titanic/test.csv\")\ngender_submission = pd.read_csv(\"../input/titanic/gender_submission.csv\")","metadata":{"papermill":{"duration":0.045887,"end_time":"2022-07-11T03:11:33.435163","exception":false,"start_time":"2022-07-11T03:11:33.389276","status":"completed"},"tags":[],"execution":{"iopub.status.busy":"2022-07-11T03:24:15.841872Z","iopub.execute_input":"2022-07-11T03:24:15.842277Z","iopub.status.idle":"2022-07-11T03:24:15.870273Z","shell.execute_reply.started":"2022-07-11T03:24:15.842243Z","shell.execute_reply":"2022-07-11T03:24:15.869158Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### train.csv - let's take a look at the DataFrames first, then types\n> 📌**Note**:\n* `df.head()`: Check to make sure the data looks correct in both train & test\n* `df.dtypes`: Check to see the data types in both train & test\n* `gender_submission.head()`: Check to see what the submission data should look like\n*  Make note that columns Sex, Cabin & Embarked may need to be converted from a `string` to an `int`","metadata":{"papermill":{"duration":0.010353,"end_time":"2022-07-11T03:11:33.455966","exception":false,"start_time":"2022-07-11T03:11:33.445613","status":"completed"},"tags":[]}},{"cell_type":"code","source":"# Explore the format\nprint(\"Train shape: {}\".format(train.shape))\nprint(\"Test shape: {}\".format(test.shape))\n\n# Explore the head of the dataframe\ntrain.head()","metadata":{"papermill":{"duration":0.03872,"end_time":"2022-07-11T03:11:33.504742","exception":false,"start_time":"2022-07-11T03:11:33.466022","status":"completed"},"tags":[],"execution":{"iopub.status.busy":"2022-07-11T03:24:15.872639Z","iopub.execute_input":"2022-07-11T03:24:15.873316Z","iopub.status.idle":"2022-07-11T03:24:15.893285Z","shell.execute_reply.started":"2022-07-11T03:24:15.873273Z","shell.execute_reply":"2022-07-11T03:24:15.892413Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# 3. Data Preprocessing 🧼\n\nThere are many techniques that can be performed during this phase. Some of them are:\n\n* checking for missing data\n* analysis of the distributions and patterns in the data\n* creating new features from the existing data (feature engineering)\n* encoding the categorical features\n* etc.\n\n### Other helpful exploratory data analysis\n* `train.columns`: Check the name of the columns\n* `train.shape`: Check the shape of the columns\n* `train.isnull().sum()`: Check to see which columns have missing data\n","metadata":{"papermill":{"duration":0.010055,"end_time":"2022-07-11T03:11:33.525248","exception":false,"start_time":"2022-07-11T03:11:33.515193","status":"completed"},"tags":[]}},{"cell_type":"code","source":"# Check for missing values\ntrain.isnull().sum()\n# Plot graphic of missing values\nmissingno.matrix(train, figsize = (30,10));","metadata":{"papermill":{"duration":0.542867,"end_time":"2022-07-11T03:11:34.078275","exception":false,"start_time":"2022-07-11T03:11:33.535408","status":"completed"},"tags":[],"execution":{"iopub.status.busy":"2022-07-11T03:24:15.894625Z","iopub.execute_input":"2022-07-11T03:24:15.895680Z","iopub.status.idle":"2022-07-11T03:24:16.510737Z","shell.execute_reply.started":"2022-07-11T03:24:15.895630Z","shell.execute_reply":"2022-07-11T03:24:16.509818Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Histogram of passenger ages.\nplt.figure(figsize=(10,6))\nsns.histplot(train['Age'].dropna(),bins=30, palette='inferno', color='#7D0A45', alpha=0.5, kde=True);","metadata":{"papermill":{"duration":0.299945,"end_time":"2022-07-11T03:11:34.390634","exception":false,"start_time":"2022-07-11T03:11:34.090689","status":"completed"},"tags":[],"execution":{"iopub.status.busy":"2022-07-11T03:24:16.512790Z","iopub.execute_input":"2022-07-11T03:24:16.513402Z","iopub.status.idle":"2022-07-11T03:24:16.820231Z","shell.execute_reply.started":"2022-07-11T03:24:16.513364Z","shell.execute_reply":"2022-07-11T03:24:16.818672Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Visualize the survival states\nplt.figure(figsize=(10,6))\ntrain['Survived'].value_counts().plot.pie(explode=[0, 0.1], autopct='%1.1f%%', shadow=True, colors=['#7D0A45', '#B11463']);","metadata":{"papermill":{"duration":0.152271,"end_time":"2022-07-11T03:11:34.555359","exception":false,"start_time":"2022-07-11T03:11:34.403088","status":"completed"},"tags":[],"execution":{"iopub.status.busy":"2022-07-11T03:24:16.822119Z","iopub.execute_input":"2022-07-11T03:24:16.822475Z","iopub.status.idle":"2022-07-11T03:24:16.972052Z","shell.execute_reply.started":"2022-07-11T03:24:16.822444Z","shell.execute_reply":"2022-07-11T03:24:16.969734Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"#### Imputation?\n\nWe know that Age and Cabin columns have some missing data, how should we think about resolving that issue? The two most common methods would be to remove the data -- or use imputation. \n\n> Imputation is a technique that replaces missing values in the data with another value (like the mean, median, mode, or other more complex operations). Be weary of the bias!\n\nFor Age, we'll use imputation","metadata":{"papermill":{"duration":0.025745,"end_time":"2022-07-11T03:11:34.607744","exception":false,"start_time":"2022-07-11T03:11:34.581999","status":"completed"},"tags":[]}},{"cell_type":"code","source":"# Age had a lot of missing values, so we will need to fill in that data. \ntrain[\"Age\"] = train[\"Age\"].fillna(-0.5)\ntest[\"Age\"] = test[\"Age\"].fillna(-0.5)\n\n# We can also sort ages into categories\nbins = [0, 5, 12, 18, 24, 35, 60, np.inf]\nlabels = ['Baby', 'Child', 'Teenager', 'Student', 'Young Adult', 'Adult', 'Senior']\ntrain['AgeGroup'] = pd.cut(train[\"Age\"], bins, labels = labels)\ntest['AgeGroup'] = pd.cut(test[\"Age\"], bins, labels = labels)\n\n# Draw a bar plot of age vs. survival\nplt.figure(figsize=(10,6))\nsns.barplot(x=\"AgeGroup\", y=\"Survived\", data=train)\nplt.show();","metadata":{"papermill":{"duration":0.376175,"end_time":"2022-07-11T03:11:35.007812","exception":false,"start_time":"2022-07-11T03:11:34.631637","status":"completed"},"tags":[],"execution":{"iopub.status.busy":"2022-07-11T03:24:16.974249Z","iopub.execute_input":"2022-07-11T03:24:16.974734Z","iopub.status.idle":"2022-07-11T03:24:17.678278Z","shell.execute_reply.started":"2022-07-11T03:24:16.974686Z","shell.execute_reply":"2022-07-11T03:24:17.677129Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# We can drop the AgeGroup column now as we were using it to view the data above\ntrain = train.drop(['AgeGroup'], axis = 1)\ntest = test.drop(['AgeGroup'], axis = 1)","metadata":{"papermill":{"duration":0.024224,"end_time":"2022-07-11T03:11:35.045532","exception":false,"start_time":"2022-07-11T03:11:35.021308","status":"completed"},"tags":[],"execution":{"iopub.status.busy":"2022-07-11T03:24:17.680090Z","iopub.execute_input":"2022-07-11T03:24:17.680926Z","iopub.status.idle":"2022-07-11T03:24:17.690130Z","shell.execute_reply.started":"2022-07-11T03:24:17.680880Z","shell.execute_reply":"2022-07-11T03:24:17.688835Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# We can see if there is a coorelation between fare price and survival rate\nfare_not_survived = train['Fare'][train['Survived'] == 0]\nfare_survived = train['Fare'][train['Survived'] == 1]\n\naverage_fare = pd.DataFrame([fare_not_survived.mean(), fare_survived.mean()])\nstd_fare = pd.DataFrame([fare_not_survived.std(), fare_survived.std()])\naverage_fare.plot(yerr=std_fare, kind='bar', legend=False)\n\nplt.show();","metadata":{"papermill":{"duration":0.178283,"end_time":"2022-07-11T03:11:35.236986","exception":false,"start_time":"2022-07-11T03:11:35.058703","status":"completed"},"tags":[],"execution":{"iopub.status.busy":"2022-07-11T03:24:17.693130Z","iopub.execute_input":"2022-07-11T03:24:17.693917Z","iopub.status.idle":"2022-07-11T03:24:17.900222Z","shell.execute_reply.started":"2022-07-11T03:24:17.693863Z","shell.execute_reply":"2022-07-11T03:24:17.898353Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# We can also view if where passengers embarked from played a role in their survival\nsns.barplot(x=\"Embarked\", y=\"Survived\", data=train)\nplt.title('Embarked and Survived rate');\nplt.show();","metadata":{"papermill":{"duration":0.25207,"end_time":"2022-07-11T03:11:35.502792","exception":false,"start_time":"2022-07-11T03:11:35.250722","status":"completed"},"tags":[],"execution":{"iopub.status.busy":"2022-07-11T03:24:17.906160Z","iopub.execute_input":"2022-07-11T03:24:17.907457Z","iopub.status.idle":"2022-07-11T03:24:18.204346Z","shell.execute_reply.started":"2022-07-11T03:24:17.907409Z","shell.execute_reply":"2022-07-11T03:24:18.203376Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"#### Fare & Embarked\n\nFrom the above data we can see some interesting points. \n\n* Passengers with a higher fare tended to have a better chance of survival\n* Passengers who embarked from Cherbourg had the highest chance of survival, and those from Southampton the least chance to survive. This is likley due to some socialeconomic differences and which part of the ship these passengers stayed in.","metadata":{"execution":{"iopub.execute_input":"2022-07-11T02:44:49.093311Z","iopub.status.busy":"2022-07-11T02:44:49.092884Z","iopub.status.idle":"2022-07-11T02:44:49.100749Z","shell.execute_reply":"2022-07-11T02:44:49.099681Z","shell.execute_reply.started":"2022-07-11T02:44:49.093276Z"},"papermill":{"duration":0.013309,"end_time":"2022-07-11T03:11:35.529717","exception":false,"start_time":"2022-07-11T03:11:35.516408","status":"completed"},"tags":[]}},{"cell_type":"code","source":"# Create a combined group of our Train and Test data\ncombine = [train, test]\n\n# Extract a title for each Name in the train and test datasets\nfor dataset in combine:\n    dataset['Title'] = dataset.Name.str.extract(' ([A-Za-z]+)\\.', expand=False)\n\npd.crosstab(train['Title'], train['Sex'])","metadata":{"papermill":{"duration":0.042677,"end_time":"2022-07-11T03:11:35.585983","exception":false,"start_time":"2022-07-11T03:11:35.543306","status":"completed"},"tags":[],"execution":{"iopub.status.busy":"2022-07-11T03:24:18.205788Z","iopub.execute_input":"2022-07-11T03:24:18.207114Z","iopub.status.idle":"2022-07-11T03:24:18.238676Z","shell.execute_reply.started":"2022-07-11T03:24:18.207052Z","shell.execute_reply":"2022-07-11T03:24:18.237336Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Replace various titles with more common names\nfor dataset in combine:\n    dataset['Title'] = dataset['Title'].replace(['Lady', 'Capt', 'Col',\n    'Don', 'Dr', 'Major', 'Rev', 'Jonkheer', 'Dona'], 'Rare')\n    \n    dataset['Title'] = dataset['Title'].replace(['Countess', 'Lady', 'Sir'], 'Royal')\n    dataset['Title'] = dataset['Title'].replace('Mlle', 'Miss')\n    dataset['Title'] = dataset['Title'].replace('Ms', 'Miss')\n    dataset['Title'] = dataset['Title'].replace('Mme', 'Mrs')\n\ntrain[['Title', 'Survived']].groupby(['Title'], as_index=False).mean()","metadata":{"papermill":{"duration":0.040556,"end_time":"2022-07-11T03:11:35.640262","exception":false,"start_time":"2022-07-11T03:11:35.599706","status":"completed"},"tags":[],"execution":{"iopub.status.busy":"2022-07-11T03:24:18.240255Z","iopub.execute_input":"2022-07-11T03:24:18.240894Z","iopub.status.idle":"2022-07-11T03:24:18.269951Z","shell.execute_reply.started":"2022-07-11T03:24:18.240837Z","shell.execute_reply":"2022-07-11T03:24:18.268834Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"for dataset in combine:\n    title_mapping = {\"Mr\": 1, \"Miss\": 2, \"Mrs\": 3, \"Master\": 4, \"Rare\": 5}\n    dataset['Title'] = dataset['Title'].map(title_mapping)\n    dataset['Title'] = dataset['Title'].fillna(0)","metadata":{"papermill":{"duration":0.025421,"end_time":"2022-07-11T03:11:35.679831","exception":false,"start_time":"2022-07-11T03:11:35.65441","status":"completed"},"tags":[],"execution":{"iopub.status.busy":"2022-07-11T03:24:18.271398Z","iopub.execute_input":"2022-07-11T03:24:18.271738Z","iopub.status.idle":"2022-07-11T03:24:18.284020Z","shell.execute_reply.started":"2022-07-11T03:24:18.271708Z","shell.execute_reply":"2022-07-11T03:24:18.282243Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# We'll drop the name feature for this analysis. While we could potentially find some data from this, for the purposes of this exercise we will not use this data\ntrain = train.drop(['Name'], axis = 1)\ntest = test.drop(['Name'], axis = 1)","metadata":{"papermill":{"duration":0.023549,"end_time":"2022-07-11T03:11:35.717288","exception":false,"start_time":"2022-07-11T03:11:35.693739","status":"completed"},"tags":[],"execution":{"iopub.status.busy":"2022-07-11T03:24:18.286300Z","iopub.execute_input":"2022-07-11T03:24:18.287132Z","iopub.status.idle":"2022-07-11T03:24:18.298254Z","shell.execute_reply.started":"2022-07-11T03:24:18.287074Z","shell.execute_reply":"2022-07-11T03:24:18.296823Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# We mentioned earlier that Sex would need to be changed to and int\nsex_mapping = {\"male\": 0, \"female\": 1}\ntrain['Sex'] = train['Sex'].map(sex_mapping)\ntest['Sex'] = test['Sex'].map(sex_mapping)\n\ntrain.head()","metadata":{"papermill":{"duration":0.034416,"end_time":"2022-07-11T03:11:35.765638","exception":false,"start_time":"2022-07-11T03:11:35.731222","status":"completed"},"tags":[],"execution":{"iopub.status.busy":"2022-07-11T03:24:18.299771Z","iopub.execute_input":"2022-07-11T03:24:18.300419Z","iopub.status.idle":"2022-07-11T03:24:18.325000Z","shell.execute_reply.started":"2022-07-11T03:24:18.300373Z","shell.execute_reply":"2022-07-11T03:24:18.323681Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"#### Predict Missing Ages\n\nNext we'll fill in the missing values in the Age feature. Since a higher percentage of values are missing, it would be illogical to fill all of them with the same value. Instead, we will try to predict the missing ages.","metadata":{"papermill":{"duration":0.013871,"end_time":"2022-07-11T03:11:35.793799","exception":false,"start_time":"2022-07-11T03:11:35.779928","status":"completed"},"tags":[]}},{"cell_type":"code","source":"# Calculate the correlation between features\ntrain_corr = train.corr().abs().unstack().sort_values(kind=\"quicksort\", ascending=False).reset_index()\ntrain_corr.rename(columns={\"level_0\": \"Feature 1\", \"level_1\": \"Feature 2\", 0: 'Correlation Coefficient'}, inplace=True)\ntrain_corr[train_corr['Feature 1'] == 'Age']","metadata":{"papermill":{"duration":0.033218,"end_time":"2022-07-11T03:11:35.841146","exception":false,"start_time":"2022-07-11T03:11:35.807928","status":"completed"},"tags":[],"execution":{"iopub.status.busy":"2022-07-11T03:24:18.326793Z","iopub.execute_input":"2022-07-11T03:24:18.327193Z","iopub.status.idle":"2022-07-11T03:24:18.348220Z","shell.execute_reply.started":"2022-07-11T03:24:18.327160Z","shell.execute_reply":"2022-07-11T03:24:18.347110Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Take the median value for Age feature based on 'Pclass' and 'Title'\ntrain['Age'] = train.groupby(['Pclass', 'Title'])['Age'].apply(lambda x: x.fillna(x.median()))\ntest['Age'] = test.groupby(['Pclass', 'Title'])['Age'].apply(lambda x: x.fillna(x.median()))","metadata":{"papermill":{"duration":0.040508,"end_time":"2022-07-11T03:11:35.896135","exception":false,"start_time":"2022-07-11T03:11:35.855627","status":"completed"},"tags":[],"execution":{"iopub.status.busy":"2022-07-11T03:24:18.349829Z","iopub.execute_input":"2022-07-11T03:24:18.350497Z","iopub.status.idle":"2022-07-11T03:24:18.378026Z","shell.execute_reply.started":"2022-07-11T03:24:18.350466Z","shell.execute_reply":"2022-07-11T03:24:18.376633Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# We can also drop the Ticket feature since it's unlikely to yield any useful information\ntrain = train.drop(['Ticket'], axis = 1)\ntest = test.drop(['Ticket'], axis = 1)","metadata":{"papermill":{"duration":0.025192,"end_time":"2022-07-11T03:11:35.935819","exception":false,"start_time":"2022-07-11T03:11:35.910627","status":"completed"},"tags":[],"execution":{"iopub.status.busy":"2022-07-11T03:24:18.379702Z","iopub.execute_input":"2022-07-11T03:24:18.380537Z","iopub.status.idle":"2022-07-11T03:24:18.391214Z","shell.execute_reply.started":"2022-07-11T03:24:18.380498Z","shell.execute_reply":"2022-07-11T03:24:18.389354Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Fill in missing Fare value in test set based on mean fare for that Pclass \nfor x in range(len(test[\"Fare\"])):\n    if pd.isnull(test[\"Fare\"][x]):\n        pclass = test[\"Pclass\"][x] #Pclass = 3\n        test[\"Fare\"][x] = round(train[train[\"Pclass\"] == pclass][\"Fare\"].mean(), 4)","metadata":{"papermill":{"duration":0.028686,"end_time":"2022-07-11T03:11:35.978924","exception":false,"start_time":"2022-07-11T03:11:35.950238","status":"completed"},"tags":[],"execution":{"iopub.status.busy":"2022-07-11T03:24:18.393701Z","iopub.execute_input":"2022-07-11T03:24:18.394334Z","iopub.status.idle":"2022-07-11T03:24:18.411339Z","shell.execute_reply.started":"2022-07-11T03:24:18.394278Z","shell.execute_reply":"2022-07-11T03:24:18.409865Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Map Fare values into groups of numerical values\ntrain['FareBand'] = pd.qcut(train['Fare'], 4, labels = [1, 2, 3, 4])\ntest['FareBand'] = pd.qcut(test['Fare'], 4, labels = [1, 2, 3, 4])\nfor dataset in combine:\n    dataset['Fare'] = dataset['Fare'].fillna(train['Fare'].median())","metadata":{"papermill":{"duration":0.028639,"end_time":"2022-07-11T03:11:36.022123","exception":false,"start_time":"2022-07-11T03:11:35.993484","status":"completed"},"tags":[],"execution":{"iopub.status.busy":"2022-07-11T03:24:18.414465Z","iopub.execute_input":"2022-07-11T03:24:18.415633Z","iopub.status.idle":"2022-07-11T03:24:18.436086Z","shell.execute_reply.started":"2022-07-11T03:24:18.415582Z","shell.execute_reply":"2022-07-11T03:24:18.434641Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"for dataset in combine:\n    dataset.loc[ dataset['Fare'] <= 7.91, 'Fare'] = 0\n    dataset.loc[(dataset['Fare'] > 7.91) & (dataset['Fare'] <= 14.454), 'Fare'] = 1\n    dataset.loc[(dataset['Fare'] > 14.454) & (dataset['Fare'] <= 31), 'Fare'] = 2\n    dataset.loc[ dataset['Fare'] > 31, 'Fare'] = 3\n    dataset['Fare'] = dataset['Fare'].astype(int)\n\n# Drop Fare values\ntrain = train.drop(['Fare'], axis = 1)\ntest = test.drop(['Fare'], axis = 1)","metadata":{"papermill":{"duration":0.033934,"end_time":"2022-07-11T03:11:36.070551","exception":false,"start_time":"2022-07-11T03:11:36.036617","status":"completed"},"tags":[],"execution":{"iopub.status.busy":"2022-07-11T03:24:18.438475Z","iopub.execute_input":"2022-07-11T03:24:18.438939Z","iopub.status.idle":"2022-07-11T03:24:18.463581Z","shell.execute_reply.started":"2022-07-11T03:24:18.438896Z","shell.execute_reply":"2022-07-11T03:24:18.461774Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Drop the Cabin feature\ntrain = train.drop(['Cabin'], axis = 1)\ntest = test.drop(['Cabin'], axis = 1)","metadata":{"papermill":{"duration":0.023551,"end_time":"2022-07-11T03:11:36.108489","exception":false,"start_time":"2022-07-11T03:11:36.084938","status":"completed"},"tags":[],"execution":{"iopub.status.busy":"2022-07-11T03:24:18.465159Z","iopub.execute_input":"2022-07-11T03:24:18.465537Z","iopub.status.idle":"2022-07-11T03:24:18.475738Z","shell.execute_reply.started":"2022-07-11T03:24:18.465506Z","shell.execute_reply":"2022-07-11T03:24:18.474203Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Fill missing Embark with \"S\"\ntrain = train.fillna({\"Embarked\": \"S\"})","metadata":{"papermill":{"duration":0.023165,"end_time":"2022-07-11T03:11:36.146126","exception":false,"start_time":"2022-07-11T03:11:36.122961","status":"completed"},"tags":[],"execution":{"iopub.status.busy":"2022-07-11T03:24:18.477664Z","iopub.execute_input":"2022-07-11T03:24:18.478173Z","iopub.status.idle":"2022-07-11T03:24:18.488417Z","shell.execute_reply.started":"2022-07-11T03:24:18.478127Z","shell.execute_reply":"2022-07-11T03:24:18.486929Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Map each Embarked value to a numerical value\nembarked_mapping = {\"S\": 1, \"C\": 2, \"Q\": 3}\ntrain['Embarked'] = train['Embarked'].map(embarked_mapping)\ntest['Embarked'] = test['Embarked'].map(embarked_mapping)\n\ntrain.head()","metadata":{"papermill":{"duration":0.035104,"end_time":"2022-07-11T03:11:36.195581","exception":false,"start_time":"2022-07-11T03:11:36.160477","status":"completed"},"tags":[],"execution":{"iopub.status.busy":"2022-07-11T03:24:18.490012Z","iopub.execute_input":"2022-07-11T03:24:18.491408Z","iopub.status.idle":"2022-07-11T03:24:18.512599Z","shell.execute_reply.started":"2022-07-11T03:24:18.491356Z","shell.execute_reply":"2022-07-11T03:24:18.511756Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Check to make sure we have no null rows\ntrain.isnull().sum()","metadata":{"papermill":{"duration":0.027295,"end_time":"2022-07-11T03:11:36.237682","exception":false,"start_time":"2022-07-11T03:11:36.210387","status":"completed"},"tags":[],"execution":{"iopub.status.busy":"2022-07-11T03:24:18.518726Z","iopub.execute_input":"2022-07-11T03:24:18.519959Z","iopub.status.idle":"2022-07-11T03:24:18.530225Z","shell.execute_reply.started":"2022-07-11T03:24:18.519901Z","shell.execute_reply":"2022-07-11T03:24:18.529404Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"target = train[\"Survived\"]\npredictors = train.drop(['Survived', 'PassengerId'], axis = 1)","metadata":{"papermill":{"duration":0.024526,"end_time":"2022-07-11T03:11:36.276987","exception":false,"start_time":"2022-07-11T03:11:36.252461","status":"completed"},"tags":[],"execution":{"iopub.status.busy":"2022-07-11T03:24:18.531391Z","iopub.execute_input":"2022-07-11T03:24:18.532225Z","iopub.status.idle":"2022-07-11T03:24:18.544050Z","shell.execute_reply.started":"2022-07-11T03:24:18.532187Z","shell.execute_reply":"2022-07-11T03:24:18.543119Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# 4. Modeling 📷💃\n\nNow to the fun part, we're going to run some models on the data. We will use the following:\n\n* Support Vector Machines\n* K-Nearst Neighbor\n* Logistic Regression\n* Random Forest\n* Naive Bayes\n* Perceptron\n* Linear SVC\n* Decision Tree\n* Stochastic Gradient Descent\n* Gradient Boosting Classifier","metadata":{"execution":{"iopub.execute_input":"2022-07-11T03:00:41.608689Z","iopub.status.busy":"2022-07-11T03:00:41.608319Z","iopub.status.idle":"2022-07-11T03:00:41.61624Z","shell.execute_reply":"2022-07-11T03:00:41.615003Z","shell.execute_reply.started":"2022-07-11T03:00:41.60866Z"},"papermill":{"duration":0.014447,"end_time":"2022-07-11T03:11:36.306121","exception":false,"start_time":"2022-07-11T03:11:36.291674","status":"completed"},"tags":[]}},{"cell_type":"code","source":"from sklearn.model_selection import train_test_split\n\nx_train, x_val, y_train, y_val = train_test_split(predictors, target, test_size = 0.20, random_state = 0)","metadata":{"papermill":{"duration":0.024633,"end_time":"2022-07-11T03:11:36.345473","exception":false,"start_time":"2022-07-11T03:11:36.32084","status":"completed"},"tags":[],"execution":{"iopub.status.busy":"2022-07-11T03:24:18.545569Z","iopub.execute_input":"2022-07-11T03:24:18.546141Z","iopub.status.idle":"2022-07-11T03:24:18.558011Z","shell.execute_reply.started":"2022-07-11T03:24:18.546108Z","shell.execute_reply":"2022-07-11T03:24:18.556770Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"predictors.head()","metadata":{"papermill":{"duration":0.032067,"end_time":"2022-07-11T03:11:36.392214","exception":false,"start_time":"2022-07-11T03:11:36.360147","status":"completed"},"tags":[],"execution":{"iopub.status.busy":"2022-07-11T03:24:18.559711Z","iopub.execute_input":"2022-07-11T03:24:18.560078Z","iopub.status.idle":"2022-07-11T03:24:18.578543Z","shell.execute_reply.started":"2022-07-11T03:24:18.560047Z","shell.execute_reply":"2022-07-11T03:24:18.577364Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Gaussian Naive Bayes\nfrom sklearn.naive_bayes import GaussianNB\nfrom sklearn.metrics import accuracy_score\n\ngaussian = GaussianNB()\ngaussian.fit(x_train, y_train)\ny_pred = gaussian.predict(x_val)\nacc_gaussian = round(accuracy_score(y_pred, y_val) * 100, 2)\nprint(acc_gaussian)","metadata":{"papermill":{"duration":0.032288,"end_time":"2022-07-11T03:11:36.439533","exception":false,"start_time":"2022-07-11T03:11:36.407245","status":"completed"},"tags":[],"execution":{"iopub.status.busy":"2022-07-11T03:24:18.580216Z","iopub.execute_input":"2022-07-11T03:24:18.580560Z","iopub.status.idle":"2022-07-11T03:24:18.596563Z","shell.execute_reply.started":"2022-07-11T03:24:18.580529Z","shell.execute_reply":"2022-07-11T03:24:18.595348Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Logistic Regression\nfrom sklearn.linear_model import LogisticRegression\n\nlogreg = LogisticRegression()\nlogreg.fit(x_train, y_train)\ny_pred = logreg.predict(x_val)\nacc_logreg = round(accuracy_score(y_pred, y_val) * 100, 2)\nprint(acc_logreg)","metadata":{"papermill":{"duration":0.045452,"end_time":"2022-07-11T03:11:36.500342","exception":false,"start_time":"2022-07-11T03:11:36.45489","status":"completed"},"tags":[],"execution":{"iopub.status.busy":"2022-07-11T03:24:18.598141Z","iopub.execute_input":"2022-07-11T03:24:18.599063Z","iopub.status.idle":"2022-07-11T03:24:18.631596Z","shell.execute_reply.started":"2022-07-11T03:24:18.599030Z","shell.execute_reply":"2022-07-11T03:24:18.630568Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Support Vector Machines\nfrom sklearn.svm import SVC\n\nsvc = SVC()\nsvc.fit(x_train, y_train)\ny_pred = svc.predict(x_val)\nacc_svc = round(accuracy_score(y_pred, y_val) * 100, 2)\nprint(acc_svc)","metadata":{"papermill":{"duration":0.057136,"end_time":"2022-07-11T03:11:36.572436","exception":false,"start_time":"2022-07-11T03:11:36.5153","status":"completed"},"tags":[],"execution":{"iopub.status.busy":"2022-07-11T03:24:18.632749Z","iopub.execute_input":"2022-07-11T03:24:18.633413Z","iopub.status.idle":"2022-07-11T03:24:18.673134Z","shell.execute_reply.started":"2022-07-11T03:24:18.633377Z","shell.execute_reply":"2022-07-11T03:24:18.672100Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Linear SVC\nfrom sklearn.svm import LinearSVC\n\nlinear_svc = LinearSVC()\nlinear_svc.fit(x_train, y_train)\ny_pred = linear_svc.predict(x_val)\nacc_linear_svc = round(accuracy_score(y_pred, y_val) * 100, 2)\nprint(acc_linear_svc)","metadata":{"papermill":{"duration":0.064757,"end_time":"2022-07-11T03:11:36.652243","exception":false,"start_time":"2022-07-11T03:11:36.587486","status":"completed"},"tags":[],"execution":{"iopub.status.busy":"2022-07-11T03:24:18.674426Z","iopub.execute_input":"2022-07-11T03:24:18.675031Z","iopub.status.idle":"2022-07-11T03:24:18.729871Z","shell.execute_reply.started":"2022-07-11T03:24:18.674997Z","shell.execute_reply":"2022-07-11T03:24:18.728660Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Perceptron\nfrom sklearn.linear_model import Perceptron\n\nperceptron = Perceptron()\nperceptron.fit(x_train, y_train)\ny_pred = perceptron.predict(x_val)\nacc_perceptron = round(accuracy_score(y_pred, y_val) * 100, 2)\nprint(acc_perceptron)","metadata":{"papermill":{"duration":0.030745,"end_time":"2022-07-11T03:11:36.698178","exception":false,"start_time":"2022-07-11T03:11:36.667433","status":"completed"},"tags":[],"execution":{"iopub.status.busy":"2022-07-11T03:24:18.731469Z","iopub.execute_input":"2022-07-11T03:24:18.732014Z","iopub.status.idle":"2022-07-11T03:24:18.746473Z","shell.execute_reply.started":"2022-07-11T03:24:18.731981Z","shell.execute_reply":"2022-07-11T03:24:18.745004Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"#Decision Tree\nfrom sklearn.tree import DecisionTreeClassifier\n\ndecisiontree = DecisionTreeClassifier()\ndecisiontree.fit(x_train, y_train)\ny_pred = decisiontree.predict(x_val)\nacc_decisiontree = round(accuracy_score(y_pred, y_val) * 100, 2)\nprint(acc_decisiontree)","metadata":{"papermill":{"duration":0.032744,"end_time":"2022-07-11T03:11:36.746721","exception":false,"start_time":"2022-07-11T03:11:36.713977","status":"completed"},"tags":[],"execution":{"iopub.status.busy":"2022-07-11T03:24:18.748468Z","iopub.execute_input":"2022-07-11T03:24:18.748824Z","iopub.status.idle":"2022-07-11T03:24:18.762959Z","shell.execute_reply.started":"2022-07-11T03:24:18.748793Z","shell.execute_reply":"2022-07-11T03:24:18.761806Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Random Forest\nfrom sklearn.ensemble import RandomForestClassifier\n\nrandomforest = RandomForestClassifier()\nrandomforest.fit(x_train, y_train)\ny_pred = randomforest.predict(x_val)\nacc_randomforest = round(accuracy_score(y_pred, y_val) * 100, 2)\nprint(acc_randomforest)","metadata":{"papermill":{"duration":0.19566,"end_time":"2022-07-11T03:11:36.957648","exception":false,"start_time":"2022-07-11T03:11:36.761988","status":"completed"},"tags":[],"execution":{"iopub.status.busy":"2022-07-11T03:24:18.764688Z","iopub.execute_input":"2022-07-11T03:24:18.765686Z","iopub.status.idle":"2022-07-11T03:24:19.031703Z","shell.execute_reply.started":"2022-07-11T03:24:18.765624Z","shell.execute_reply":"2022-07-11T03:24:19.030525Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# KNN or k-Nearest Neighbors\nfrom sklearn.neighbors import KNeighborsClassifier\n\nknn = KNeighborsClassifier()\nknn.fit(x_train, y_train)\ny_pred = knn.predict(x_val)\nacc_knn = round(accuracy_score(y_pred, y_val) * 100, 2)\nprint(acc_knn)","metadata":{"papermill":{"duration":0.03642,"end_time":"2022-07-11T03:11:37.009368","exception":false,"start_time":"2022-07-11T03:11:36.972948","status":"completed"},"tags":[],"execution":{"iopub.status.busy":"2022-07-11T03:24:19.033191Z","iopub.execute_input":"2022-07-11T03:24:19.033504Z","iopub.status.idle":"2022-07-11T03:24:19.057260Z","shell.execute_reply.started":"2022-07-11T03:24:19.033475Z","shell.execute_reply":"2022-07-11T03:24:19.056069Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Stochastic Gradient Descent\nfrom sklearn.linear_model import SGDClassifier\n\nsgd = SGDClassifier()\nsgd.fit(x_train, y_train)\ny_pred = sgd.predict(x_val)\nacc_sgd = round(accuracy_score(y_pred, y_val) * 100, 2)\nprint(acc_sgd)","metadata":{"papermill":{"duration":0.03133,"end_time":"2022-07-11T03:11:37.056217","exception":false,"start_time":"2022-07-11T03:11:37.024887","status":"completed"},"tags":[],"execution":{"iopub.status.busy":"2022-07-11T03:24:19.059231Z","iopub.execute_input":"2022-07-11T03:24:19.060009Z","iopub.status.idle":"2022-07-11T03:24:19.076247Z","shell.execute_reply.started":"2022-07-11T03:24:19.059959Z","shell.execute_reply":"2022-07-11T03:24:19.074948Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Gradient Boosting Classifier\nfrom sklearn.ensemble import GradientBoostingClassifier\n\ngbk = GradientBoostingClassifier()\ngbk.fit(x_train, y_train)\ny_pred = gbk.predict(x_val)\nacc_gbk = round(accuracy_score(y_pred, y_val) * 100, 2)\nprint(acc_gbk)","metadata":{"papermill":{"duration":0.115479,"end_time":"2022-07-11T03:11:37.18748","exception":false,"start_time":"2022-07-11T03:11:37.072001","status":"completed"},"tags":[],"execution":{"iopub.status.busy":"2022-07-11T03:24:19.078088Z","iopub.execute_input":"2022-07-11T03:24:19.078496Z","iopub.status.idle":"2022-07-11T03:24:19.203480Z","shell.execute_reply.started":"2022-07-11T03:24:19.078464Z","shell.execute_reply":"2022-07-11T03:24:19.202106Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"models = pd.DataFrame({\n    'Model': ['Support Vector Machines', 'KNN', 'Logistic Regression', \n              'Random Forest', 'Naive Bayes', 'Perceptron', 'Linear SVC', \n              'Decision Tree', 'Stochastic Gradient Descent', 'Gradient Boosting Classifier'],\n    'Score': [acc_svc, acc_knn, acc_logreg, \n              acc_randomforest, acc_gaussian, acc_perceptron,acc_linear_svc, acc_decisiontree,\n              acc_sgd, acc_gbk]})\nmodels.sort_values(by='Score', ascending=False)","metadata":{"papermill":{"duration":0.031091,"end_time":"2022-07-11T03:11:37.234217","exception":false,"start_time":"2022-07-11T03:11:37.203126","status":"completed"},"tags":[],"execution":{"iopub.status.busy":"2022-07-11T03:24:19.205728Z","iopub.execute_input":"2022-07-11T03:24:19.206580Z","iopub.status.idle":"2022-07-11T03:24:19.224368Z","shell.execute_reply.started":"2022-07-11T03:24:19.206520Z","shell.execute_reply":"2022-07-11T03:24:19.223004Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"from sklearn.model_selection import KFold #for K-fold cross validation\nfrom sklearn.model_selection import cross_val_score #score evaluation\nfrom sklearn.model_selection import cross_val_predict #prediction\nfrom xgboost import XGBClassifier\n\nX = train.drop(['Survived', 'PassengerId'], axis=1)\nY = train[\"Survived\"]\n\nkfold = KFold(n_splits=10) # k=10, split the data into 10 equal parts\nxyz=[]\naccuracy=[]\nstd=[]\nclassifiers=['Support Vector Machines', 'K-Nearst Neighbor', 'Logistic Regression', \n              'Random Forest', 'Naive Bayes', 'Perceptron', 'Linear SVC', \n              'Decision Tree', 'Stochastic Gradient Descent', 'Gradient Boosting Classifier']\nmodels=[SVC(), KNeighborsClassifier(), LogisticRegression(), RandomForestClassifier(), GaussianNB(), \n        Perceptron(), LinearSVC(), DecisionTreeClassifier(), SGDClassifier(), \n        GradientBoostingClassifier()]\nfor i in models:\n    model = i\n    cv_result = cross_val_score(model,X,Y, cv = kfold,scoring = \"accuracy\")\n    xyz.append(cv_result.mean())\n    std.append(cv_result.std())\n    accuracy.append(cv_result)\nnew_models_dataframe2=pd.DataFrame({'CV Mean':xyz,'Std':std},index=classifiers)       \nnew_models_dataframe2","metadata":{"papermill":{"duration":3.97785,"end_time":"2022-07-11T03:11:41.227855","exception":false,"start_time":"2022-07-11T03:11:37.250005","status":"completed"},"tags":[],"execution":{"iopub.status.busy":"2022-07-11T03:24:19.226294Z","iopub.execute_input":"2022-07-11T03:24:19.227373Z","iopub.status.idle":"2022-07-11T03:24:24.119739Z","shell.execute_reply.started":"2022-07-11T03:24:19.227323Z","shell.execute_reply":"2022-07-11T03:24:24.118448Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# 5. Modeling Results ⌛\n\nWe've run the data through several models and we can see that Gradient Boosting Classifier was the best result at ~84% accuracy. ","metadata":{"papermill":{"duration":0.015829,"end_time":"2022-07-11T03:11:41.25999","exception":false,"start_time":"2022-07-11T03:11:41.244161","status":"completed"},"tags":[]}},{"cell_type":"code","source":"test.head()","metadata":{"papermill":{"duration":0.031173,"end_time":"2022-07-11T03:11:41.307188","exception":false,"start_time":"2022-07-11T03:11:41.276015","status":"completed"},"tags":[],"execution":{"iopub.status.busy":"2022-07-11T03:24:24.122098Z","iopub.execute_input":"2022-07-11T03:24:24.122593Z","iopub.status.idle":"2022-07-11T03:24:24.139972Z","shell.execute_reply.started":"2022-07-11T03:24:24.122546Z","shell.execute_reply":"2022-07-11T03:24:24.138105Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"train.head()","metadata":{"papermill":{"duration":0.03271,"end_time":"2022-07-11T03:11:41.356061","exception":false,"start_time":"2022-07-11T03:11:41.323351","status":"completed"},"tags":[],"execution":{"iopub.status.busy":"2022-07-11T03:24:24.141514Z","iopub.execute_input":"2022-07-11T03:24:24.141875Z","iopub.status.idle":"2022-07-11T03:24:24.159591Z","shell.execute_reply.started":"2022-07-11T03:24:24.141827Z","shell.execute_reply":"2022-07-11T03:24:24.158779Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Set ids as PassengerId and predict survival \nids = test['PassengerId']\npredictions = gbk.predict(test.drop('PassengerId', axis=1))\n\n# Set the output as a dataframe and convert to csv file named submission.csv\noutput = pd.DataFrame({ 'PassengerId' : ids, 'Survived': predictions })\noutput.to_csv('submission.csv', index=False)\nprint(\"The submission was successfully saved!\")","metadata":{"papermill":{"duration":0.034482,"end_time":"2022-07-11T03:11:41.40696","exception":false,"start_time":"2022-07-11T03:11:41.372478","status":"completed"},"tags":[],"execution":{"iopub.status.busy":"2022-07-11T03:24:24.160987Z","iopub.execute_input":"2022-07-11T03:24:24.161477Z","iopub.status.idle":"2022-07-11T03:24:24.181813Z","shell.execute_reply.started":"2022-07-11T03:24:24.161433Z","shell.execute_reply":"2022-07-11T03:24:24.180207Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"test.head()","metadata":{"papermill":{"duration":0.030597,"end_time":"2022-07-11T03:11:41.454065","exception":false,"start_time":"2022-07-11T03:11:41.423468","status":"completed"},"tags":[],"execution":{"iopub.status.busy":"2022-07-11T03:24:24.183926Z","iopub.execute_input":"2022-07-11T03:24:24.185082Z","iopub.status.idle":"2022-07-11T03:24:24.201829Z","shell.execute_reply.started":"2022-07-11T03:24:24.185039Z","shell.execute_reply":"2022-07-11T03:24:24.200209Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"output.head()","metadata":{"papermill":{"duration":0.027516,"end_time":"2022-07-11T03:11:41.498382","exception":false,"start_time":"2022-07-11T03:11:41.470866","status":"completed"},"tags":[],"execution":{"iopub.status.busy":"2022-07-11T03:24:24.203453Z","iopub.execute_input":"2022-07-11T03:24:24.203790Z","iopub.status.idle":"2022-07-11T03:24:24.214803Z","shell.execute_reply.started":"2022-07-11T03:24:24.203760Z","shell.execute_reply":"2022-07-11T03:24:24.213511Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# 6. Submit the Data 📤\n\nWe have finished up by ensuring that our output matches the format we need to submit the data for this competition! ","metadata":{"papermill":{"duration":0.016209,"end_time":"2022-07-11T03:11:41.531461","exception":false,"start_time":"2022-07-11T03:11:41.515252","status":"completed"},"tags":[]}},{"cell_type":"markdown","source":"# 7. How Can We Improve? 📈\n\nNow that we have submitted out data, what could we have done to better tune the hyperparameters? Leave comments on this submission with your ideas!","metadata":{"papermill":{"duration":0.016274,"end_time":"2022-07-11T03:11:41.566445","exception":false,"start_time":"2022-07-11T03:11:41.550171","status":"completed"},"tags":[]}},{"cell_type":"code","source":"","metadata":{"papermill":{"duration":0.015996,"end_time":"2022-07-11T03:11:41.598853","exception":false,"start_time":"2022-07-11T03:11:41.582857","status":"completed"},"tags":[]},"execution_count":null,"outputs":[]}]}