{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"pygments_lexer":"ipython3","nbconvert_exporter":"python","version":"3.6.4","file_extension":".py","codemirror_mode":{"name":"ipython","version":3},"name":"python","mimetype":"text/x-python"}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"# Spaceship Titanic - SVM|DecisionTree|RandomForest\nThe aim of this notebook is to compare SVM, Decision Trees and Random Forest Classifier algorithms without much hyperparameter tuning.","metadata":{}},{"cell_type":"markdown","source":"**Problem Statement :**\n<br><br>\nWelcome to the year 2912, where your data science skills are needed to solve a cosmic mystery. We've received a transmission from four lightyears away and things aren't looking good.\n\nThe Spaceship Titanic was an interstellar passenger liner launched a month ago. With almost 13,000 passengers on board, the vessel set out on its maiden voyage transporting emigrants from our solar system to three newly habitable exoplanets orbiting nearby stars.\n\nWhile rounding Alpha Centauri en route to its first destination—the torrid 55 Cancri E—the unwary Spaceship Titanic collided with a spacetime anomaly hidden within a dust cloud. Sadly, it met a similar fate as its namesake from 1000 years before. Though the ship stayed intact, almost half of the passengers were transported to an alternate dimension!\n\nTo help rescue crews and retrieve the lost passengers, you are challenged to predict which passengers were transported by the anomaly using records recovered from the spaceship’s damaged computer system.\n\nHelp save them and change history!","metadata":{}},{"cell_type":"markdown","source":"**Importing necessary libraries.**","metadata":{}},{"cell_type":"code","source":"import pandas as pd\nimport numpy as np\nimport matplotlib.pyplot as plt\n%matplotlib inline\nimport seaborn as sns\n \nfrom sklearn.model_selection import train_test_split\nfrom sklearn.metrics import accuracy_score","metadata":{"execution":{"iopub.status.busy":"2022-08-13T14:17:04.678188Z","iopub.execute_input":"2022-08-13T14:17:04.678637Z","iopub.status.idle":"2022-08-13T14:17:04.687008Z","shell.execute_reply.started":"2022-08-13T14:17:04.678601Z","shell.execute_reply":"2022-08-13T14:17:04.685941Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"**Loading the data**","metadata":{}},{"cell_type":"code","source":"df_train = pd.read_csv('../input/spaceship-titanic/train.csv')\ndf_test = pd.read_csv('../input/spaceship-titanic/test.csv')","metadata":{"execution":{"iopub.status.busy":"2022-08-13T14:17:04.693357Z","iopub.execute_input":"2022-08-13T14:17:04.693941Z","iopub.status.idle":"2022-08-13T14:17:04.747213Z","shell.execute_reply.started":"2022-08-13T14:17:04.693869Z","shell.execute_reply":"2022-08-13T14:17:04.746069Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"**Exploring how our data is distributed**","metadata":{}},{"cell_type":"code","source":"df_train.head()","metadata":{"execution":{"iopub.status.busy":"2022-08-13T14:17:04.748971Z","iopub.execute_input":"2022-08-13T14:17:04.749296Z","iopub.status.idle":"2022-08-13T14:17:04.770381Z","shell.execute_reply.started":"2022-08-13T14:17:04.749266Z","shell.execute_reply":"2022-08-13T14:17:04.769611Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"df_test.head()","metadata":{"execution":{"iopub.status.busy":"2022-08-13T14:17:04.771666Z","iopub.execute_input":"2022-08-13T14:17:04.772161Z","iopub.status.idle":"2022-08-13T14:17:04.796459Z","shell.execute_reply.started":"2022-08-13T14:17:04.772130Z","shell.execute_reply":"2022-08-13T14:17:04.795677Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"**Columns Description :**\n<ul>\n<li>PassengerId - A unique Id for each passenger. Each Id takes the form gggg_pp where gggg indicates a group the passenger is travelling with and pp is their number within the group. People in a group are often family members, but not always.</li>\n<li>HomePlanet - The planet the passenger departed from, typically their planet of permanent residence.</li>\n<li>CryoSleep - Indicates whether the passenger elected to be put into suspended animation for the duration of the voyage. Passengers in cryosleep are confined to their cabins.</li>\n<li>Cabin - The cabin number where the passenger is staying. Takes the form deck/num/side, where side can be either P for Port or S for Starboard.</li>\n<li>Destination - The planet the passenger will be debarking to.</li>\n    <li>Age - The age of the passenger.</li>\n    <li>VIP - Whether the passenger has paid for special VIP service during the voyage.</li>\n<li>RoomService, FoodCourt, ShoppingMall, Spa, VRDeck - Amount the passenger has billed at each of the Spaceship Titanic's many luxury amenities.</li>\n    <li>Name - The first and last names of the passenger.</li>\n<li>Transported - Whether the passenger was transported to another dimension. This is the target, the column you are trying to predict.</li>\n</ul>","metadata":{}},{"cell_type":"code","source":"df_train.shape","metadata":{"execution":{"iopub.status.busy":"2022-08-13T14:17:04.799122Z","iopub.execute_input":"2022-08-13T14:17:04.799447Z","iopub.status.idle":"2022-08-13T14:17:04.809132Z","shell.execute_reply.started":"2022-08-13T14:17:04.799411Z","shell.execute_reply":"2022-08-13T14:17:04.807971Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"df_test.shape","metadata":{"execution":{"iopub.status.busy":"2022-08-13T14:17:04.810517Z","iopub.execute_input":"2022-08-13T14:17:04.811050Z","iopub.status.idle":"2022-08-13T14:17:04.821146Z","shell.execute_reply.started":"2022-08-13T14:17:04.810998Z","shell.execute_reply":"2022-08-13T14:17:04.820089Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<h2>Performing EDA</h2>","metadata":{}},{"cell_type":"code","source":"df_train.info()","metadata":{"execution":{"iopub.status.busy":"2022-08-13T14:17:04.822954Z","iopub.execute_input":"2022-08-13T14:17:04.823269Z","iopub.status.idle":"2022-08-13T14:17:04.842950Z","shell.execute_reply.started":"2022-08-13T14:17:04.823239Z","shell.execute_reply":"2022-08-13T14:17:04.842149Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"df_train.describe()","metadata":{"execution":{"iopub.status.busy":"2022-08-13T14:17:04.844124Z","iopub.execute_input":"2022-08-13T14:17:04.844923Z","iopub.status.idle":"2022-08-13T14:17:04.875318Z","shell.execute_reply.started":"2022-08-13T14:17:04.844861Z","shell.execute_reply":"2022-08-13T14:17:04.874333Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"df_test.describe()","metadata":{"execution":{"iopub.status.busy":"2022-08-13T14:17:04.876540Z","iopub.execute_input":"2022-08-13T14:17:04.876859Z","iopub.status.idle":"2022-08-13T14:17:04.907960Z","shell.execute_reply.started":"2022-08-13T14:17:04.876830Z","shell.execute_reply":"2022-08-13T14:17:04.906676Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"df_train.set_index('PassengerId',inplace=True)\ndf_test.set_index('PassengerId',inplace=True)","metadata":{"execution":{"iopub.status.busy":"2022-08-13T14:17:04.909593Z","iopub.execute_input":"2022-08-13T14:17:04.910074Z","iopub.status.idle":"2022-08-13T14:17:04.918876Z","shell.execute_reply.started":"2022-08-13T14:17:04.910025Z","shell.execute_reply":"2022-08-13T14:17:04.917683Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<h4>Dealing with Missing Values</h4>","metadata":{}},{"cell_type":"code","source":"print(\"Missing Values in Training Dataset\")\nprint(df_train.isnull().sum())\nprint(\" \")\nprint(\"Missing Values in Test Dataset\")\nprint(df_test.isnull().sum())\n","metadata":{"execution":{"iopub.status.busy":"2022-08-13T14:17:04.922921Z","iopub.execute_input":"2022-08-13T14:17:04.923301Z","iopub.status.idle":"2022-08-13T14:17:04.940636Z","shell.execute_reply.started":"2022-08-13T14:17:04.923267Z","shell.execute_reply":"2022-08-13T14:17:04.939504Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"In columns with data type as **float** we just replace the missing values with 0","metadata":{}},{"cell_type":"code","source":"df_train[['RoomService', 'FoodCourt', 'ShoppingMall', 'Spa', 'VRDeck']] = df_train[['RoomService', 'FoodCourt', 'ShoppingMall', 'Spa', 'VRDeck']].fillna(0)\ndf_test[['RoomService', 'FoodCourt', 'ShoppingMall', 'Spa', 'VRDeck']] = df_test[['RoomService', 'FoodCourt', 'ShoppingMall', 'Spa', 'VRDeck']].fillna(0)","metadata":{"execution":{"iopub.status.busy":"2022-08-13T14:17:04.942165Z","iopub.execute_input":"2022-08-13T14:17:04.943262Z","iopub.status.idle":"2022-08-13T14:17:04.955831Z","shell.execute_reply.started":"2022-08-13T14:17:04.943217Z","shell.execute_reply":"2022-08-13T14:17:04.955004Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"df_train['Age'] = df_train['Age'].fillna(df_train['Age'].median())\ndf_test['Age'] = df_test['Age'].fillna(df_test['Age'].median())","metadata":{"execution":{"iopub.status.busy":"2022-08-13T14:17:04.957302Z","iopub.execute_input":"2022-08-13T14:17:04.958476Z","iopub.status.idle":"2022-08-13T14:17:04.966778Z","shell.execute_reply.started":"2022-08-13T14:17:04.958432Z","shell.execute_reply":"2022-08-13T14:17:04.965982Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"In columns with data type as **object** we fill the missing values with the most repeated value.","metadata":{}},{"cell_type":"code","source":"df_train['HomePlanet'].unique()","metadata":{"scrolled":true,"execution":{"iopub.status.busy":"2022-08-13T14:17:04.968486Z","iopub.execute_input":"2022-08-13T14:17:04.969583Z","iopub.status.idle":"2022-08-13T14:17:04.979679Z","shell.execute_reply.started":"2022-08-13T14:17:04.969539Z","shell.execute_reply":"2022-08-13T14:17:04.978835Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"df_train['HomePlanet'].value_counts()","metadata":{"scrolled":true,"execution":{"iopub.status.busy":"2022-08-13T14:17:04.980858Z","iopub.execute_input":"2022-08-13T14:17:04.981235Z","iopub.status.idle":"2022-08-13T14:17:04.995502Z","shell.execute_reply.started":"2022-08-13T14:17:04.981204Z","shell.execute_reply":"2022-08-13T14:17:04.994506Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"df_train['HomePlanet'] = df_train['HomePlanet'].fillna('Earth')\ndf_test['HomePlanet'] = df_test['HomePlanet'].fillna('Earth')","metadata":{"execution":{"iopub.status.busy":"2022-08-13T14:17:04.997110Z","iopub.execute_input":"2022-08-13T14:17:04.997833Z","iopub.status.idle":"2022-08-13T14:17:05.007165Z","shell.execute_reply.started":"2022-08-13T14:17:04.997790Z","shell.execute_reply":"2022-08-13T14:17:05.006108Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"df_train['CryoSleep'].value_counts()","metadata":{"execution":{"iopub.status.busy":"2022-08-13T14:17:05.008649Z","iopub.execute_input":"2022-08-13T14:17:05.009660Z","iopub.status.idle":"2022-08-13T14:17:05.021711Z","shell.execute_reply.started":"2022-08-13T14:17:05.009615Z","shell.execute_reply":"2022-08-13T14:17:05.020983Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"df_train['CryoSleep'] = df_train['CryoSleep'].fillna(False)\ndf_test['CryoSleep'] = df_test['CryoSleep'].fillna(False)","metadata":{"execution":{"iopub.status.busy":"2022-08-13T14:17:05.022703Z","iopub.execute_input":"2022-08-13T14:17:05.023703Z","iopub.status.idle":"2022-08-13T14:17:05.037082Z","shell.execute_reply.started":"2022-08-13T14:17:05.023668Z","shell.execute_reply":"2022-08-13T14:17:05.035837Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"df_train['Cabin'].value_counts()","metadata":{"execution":{"iopub.status.busy":"2022-08-13T14:17:05.038360Z","iopub.execute_input":"2022-08-13T14:17:05.039214Z","iopub.status.idle":"2022-08-13T14:17:05.051604Z","shell.execute_reply.started":"2022-08-13T14:17:05.039181Z","shell.execute_reply":"2022-08-13T14:17:05.050411Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"df_train['Cabin'] = df_train['Cabin'].fillna('T/0/P')\ndf_test['Cabin'] = df_test['Cabin'].fillna('T/0/P')","metadata":{"execution":{"iopub.status.busy":"2022-08-13T14:17:05.052822Z","iopub.execute_input":"2022-08-13T14:17:05.053557Z","iopub.status.idle":"2022-08-13T14:17:05.062430Z","shell.execute_reply.started":"2022-08-13T14:17:05.053524Z","shell.execute_reply":"2022-08-13T14:17:05.061541Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"df_train['Destination'].value_counts()","metadata":{"execution":{"iopub.status.busy":"2022-08-13T14:17:05.063666Z","iopub.execute_input":"2022-08-13T14:17:05.064240Z","iopub.status.idle":"2022-08-13T14:17:05.076022Z","shell.execute_reply.started":"2022-08-13T14:17:05.064209Z","shell.execute_reply":"2022-08-13T14:17:05.075180Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"df_train['Destination'] = df_train['Destination'].fillna('TRAPPIST-1e')\ndf_test['Destination'] = df_test['Destination'].fillna('TRAPPIST-1e')","metadata":{"execution":{"iopub.status.busy":"2022-08-13T14:17:05.077551Z","iopub.execute_input":"2022-08-13T14:17:05.077857Z","iopub.status.idle":"2022-08-13T14:17:05.087101Z","shell.execute_reply.started":"2022-08-13T14:17:05.077829Z","shell.execute_reply":"2022-08-13T14:17:05.085928Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"df_train['VIP'].value_counts()","metadata":{"execution":{"iopub.status.busy":"2022-08-13T14:17:05.088431Z","iopub.execute_input":"2022-08-13T14:17:05.089113Z","iopub.status.idle":"2022-08-13T14:17:05.107108Z","shell.execute_reply.started":"2022-08-13T14:17:05.089081Z","shell.execute_reply":"2022-08-13T14:17:05.105970Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"df_train['VIP'] = df_train['VIP'].fillna(False)\ndf_test['VIP'] = df_test['VIP'].fillna(False)","metadata":{"execution":{"iopub.status.busy":"2022-08-13T14:17:05.114959Z","iopub.execute_input":"2022-08-13T14:17:05.115911Z","iopub.status.idle":"2022-08-13T14:17:05.125298Z","shell.execute_reply.started":"2022-08-13T14:17:05.115847Z","shell.execute_reply":"2022-08-13T14:17:05.124303Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"df_train.head()","metadata":{"execution":{"iopub.status.busy":"2022-08-13T14:17:05.126874Z","iopub.execute_input":"2022-08-13T14:17:05.127559Z","iopub.status.idle":"2022-08-13T14:17:05.153535Z","shell.execute_reply.started":"2022-08-13T14:17:05.127517Z","shell.execute_reply":"2022-08-13T14:17:05.152318Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<h3>Visualizing our data<h3>","metadata":{}},{"cell_type":"code","source":"plt.figure(figsize=(10,8))\nsns.heatmap(df_train.corr(), annot=True);","metadata":{"execution":{"iopub.status.busy":"2022-08-13T14:17:05.154874Z","iopub.execute_input":"2022-08-13T14:17:05.155281Z","iopub.status.idle":"2022-08-13T14:17:05.788878Z","shell.execute_reply.started":"2022-08-13T14:17:05.155250Z","shell.execute_reply":"2022-08-13T14:17:05.788086Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"sns.countplot(df_train.Transported)","metadata":{"execution":{"iopub.status.busy":"2022-08-13T14:17:05.790264Z","iopub.execute_input":"2022-08-13T14:17:05.790806Z","iopub.status.idle":"2022-08-13T14:17:05.956909Z","shell.execute_reply.started":"2022-08-13T14:17:05.790773Z","shell.execute_reply":"2022-08-13T14:17:05.955865Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"sns.countplot(df_train.HomePlanet, hue=df_train.Transported)","metadata":{"execution":{"iopub.status.busy":"2022-08-13T14:17:05.958278Z","iopub.execute_input":"2022-08-13T14:17:05.958586Z","iopub.status.idle":"2022-08-13T14:17:06.179257Z","shell.execute_reply.started":"2022-08-13T14:17:05.958558Z","shell.execute_reply":"2022-08-13T14:17:06.178154Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"sns.countplot(df_train.Destination, hue=df_train.Transported)","metadata":{"execution":{"iopub.status.busy":"2022-08-13T14:17:06.180830Z","iopub.execute_input":"2022-08-13T14:17:06.181208Z","iopub.status.idle":"2022-08-13T14:17:06.405295Z","shell.execute_reply.started":"2022-08-13T14:17:06.181175Z","shell.execute_reply":"2022-08-13T14:17:06.403954Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"sns.countplot(df_train.CryoSleep, hue=df_train.Transported)","metadata":{"execution":{"iopub.status.busy":"2022-08-13T14:17:06.406997Z","iopub.execute_input":"2022-08-13T14:17:06.407845Z","iopub.status.idle":"2022-08-13T14:17:06.612672Z","shell.execute_reply.started":"2022-08-13T14:17:06.407801Z","shell.execute_reply":"2022-08-13T14:17:06.611555Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"sns.boxplot(y = df_train.Age, x = df_train.Transported)","metadata":{"execution":{"iopub.status.busy":"2022-08-13T14:17:06.614411Z","iopub.execute_input":"2022-08-13T14:17:06.615094Z","iopub.status.idle":"2022-08-13T14:17:06.814839Z","shell.execute_reply.started":"2022-08-13T14:17:06.615059Z","shell.execute_reply":"2022-08-13T14:17:06.814116Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"df_train[['Deck', 'Num', 'Side']] = df_train.Cabin.str.split('/', expand=True)\ndf_test[['Deck', 'Num', 'Side']] = df_test.Cabin.str.split('/', expand=True)","metadata":{"execution":{"iopub.status.busy":"2022-08-13T14:17:06.816275Z","iopub.execute_input":"2022-08-13T14:17:06.816688Z","iopub.status.idle":"2022-08-13T14:17:06.847859Z","shell.execute_reply.started":"2022-08-13T14:17:06.816646Z","shell.execute_reply":"2022-08-13T14:17:06.846938Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"sns.countplot(df_train.Deck, hue=df_train.Transported)","metadata":{"execution":{"iopub.status.busy":"2022-08-13T14:17:06.849186Z","iopub.execute_input":"2022-08-13T14:17:06.849570Z","iopub.status.idle":"2022-08-13T14:17:07.126417Z","shell.execute_reply.started":"2022-08-13T14:17:06.849542Z","shell.execute_reply":"2022-08-13T14:17:07.125261Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"sns.countplot(df_train.Side, hue=df_train.Transported)","metadata":{"execution":{"iopub.status.busy":"2022-08-13T14:17:07.127840Z","iopub.execute_input":"2022-08-13T14:17:07.128210Z","iopub.status.idle":"2022-08-13T14:17:07.332020Z","shell.execute_reply.started":"2022-08-13T14:17:07.128178Z","shell.execute_reply":"2022-08-13T14:17:07.331197Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<h3>Performing Feature Engineering</h3>","metadata":{}},{"cell_type":"code","source":"df_train['TotalSpent'] = df_train['RoomService'] + df_train['FoodCourt'] + df_train['ShoppingMall'] + df_train['Spa'] +df_train['VRDeck']\n\ndf_test['TotalSpent'] = df_test['RoomService'] + df_test['FoodCourt'] + df_test['ShoppingMall'] + df_test['Spa'] +df_test['VRDeck']","metadata":{"execution":{"iopub.status.busy":"2022-08-13T14:17:07.333043Z","iopub.execute_input":"2022-08-13T14:17:07.334240Z","iopub.status.idle":"2022-08-13T14:17:07.344685Z","shell.execute_reply.started":"2022-08-13T14:17:07.334197Z","shell.execute_reply":"2022-08-13T14:17:07.343644Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"df_train['AgeGroup'] = 0\nfor i in range(6):\n    df_train.loc[(df_train.Age >= 10*i) & (df_train.Age < 10*(i+1)),'AgeGroup'] = i\n    \ndf_test['AgeGroup'] = 0\nfor i in range(6):\n    df_test.loc[(df_test.Age >= 10*i) & (df_test.Age < 10*(i+1)),'AgeGroup'] = i","metadata":{"execution":{"iopub.status.busy":"2022-08-13T14:17:07.347976Z","iopub.execute_input":"2022-08-13T14:17:07.348679Z","iopub.status.idle":"2022-08-13T14:17:07.367990Z","shell.execute_reply.started":"2022-08-13T14:17:07.348638Z","shell.execute_reply":"2022-08-13T14:17:07.367165Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"df_train.head()","metadata":{"execution":{"iopub.status.busy":"2022-08-13T14:17:07.369635Z","iopub.execute_input":"2022-08-13T14:17:07.370085Z","iopub.status.idle":"2022-08-13T14:17:07.394375Z","shell.execute_reply.started":"2022-08-13T14:17:07.370043Z","shell.execute_reply":"2022-08-13T14:17:07.393310Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"sns.countplot(df_train.AgeGroup, hue=df_train.Transported)","metadata":{"execution":{"iopub.status.busy":"2022-08-13T14:17:07.395851Z","iopub.execute_input":"2022-08-13T14:17:07.396941Z","iopub.status.idle":"2022-08-13T14:17:07.646376Z","shell.execute_reply.started":"2022-08-13T14:17:07.396879Z","shell.execute_reply":"2022-08-13T14:17:07.645628Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<h2>Preprocessing and Modelling</h2>","metadata":{}},{"cell_type":"markdown","source":"Using Label Encoding and One Hot Encoding to convert categorical values into numerical values.","metadata":{}},{"cell_type":"code","source":"from sklearn.preprocessing import LabelEncoder\n\ncategorical_cols = ['Deck', 'Side', 'Num']\n\nfor i in categorical_cols:\n    print(i)\n    le = LabelEncoder()\n    arr = np.concatenate((df_train[i], df_test[i])).astype(str)\n    le.fit(arr)\n    df_train[i] = le.transform(df_train[i].astype(str))\n    df_test[i] = le.transform(df_test[i].astype(str))","metadata":{"execution":{"iopub.status.busy":"2022-08-13T14:17:07.647595Z","iopub.execute_input":"2022-08-13T14:17:07.648240Z","iopub.status.idle":"2022-08-13T14:17:07.685776Z","shell.execute_reply.started":"2022-08-13T14:17:07.648208Z","shell.execute_reply":"2022-08-13T14:17:07.684654Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"df_train = pd.get_dummies(df_train, columns =  ['HomePlanet', 'CryoSleep', 'Destination', 'VIP'])\n\ndf_test = pd.get_dummies(df_test, columns =  ['HomePlanet', 'CryoSleep', 'Destination', 'VIP'])","metadata":{"execution":{"iopub.status.busy":"2022-08-13T14:17:07.687483Z","iopub.execute_input":"2022-08-13T14:17:07.687835Z","iopub.status.idle":"2022-08-13T14:17:07.716691Z","shell.execute_reply.started":"2022-08-13T14:17:07.687803Z","shell.execute_reply":"2022-08-13T14:17:07.715526Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"df_train.head()","metadata":{"scrolled":true,"execution":{"iopub.status.busy":"2022-08-13T14:17:07.717999Z","iopub.execute_input":"2022-08-13T14:17:07.718312Z","iopub.status.idle":"2022-08-13T14:17:07.744212Z","shell.execute_reply.started":"2022-08-13T14:17:07.718283Z","shell.execute_reply":"2022-08-13T14:17:07.743136Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Dropping the **Name** and **Cabin** columns from our dataset.","metadata":{}},{"cell_type":"code","source":"df_train= df_train.drop(['Name','Cabin'],axis=1)\ndf_test= df_test.drop(['Name','Cabin'],axis=1)","metadata":{"execution":{"iopub.status.busy":"2022-08-13T14:17:07.768859Z","iopub.execute_input":"2022-08-13T14:17:07.769216Z","iopub.status.idle":"2022-08-13T14:17:07.778843Z","shell.execute_reply.started":"2022-08-13T14:17:07.769183Z","shell.execute_reply":"2022-08-13T14:17:07.777816Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"df_train['Transported']=df_train['Transported'].replace({True:1,False:0})\n","metadata":{"execution":{"iopub.status.busy":"2022-08-13T14:17:07.800796Z","iopub.execute_input":"2022-08-13T14:17:07.801151Z","iopub.status.idle":"2022-08-13T14:17:07.809067Z","shell.execute_reply.started":"2022-08-13T14:17:07.801119Z","shell.execute_reply":"2022-08-13T14:17:07.808200Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"plt.figure(figsize=(20,15))\nsns.heatmap(df_train.corr(), annot=True)","metadata":{"execution":{"iopub.status.busy":"2022-08-13T14:17:07.810199Z","iopub.execute_input":"2022-08-13T14:17:07.810490Z","iopub.status.idle":"2022-08-13T14:17:10.300798Z","shell.execute_reply.started":"2022-08-13T14:17:07.810462Z","shell.execute_reply":"2022-08-13T14:17:10.299650Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<h4>Splitting our dataset into Test and Train datasets<h4>","metadata":{}},{"cell_type":"code","source":"X = df_train.drop(['Transported'], axis=1)\ny = df_train['Transported']","metadata":{"execution":{"iopub.status.busy":"2022-08-13T14:17:10.302222Z","iopub.execute_input":"2022-08-13T14:17:10.302565Z","iopub.status.idle":"2022-08-13T14:17:10.310360Z","shell.execute_reply.started":"2022-08-13T14:17:10.302534Z","shell.execute_reply":"2022-08-13T14:17:10.309261Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"X.columns","metadata":{"execution":{"iopub.status.busy":"2022-08-13T14:17:10.311629Z","iopub.execute_input":"2022-08-13T14:17:10.311990Z","iopub.status.idle":"2022-08-13T14:17:10.322574Z","shell.execute_reply.started":"2022-08-13T14:17:10.311952Z","shell.execute_reply":"2022-08-13T14:17:10.321790Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"X_train, X_test, y_train, y_test = train_test_split(X, y, test_size=0.2, random_state=2)","metadata":{"execution":{"iopub.status.busy":"2022-08-13T14:17:10.323968Z","iopub.execute_input":"2022-08-13T14:17:10.324630Z","iopub.status.idle":"2022-08-13T14:17:10.337107Z","shell.execute_reply.started":"2022-08-13T14:17:10.324599Z","shell.execute_reply":"2022-08-13T14:17:10.336166Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"X_train.shape","metadata":{"execution":{"iopub.status.busy":"2022-08-13T14:17:10.338920Z","iopub.execute_input":"2022-08-13T14:17:10.339666Z","iopub.status.idle":"2022-08-13T14:17:10.347249Z","shell.execute_reply.started":"2022-08-13T14:17:10.339623Z","shell.execute_reply":"2022-08-13T14:17:10.346106Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"X_test.shape","metadata":{"execution":{"iopub.status.busy":"2022-08-13T14:17:10.349367Z","iopub.execute_input":"2022-08-13T14:17:10.350170Z","iopub.status.idle":"2022-08-13T14:17:10.358094Z","shell.execute_reply.started":"2022-08-13T14:17:10.350128Z","shell.execute_reply":"2022-08-13T14:17:10.357366Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<h3>Using SVM</h3>","metadata":{}},{"cell_type":"code","source":"from sklearn import svm\nsvm_model = svm.SVC(kernel='linear', C=0.01)\nsvm_model.fit(X_train,y_train)","metadata":{"execution":{"iopub.status.busy":"2022-08-13T14:17:10.359275Z","iopub.execute_input":"2022-08-13T14:17:10.359934Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"pred_y = svm_model.predict(X_test)\npred_X = svm_model.predict(X_train)\n\nprint(\"Accuracy Score for Training Dataset using SVM: \", accuracy_score(y_train.values, pred_X))\nprint(\"Accuracy Score for Test Dataset using SVM: \", accuracy_score(y_test.values, pred_y))","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"The **SVM** algorithm produces a good accuracy score of however it takes a lot of time in running, hence is not suitable.","metadata":{}},{"cell_type":"markdown","source":"<h3>Using Decision Tree Classifier</h3>","metadata":{}},{"cell_type":"code","source":"from sklearn.tree import DecisionTreeClassifier\ndt_model = DecisionTreeClassifier(max_depth=5)\ndt_model.fit(X_train, y_train)","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"pred_y = dt_model.predict(X_test)\npred_X = dt_model.predict(X_train)\n\nprint(\"Accuracy Score for Training Dataset using Decision Tree Classifier: \", accuracy_score(y_train.values, pred_X))\nprint(\"Accuracy Score for Test Dataset using Decision Tree Classifier: \", accuracy_score(y_test.values, pred_y))","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"The **Decision Tree** algorithm produces a similar accuracy score as produced by **SVM** algorithm, however it takes much lesser time to produce this result. ","metadata":{}},{"cell_type":"markdown","source":"<h3>Using Random Forest Classifier</h3>","metadata":{}},{"cell_type":"code","source":"from sklearn.ensemble import RandomForestClassifier\n\nrf_model = RandomForestClassifier(n_estimators=12, max_depth = 10)\nrf_model.fit(X_train, y_train)","metadata":{"execution":{"iopub.status.busy":"2022-08-13T14:35:55.864727Z","iopub.execute_input":"2022-08-13T14:35:55.865169Z","iopub.status.idle":"2022-08-13T14:35:56.777284Z","shell.execute_reply.started":"2022-08-13T14:35:55.865135Z","shell.execute_reply":"2022-08-13T14:35:56.775372Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"pred_y = rf_model.predict(X_test)\npred_X = rf_model.predict(X_train)\n\nprint(\"Accuracy Score for Training Dataset using Random Forrest Classifier: \", accuracy_score(y_train.values, pred_X))\nprint(\"Accuracy Score for Test Dataset using Random Forrest Classifier: \", accuracy_score(y_test.values, pred_y))","metadata":{"execution":{"iopub.status.busy":"2022-08-13T14:35:51.607916Z","iopub.execute_input":"2022-08-13T14:35:51.609312Z","iopub.status.idle":"2022-08-13T14:35:51.720218Z","shell.execute_reply.started":"2022-08-13T14:35:51.609148Z","shell.execute_reply":"2022-08-13T14:35:51.718428Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"The **Random Forest** algorithm produces the highest accuracy score among the three algorithms we've used, hence it is the most suitable algorithm.","metadata":{}},{"cell_type":"markdown","source":"<h2>Submitting our Results !</h2>","metadata":{}},{"cell_type":"code","source":"y_pred = rf_model.predict(df_test)\n\nsub = pd.DataFrame({'Transported': y_pred.astype(bool)}, index = df_test.index)","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"sub.head()","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"sub.to_csv('submission')\n","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<h3>Thanks for reading !!</h3>\n    <h3>Upvote and leave some suggestions.<h3>","metadata":{}},{"cell_type":"code","source":"","metadata":{},"execution_count":null,"outputs":[]}]}