{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"pygments_lexer":"ipython3","nbconvert_exporter":"python","version":"3.6.4","file_extension":".py","codemirror_mode":{"name":"ipython","version":3},"name":"python","mimetype":"text/x-python"}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"<div style=\"border-radius:20px;\n            border : black solid;\n            background-color: ##FFFFFF;\n            font-size:200%;\n            text-align: left\">\n\n<h1 style='; border:0; border-radius: 15px; text-shadow: 1px 1px black; font-weight: bold; color:green'><center>  Spaceship Titanic Dataset </center></h1>","metadata":{}},{"cell_type":"markdown","source":"![](https://miro.medium.com/max/1080/1*OmoqdehwIUtvm2MakA5Xmw.jpeg)","metadata":{}},{"cell_type":"markdown","source":"- - -\n\n<div style='font-size:200%;'>\n    <a id='nan'></a>\n    <h1 style='color: chartreuse; font-weight: bold; font-family: Cascadia code;'> Contents </h1>\n</div>\n\n- [Importing necessary libraries](#import)\n- [Loading Train and Test Datasets](#data)\n- [Exploratory Data Analysis](#eda)\n    - [NaN values heat-map](#heatmap)\n    - [Typecasting and dropping columns](#typecast)\n    - [Correlation between different features and our target variable](#corr)\n    - [Distribution of Individuals based on HomePlanet](#planet)\n- [Distribution of transported individuals](#trans)\n- [Data Pre-processing](#preprocess)\n    - [Imputing missing data](#impute)\n    - [Typecasting and dropping columns](#typecast)\n    - [One-Hot Encoding](#ohe)\n    - [Splitting data into x (Values) and y (labels)](#split)\n- [Classifying](#classify)\n    - [Building and fitting the models](#build)\n    - [Performance Analysis of the different models](#anal)\n- [Submission](#submit)\n\n- - -","metadata":{}},{"cell_type":"markdown","source":"<div style=\"border-radius:10px;\n            border : black solid;\n            background-color:  #FFA07A;\n            font-size:110%;\n            text-align: left\">\n\n<h2 style='; border:0; border-radius: 15px; text-shadow: 1px 1px black; font-weight: bold; color:black'><center> PROBLEM STATMENT </center></h2>\n","metadata":{}},{"cell_type":"markdown","source":"**Data Description**\n\nYou must decide whether a contestant was transported to a different dimension when the Spaceship Titanic collided with the anomaly in spacetime. You are handed a set of personal documents that were salvaged from the ship's damaged computer system to aid in making these predictions.\n\n**File and Data Field Descriptions**\n\ntrain.csv : Approximately 8700 passengers' personal records from the train.csv file will be used as training data.\n\nPassengerId : Each passenger's individual ID. Each Id has the format gggg pp, where gggg stands for the group the passenger is travelling in and pp stands for their position inside the group. Family members are frequently, but not always, present in a gathering.\n\nHomePlanet : Usually the planet of the passenger's permanent domicile, from which they had just departed.\n\nCryoSleep : If the passenger chose to be put in suspended animation for the duration of the trip, it will be indicated. When in cryosleep, travellers are confined to their cabins.\n\nCabin : The passenger's accommodation's cabin number. consists of the elements deck/num/side, where side can either be P for Port or S for Starboard.\n\nDestination : The planet the passenger will be debarking to.\nAge : The age of the passenger.\n\nVIP : Whether the passenger has paid for special VIP service during the voyage.\n\nRoomService, FoodCourt, ShoppingMall, Spa, VRDeck : Amount the passenger has billed at each of the Spaceship Titanic's many luxury amenities.\n\nName : The first and last names of the passenger.\n\nTransported : Whether the passenger was transported to another dimension. This is the target, the column you are trying to predict.\n\ntest.csv : Records for the remaining 4300 passengers, or the remaining third of the passengers, to be utilised as test data. It is your responsibility to determine the value of Transported for each passenger in this set.\n\nsample_submission.csv : A submission file in the correct format.\n\nPassengerId : Id for each passenger in the test set.\n\nTransported : The target. For each passenger, predict either True or False.","metadata":{}},{"cell_type":"markdown","source":"<div style=\"border-radius:10px;\n            border : black solid;\n            background-color:  #FFA07A;\n            font-size:110%;\n            text-align: left\">\n    <h2 style='; border:0; border-radius: 15px; text-shadow: 1px 1px black; font-weight: bold; color:black'><center> Importing relevant Libraries </center></h2><a id=\"import\"></a>\n","metadata":{}},{"cell_type":"code","source":"import numpy as np\nimport pandas as pd\nimport sklearn\nimport matplotlib.pyplot as plt\nimport seaborn as sns","metadata":{"execution":{"iopub.status.busy":"2022-07-15T06:29:06.863986Z","iopub.execute_input":"2022-07-15T06:29:06.865044Z","iopub.status.idle":"2022-07-15T06:29:06.870613Z","shell.execute_reply.started":"2022-07-15T06:29:06.864988Z","shell.execute_reply":"2022-07-15T06:29:06.869528Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"import os\nfor dirname, _, filenames in os.walk('/kaggle/input'):\n    for filename in filenames:\n        print(os.path.join(dirname, filename))","metadata":{"execution":{"iopub.status.busy":"2022-07-15T06:29:06.953396Z","iopub.execute_input":"2022-07-15T06:29:06.954149Z","iopub.status.idle":"2022-07-15T06:29:06.963456Z","shell.execute_reply.started":"2022-07-15T06:29:06.954105Z","shell.execute_reply":"2022-07-15T06:29:06.962243Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<div style=\"border-radius:10px;\n            border : black solid;\n            background-color:  #FFA07A;\n            font-size:110%;\n            text-align: left\">\n    <h2 style='; border:0; border-radius: 15px; text-shadow: 1px 1px black; font-weight: bold; color:black'><center> Loading Train and Test Datasets </center></h2><a id=\"data\"></a>\n","metadata":{}},{"cell_type":"code","source":"Train = pd.read_csv('/kaggle/input/spaceship-titanic/train.csv')\nTest = pd.read_csv('/kaggle/input/spaceship-titanic/test.csv')","metadata":{"execution":{"iopub.status.busy":"2022-07-15T06:29:06.965097Z","iopub.execute_input":"2022-07-15T06:29:06.965782Z","iopub.status.idle":"2022-07-15T06:29:07.013048Z","shell.execute_reply.started":"2022-07-15T06:29:06.965747Z","shell.execute_reply":"2022-07-15T06:29:07.011777Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<div style=\"border-radius:10px;\n            border : black solid;\n            background-color:  #FFA07A;\n            font-size:110%;\n            text-align: left\">\n    <h2 style='; border:0; border-radius: 15px; text-shadow: 1px 1px black; font-weight: bold; color:black'><center> Explore Data Analysis(EDA) </center></h2><a id=\"eda\"></a>","metadata":{}},{"cell_type":"markdown","source":"![](https://cdn-images-1.medium.com/max/1000/1*Owa2rsDG6Rwv1IM_RdsL3A.gif)","metadata":{}},{"cell_type":"markdown","source":"<h1 align=\"center\" ><a id='heatmap'><b>Null values heat-map<b></a></h1>","metadata":{}},{"cell_type":"code","source":"fig, axes = plt.subplots(1, 2, sharex=True, figsize=(20,10))\nsns.heatmap(ax=axes[0], yticklabels=False, data=Train.isnull(), cbar=False, cmap=\"viridis\")\nsns.heatmap(ax=axes[1], yticklabels=False, data=Test.isnull(), cbar=False, cmap=\"tab20c\")\naxes[0].set_title('Heatmap of missing values in training data')\naxes[1].set_title('Heatmap of missing values in testing data')\nplt.show()","metadata":{"execution":{"iopub.status.busy":"2022-07-15T06:29:07.014941Z","iopub.execute_input":"2022-07-15T06:29:07.015231Z","iopub.status.idle":"2022-07-15T06:29:07.613330Z","shell.execute_reply.started":"2022-07-15T06:29:07.015204Z","shell.execute_reply":"2022-07-15T06:29:07.611990Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"print('Unique HomePlanet:', Train.HomePlanet.unique(), '\\nUnique Destination:', Train.Destination.unique())","metadata":{"execution":{"iopub.status.busy":"2022-07-15T06:29:07.614861Z","iopub.execute_input":"2022-07-15T06:29:07.615293Z","iopub.status.idle":"2022-07-15T06:29:07.623267Z","shell.execute_reply.started":"2022-07-15T06:29:07.615250Z","shell.execute_reply":"2022-07-15T06:29:07.622170Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<h1 align=\"center\" ><a id='corr'><b>Correlation between different features and our target variable<b></a></h1>","metadata":{}},{"cell_type":"markdown","source":"## **Correlation coefficient**\n\nThe correlation coefficient is a statistical measure of the strength of the relationship between the relative movements of two variables. The values range between -1.0 and 1.0. A calculated number greater than 1.0 or less than -1.0 means that there was an error in the correlation measurement. A correlation of -1.0 shows a perfect negative correlation, while a correlation of 1.0 shows a perfect positive correlation. A correlation of 0.0 shows no linear relationship between the movement of the two variables.","metadata":{}},{"cell_type":"code","source":"plt.figure(figsize=(12,8))\ndata = Train.corr()[\"Transported\"].sort_values(ascending=False)\nindices = data.index\nlabels = []\ncorr = []\nfor i in range(1, len(indices)):\n    labels.append(indices[i])\n    corr.append(data[i])\nsns.barplot(x=corr, y=labels, palette='viridis')\nplt.title('Correlation coefficient between different features and Transported')","metadata":{"_kg_hide-input":true,"execution":{"iopub.status.busy":"2022-07-15T06:29:07.626009Z","iopub.execute_input":"2022-07-15T06:29:07.626979Z","iopub.status.idle":"2022-07-15T06:29:08.266073Z","shell.execute_reply.started":"2022-07-15T06:29:07.626933Z","shell.execute_reply":"2022-07-15T06:29:08.265071Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<h1 align=\"center\" ><a id='planet'><b>Distribution of Individuals based on HomePlanet<b></a></h1>","metadata":{}},{"cell_type":"code","source":"tPlanet = pd.crosstab(Train['Transported'], Train['HomePlanet'])\ntDest = pd.crosstab(Train['Transported'], Train['Destination'])","metadata":{"execution":{"iopub.status.busy":"2022-07-15T06:29:08.267865Z","iopub.execute_input":"2022-07-15T06:29:08.268290Z","iopub.status.idle":"2022-07-15T06:29:08.297165Z","shell.execute_reply.started":"2022-07-15T06:29:08.268250Z","shell.execute_reply":"2022-07-15T06:29:08.296156Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"plt.figure(figsize=(12,8))\ncolors = sns.color_palette('pastel')\nplt.pie([item/len(Train.HomePlanet) for item in Train.HomePlanet.value_counts()], labels=['Earth', 'Europa', 'Mars'], colors=colors, autopct='%.0f%%')\nplt.show()","metadata":{"_kg_hide-input":true,"execution":{"iopub.status.busy":"2022-07-15T06:29:08.299268Z","iopub.execute_input":"2022-07-15T06:29:08.300344Z","iopub.status.idle":"2022-07-15T06:29:08.412741Z","shell.execute_reply.started":"2022-07-15T06:29:08.300294Z","shell.execute_reply":"2022-07-15T06:29:08.411414Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<h1 align=\"center\" ><a id='trans'><b>Distribution of transported individuals<b></a></h1>","metadata":{}},{"cell_type":"code","source":"plt.figure(figsize=(20,15))\nplt.subplot(2,2,1)\nsns.countplot(x = Train.HomePlanet, hue = Train.Transported, palette=\"viridis\")\nplt.title('Transported individuals - Home Planets', fontsize=15)\nplt.xlabel('HomePlanet', fontsize=15)\nplt.ylabel('Number of Individuals', fontsize=15)\n\nplt.subplot(2,2,2)\nsns.countplot(x = Train.HomePlanet, hue = Train.CryoSleep, palette=\"viridis\")\nplt.title('Transported individuals - Cryosleep', fontsize=14)\nplt.xlabel('HomePlanet', fontsize=15)\nplt.ylabel('Number of passengers', fontsize=15)","metadata":{"execution":{"iopub.status.busy":"2022-07-15T06:29:08.415159Z","iopub.execute_input":"2022-07-15T06:29:08.416155Z","iopub.status.idle":"2022-07-15T06:29:08.794711Z","shell.execute_reply.started":"2022-07-15T06:29:08.416103Z","shell.execute_reply":"2022-07-15T06:29:08.793900Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<h1 align=\"center\" ><a id='age'><b>Age distribution of the passengers<b></a></h1>","metadata":{}},{"cell_type":"code","source":"plt.figure(figsize=(20,8))\nsns.histplot(Train.Age, color=sns.color_palette('magma')[2])\nplt.show()","metadata":{"execution":{"iopub.status.busy":"2022-07-15T06:29:08.796059Z","iopub.execute_input":"2022-07-15T06:29:08.796585Z","iopub.status.idle":"2022-07-15T06:29:09.096313Z","shell.execute_reply.started":"2022-07-15T06:29:08.796557Z","shell.execute_reply":"2022-07-15T06:29:09.094907Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<div style=\"border-radius:10px;\n            border : black solid;\n            background-color:  #FFA07A;\n            font-size:110%;\n            text-align: left\">\n    <h1 style='; border:0; border-radius: 15px; text-shadow: 1px 1px black; font-weight: bold; color:black'><center>  Data Pre-processing ⌛</center></h1> <a id='preprocess'></a>","metadata":{}},{"cell_type":"markdown","source":"![](https://cdn.dribbble.com/users/2017910/screenshots/5102683/ai_trends_dribbble_shot.gif)","metadata":{}},{"cell_type":"markdown","source":"<h1 align=\"center\" ><a id='impute'><b>Imputing missing data<b></a></h1>","metadata":{}},{"cell_type":"markdown","source":"## **Data imputation**\n\nImputation is a technique used for replacing the missing data with some substitute value to retain most of the data/information of the dataset. These techniques are used because removing the data from the dataset every time is not feasible and can lead to a reduction in the size of the dataset to a large extend, which not only raises concerns for biasing the dataset but also leads to incorrect analysis.","metadata":{}},{"cell_type":"code","source":"Test.head()","metadata":{"execution":{"iopub.status.busy":"2022-07-15T06:29:09.097913Z","iopub.execute_input":"2022-07-15T06:29:09.098286Z","iopub.status.idle":"2022-07-15T06:29:09.116647Z","shell.execute_reply.started":"2022-07-15T06:29:09.098255Z","shell.execute_reply":"2022-07-15T06:29:09.115832Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"idCol=Test.PassengerId.to_numpy()\nTrain.set_index('PassengerId', inplace=True)\nTest.set_index('PassengerId', inplace=True)","metadata":{"execution":{"iopub.status.busy":"2022-07-15T06:29:09.119475Z","iopub.execute_input":"2022-07-15T06:29:09.119943Z","iopub.status.idle":"2022-07-15T06:29:09.129713Z","shell.execute_reply.started":"2022-07-15T06:29:09.119914Z","shell.execute_reply":"2022-07-15T06:29:09.128970Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"from sklearn.impute import SimpleImputer\nimputer = SimpleImputer(missing_values=np.nan, strategy='most_frequent')\nTrain = pd.DataFrame(imputer.fit_transform(Train), columns=Train.columns, index=Train.index)\nTest = pd.DataFrame(imputer.fit_transform(Test), columns=Test.columns, index=Test.index)\nTrain = Train.reset_index(drop=True)\nTest = Test.reset_index(drop=True)","metadata":{"execution":{"iopub.status.busy":"2022-07-15T06:29:09.131366Z","iopub.execute_input":"2022-07-15T06:29:09.131960Z","iopub.status.idle":"2022-07-15T06:29:09.203217Z","shell.execute_reply.started":"2022-07-15T06:29:09.131918Z","shell.execute_reply":"2022-07-15T06:29:09.202393Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<h1 align=\"center\" ><a id='typecast'><b>Typecasting and dropping columns<b></a></h1>","metadata":{}},{"cell_type":"code","source":"Train.Transported = Train.Transported.astype('int')\nTrain.VIP = Train.VIP.astype('int')\nTrain.CryoSleep = Train.CryoSleep.astype('int')\nTrain.drop(columns=['Cabin', 'Name'], inplace=True)\nTest.drop(columns=['Cabin', 'Name'], inplace=True)\nTrain.head()","metadata":{"execution":{"iopub.status.busy":"2022-07-15T06:29:09.204950Z","iopub.execute_input":"2022-07-15T06:29:09.205515Z","iopub.status.idle":"2022-07-15T06:29:09.233827Z","shell.execute_reply.started":"2022-07-15T06:29:09.205474Z","shell.execute_reply":"2022-07-15T06:29:09.232615Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<h1 align=\"center\" ><a id='ohe'><b>One-Hot Encoding 🔥<b></a></h1>","metadata":{}},{"cell_type":"markdown","source":"## **One-Hot Encoding**\n\nOne hot encoding is one method of converting data to prepare it for an algorithm and get a better prediction. With one-hot, we convert each categorical value into a new categorical column and assign a binary value of 1 or 0 to those columns. Each integer value is represented as a binary vector. All the values are zero, and the index is marked with a 1.","metadata":{}},{"cell_type":"code","source":"Train = pd.get_dummies(Train, columns=['HomePlanet', 'CryoSleep', 'Destination'])\nTest = pd.get_dummies(Test, columns=['HomePlanet', 'CryoSleep', 'Destination'])\nTrain.head()","metadata":{"execution":{"iopub.status.busy":"2022-07-15T06:29:09.235475Z","iopub.execute_input":"2022-07-15T06:29:09.235822Z","iopub.status.idle":"2022-07-15T06:29:09.265161Z","shell.execute_reply.started":"2022-07-15T06:29:09.235785Z","shell.execute_reply":"2022-07-15T06:29:09.263752Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<h1 align=\"center\" ><a id='split'><b>Splitting data into x (Values) and y (labels) 🪓<b></a</h1>","metadata":{}},{"cell_type":"code","source":"y_Train = Train.pop('Transported').to_numpy()\nx_Train = Train.to_numpy()\nx_Test = Test.to_numpy()\ny_Test = Test.to_numpy()\nx_Train.shape, y_Train.shape, x_Test.shape,y_Test.shape","metadata":{"execution":{"iopub.status.busy":"2022-07-15T06:29:09.266583Z","iopub.execute_input":"2022-07-15T06:29:09.266929Z","iopub.status.idle":"2022-07-15T06:29:09.284094Z","shell.execute_reply.started":"2022-07-15T06:29:09.266896Z","shell.execute_reply":"2022-07-15T06:29:09.283171Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<div style=\"border-radius:10px;\n            border : black solid;\n            background-color:  #FFA07A;\n            font-size:110%;\n            text-align: left\">\n    <h2 style='; border:0; border-radius: 15px; text-shadow: 1px 1px black; font-weight: bold; color:black'><center> Model Creation and Evalutation </center></h2><a id=\"model\"></a>\n","metadata":{}},{"cell_type":"markdown","source":"![](https://miro.medium.com/max/1280/1*czcdGNhz6jvyxSRvmuxlSQ.gif)","metadata":{}},{"cell_type":"markdown","source":"<h1 align=\"center\" ><a id='build'><b>Building and fitting the models 🏗️<b></a></h1>","metadata":{}},{"cell_type":"markdown","source":"**K-nearest neighbours classifier**<a id=\"knn\"></a>","metadata":{}},{"cell_type":"markdown","source":"\nThe K-nearest neighbours (KNN) classifier uses proximity to make classifications or predictions about independent data points. This technique may be used for both classification and regression scenarios and the output will vary. In classification instances, a decision is made based on majority vote, i.e., the class assigned to the new data point is taken to be the one that is most frequently seen in the vicinity of the point. KNN is also known as a lazy learner technique since a model is not learned. Instead, the raw data is stored and used everytime a prediction must be made.\n![](https://analyticsjobs.in/wp-content/uploads/2020/02/Calculate-the-Euclidean-distance-between-the-data-points.jpg)","metadata":{}},{"cell_type":"code","source":"from sklearn.neighbors import KNeighborsClassifier\n\nknnClassifier = KNeighborsClassifier(3)\nknnClassifier.fit(x_Train, y_Train)\nknnClassifier.score(x_Train, y_Train)","metadata":{"execution":{"iopub.status.busy":"2022-07-15T06:29:09.285147Z","iopub.execute_input":"2022-07-15T06:29:09.286007Z","iopub.status.idle":"2022-07-15T06:29:10.030223Z","shell.execute_reply.started":"2022-07-15T06:29:09.285976Z","shell.execute_reply":"2022-07-15T06:29:10.029005Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"**Support vector machines**<a id=\"svm\"></a>","metadata":{}},{"cell_type":"markdown","source":"\n\nSupport vector machines (SVM) are a class of machine learning algorithms that map data to a high dimensionality feature space in such a manner that the data points can be categorized, even when they are not otherwise linearly seperable. A seperator between the categories is found and a hyperplane is drawn accordingly. Support vectors are those data points that are close to the hyperplane and influence the position and orientation of the hyperplane. The function used to map data to a high dimensionality feature space is called a kernel function.\n![](https://www.analyticssteps.com/backend/media/thumbnail/338466/8680904_1588569086_SVM.jpg)","metadata":{}},{"cell_type":"code","source":"from sklearn.svm import SVC\n\nsvClassifier = SVC()\nsvClassifier.fit(x_Train, y_Train)\nsvClassifier.score(x_Train, y_Train)","metadata":{"execution":{"iopub.status.busy":"2022-07-15T06:29:10.031985Z","iopub.execute_input":"2022-07-15T06:29:10.032429Z","iopub.status.idle":"2022-07-15T06:29:15.933692Z","shell.execute_reply.started":"2022-07-15T06:29:10.032387Z","shell.execute_reply":"2022-07-15T06:29:15.932473Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## **Random forest classifier**\n\nThe random forest classifier is an improvement over decision tree classifiers. Based on ensemble learning, a random forest classifier contains a number of decision trees on various subsets of the given dataset and takes the average to improve the predictive accuracy of that dataset. In general, a greater number of trees in the forest leads to higher accuracy and prevents the problem of overfitting.","metadata":{}},{"cell_type":"markdown","source":"![](https://miro.medium.com/max/1200/1*hmtbIgxoflflJqMJ_UHwXw.jpeg)","metadata":{}},{"cell_type":"code","source":"from sklearn.ensemble import RandomForestClassifier\n\nrfClassifier = RandomForestClassifier()\nrfClassifier.fit(x_Train, y_Train)\nrfClassifier.score(x_Train, y_Train)","metadata":{"execution":{"iopub.status.busy":"2022-07-15T06:29:15.936002Z","iopub.execute_input":"2022-07-15T06:29:15.936341Z","iopub.status.idle":"2022-07-15T06:29:16.939622Z","shell.execute_reply.started":"2022-07-15T06:29:15.936293Z","shell.execute_reply":"2022-07-15T06:29:16.938357Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## **Naive Bayes classifier**\n\nNaive Bayes makes the assumption that the features are independent. This means that we are still assuming class-specific covariance matrices (as in QDA), but the covariance matrices are diagonal matrices. This is due to the assumption that the features are independent.\n\nSo, given a training dataset of N input variables x with corresponding target variables t, (Gaussian) Naive Bayes assumes that the class-conditional densities are normally distributed.","metadata":{}},{"cell_type":"markdown","source":"![](https://cdn-images-1.medium.com/max/600/1*aFhOj7TdBIZir4keHMgHOw.png)","metadata":{}},{"cell_type":"code","source":"from sklearn.naive_bayes import GaussianNB\n\nnbClassifier = GaussianNB()\nnbClassifier.fit(x_Train, y_Train)\nnbClassifier.score(x_Train, y_Train)","metadata":{"execution":{"iopub.status.busy":"2022-07-15T06:29:16.941107Z","iopub.execute_input":"2022-07-15T06:29:16.941880Z","iopub.status.idle":"2022-07-15T06:29:16.967895Z","shell.execute_reply.started":"2022-07-15T06:29:16.941834Z","shell.execute_reply":"2022-07-15T06:29:16.966834Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<h1 align=\"center\" ><a id='anal'><b>Performance Analysis of the different models 📈<b></a></h1>","metadata":{}},{"cell_type":"code","source":"dataPerf = pd.DataFrame(data={'Model': ['SVM', 'RandomForest', 'Naive-Bayes','KNN'], 'Score': [svClassifier.score(x_Train, y_Train), rfClassifier.score(x_Train, y_Train), nbClassifier.score(x_Train, y_Train), knnClassifier.score(x_Train, y_Train)]})\n\nplt.figure(figsize=(12, 8))\nsns.barplot(x=\"Model\", y=\"Score\", data=dataPerf, palette=\"magma\")\nplt.title('Performance analysis of different classifiers')\nplt.show()","metadata":{"execution":{"iopub.status.busy":"2022-07-15T06:29:16.969554Z","iopub.execute_input":"2022-07-15T06:29:16.969844Z","iopub.status.idle":"2022-07-15T06:29:20.613724Z","shell.execute_reply.started":"2022-07-15T06:29:16.969818Z","shell.execute_reply":"2022-07-15T06:29:20.612604Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"----\n#### Based on the above findings, we can conclude that the RandomForestClassifier is best suited for this classification.\n#### Hence we will use this model for our final submission.\n----","metadata":{}},{"cell_type":"markdown","source":"<div style='font-size:200%;'>\n    <a id='submit'></a>\n    <h1 style='color: lightskyblue; font-weight: bold; font-family: Cascadia code;'>\n        <center> Submission ✅ </center>\n    </h1>\n</div>","metadata":{}},{"cell_type":"code","source":"submission = pd.DataFrame(columns=[\"PassengerId\",\"Transported\"])\nsubmission[\"PassengerId\"] = idCol\nsubmission.set_index('PassengerId')\nsubmission[\"Transported\"] = rfClassifier.predict(x_Test).astype(bool)\nsubmission","metadata":{"execution":{"iopub.status.busy":"2022-07-15T06:29:20.615535Z","iopub.execute_input":"2022-07-15T06:29:20.616251Z","iopub.status.idle":"2022-07-15T06:29:20.727092Z","shell.execute_reply.started":"2022-07-15T06:29:20.616215Z","shell.execute_reply":"2022-07-15T06:29:20.726024Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"submission.to_csv('submission.csv', index=False)\nprint('Submission succesful!')","metadata":{"execution":{"iopub.status.busy":"2022-07-15T06:29:20.728531Z","iopub.execute_input":"2022-07-15T06:29:20.729446Z","iopub.status.idle":"2022-07-15T06:29:20.739961Z","shell.execute_reply.started":"2022-07-15T06:29:20.729412Z","shell.execute_reply":"2022-07-15T06:29:20.738886Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<div style=\"border-radius:10px;\n            border : black solid;\n            background-color: #DA70D6;\n            font-size:200%;\n            text-align: left\">\n\n<h1 style='; border:0; border-radius: 10px; text-shadow: 1px 1px black; font-weight: bold; color:black'><center> YOUR FEEDBACKS IS SO VALUABLE FOR ME </center></h1>","metadata":{}},{"cell_type":"markdown","source":"![](https://st2.depositphotos.com/1006899/7664/i/600/depositphotos_76643019-stock-photo-thank-you-words.jpg)","metadata":{}}]}