{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"pygments_lexer":"ipython3","nbconvert_exporter":"python","version":"3.6.4","file_extension":".py","codemirror_mode":{"name":"ipython","version":3},"name":"python","mimetype":"text/x-python"}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"# Data Description\n* Welcome to the year 2912, where your data science skills are needed to solve a cosmic mystery. We've received a transmission from four lightyears away and things aren't looking good.\n\n* The Spaceship Titanic was an interstellar passenger liner launched a month ago. With almost 13,000 passengers on board, the vessel set out on its maiden voyage transporting emigrants from our solar system to three newly habitable exoplanets orbiting nearby stars.\n\n* While rounding Alpha Centauri en route to its first destination—the torrid 55 Cancri E—the unwary Spaceship Titanic collided with a spacetime anomaly hidden within a dust cloud. Sadly, it met a similar fate as its namesake from 1000 years before. Though the ship stayed intact, almost half of the passengers were transported to an alternate dimension!\n\n* To help rescue crews and retrieve the lost passengers, you are challenged to predict which passengers were transported by the anomaly using records recovered from the spaceship’s damaged computer system.","metadata":{"execution":{"iopub.status.busy":"2022-07-05T16:36:55.265924Z","iopub.execute_input":"2022-07-05T16:36:55.266334Z","iopub.status.idle":"2022-07-05T16:36:55.270914Z","shell.execute_reply.started":"2022-07-05T16:36:55.266293Z","shell.execute_reply":"2022-07-05T16:36:55.269873Z"}}},{"cell_type":"markdown","source":"# Field and Data Field Description\n* **train.csv** - Personal records for about two-thirds (~8700) of the passengers, to be used as training data.\n* **PassengerId** - A unique Id for each passenger. Each Id takes the form gggg_pp where gggg indicates a group the passenger is travelling with and pp is their number within the group. People in a group are often family members, but not always.\n* **HomePlanet** - The planet the passenger departed from, typically their planet of permanent residence.\n* **CryoSleep** - Indicates whether the passenger elected to be put into suspended animation for the duration of the voyage. Passengers in cryosleep are confined to their cabins.\n* **Cabin** - The cabin number where the passenger is staying. Takes the form deck/num/side, where side can be either P for Port or S for Starboard.\n* **Destination** - The planet the passenger will be debarking to.\n* **Age** - The age of the passenger.\n* **VIP** - Whether the passenger has paid for special VIP service during the voyage.\n* **RoomService, FoodCourt, ShoppingMall, Spa, VRDeck** - Amount the passenger has billed at each of the Spaceship Titanic's many luxury amenities.\n* **Name** - The first and last names of the passenger.\n* **Transported** - Whether the passenger was transported to another dimension. This is the target, the column you are trying to predict.\n* **test.csv** - Personal records for the remaining one-third (~4300) of the passengers, to be used as test data. Your task is to predict the value of Transported for the passengers in this set.\n* **sample_submission.csv** - A submission file in the correct format.\n* **PassengerId** - Id for each passenger in the test set.\n* **Transported** - The target. For each passenger, predict either True or False.","metadata":{}},{"cell_type":"markdown","source":"# Importing libraries","metadata":{}},{"cell_type":"code","source":"import pandas as pd\nimport numpy as np\nimport matplotlib.pyplot as plt\nimport seaborn as sns\nimport warnings\nwarnings.filterwarnings(\"ignore\")\nfrom sklearn import preprocessing\nfrom sklearn.model_selection import train_test_split\nfrom sklearn.ensemble import RandomForestClassifier\nfrom sklearn.preprocessing import LabelEncoder","metadata":{"execution":{"iopub.status.busy":"2022-07-20T07:22:37.590051Z","iopub.execute_input":"2022-07-20T07:22:37.591076Z","iopub.status.idle":"2022-07-20T07:22:37.598077Z","shell.execute_reply.started":"2022-07-20T07:22:37.591034Z","shell.execute_reply":"2022-07-20T07:22:37.596891Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Loading Train and Test datasets","metadata":{}},{"cell_type":"code","source":"df_train = pd.read_csv('../input/spaceship-titanic/train.csv')\ndf_train.head()","metadata":{"execution":{"iopub.status.busy":"2022-07-20T07:22:37.655355Z","iopub.execute_input":"2022-07-20T07:22:37.656814Z","iopub.status.idle":"2022-07-20T07:22:37.712563Z","shell.execute_reply.started":"2022-07-20T07:22:37.656758Z","shell.execute_reply":"2022-07-20T07:22:37.711262Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"df_test =  pd.read_csv('../input/spaceship-titanic/test.csv')\ndf_test.head()","metadata":{"execution":{"iopub.status.busy":"2022-07-20T07:22:37.761228Z","iopub.execute_input":"2022-07-20T07:22:37.761779Z","iopub.status.idle":"2022-07-20T07:22:37.800828Z","shell.execute_reply.started":"2022-07-20T07:22:37.761740Z","shell.execute_reply":"2022-07-20T07:22:37.799427Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## Exploratory Data Analysis(EDA)\n* In statistics, exploratory data analysis is an approach of analyzing data sets to summarize their main characteristics, often using statistical graphics and other data visualization methods. A statistical model can be used or not, but primarily EDA is for seeing what the data can tell us beyond the formal modeling and thereby contrasts traditional hypothesis testing. \n* Exploratory data analysis has been promoted by John Tukey since 1970 to encourage statisticians to explore the data, and possibly formulate hypotheses that could lead to new data collection and experiments. EDA is different from initial data analysis (IDA),[1][2] which focuses more narrowly on checking assumptions required for model fitting and hypothesis testing, and handling missing values and making transformations of variables as needed. EDA encompasses IDA.","metadata":{}},{"cell_type":"code","source":"df_train.isnull().sum()","metadata":{"execution":{"iopub.status.busy":"2022-07-20T07:22:37.865430Z","iopub.execute_input":"2022-07-20T07:22:37.865839Z","iopub.status.idle":"2022-07-20T07:22:37.885222Z","shell.execute_reply.started":"2022-07-20T07:22:37.865799Z","shell.execute_reply":"2022-07-20T07:22:37.883543Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"df_train.describe()","metadata":{"execution":{"iopub.status.busy":"2022-07-20T07:22:37.915685Z","iopub.execute_input":"2022-07-20T07:22:37.916484Z","iopub.status.idle":"2022-07-20T07:22:37.954383Z","shell.execute_reply.started":"2022-07-20T07:22:37.916432Z","shell.execute_reply":"2022-07-20T07:22:37.952927Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"df_train.info()","metadata":{"execution":{"iopub.status.busy":"2022-07-20T07:22:37.969071Z","iopub.execute_input":"2022-07-20T07:22:37.969456Z","iopub.status.idle":"2022-07-20T07:22:37.993544Z","shell.execute_reply.started":"2022-07-20T07:22:37.969428Z","shell.execute_reply":"2022-07-20T07:22:37.992202Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"df_train.dtypes","metadata":{"execution":{"iopub.status.busy":"2022-07-20T07:22:38.013326Z","iopub.execute_input":"2022-07-20T07:22:38.013822Z","iopub.status.idle":"2022-07-20T07:22:38.023846Z","shell.execute_reply.started":"2022-07-20T07:22:38.013783Z","shell.execute_reply":"2022-07-20T07:22:38.022450Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"df_train['PassengerId'].unique()","metadata":{"execution":{"iopub.status.busy":"2022-07-20T07:22:38.066126Z","iopub.execute_input":"2022-07-20T07:22:38.066542Z","iopub.status.idle":"2022-07-20T07:22:38.077517Z","shell.execute_reply.started":"2022-07-20T07:22:38.066510Z","shell.execute_reply":"2022-07-20T07:22:38.075808Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"a= df_train['PassengerId']\na","metadata":{"execution":{"iopub.status.busy":"2022-07-20T07:22:38.131128Z","iopub.execute_input":"2022-07-20T07:22:38.132378Z","iopub.status.idle":"2022-07-20T07:22:38.142358Z","shell.execute_reply.started":"2022-07-20T07:22:38.132324Z","shell.execute_reply":"2022-07-20T07:22:38.140538Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"df_train.drop('PassengerId',axis=1,inplace=True)","metadata":{"execution":{"iopub.status.busy":"2022-07-20T07:22:38.144692Z","iopub.execute_input":"2022-07-20T07:22:38.145140Z","iopub.status.idle":"2022-07-20T07:22:38.153233Z","shell.execute_reply.started":"2022-07-20T07:22:38.145105Z","shell.execute_reply":"2022-07-20T07:22:38.151623Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"df_train['Transported']=df_train['Transported'].astype('str')\ndf_train['CryoSleep']=df_train['CryoSleep'].astype('str')\ndf_train['VIP']=df_train['VIP'].astype('str')","metadata":{"execution":{"iopub.status.busy":"2022-07-20T07:22:38.154811Z","iopub.execute_input":"2022-07-20T07:22:38.155369Z","iopub.status.idle":"2022-07-20T07:22:38.181660Z","shell.execute_reply.started":"2022-07-20T07:22:38.155336Z","shell.execute_reply":"2022-07-20T07:22:38.180373Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### Data Visualizations\n\n> For visualizing data we can use different types of plots like categorical plots,relational plots and distribution plots\n\n* **Distribution plots**: It shows how data is distributed,by this plot you can define wheather data is in normal distribution or it is skewed data,for ex: displot,distplot\n* **Categorical plots** : it is by name used to find count,kernel density of the categorical data,for ex:boxplot,violinplot,boxenplot\n","metadata":{}},{"cell_type":"code","source":"sns.countplot(df_train.dtypes.map(str))\nplt.show()","metadata":{"execution":{"iopub.status.busy":"2022-07-20T07:22:38.202064Z","iopub.execute_input":"2022-07-20T07:22:38.202494Z","iopub.status.idle":"2022-07-20T07:22:38.413237Z","shell.execute_reply.started":"2022-07-20T07:22:38.202462Z","shell.execute_reply":"2022-07-20T07:22:38.411769Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"df_train.dtypes.value_counts()","metadata":{"execution":{"iopub.status.busy":"2022-07-20T07:22:38.416065Z","iopub.execute_input":"2022-07-20T07:22:38.416436Z","iopub.status.idle":"2022-07-20T07:22:38.427230Z","shell.execute_reply.started":"2022-07-20T07:22:38.416405Z","shell.execute_reply":"2022-07-20T07:22:38.425788Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Data prep\n* DataPrep lets you prepare your data using a single library with a few lines of code.DataPrep is an open-source library available for python that lets you prepare your data using a single library with only a few lines of code. DataPrep can be used to address multiple data-related problems, and the library provides numerous features through which every problem can be solved and taken care of\n\n* Collect data from common data sources (through Connector)\n* Do your exploratory data analysis (through EDA)\n\n* Clean and standardize data (through Clean)\n\n…more modules are coming","metadata":{}},{"cell_type":"code","source":"!pip install dataprep","metadata":{"execution":{"iopub.status.busy":"2022-07-20T07:22:38.429622Z","iopub.execute_input":"2022-07-20T07:22:38.430320Z","iopub.status.idle":"2022-07-20T07:22:50.479576Z","shell.execute_reply.started":"2022-07-20T07:22:38.430266Z","shell.execute_reply":"2022-07-20T07:22:50.478281Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"from dataprep.eda import plot, plot_correlation, create_report, plot_missing","metadata":{"execution":{"iopub.status.busy":"2022-07-20T07:22:50.482427Z","iopub.execute_input":"2022-07-20T07:22:50.482788Z","iopub.status.idle":"2022-07-20T07:22:50.489085Z","shell.execute_reply.started":"2022-07-20T07:22:50.482755Z","shell.execute_reply":"2022-07-20T07:22:50.487345Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"plot(df_train)","metadata":{"execution":{"iopub.status.busy":"2022-07-20T07:22:50.491272Z","iopub.execute_input":"2022-07-20T07:22:50.491758Z","iopub.status.idle":"2022-07-20T07:22:53.278283Z","shell.execute_reply.started":"2022-07-20T07:22:50.491703Z","shell.execute_reply":"2022-07-20T07:22:53.275767Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"count=1\nplt.subplots(figsize=(30,25))\nfor i in df_train.columns:\n    if df_train[i].dtypes!='object':\n        plt.subplot(4,2,count)\n        sns.distplot(df_train[i])\n        count+=1","metadata":{"execution":{"iopub.status.busy":"2022-07-20T07:22:53.279818Z","iopub.execute_input":"2022-07-20T07:22:53.280172Z","iopub.status.idle":"2022-07-20T07:22:55.392661Z","shell.execute_reply.started":"2022-07-20T07:22:53.280143Z","shell.execute_reply":"2022-07-20T07:22:55.391361Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"df_train.dtypes","metadata":{"execution":{"iopub.status.busy":"2022-07-20T07:22:55.395265Z","iopub.execute_input":"2022-07-20T07:22:55.395709Z","iopub.status.idle":"2022-07-20T07:22:55.404400Z","shell.execute_reply.started":"2022-07-20T07:22:55.395673Z","shell.execute_reply":"2022-07-20T07:22:55.403346Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"count=1\nplt.subplots(figsize=(30,25))\nfor i in df_train.columns:\n    if df_train[i].dtypes!='object':\n        plt.subplot(4,2,count)\n        sns.boxplot(df_train[i])\n        count+=1","metadata":{"execution":{"iopub.status.busy":"2022-07-20T07:22:55.407869Z","iopub.execute_input":"2022-07-20T07:22:55.409258Z","iopub.status.idle":"2022-07-20T07:22:56.310820Z","shell.execute_reply.started":"2022-07-20T07:22:55.409218Z","shell.execute_reply":"2022-07-20T07:22:56.309280Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"count=1\nplt.subplots(figsize=(30,25))\nfor i in df_train.columns:\n    if df_train[i].dtypes!='object':\n        plt.subplot(4,2,count)\n        sns.violinplot(df_train[i])\n        count+=1","metadata":{"execution":{"iopub.status.busy":"2022-07-20T07:22:56.313133Z","iopub.execute_input":"2022-07-20T07:22:56.313685Z","iopub.status.idle":"2022-07-20T07:22:57.363684Z","shell.execute_reply.started":"2022-07-20T07:22:56.313631Z","shell.execute_reply":"2022-07-20T07:22:57.362424Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## Filling Null values\n* quantitative data null values filling with median\n* qualitative data null values filling with mode","metadata":{}},{"cell_type":"code","source":"for col in df_train.columns:\n    if df_train[col].dtypes=='object':\n        df_train[col]=df_train[col].fillna(df_train[col].mode()[0])\n    else:\n        df_train[col]=df_train[col].fillna(df_train[col].median())","metadata":{"execution":{"iopub.status.busy":"2022-07-20T07:22:57.365246Z","iopub.execute_input":"2022-07-20T07:22:57.365724Z","iopub.status.idle":"2022-07-20T07:22:57.407251Z","shell.execute_reply.started":"2022-07-20T07:22:57.365563Z","shell.execute_reply":"2022-07-20T07:22:57.406050Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"df_train.head(2)","metadata":{"execution":{"iopub.status.busy":"2022-07-20T07:22:57.409119Z","iopub.execute_input":"2022-07-20T07:22:57.409798Z","iopub.status.idle":"2022-07-20T07:22:57.429175Z","shell.execute_reply.started":"2022-07-20T07:22:57.409760Z","shell.execute_reply":"2022-07-20T07:22:57.427736Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"df_train.head(2)","metadata":{"execution":{"iopub.status.busy":"2022-07-20T07:22:57.430683Z","iopub.execute_input":"2022-07-20T07:22:57.431770Z","iopub.status.idle":"2022-07-20T07:22:57.452571Z","shell.execute_reply.started":"2022-07-20T07:22:57.431732Z","shell.execute_reply":"2022-07-20T07:22:57.451248Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### Converting Categorical and boolean values to numerical using label Encoding","metadata":{}},{"cell_type":"code","source":"for i in df_train.columns:\n    if df_train[i].dtype=='object':\n        label_encoder=preprocessing.LabelEncoder()\n        df_train[i]=label_encoder.fit_transform(df_train[i])","metadata":{"execution":{"iopub.status.busy":"2022-07-20T07:22:57.454809Z","iopub.execute_input":"2022-07-20T07:22:57.455723Z","iopub.status.idle":"2022-07-20T07:22:57.523004Z","shell.execute_reply.started":"2022-07-20T07:22:57.455664Z","shell.execute_reply":"2022-07-20T07:22:57.521611Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"df_train.dtypes","metadata":{"execution":{"iopub.status.busy":"2022-07-20T07:22:57.524516Z","iopub.execute_input":"2022-07-20T07:22:57.524825Z","iopub.status.idle":"2022-07-20T07:22:57.532510Z","shell.execute_reply.started":"2022-07-20T07:22:57.524797Z","shell.execute_reply":"2022-07-20T07:22:57.531627Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Feature Selection","metadata":{}},{"cell_type":"code","source":"x = df_train.drop('Transported',axis=1)\ny = df_train['Transported']","metadata":{"execution":{"iopub.status.busy":"2022-07-20T07:22:57.533538Z","iopub.execute_input":"2022-07-20T07:22:57.534460Z","iopub.status.idle":"2022-07-20T07:22:57.549735Z","shell.execute_reply.started":"2022-07-20T07:22:57.534413Z","shell.execute_reply":"2022-07-20T07:22:57.548125Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## Splitting the data into train and test","metadata":{}},{"cell_type":"code","source":"x_train,x_test,y_train,y_test = train_test_split(x,y,test_size=0.25,random_state=0)","metadata":{"execution":{"iopub.status.busy":"2022-07-20T07:22:57.552185Z","iopub.execute_input":"2022-07-20T07:22:57.553225Z","iopub.status.idle":"2022-07-20T07:22:57.564355Z","shell.execute_reply.started":"2022-07-20T07:22:57.553180Z","shell.execute_reply":"2022-07-20T07:22:57.562749Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# RandomForestClassifier ","metadata":{}},{"cell_type":"code","source":"rf = RandomForestClassifier(n_estimators = 150,random_state=42,max_depth=12,max_features=12,oob_score=True)\nrf.fit(x_train, y_train)","metadata":{"execution":{"iopub.status.busy":"2022-07-20T07:22:57.566364Z","iopub.execute_input":"2022-07-20T07:22:57.567155Z","iopub.status.idle":"2022-07-20T07:23:01.481580Z","shell.execute_reply.started":"2022-07-20T07:22:57.567110Z","shell.execute_reply":"2022-07-20T07:23:01.480144Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"y_pred_rf= rf.predict(x_test)\ny_pred_rf","metadata":{"execution":{"iopub.status.busy":"2022-07-20T07:23:01.483293Z","iopub.execute_input":"2022-07-20T07:23:01.483637Z","iopub.status.idle":"2022-07-20T07:23:01.570605Z","shell.execute_reply.started":"2022-07-20T07:23:01.483600Z","shell.execute_reply":"2022-07-20T07:23:01.569422Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"from sklearn.metrics import confusion_matrix,classification_report,accuracy_score\ncm = confusion_matrix(y_test,y_pred_rf)\nprint(classification_report(y_test,y_pred_rf))\nprint(cm)","metadata":{"execution":{"iopub.status.busy":"2022-07-20T07:23:01.572441Z","iopub.execute_input":"2022-07-20T07:23:01.572833Z","iopub.status.idle":"2022-07-20T07:23:01.592494Z","shell.execute_reply.started":"2022-07-20T07:23:01.572798Z","shell.execute_reply":"2022-07-20T07:23:01.590774Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"rf_acc = round(accuracy_score(y_test,y_pred_rf)*100)# it defines how much correctly the model is predicting the actual value\nrf_acc","metadata":{"execution":{"iopub.status.busy":"2022-07-20T07:23:01.599100Z","iopub.execute_input":"2022-07-20T07:23:01.599725Z","iopub.status.idle":"2022-07-20T07:23:01.610463Z","shell.execute_reply.started":"2022-07-20T07:23:01.599539Z","shell.execute_reply":"2022-07-20T07:23:01.608807Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## K nearest Neighbor Classifier","metadata":{}},{"cell_type":"code","source":"from sklearn import neighbors\nfrom math import sqrt\n\nknn = neighbors.KNeighborsClassifier(n_neighbors=5,p=2)\nknn.fit(x_train, y_train)","metadata":{"execution":{"iopub.status.busy":"2022-07-20T07:23:01.612531Z","iopub.execute_input":"2022-07-20T07:23:01.612991Z","iopub.status.idle":"2022-07-20T07:23:01.633204Z","shell.execute_reply.started":"2022-07-20T07:23:01.612907Z","shell.execute_reply":"2022-07-20T07:23:01.632282Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"y_pred_knn = knn.predict(x_test)\ny_pred_knn","metadata":{"execution":{"iopub.status.busy":"2022-07-20T07:23:01.634568Z","iopub.execute_input":"2022-07-20T07:23:01.634916Z","iopub.status.idle":"2022-07-20T07:23:01.808647Z","shell.execute_reply.started":"2022-07-20T07:23:01.634886Z","shell.execute_reply":"2022-07-20T07:23:01.807225Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Logistic Regression","metadata":{}},{"cell_type":"code","source":"from sklearn.metrics import confusion_matrix,classification_report,accuracy_score\ncm = confusion_matrix(y_test,y_pred_knn)\nprint(classification_report(y_test,y_pred_knn))\nprint(cm)","metadata":{"execution":{"iopub.status.busy":"2022-07-20T07:23:01.810577Z","iopub.execute_input":"2022-07-20T07:23:01.811117Z","iopub.status.idle":"2022-07-20T07:23:01.831287Z","shell.execute_reply.started":"2022-07-20T07:23:01.811069Z","shell.execute_reply":"2022-07-20T07:23:01.830086Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"knn_acc = accuracy_score(y_test,y_pred_knn)*100# it defines how much correctly the model is predicting the actual value\nknn_acc","metadata":{"execution":{"iopub.status.busy":"2022-07-20T07:23:01.832814Z","iopub.execute_input":"2022-07-20T07:23:01.833171Z","iopub.status.idle":"2022-07-20T07:23:01.843819Z","shell.execute_reply.started":"2022-07-20T07:23:01.833139Z","shell.execute_reply":"2022-07-20T07:23:01.842328Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Support Vector Classifier","metadata":{}},{"cell_type":"code","source":"from sklearn import svm\nsvm = svm.SVC(kernel='rbf',C = 0.01)\nsvm.fit(x_train,y_train)","metadata":{"execution":{"iopub.status.busy":"2022-07-20T07:23:01.846163Z","iopub.execute_input":"2022-07-20T07:23:01.846713Z","iopub.status.idle":"2022-07-20T07:23:04.635395Z","shell.execute_reply.started":"2022-07-20T07:23:01.846666Z","shell.execute_reply":"2022-07-20T07:23:04.634117Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"y_pred_svm = svm.predict(x_test)\ny_pred_svm","metadata":{"execution":{"iopub.status.busy":"2022-07-20T07:23:04.636888Z","iopub.execute_input":"2022-07-20T07:23:04.638153Z","iopub.status.idle":"2022-07-20T07:23:05.580602Z","shell.execute_reply.started":"2022-07-20T07:23:04.638089Z","shell.execute_reply":"2022-07-20T07:23:05.578798Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"from sklearn.metrics import confusion_matrix,classification_report,accuracy_score\ncm = confusion_matrix(y_test,y_pred_svm)\nprint(classification_report(y_test,y_pred_svm))\nprint(cm)","metadata":{"execution":{"iopub.status.busy":"2022-07-20T07:23:05.582594Z","iopub.execute_input":"2022-07-20T07:23:05.583514Z","iopub.status.idle":"2022-07-20T07:23:05.604180Z","shell.execute_reply.started":"2022-07-20T07:23:05.583476Z","shell.execute_reply":"2022-07-20T07:23:05.602457Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"svm_acc = accuracy_score(y_test,y_pred_svm)*100# it defines how much correctly the model is predicting the actual value\nsvm_acc","metadata":{"execution":{"iopub.status.busy":"2022-07-20T07:23:05.606217Z","iopub.execute_input":"2022-07-20T07:23:05.606616Z","iopub.status.idle":"2022-07-20T07:23:05.620383Z","shell.execute_reply.started":"2022-07-20T07:23:05.606582Z","shell.execute_reply":"2022-07-20T07:23:05.618499Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"from sklearn.linear_model import LogisticRegression\nclassifier=LogisticRegression(random_state=42)\nclassifier.fit(x_train,y_train)","metadata":{"execution":{"iopub.status.busy":"2022-07-20T07:23:05.622176Z","iopub.execute_input":"2022-07-20T07:23:05.622885Z","iopub.status.idle":"2022-07-20T07:23:05.778893Z","shell.execute_reply.started":"2022-07-20T07:23:05.622844Z","shell.execute_reply":"2022-07-20T07:23:05.777664Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"y_pred = classifier.predict(x_test)\ny_pred","metadata":{"execution":{"iopub.status.busy":"2022-07-20T07:23:05.780983Z","iopub.execute_input":"2022-07-20T07:23:05.781804Z","iopub.status.idle":"2022-07-20T07:23:05.793806Z","shell.execute_reply.started":"2022-07-20T07:23:05.781758Z","shell.execute_reply":"2022-07-20T07:23:05.792537Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"from sklearn.metrics import confusion_matrix,classification_report,accuracy_score\ncm = confusion_matrix(y_test,y_pred)\naccuracy_score(y_test,y_pred)\nprint(classification_report(y_test,y_pred))\nprint(cm)","metadata":{"scrolled":true,"execution":{"iopub.status.busy":"2022-07-20T07:23:05.796236Z","iopub.execute_input":"2022-07-20T07:23:05.797276Z","iopub.status.idle":"2022-07-20T07:23:05.823976Z","shell.execute_reply.started":"2022-07-20T07:23:05.797230Z","shell.execute_reply":"2022-07-20T07:23:05.822806Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"log_acc=accuracy_score(y_test,y_pred)*100\nlog_acc","metadata":{"execution":{"iopub.status.busy":"2022-07-20T07:23:05.826005Z","iopub.execute_input":"2022-07-20T07:23:05.826798Z","iopub.status.idle":"2022-07-20T07:23:05.838329Z","shell.execute_reply.started":"2022-07-20T07:23:05.826752Z","shell.execute_reply":"2022-07-20T07:23:05.836727Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"for col in df_test.columns:\n    if df_test[col].dtypes=='object':\n        df_test[col]=df_test[col].fillna(df_test[col].mode()[0])\n    else:\n        df_test[col]=df_test[col].fillna(df_test[col].median())","metadata":{"execution":{"iopub.status.busy":"2022-07-20T07:23:05.840417Z","iopub.execute_input":"2022-07-20T07:23:05.841505Z","iopub.status.idle":"2022-07-20T07:23:05.891687Z","shell.execute_reply.started":"2022-07-20T07:23:05.841458Z","shell.execute_reply":"2022-07-20T07:23:05.890323Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"b = df_test['PassengerId']","metadata":{"execution":{"iopub.status.busy":"2022-07-20T07:23:05.893718Z","iopub.execute_input":"2022-07-20T07:23:05.894555Z","iopub.status.idle":"2022-07-20T07:23:05.900770Z","shell.execute_reply.started":"2022-07-20T07:23:05.894512Z","shell.execute_reply":"2022-07-20T07:23:05.899501Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"df_test.drop('PassengerId',axis=1,inplace=True)","metadata":{"execution":{"iopub.status.busy":"2022-07-20T07:23:05.906461Z","iopub.execute_input":"2022-07-20T07:23:05.909552Z","iopub.status.idle":"2022-07-20T07:23:05.916742Z","shell.execute_reply.started":"2022-07-20T07:23:05.909511Z","shell.execute_reply":"2022-07-20T07:23:05.915067Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"from sklearn import preprocessing\nfor i in df_test.columns:\n    if df_test[i].dtype=='object' or df_test[i].dtype=='bool':\n        label_encoder=preprocessing.LabelEncoder()\n        df_test[i]=label_encoder.fit_transform(df_test[i])","metadata":{"execution":{"iopub.status.busy":"2022-07-20T07:23:05.918928Z","iopub.execute_input":"2022-07-20T07:23:05.919457Z","iopub.status.idle":"2022-07-20T07:23:05.963406Z","shell.execute_reply.started":"2022-07-20T07:23:05.919421Z","shell.execute_reply":"2022-07-20T07:23:05.961711Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"y_valid = rf.predict(df_test)\ny_valid","metadata":{"execution":{"iopub.status.busy":"2022-07-20T07:23:05.965489Z","iopub.execute_input":"2022-07-20T07:23:05.966676Z","iopub.status.idle":"2022-07-20T07:23:06.094733Z","shell.execute_reply.started":"2022-07-20T07:23:05.966637Z","shell.execute_reply":"2022-07-20T07:23:06.093367Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"submission_spaceship = pd.DataFrame({\"PassengerId\": b, \"Transported\":y_valid})\nsubmission_spaceship","metadata":{"execution":{"iopub.status.busy":"2022-07-20T07:23:06.096624Z","iopub.execute_input":"2022-07-20T07:23:06.097337Z","iopub.status.idle":"2022-07-20T07:23:06.113617Z","shell.execute_reply.started":"2022-07-20T07:23:06.097302Z","shell.execute_reply":"2022-07-20T07:23:06.112218Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"submission_spaceship['Transported'] =submission_spaceship['Transported'].apply(lambda x:True if x==1 else False)","metadata":{"execution":{"iopub.status.busy":"2022-07-20T07:23:06.115141Z","iopub.execute_input":"2022-07-20T07:23:06.116478Z","iopub.status.idle":"2022-07-20T07:23:06.126789Z","shell.execute_reply.started":"2022-07-20T07:23:06.116423Z","shell.execute_reply":"2022-07-20T07:23:06.125375Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"submission_spaceship","metadata":{"execution":{"iopub.status.busy":"2022-07-20T07:23:06.128657Z","iopub.execute_input":"2022-07-20T07:23:06.129155Z","iopub.status.idle":"2022-07-20T07:23:06.150716Z","shell.execute_reply.started":"2022-07-20T07:23:06.129119Z","shell.execute_reply":"2022-07-20T07:23:06.149175Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# XGBoost Algorithm","metadata":{}},{"cell_type":"code","source":"from xgboost import XGBClassifier\nfrom sklearn.metrics import accuracy_score\nfrom sklearn.metrics import roc_auc_score\n\nboost_model = XGBClassifier(n_jobs=-1, random_state=42,max_depth = 5)\n\n#Fitting the model\nboost_model.fit(x_train,y_train)","metadata":{"execution":{"iopub.status.busy":"2022-07-20T07:23:06.152282Z","iopub.execute_input":"2022-07-20T07:23:06.152785Z","iopub.status.idle":"2022-07-20T07:23:07.006691Z","shell.execute_reply.started":"2022-07-20T07:23:06.152751Z","shell.execute_reply":"2022-07-20T07:23:07.005564Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"pred = boost_model.predict(x_test)","metadata":{"execution":{"iopub.status.busy":"2022-07-20T07:23:07.008411Z","iopub.execute_input":"2022-07-20T07:23:07.010181Z","iopub.status.idle":"2022-07-20T07:23:07.031077Z","shell.execute_reply.started":"2022-07-20T07:23:07.010110Z","shell.execute_reply":"2022-07-20T07:23:07.030033Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"from sklearn.metrics import confusion_matrix,classification_report,accuracy_score\ncm = confusion_matrix(y_test,pred)\nxg_acc=accuracy_score(y_test,pred)\nprint(classification_report(y_test,pred))\nprint(cm)\nprint(xg_acc)","metadata":{"execution":{"iopub.status.busy":"2022-07-20T07:23:07.032630Z","iopub.execute_input":"2022-07-20T07:23:07.033827Z","iopub.status.idle":"2022-07-20T07:23:07.054057Z","shell.execute_reply.started":"2022-07-20T07:23:07.033790Z","shell.execute_reply":"2022-07-20T07:23:07.053147Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"y_valid_xg = boost_model.predict(df_test)\ny_valid_xg","metadata":{"execution":{"iopub.status.busy":"2022-07-20T07:23:07.055582Z","iopub.execute_input":"2022-07-20T07:23:07.056667Z","iopub.status.idle":"2022-07-20T07:23:07.081774Z","shell.execute_reply.started":"2022-07-20T07:23:07.056621Z","shell.execute_reply":"2022-07-20T07:23:07.080292Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"spaceship_xg = pd.DataFrame({\"PassengerId\": b, \"Transported\":y_valid_xg})\nspaceship_xg","metadata":{"execution":{"iopub.status.busy":"2022-07-20T07:23:07.086444Z","iopub.execute_input":"2022-07-20T07:23:07.087345Z","iopub.status.idle":"2022-07-20T07:23:07.103660Z","shell.execute_reply.started":"2022-07-20T07:23:07.087302Z","shell.execute_reply":"2022-07-20T07:23:07.102297Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"spaceship_xg['Transported'] =spaceship_xg['Transported'].apply(lambda x:True if x==1 else False)","metadata":{"execution":{"iopub.status.busy":"2022-07-20T07:23:07.105518Z","iopub.execute_input":"2022-07-20T07:23:07.106278Z","iopub.status.idle":"2022-07-20T07:23:07.115243Z","shell.execute_reply.started":"2022-07-20T07:23:07.106243Z","shell.execute_reply":"2022-07-20T07:23:07.113382Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"spaceship_xg ","metadata":{"execution":{"iopub.status.busy":"2022-07-20T07:23:07.117097Z","iopub.execute_input":"2022-07-20T07:23:07.118627Z","iopub.status.idle":"2022-07-20T07:23:07.134985Z","shell.execute_reply.started":"2022-07-20T07:23:07.118570Z","shell.execute_reply":"2022-07-20T07:23:07.133569Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Accuracy of models ","metadata":{}},{"cell_type":"code","source":"import numpy as np\n\nplt.rcParams[\"figure.figsize\"] = [12, 5]\nplt.rcParams[\"figure.autolayout\"] = True\n\nx=['RandomForest','KNN','LogisticRegression','SVC','XGBClassifier']\ny=[rf_acc,knn_acc,log_acc,svm_acc,xg_acc*100]\n\nwidth = 0.75\nfig, ax = plt.subplots()\n\npps = ax.bar(x, y, width, align='center')\n\nfor p in pps:\n   height = p.get_height()\n   ax.text(x=p.get_x() + p.get_width() / 2, y=height+1,\n      s=\"{}%\".format(height),\n      ha='center')\nplt.title('Accuracy of models')\nplt.show()","metadata":{"execution":{"iopub.status.busy":"2022-07-20T07:23:07.137075Z","iopub.execute_input":"2022-07-20T07:23:07.137835Z","iopub.status.idle":"2022-07-20T07:23:07.334688Z","shell.execute_reply.started":"2022-07-20T07:23:07.137783Z","shell.execute_reply":"2022-07-20T07:23:07.333648Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"","metadata":{},"execution_count":null,"outputs":[]}]}