{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"pygments_lexer":"ipython3","nbconvert_exporter":"python","version":"3.6.4","file_extension":".py","codemirror_mode":{"name":"ipython","version":3},"name":"python","mimetype":"text/x-python"}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"You should construct your machine learning models using the training set. We give each passengers outcome (often referred to as the ground reality) for the training set. Your model will be based on \"features\" like the class and gender of the passengers. In order to develop new features, feature engineering can also be used.\n\nYou should evaluate your models performance on unobserved data using the test set. We do not give each passengers ground truth for the test set. Your responsibility is to foresee these outcomes. Use the model you trained to forecast whether each test set passenger survived the Titanics sinking for each passenger in the test set.\n\nVariable Notes\n\npclass: A proxy for socio-economic status (SES) 1st = Upper 2nd = Middle 3rd = Lower\n\nage: Age is fractional if less than 1. If the age is estimated, is it in the form of xx.5\n\nsibsp: The dataset defines family relations in this way...\n\nSibling = brother, sister, stepbrother, stepsister\n\nSpouse = husband, wife (mistresses and fiancés were ignored)\n\nparch: The dataset defines family relations in this way...\n\nParent = mother, father\n\nChild = daughter, son, stepdaughter, stepson\n\nSome children travelled only with a nanny, therefore parch=0 for them.","metadata":{}},{"cell_type":"code","source":"# This Python 3 environment comes with many helpful analytics libraries installed\n# It is defined by the kaggle/python Docker image: https://github.com/kaggle/docker-python\n# For example, here's several helpful packages to load\n\nimport numpy as np # linear algebra\nimport pandas as pd # data processing, CSV file I/O (e.g. pd.read_csv)\n\n# Input data files are available in the read-only \"../input/\" directory\n# For example, running this (by clicking run or pressing Shift+Enter) will list all files under the input directory\n\nimport os\nfor dirname, _, filenames in os.walk('../input/titanic'):\n    for filename in filenames:\n        print(os.path.join(dirname, filename))\n\n# You can write up to 20GB to the current directory (/kaggle/working/) that gets preserved as output when you create a version using \"Save & Run All\" \n# You can also write temporary files to /kaggle/temp/, but they won't be saved outside of the current session","metadata":{"_uuid":"8f2839f25d086af736a60e9eeb907d3b93b6e0e5","_cell_guid":"b1076dfc-b9ad-4769-8c92-a6c4dae69d19","execution":{"iopub.status.busy":"2022-07-21T04:36:42.809161Z","iopub.execute_input":"2022-07-21T04:36:42.809867Z","iopub.status.idle":"2022-07-21T04:36:42.818072Z","shell.execute_reply.started":"2022-07-21T04:36:42.809822Z","shell.execute_reply":"2022-07-21T04:36:42.817188Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# **Importing the essential modules**","metadata":{}},{"cell_type":"code","source":"import pandas as pd\nimport numpy as np \nimport matplotlib.pyplot as plt\nimport seaborn as sns\nfrom sklearn.ensemble import ExtraTreesClassifier\nfrom sklearn.linear_model import LogisticRegression\nfrom sklearn.model_selection import train_test_split, GridSearchCV\nfrom sklearn.preprocessing import StandardScaler\nfrom sklearn.metrics import accuracy_score\nfrom sklearn.tree import DecisionTreeClassifier\nfrom sklearn import svm\nfrom sklearn.neighbors import KNeighborsClassifier\nfrom IPython.display import display","metadata":{"execution":{"iopub.status.busy":"2022-07-21T04:36:47.656608Z","iopub.execute_input":"2022-07-21T04:36:47.656960Z","iopub.status.idle":"2022-07-21T04:36:48.687763Z","shell.execute_reply.started":"2022-07-21T04:36:47.656931Z","shell.execute_reply":"2022-07-21T04:36:48.686485Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"train = pd.read_csv(\"../input/titanic/train.csv\") \ntest = pd.read_csv(\"../input/titanic/test.csv\")","metadata":{"execution":{"iopub.status.busy":"2022-07-21T04:37:18.119082Z","iopub.execute_input":"2022-07-21T04:37:18.119486Z","iopub.status.idle":"2022-07-21T04:37:18.135297Z","shell.execute_reply.started":"2022-07-21T04:37:18.119454Z","shell.execute_reply":"2022-07-21T04:37:18.134355Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# **Display Top 10 data of train and test**","metadata":{}},{"cell_type":"code","source":"display(train.head(10))","metadata":{"execution":{"iopub.status.busy":"2022-07-21T04:37:29.940558Z","iopub.execute_input":"2022-07-21T04:37:29.941328Z","iopub.status.idle":"2022-07-21T04:37:29.961570Z","shell.execute_reply.started":"2022-07-21T04:37:29.941288Z","shell.execute_reply":"2022-07-21T04:37:29.960619Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"display(test.head(10))","metadata":{"execution":{"iopub.status.busy":"2022-07-21T04:37:30.077669Z","iopub.execute_input":"2022-07-21T04:37:30.080046Z","iopub.status.idle":"2022-07-21T04:37:30.101148Z","shell.execute_reply.started":"2022-07-21T04:37:30.080005Z","shell.execute_reply":"2022-07-21T04:37:30.100329Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"train.info()","metadata":{"execution":{"iopub.status.busy":"2022-07-21T04:37:30.253378Z","iopub.execute_input":"2022-07-21T04:37:30.253732Z","iopub.status.idle":"2022-07-21T04:37:30.281267Z","shell.execute_reply.started":"2022-07-21T04:37:30.253703Z","shell.execute_reply":"2022-07-21T04:37:30.280037Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"test.info()","metadata":{"execution":{"iopub.status.busy":"2022-07-21T04:37:30.411939Z","iopub.execute_input":"2022-07-21T04:37:30.412571Z","iopub.status.idle":"2022-07-21T04:37:30.426946Z","shell.execute_reply.started":"2022-07-21T04:37:30.412521Z","shell.execute_reply":"2022-07-21T04:37:30.425396Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"train.describe(include = \"all\")","metadata":{"execution":{"iopub.status.busy":"2022-07-21T04:37:30.587860Z","iopub.execute_input":"2022-07-21T04:37:30.588454Z","iopub.status.idle":"2022-07-21T04:37:30.637874Z","shell.execute_reply.started":"2022-07-21T04:37:30.588420Z","shell.execute_reply":"2022-07-21T04:37:30.636767Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"test.describe(include = \"all\")","metadata":{"execution":{"iopub.status.busy":"2022-07-21T04:37:30.767856Z","iopub.execute_input":"2022-07-21T04:37:30.768877Z","iopub.status.idle":"2022-07-21T04:37:30.831647Z","shell.execute_reply.started":"2022-07-21T04:37:30.768818Z","shell.execute_reply":"2022-07-21T04:37:30.830257Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"#number of column in test and train\nprint(\"Number of column in train:\",train.shape[1], \" Number of column in test:\", test.shape[1])","metadata":{"execution":{"iopub.status.busy":"2022-07-21T04:37:30.932799Z","iopub.execute_input":"2022-07-21T04:37:30.933232Z","iopub.status.idle":"2022-07-21T04:37:30.939712Z","shell.execute_reply.started":"2022-07-21T04:37:30.933196Z","shell.execute_reply":"2022-07-21T04:37:30.938484Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"#number of rows in train and test\nprint(\"Number of rows in train:\",train.shape[0], \"Number of rows in test:\",test.shape[0])","metadata":{"execution":{"iopub.status.busy":"2022-07-21T04:37:31.083035Z","iopub.execute_input":"2022-07-21T04:37:31.083471Z","iopub.status.idle":"2022-07-21T04:37:31.089534Z","shell.execute_reply.started":"2022-07-21T04:37:31.083433Z","shell.execute_reply":"2022-07-21T04:37:31.088334Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# **Data Cleaning**\n1. Removing null values\n2. Removing unwanted column\n3. Formatting the column\n4. Added new column according to need","metadata":{}},{"cell_type":"code","source":"def null_value(df):\n    percentage = ((df.isna().sum()/df.isna().count())*100).sort_values(ascending = False)\n    count = (df.isna().sum()).sort_values(ascending = False)\n    diff= pd.concat([count,percentage],axis = 1, keys = [\"Count\", \"Percentage\"])\n    return diff","metadata":{"execution":{"iopub.status.busy":"2022-07-21T04:37:31.434201Z","iopub.execute_input":"2022-07-21T04:37:31.434597Z","iopub.status.idle":"2022-07-21T04:37:31.440972Z","shell.execute_reply.started":"2022-07-21T04:37:31.434561Z","shell.execute_reply":"2022-07-21T04:37:31.440078Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"null_value(train)","metadata":{"execution":{"iopub.status.busy":"2022-07-21T04:37:31.590719Z","iopub.execute_input":"2022-07-21T04:37:31.591163Z","iopub.status.idle":"2022-07-21T04:37:31.614616Z","shell.execute_reply.started":"2022-07-21T04:37:31.591129Z","shell.execute_reply":"2022-07-21T04:37:31.613440Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"null_value(test)","metadata":{"execution":{"iopub.status.busy":"2022-07-21T04:37:31.739984Z","iopub.execute_input":"2022-07-21T04:37:31.741287Z","iopub.status.idle":"2022-07-21T04:37:31.759885Z","shell.execute_reply.started":"2022-07-21T04:37:31.741233Z","shell.execute_reply":"2022-07-21T04:37:31.758716Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"#dropping the column cabin because it has more percentage of null and the cabin can be found from the fare rate and pclass\ntrain.drop(\"Cabin\",axis = 1,inplace = True)\ntest.drop(\"Cabin\",axis=1, inplace = True)","metadata":{"execution":{"iopub.status.busy":"2022-07-21T04:37:31.910176Z","iopub.execute_input":"2022-07-21T04:37:31.910587Z","iopub.status.idle":"2022-07-21T04:37:31.921281Z","shell.execute_reply.started":"2022-07-21T04:37:31.910550Z","shell.execute_reply":"2022-07-21T04:37:31.919937Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"#fill up or replace the null in the age with the mean of age\ntrain[\"Age\"].replace(np.nan, train[\"Age\"].mean(), inplace = True)\ntrain[\"Embarked\"].replace(np.nan, train[\"Embarked\"].mode()[0], inplace = True)\ntrain[\"Age\"].isna().sum(), train[\"Embarked\"].isna().sum()","metadata":{"execution":{"iopub.status.busy":"2022-07-21T04:37:32.083204Z","iopub.execute_input":"2022-07-21T04:37:32.083596Z","iopub.status.idle":"2022-07-21T04:37:32.097238Z","shell.execute_reply.started":"2022-07-21T04:37:32.083563Z","shell.execute_reply":"2022-07-21T04:37:32.096209Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"test[\"Age\"].replace(np.nan, test[\"Age\"].mean(),inplace = True)\ntest[\"Fare\"].replace(np.nan, test[\"Fare\"].mode()[0],inplace = True)\ntest[\"Age\"].isna().sum(), test[\"Fare\"].isna().sum()","metadata":{"execution":{"iopub.status.busy":"2022-07-21T04:37:32.252146Z","iopub.execute_input":"2022-07-21T04:37:32.252549Z","iopub.status.idle":"2022-07-21T04:37:32.267805Z","shell.execute_reply.started":"2022-07-21T04:37:32.252515Z","shell.execute_reply":"2022-07-21T04:37:32.266627Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"#Dropping of the Name and the passenger ticket because both has no recurring pattern\ntrain.drop([\"Name\",\"Ticket\"], axis =1, inplace = True)\ntest.drop([\"Name\",\"Ticket\"],axis = 1, inplace = True)","metadata":{"execution":{"iopub.status.busy":"2022-07-21T04:37:32.422755Z","iopub.execute_input":"2022-07-21T04:37:32.425086Z","iopub.status.idle":"2022-07-21T04:37:32.432847Z","shell.execute_reply.started":"2022-07-21T04:37:32.425047Z","shell.execute_reply":"2022-07-21T04:37:32.431568Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"**Visualization of Sex and the Survival count**","metadata":{}},{"cell_type":"code","source":"dd = train[[\"Sex\",\"Survived\"]].groupby([\"Survived\"], as_index = False).count()\ndd.plot(kind = \"bar\",figsize = (10,7))\nplt.xlabel(\"Sex\")\nplt.ylabel(\"Survived\")\nplt.show()","metadata":{"execution":{"iopub.status.busy":"2022-07-21T04:37:32.782583Z","iopub.execute_input":"2022-07-21T04:37:32.783720Z","iopub.status.idle":"2022-07-21T04:37:33.035284Z","shell.execute_reply.started":"2022-07-21T04:37:32.783675Z","shell.execute_reply":"2022-07-21T04:37:33.034088Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"dd = train[[\"Embarked\",\"Survived\"]].groupby([\"Survived\"], as_index = False).count()\ndd.plot(kind = \"bar\",figsize = (10,7))\nplt.xlabel(\"Embarked\")\nplt.ylabel(\"Survived\")\nplt.show()","metadata":{"execution":{"iopub.status.busy":"2022-07-21T04:37:33.039166Z","iopub.execute_input":"2022-07-21T04:37:33.039507Z","iopub.status.idle":"2022-07-21T04:37:33.246380Z","shell.execute_reply.started":"2022-07-21T04:37:33.039475Z","shell.execute_reply":"2022-07-21T04:37:33.245194Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"As we can see that both **Sex** and **Embarked** shows a great dependacy/relationship to the target.","metadata":{}},{"cell_type":"code","source":"#formating the columns of Sex and Embarked\ntrain = pd.get_dummies(train, prefix = [\"Sex\",\"Embarked\"])\ntest = pd.get_dummies(test, prefix = [\"Sex\",\"Embarked\"])","metadata":{"execution":{"iopub.status.busy":"2022-07-21T04:37:33.285065Z","iopub.execute_input":"2022-07-21T04:37:33.285504Z","iopub.status.idle":"2022-07-21T04:37:33.302765Z","shell.execute_reply.started":"2022-07-21T04:37:33.285467Z","shell.execute_reply":"2022-07-21T04:37:33.301594Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"train","metadata":{"execution":{"iopub.status.busy":"2022-07-21T04:37:33.448460Z","iopub.execute_input":"2022-07-21T04:37:33.448866Z","iopub.status.idle":"2022-07-21T04:37:33.470825Z","shell.execute_reply.started":"2022-07-21T04:37:33.448830Z","shell.execute_reply":"2022-07-21T04:37:33.469709Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"#combining the SibSp and Parch\ndef merge(df):\n    rel = []\n    for i in df[\"SibSp\"].values.tolist():\n        rel.append(i)\n    rel1 = []\n    for i in df[\"Parch\"].values.tolist():\n        rel1.append(i)\n        \n    concat = []\n    for index in range(len(rel)):\n        concat.append(rel[index] + rel1[index])\n        \n    df1 = pd.DataFrame(concat, columns = [\"Relatives\"])\n    df = pd.concat([df,df1], axis = 1)\n    return df","metadata":{"execution":{"iopub.status.busy":"2022-07-21T04:37:33.620198Z","iopub.execute_input":"2022-07-21T04:37:33.620610Z","iopub.status.idle":"2022-07-21T04:37:33.629168Z","shell.execute_reply.started":"2022-07-21T04:37:33.620575Z","shell.execute_reply":"2022-07-21T04:37:33.627990Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"train = merge(train)","metadata":{"execution":{"iopub.status.busy":"2022-07-21T04:37:33.793339Z","iopub.execute_input":"2022-07-21T04:37:33.794432Z","iopub.status.idle":"2022-07-21T04:37:33.802156Z","shell.execute_reply.started":"2022-07-21T04:37:33.794365Z","shell.execute_reply":"2022-07-21T04:37:33.800943Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"train.drop([\"SibSp\",\"Parch\"],axis=1, inplace = True)","metadata":{"execution":{"iopub.status.busy":"2022-07-21T04:37:33.969004Z","iopub.execute_input":"2022-07-21T04:37:33.970036Z","iopub.status.idle":"2022-07-21T04:37:33.979404Z","shell.execute_reply.started":"2022-07-21T04:37:33.969978Z","shell.execute_reply":"2022-07-21T04:37:33.977884Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"train.loc[888]","metadata":{"execution":{"iopub.status.busy":"2022-07-21T04:37:34.180070Z","iopub.execute_input":"2022-07-21T04:37:34.182430Z","iopub.status.idle":"2022-07-21T04:37:34.192309Z","shell.execute_reply.started":"2022-07-21T04:37:34.182392Z","shell.execute_reply":"2022-07-21T04:37:34.191076Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"test = merge(test)\ntest.drop([\"SibSp\",\"Parch\"],axis = 1, inplace = True)","metadata":{"execution":{"iopub.status.busy":"2022-07-21T04:37:34.604580Z","iopub.execute_input":"2022-07-21T04:37:34.605515Z","iopub.status.idle":"2022-07-21T04:37:34.613757Z","shell.execute_reply.started":"2022-07-21T04:37:34.605476Z","shell.execute_reply":"2022-07-21T04:37:34.612747Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"test.head()","metadata":{"execution":{"iopub.status.busy":"2022-07-21T04:37:35.040477Z","iopub.execute_input":"2022-07-21T04:37:35.041434Z","iopub.status.idle":"2022-07-21T04:37:35.055914Z","shell.execute_reply.started":"2022-07-21T04:37:35.041395Z","shell.execute_reply":"2022-07-21T04:37:35.054854Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# **Feature Importance using ExtraTreeClassifier**","metadata":{}},{"cell_type":"code","source":"x = train.drop(\"Survived\",axis = 1)\ny = train[\"Survived\"]","metadata":{"execution":{"iopub.status.busy":"2022-07-21T04:37:35.573578Z","iopub.execute_input":"2022-07-21T04:37:35.573954Z","iopub.status.idle":"2022-07-21T04:37:35.580332Z","shell.execute_reply.started":"2022-07-21T04:37:35.573923Z","shell.execute_reply":"2022-07-21T04:37:35.579155Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"#model object creation and fitting\nmodel=ExtraTreesClassifier()\nmodel.fit(x,y)","metadata":{"execution":{"iopub.status.busy":"2022-07-21T04:37:35.825419Z","iopub.execute_input":"2022-07-21T04:37:35.825835Z","iopub.status.idle":"2022-07-21T04:37:36.014018Z","shell.execute_reply.started":"2022-07-21T04:37:35.825797Z","shell.execute_reply":"2022-07-21T04:37:36.012917Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"ft = pd.Series(model.feature_importances_, index = x.columns)\nft.plot(kind = \"barh\")\nplt.show()","metadata":{"execution":{"iopub.status.busy":"2022-07-21T04:37:36.083247Z","iopub.execute_input":"2022-07-21T04:37:36.083658Z","iopub.status.idle":"2022-07-21T04:37:36.303475Z","shell.execute_reply.started":"2022-07-21T04:37:36.083624Z","shell.execute_reply":"2022-07-21T04:37:36.302342Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"#heatmap for correlation between feature\nplt.figure(figsize = (20,15))\nsns.heatmap(train.corr(),annot = True)\nplt.show()","metadata":{"execution":{"iopub.status.busy":"2022-07-21T04:37:36.347133Z","iopub.execute_input":"2022-07-21T04:37:36.348304Z","iopub.status.idle":"2022-07-21T04:37:37.165171Z","shell.execute_reply.started":"2022-07-21T04:37:36.348262Z","shell.execute_reply":"2022-07-21T04:37:37.164277Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"#Dropping passengerId as we can see the correlation of passengerId is near to zero with respect to Survived\ntrain.drop(\"PassengerId\",axis = 1,inplace = True)","metadata":{"execution":{"iopub.status.busy":"2022-07-21T04:37:37.166390Z","iopub.execute_input":"2022-07-21T04:37:37.166682Z","iopub.status.idle":"2022-07-21T04:37:37.173238Z","shell.execute_reply.started":"2022-07-21T04:37:37.166654Z","shell.execute_reply":"2022-07-21T04:37:37.171899Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# **Data Preprocessing**","metadata":{}},{"cell_type":"code","source":"X_train, X_test, y_train, y_test = train_test_split(x,y, test_size = 0.2, random_state = 10)\nX_train = StandardScaler().fit_transform(X_train)\nX_test = StandardScaler().fit_transform(X_test)\ntest = StandardScaler().fit_transform(test)","metadata":{"execution":{"iopub.status.busy":"2022-07-21T04:37:37.177283Z","iopub.execute_input":"2022-07-21T04:37:37.177845Z","iopub.status.idle":"2022-07-21T04:37:37.203211Z","shell.execute_reply.started":"2022-07-21T04:37:37.177804Z","shell.execute_reply":"2022-07-21T04:37:37.202022Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"**Logistic Regression**","metadata":{}},{"cell_type":"code","source":"lr = LogisticRegression()\nparams = { \"penalty\": (\"l1\", \"l2\", \"elasticnet\"), \"tol\": (0.1, 0.01, 0.001, 0.0001), \"C\": (10.0, 1.0, 0.1, 0.01)}\nmodelLR = GridSearchCV(lr, params, cv=10)\nmodelLR.fit(X_train, y_train)","metadata":{"execution":{"iopub.status.busy":"2022-07-21T04:37:37.707955Z","iopub.execute_input":"2022-07-21T04:37:37.708651Z","iopub.status.idle":"2022-07-21T04:37:38.804985Z","shell.execute_reply.started":"2022-07-21T04:37:37.708599Z","shell.execute_reply":"2022-07-21T04:37:38.803823Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"print(accuracy_score(modelLR.predict(X_test),y_test))","metadata":{"execution":{"iopub.status.busy":"2022-07-21T04:37:38.806812Z","iopub.execute_input":"2022-07-21T04:37:38.807143Z","iopub.status.idle":"2022-07-21T04:37:38.814488Z","shell.execute_reply.started":"2022-07-21T04:37:38.807099Z","shell.execute_reply":"2022-07-21T04:37:38.813280Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"**Support Vector Machine(SVM)**","metadata":{}},{"cell_type":"code","source":"modelSVM = svm.SVC(kernel = \"rbf\")\nmodelSVM.fit(X_train,y_train)","metadata":{"execution":{"iopub.status.busy":"2022-07-21T04:37:38.865339Z","iopub.execute_input":"2022-07-21T04:37:38.865762Z","iopub.status.idle":"2022-07-21T04:37:38.893860Z","shell.execute_reply.started":"2022-07-21T04:37:38.865729Z","shell.execute_reply":"2022-07-21T04:37:38.892683Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"print(accuracy_score(modelSVM.predict(X_test),y_test))","metadata":{"execution":{"iopub.status.busy":"2022-07-21T04:37:39.040341Z","iopub.execute_input":"2022-07-21T04:37:39.040755Z","iopub.status.idle":"2022-07-21T04:37:39.052691Z","shell.execute_reply.started":"2022-07-21T04:37:39.040720Z","shell.execute_reply":"2022-07-21T04:37:39.051363Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"**DecisionTreeClassifier**","metadata":{}},{"cell_type":"code","source":"modelDTC = DecisionTreeClassifier(criterion=\"entropy\")\nmodelDTC.fit(X_train,y_train)","metadata":{"execution":{"iopub.status.busy":"2022-07-21T04:37:39.511054Z","iopub.execute_input":"2022-07-21T04:37:39.511455Z","iopub.status.idle":"2022-07-21T04:37:39.523501Z","shell.execute_reply.started":"2022-07-21T04:37:39.511424Z","shell.execute_reply":"2022-07-21T04:37:39.522150Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"print(accuracy_score(modelDTC.predict(X_test),y_test))","metadata":{"execution":{"iopub.status.busy":"2022-07-21T04:37:39.759076Z","iopub.execute_input":"2022-07-21T04:37:39.759804Z","iopub.status.idle":"2022-07-21T04:37:39.766217Z","shell.execute_reply.started":"2022-07-21T04:37:39.759763Z","shell.execute_reply":"2022-07-21T04:37:39.765394Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"**KNeighborsClassifier**","metadata":{}},{"cell_type":"code","source":"modelKNC = KNeighborsClassifier(n_neighbors=4)\nmodelKNC.fit(X_train, y_train)","metadata":{"execution":{"iopub.status.busy":"2022-07-21T04:37:40.225897Z","iopub.execute_input":"2022-07-21T04:37:40.226342Z","iopub.status.idle":"2022-07-21T04:37:40.234592Z","shell.execute_reply.started":"2022-07-21T04:37:40.226305Z","shell.execute_reply":"2022-07-21T04:37:40.233738Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"print(accuracy_score(modelKNC.predict(X_test),y_test))","metadata":{"execution":{"iopub.status.busy":"2022-07-21T04:37:40.482275Z","iopub.execute_input":"2022-07-21T04:37:40.482649Z","iopub.status.idle":"2022-07-21T04:37:40.498919Z","shell.execute_reply.started":"2022-07-21T04:37:40.482619Z","shell.execute_reply":"2022-07-21T04:37:40.498121Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"report = pd.DataFrame({\n    \"Model\": [\"LogisticRegression\",\"SVM\",\"DecisionTreeClassifier\",\"KNeighborsClassifier\"],\n    \"Accuracy\": [accuracy_score(modelLR.predict(X_test),y_test),accuracy_score(modelSVM.predict(X_test),y_test),\n                  accuracy_score(modelDTC.predict(X_test),y_test),accuracy_score(modelKNC.predict(X_test),y_test)]\n})","metadata":{"execution":{"iopub.status.busy":"2022-07-21T04:37:40.843498Z","iopub.execute_input":"2022-07-21T04:37:40.844091Z","iopub.status.idle":"2022-07-21T04:37:40.866214Z","shell.execute_reply.started":"2022-07-21T04:37:40.844058Z","shell.execute_reply":"2022-07-21T04:37:40.865351Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"report.sort_values(by = \"Accuracy\")","metadata":{"execution":{"iopub.status.busy":"2022-07-21T04:37:41.088034Z","iopub.execute_input":"2022-07-21T04:37:41.088636Z","iopub.status.idle":"2022-07-21T04:37:41.099809Z","shell.execute_reply.started":"2022-07-21T04:37:41.088603Z","shell.execute_reply":"2022-07-21T04:37:41.098877Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Test Data Prediction","metadata":{}},{"cell_type":"markdown","source":"**Using SVM as it has a better accuracy than other**","metadata":{}},{"cell_type":"code","source":"pred = modelSVM.predict(test)","metadata":{"execution":{"iopub.status.busy":"2022-07-21T04:37:42.222943Z","iopub.execute_input":"2022-07-21T04:37:42.223336Z","iopub.status.idle":"2022-07-21T04:37:42.237844Z","shell.execute_reply.started":"2022-07-21T04:37:42.223292Z","shell.execute_reply":"2022-07-21T04:37:42.236860Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"x = pd.read_csv(\"../input/titanic/test.csv\")\nprediction = pd.DataFrame({\"PassengerId\":x[\"PassengerId\"], \"Survived\": pred})","metadata":{"execution":{"iopub.status.busy":"2022-07-21T04:37:42.436069Z","iopub.execute_input":"2022-07-21T04:37:42.436470Z","iopub.status.idle":"2022-07-21T04:37:42.448145Z","shell.execute_reply.started":"2022-07-21T04:37:42.436437Z","shell.execute_reply":"2022-07-21T04:37:42.447342Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"prediction.sample(10)","metadata":{"execution":{"iopub.status.busy":"2022-07-21T04:37:42.821665Z","iopub.execute_input":"2022-07-21T04:37:42.822787Z","iopub.status.idle":"2022-07-21T04:37:42.834376Z","shell.execute_reply.started":"2022-07-21T04:37:42.822742Z","shell.execute_reply":"2022-07-21T04:37:42.833287Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### You can also use feature selection module from [scikit-learn](https://scikit-learn.org/stable/modules/feature_selection.html) to know the best feature for better accuracy.\n### Such as: SelectKBest -> Univariate feature selection works by selecting the best features based on univariate statistical tests. It can be seen as a preprocessing step to an estimator","metadata":{}},{"cell_type":"code","source":"prediction.to_csv(\"submission.csv\")","metadata":{"execution":{"iopub.status.busy":"2022-07-16T16:35:26.357341Z","iopub.execute_input":"2022-07-16T16:35:26.358177Z","iopub.status.idle":"2022-07-16T16:35:26.369683Z","shell.execute_reply.started":"2022-07-16T16:35:26.358113Z","shell.execute_reply":"2022-07-16T16:35:26.368359Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":">**Inspired by one of the contributor of this competition**","metadata":{}}]}