{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"pygments_lexer":"ipython3","nbconvert_exporter":"python","version":"3.6.4","file_extension":".py","codemirror_mode":{"name":"ipython","version":3},"name":"python","mimetype":"text/x-python"}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"**Introduction**\n\n**Titanic disaster survivals dataset**\n> *On April 15, 1912, during her maiden voyage, the widely considered “unsinkable” RMS Titanic sank after colliding with an iceberg. Unfortunately, there weren’t enough lifeboats for everyone onboard, resulting in the death of 1502 out of 2224 passengers and crew.*\n\n**Mission**\n\nBuild a predictive model that answers the question: “what sorts of people were more likely to survive?” using passenger data (ie name, age, gender, socio-economic class, etc).\n\n**Work plan**\n* Analyze and explore data\n* Building a Machine Learning Model","metadata":{}},{"cell_type":"code","source":"#Import libraries\nimport pandas as pd\nimport numpy as np\nimport matplotlib.pyplot as plt\nimport seaborn as sn\nfrom sklearn.preprocessing import StandardScaler\nfrom sklearn.linear_model import LogisticRegression\nfrom sklearn.naive_bayes import GaussianNB\nfrom sklearn.ensemble import RandomForestClassifier\nfrom sklearn.tree import DecisionTreeClassifier\nfrom sklearn.svm import SVC\nfrom sklearn.neighbors import KNeighborsClassifier\nfrom sklearn.model_selection import cross_val_score\nfrom sklearn.model_selection import GridSearchCV","metadata":{"execution":{"iopub.status.busy":"2022-07-29T08:55:02.553023Z","iopub.execute_input":"2022-07-29T08:55:02.554132Z","iopub.status.idle":"2022-07-29T08:55:02.563090Z","shell.execute_reply.started":"2022-07-29T08:55:02.554080Z","shell.execute_reply":"2022-07-29T08:55:02.562054Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"#Load the data\ndf = pd.read_csv('../input/titanic/train.csv')\ndf.head()","metadata":{"execution":{"iopub.status.busy":"2022-07-29T08:55:02.663964Z","iopub.execute_input":"2022-07-29T08:55:02.666683Z","iopub.status.idle":"2022-07-29T08:55:02.693894Z","shell.execute_reply.started":"2022-07-29T08:55:02.666629Z","shell.execute_reply":"2022-07-29T08:55:02.692571Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Data visualization","metadata":{}},{"cell_type":"code","source":"#lets look at how many entries are there\ndf.shape","metadata":{"execution":{"iopub.status.busy":"2022-07-29T08:55:04.035709Z","iopub.execute_input":"2022-07-29T08:55:04.036084Z","iopub.status.idle":"2022-07-29T08:55:04.043009Z","shell.execute_reply.started":"2022-07-29T08:55:04.036054Z","shell.execute_reply":"2022-07-29T08:55:04.041666Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"#### This dataset contain Numerical and Categorical data. So lets separate them into numerical and categorical. It's easy to us work with further\n\n* Numerical\n    1. Age\n    2. SibSp\n    3. Parch\n    4. Fare\n    \n* Categorical\n    1. Pclass\n    2. Sex\n    3. Ticket\n    4. Embarked\n    5. Cabin","metadata":{}},{"cell_type":"code","source":"#separate these data into numerical and categorical \nnumerical = df[['Age','SibSp','Parch','Fare']]\ncategorical = df[['Pclass','Sex','Ticket','Embarked','Cabin']]","metadata":{"execution":{"iopub.status.busy":"2022-07-29T08:55:04.215416Z","iopub.execute_input":"2022-07-29T08:55:04.215838Z","iopub.status.idle":"2022-07-29T08:55:04.226138Z","shell.execute_reply.started":"2022-07-29T08:55:04.215805Z","shell.execute_reply":"2022-07-29T08:55:04.224961Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"**First look at the numerical data visualization**","metadata":{}},{"cell_type":"code","source":"#Numerical data visualization\nfor a in numerical.columns:\n    plt.hist(numerical[a])\n    plt.title(a)\n    plt.show()","metadata":{"execution":{"iopub.status.busy":"2022-07-29T08:55:04.269203Z","iopub.execute_input":"2022-07-29T08:55:04.269843Z","iopub.status.idle":"2022-07-29T08:55:05.135558Z","shell.execute_reply.started":"2022-07-29T08:55:04.269800Z","shell.execute_reply":"2022-07-29T08:55:05.133801Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"So lets look at the survival rate with above values","metadata":{}},{"cell_type":"code","source":"#Lets look relationship between survival and numerical values\npd.pivot_table(df, index = 'Survived', values = ['Age','SibSp','Parch','Fare'])","metadata":{"execution":{"iopub.status.busy":"2022-07-29T08:55:05.138380Z","iopub.execute_input":"2022-07-29T08:55:05.139275Z","iopub.status.idle":"2022-07-29T08:55:05.163116Z","shell.execute_reply.started":"2022-07-29T08:55:05.139237Z","shell.execute_reply":"2022-07-29T08:55:05.161965Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"**Ok we saw relationship between numerical data. Lets look at the categorical data visualization**","metadata":{}},{"cell_type":"code","source":"#Categorical data visualization\nfor a in categorical.columns:\n    sn.barplot(categorical[a].value_counts().index,categorical[a].value_counts()).set_title(a)\n    plt.show()","metadata":{"execution":{"iopub.status.busy":"2022-07-29T08:55:05.164621Z","iopub.execute_input":"2022-07-29T08:55:05.165359Z","iopub.status.idle":"2022-07-29T08:55:16.881868Z","shell.execute_reply.started":"2022-07-29T08:55:05.165318Z","shell.execute_reply":"2022-07-29T08:55:16.880583Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Survived rate with categorical data","metadata":{}},{"cell_type":"code","source":"#Lets look relationship between survival and categorical values\nprint(pd.pivot_table(df, index = 'Survived', columns = 'Pclass', values = 'Ticket' , aggfunc ='count'))\nprint('--------------------')\nprint(pd.pivot_table(df, index = 'Survived', columns = 'Sex', values = 'Ticket' ,aggfunc ='count'))\nprint('--------------------')\nprint(pd.pivot_table(df, index = 'Survived', columns = 'Embarked', values = 'Ticket' ,aggfunc ='count'))","metadata":{"execution":{"iopub.status.busy":"2022-07-29T08:55:16.883329Z","iopub.execute_input":"2022-07-29T08:55:16.883740Z","iopub.status.idle":"2022-07-29T08:55:16.927406Z","shell.execute_reply.started":"2022-07-29T08:55:16.883708Z","shell.execute_reply":"2022-07-29T08:55:16.926191Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Feature Engineering","metadata":{}},{"cell_type":"markdown","source":"Lets look at some information about our dataframe","metadata":{}},{"cell_type":"code","source":"df.describe()","metadata":{"execution":{"iopub.status.busy":"2022-07-29T08:55:16.930011Z","iopub.execute_input":"2022-07-29T08:55:16.930448Z","iopub.status.idle":"2022-07-29T08:55:16.968020Z","shell.execute_reply.started":"2022-07-29T08:55:16.930418Z","shell.execute_reply":"2022-07-29T08:55:16.966767Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### Data Cleaning","metadata":{}},{"cell_type":"markdown","source":"**Ok according to above chart. There are some outliers. look max value and min value of fare. I think that is a outliers. If I train my model with this outliers, It will be a problem.**\n* **So first remove this outliers**","metadata":{}},{"cell_type":"code","source":"def outliers_remove(df):\n    outliers_list = ['Age','SibSp','Parch','Fare']\n    for entry in outliers_list:\n        max_limit = df[entry].mean() + 4*df[entry].std()\n        min_limit = df[entry].mean() - 4*df[entry].std()\n        df = df[(df[entry]<max_limit) & (df[entry]>min_limit)]\n    return df","metadata":{"execution":{"iopub.status.busy":"2022-07-29T08:55:16.969563Z","iopub.execute_input":"2022-07-29T08:55:16.969903Z","iopub.status.idle":"2022-07-29T08:55:16.978204Z","shell.execute_reply.started":"2022-07-29T08:55:16.969873Z","shell.execute_reply":"2022-07-29T08:55:16.976584Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"df = outliers_remove(df)\ndf.describe()","metadata":{"execution":{"iopub.status.busy":"2022-07-29T08:55:16.980500Z","iopub.execute_input":"2022-07-29T08:55:16.980945Z","iopub.status.idle":"2022-07-29T08:55:17.036642Z","shell.execute_reply.started":"2022-07-29T08:55:16.980908Z","shell.execute_reply":"2022-07-29T08:55:17.035494Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"*** I used 4 standrad Deviation to remove these outliers.***","metadata":{}},{"cell_type":"code","source":"df.shape","metadata":{"execution":{"iopub.status.busy":"2022-07-29T08:55:17.038116Z","iopub.execute_input":"2022-07-29T08:55:17.038461Z","iopub.status.idle":"2022-07-29T08:55:17.045799Z","shell.execute_reply.started":"2022-07-29T08:55:17.038433Z","shell.execute_reply":"2022-07-29T08:55:17.044546Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"**I think I lost about 202 entries from outliers removal (891-689)**","metadata":{}},{"cell_type":"markdown","source":"**Lets look unique values in each column and is there any invalid entirs?**","metadata":{}},{"cell_type":"code","source":"df.head()","metadata":{"execution":{"iopub.status.busy":"2022-07-29T08:55:17.047652Z","iopub.execute_input":"2022-07-29T08:55:17.048346Z","iopub.status.idle":"2022-07-29T08:55:17.072515Z","shell.execute_reply.started":"2022-07-29T08:55:17.048303Z","shell.execute_reply":"2022-07-29T08:55:17.070808Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"print('Survived :',df.Survived.unique())\nprint('Pclass :',df.Pclass.unique())\nprint('Sex :',df.Sex.unique())\nprint('SibSp :',df.SibSp.unique())\nprint('Parch :',df.Parch.unique())\nprint('Cabin :',df.Cabin.unique())\nprint('Embarked :',df.Embarked.unique())","metadata":{"execution":{"iopub.status.busy":"2022-07-29T08:55:17.073864Z","iopub.execute_input":"2022-07-29T08:55:17.074171Z","iopub.status.idle":"2022-07-29T08:55:17.086191Z","shell.execute_reply.started":"2022-07-29T08:55:17.074144Z","shell.execute_reply":"2022-07-29T08:55:17.084930Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"**There are some NaN values. Lets look at how much**","metadata":{}},{"cell_type":"code","source":"#Lets look how many NaN values are there\ndf.isnull().sum()","metadata":{"execution":{"iopub.status.busy":"2022-07-29T08:55:17.087784Z","iopub.execute_input":"2022-07-29T08:55:17.088185Z","iopub.status.idle":"2022-07-29T08:55:17.104196Z","shell.execute_reply.started":"2022-07-29T08:55:17.088151Z","shell.execute_reply":"2022-07-29T08:55:17.103028Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"* There are 687 NaN Cabin entries out of 891. So we can ignore that column\n* Also drop 2 embarked entries\n* And we can fill other NaN values with there mean","metadata":{}},{"cell_type":"code","source":"#There are 687 NaN Cabin entries out of 891. So we can ignore that column\n#Also drop 2 embarked entries\n#And we can fill other NaN values with there mean\ndef fix_nan_values(df):\n    df.drop('Cabin',axis='columns',inplace=True)\n    df.Age.fillna(df.Age.mean(),inplace=True)\n    df.Fare.fillna(df.Fare.mean(),inplace=True)\n    df.dropna(subset=['Embarked'],inplace=True)\n    return df\n","metadata":{"execution":{"iopub.status.busy":"2022-07-29T08:55:17.106341Z","iopub.execute_input":"2022-07-29T08:55:17.106846Z","iopub.status.idle":"2022-07-29T08:55:17.118407Z","shell.execute_reply.started":"2022-07-29T08:55:17.106801Z","shell.execute_reply":"2022-07-29T08:55:17.117209Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"fix_nan_values(df)\ndf.isnull().sum()","metadata":{"execution":{"iopub.status.busy":"2022-07-29T08:55:17.120024Z","iopub.execute_input":"2022-07-29T08:55:17.120367Z","iopub.status.idle":"2022-07-29T08:55:17.141593Z","shell.execute_reply.started":"2022-07-29T08:55:17.120338Z","shell.execute_reply":"2022-07-29T08:55:17.140638Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"**So there isn't any NaN values further. Lets apply dummies for categorical variables.**","metadata":{}},{"cell_type":"markdown","source":"> ***Every time wrotes functions because it helpful when do same process to test data set.***","metadata":{}},{"cell_type":"code","source":"def apply_dummies(df):\n    dummies=pd.get_dummies(df.Sex)\n    df1 = pd.concat([df,dummies.drop(['female'],axis='columns')],axis='columns')\n    df1.drop('Sex',axis='columns',inplace=True)\n    dummies=pd.get_dummies(df1.Embarked)\n    df2 = pd.concat([df1,dummies.drop(['Q'],axis='columns')],axis='columns')\n    df2.drop('Embarked', axis='columns',inplace=True)\n    return df2","metadata":{"execution":{"iopub.status.busy":"2022-07-29T08:55:17.146541Z","iopub.execute_input":"2022-07-29T08:55:17.147104Z","iopub.status.idle":"2022-07-29T08:55:17.153862Z","shell.execute_reply.started":"2022-07-29T08:55:17.147073Z","shell.execute_reply":"2022-07-29T08:55:17.152828Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"df = apply_dummies(df)\ndf.head()","metadata":{"execution":{"iopub.status.busy":"2022-07-29T08:55:17.155185Z","iopub.execute_input":"2022-07-29T08:55:17.156102Z","iopub.status.idle":"2022-07-29T08:55:17.188306Z","shell.execute_reply.started":"2022-07-29T08:55:17.156064Z","shell.execute_reply":"2022-07-29T08:55:17.187339Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"*** Name and Ticket columns are no need. It looks there isn't any relationship between servival rate and that Name/Ticket columns. I don't know exactly but I think so. ***","metadata":{}},{"cell_type":"markdown","source":"**Therefore I drop these columns**","metadata":{}},{"cell_type":"code","source":"#Name and Ticket column is no need. So drop them\ndef drop_columns(df):\n    df.drop('Name',axis='columns',inplace=True)\n    df.drop('Ticket',axis='columns',inplace=True)\n    return df\n    ","metadata":{"execution":{"iopub.status.busy":"2022-07-29T08:55:17.190014Z","iopub.execute_input":"2022-07-29T08:55:17.190629Z","iopub.status.idle":"2022-07-29T08:55:17.195678Z","shell.execute_reply.started":"2022-07-29T08:55:17.190589Z","shell.execute_reply":"2022-07-29T08:55:17.194741Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"drop_columns(df)\ndf.head()","metadata":{"execution":{"iopub.status.busy":"2022-07-29T08:55:17.197340Z","iopub.execute_input":"2022-07-29T08:55:17.197810Z","iopub.status.idle":"2022-07-29T08:55:17.220807Z","shell.execute_reply.started":"2022-07-29T08:55:17.197778Z","shell.execute_reply":"2022-07-29T08:55:17.219420Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"**Now our dataset is looking good. Let's apply a scaller for Age and Fare data columns**","metadata":{}},{"cell_type":"markdown","source":"I  use StandardScaller","metadata":{}},{"cell_type":"code","source":"def scale_data(df):\n    scaler = StandardScaler()\n    df[['Age','Fare']] = scaler.fit_transform(df[['Age','Fare']])\n    df.head()\n    return df","metadata":{"execution":{"iopub.status.busy":"2022-07-29T08:55:17.222170Z","iopub.execute_input":"2022-07-29T08:55:17.223073Z","iopub.status.idle":"2022-07-29T08:55:17.234492Z","shell.execute_reply.started":"2022-07-29T08:55:17.223038Z","shell.execute_reply":"2022-07-29T08:55:17.233483Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"scale_data(df)\ndf.head()","metadata":{"execution":{"iopub.status.busy":"2022-07-29T08:55:17.235935Z","iopub.execute_input":"2022-07-29T08:55:17.236531Z","iopub.status.idle":"2022-07-29T08:55:17.264941Z","shell.execute_reply.started":"2022-07-29T08:55:17.236490Z","shell.execute_reply":"2022-07-29T08:55:17.263567Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Model Training ","metadata":{}},{"cell_type":"markdown","source":"**Now we can take this data to train our model**","metadata":{}},{"cell_type":"markdown","source":"**Split data into x train and y train**","metadata":{}},{"cell_type":"code","source":"X_train, y_train = df.drop(['Survived','PassengerId'],axis='columns'),df['Survived']\nX_train.shape, y_train.shape","metadata":{"execution":{"iopub.status.busy":"2022-07-29T08:55:17.266421Z","iopub.execute_input":"2022-07-29T08:55:17.267320Z","iopub.status.idle":"2022-07-29T08:55:17.275769Z","shell.execute_reply.started":"2022-07-29T08:55:17.267278Z","shell.execute_reply":"2022-07-29T08:55:17.274833Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"**Then load test data and apply that previous data cleaning and scalling functions**","metadata":{}},{"cell_type":"code","source":"test_data = pd.read_csv('../input/titanic/test.csv')\nfix_nan_values(test_data)\ntest_data = apply_dummies(test_data)\ndrop_columns(test_data)\nscale_data(test_data)\ntest_data.head()\n","metadata":{"execution":{"iopub.status.busy":"2022-07-29T08:55:17.277631Z","iopub.execute_input":"2022-07-29T08:55:17.278595Z","iopub.status.idle":"2022-07-29T08:55:17.325848Z","shell.execute_reply.started":"2022-07-29T08:55:17.278545Z","shell.execute_reply":"2022-07-29T08:55:17.324545Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"X_test = test_data.drop('PassengerId',axis='columns')\nX_test.head()","metadata":{"execution":{"iopub.status.busy":"2022-07-29T08:55:17.327670Z","iopub.execute_input":"2022-07-29T08:55:17.328674Z","iopub.status.idle":"2022-07-29T08:55:17.344742Z","shell.execute_reply.started":"2022-07-29T08:55:17.328625Z","shell.execute_reply":"2022-07-29T08:55:17.343266Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"**Lets look how many entries with train and test data**","metadata":{}},{"cell_type":"code","source":"X_train.shape, X_test.shape, y_train.shape","metadata":{"execution":{"iopub.status.busy":"2022-07-29T08:55:17.346303Z","iopub.execute_input":"2022-07-29T08:55:17.347239Z","iopub.status.idle":"2022-07-29T08:55:17.356036Z","shell.execute_reply.started":"2022-07-29T08:55:17.347171Z","shell.execute_reply":"2022-07-29T08:55:17.354856Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"**Lets try to train different models and look their avg scores.**","metadata":{}},{"cell_type":"code","source":"m_gnb = GaussianNB()\nscore = cross_val_score(m_gnb,X_train,y_train,cv=5)\nprint(score.mean())","metadata":{"execution":{"iopub.status.busy":"2022-07-29T08:55:17.358005Z","iopub.execute_input":"2022-07-29T08:55:17.358761Z","iopub.status.idle":"2022-07-29T08:55:17.400408Z","shell.execute_reply.started":"2022-07-29T08:55:17.358715Z","shell.execute_reply":"2022-07-29T08:55:17.399472Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"m_lr = LogisticRegression()\nscore = cross_val_score(m_lr,X_train,y_train,cv=5)\nprint(score.mean())","metadata":{"execution":{"iopub.status.busy":"2022-07-29T08:55:17.401731Z","iopub.execute_input":"2022-07-29T08:55:17.402256Z","iopub.status.idle":"2022-07-29T08:55:17.469181Z","shell.execute_reply.started":"2022-07-29T08:55:17.402215Z","shell.execute_reply":"2022-07-29T08:55:17.468034Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"m_rfc = RandomForestClassifier()\nscore = cross_val_score(m_rfc,X_train,y_train,cv=5)\nprint(score.mean())","metadata":{"execution":{"iopub.status.busy":"2022-07-29T08:55:17.470775Z","iopub.execute_input":"2022-07-29T08:55:17.471477Z","iopub.status.idle":"2022-07-29T08:55:18.623784Z","shell.execute_reply.started":"2022-07-29T08:55:17.471431Z","shell.execute_reply":"2022-07-29T08:55:18.622492Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"m_dt = DecisionTreeClassifier()\nscore = cross_val_score(m_dt,X_train,y_train,cv=5)\nprint(score.mean())","metadata":{"execution":{"iopub.status.busy":"2022-07-29T08:55:18.625073Z","iopub.execute_input":"2022-07-29T08:55:18.625772Z","iopub.status.idle":"2022-07-29T08:55:18.667125Z","shell.execute_reply.started":"2022-07-29T08:55:18.625733Z","shell.execute_reply":"2022-07-29T08:55:18.665978Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"m_svc = SVC()\nscore = cross_val_score(m_svc,X_train,y_train,cv=5)\nprint(score.mean())","metadata":{"execution":{"iopub.status.busy":"2022-07-29T08:55:18.668674Z","iopub.execute_input":"2022-07-29T08:55:18.669661Z","iopub.status.idle":"2022-07-29T08:55:18.774082Z","shell.execute_reply.started":"2022-07-29T08:55:18.669620Z","shell.execute_reply":"2022-07-29T08:55:18.772724Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"m_knc = KNeighborsClassifier()\nscore = cross_val_score(m_knc,X_train,y_train,cv=5)\nprint(score.mean())","metadata":{"execution":{"iopub.status.busy":"2022-07-29T08:55:18.775469Z","iopub.execute_input":"2022-07-29T08:55:18.775795Z","iopub.status.idle":"2022-07-29T08:55:18.849582Z","shell.execute_reply.started":"2022-07-29T08:55:18.775767Z","shell.execute_reply":"2022-07-29T08:55:18.848100Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## Hyper Parameter Tuning","metadata":{}},{"cell_type":"markdown","source":"**Lets try different parameters in different models**","metadata":{}},{"cell_type":"code","source":"model_param = {\n    'svc': {\n        'model': SVC(gamma='auto'),\n        'params' : {\n            'C': [1,10,20],\n            'kernel': ['rbf','linear']\n        }  \n    },\n    'random_forest': {\n        'model': RandomForestClassifier(),\n        'params' : {\n            'n_estimators': [1,5,10],\n            'criterion':['gini','entropy']\n        }\n    },\n    'logistic_regression' : {\n        'model': LogisticRegression(solver='liblinear',multi_class='auto'),\n        'params': {\n            'C': [1,5,10]\n        }\n    },\n    'gaussian' :{\n        'model' : GaussianNB(),\n        'params' : {\n            \n        }\n    },\n    'kneighbors' : {\n        'model' : KNeighborsClassifier(),\n        'params' : {\n            'n_neighbors' : [3,5,7,9],\n            'weights' : ['uniform', 'distance'],\n            'algorithm' : ['auto', 'ball_tree','kd_tree']\n        }\n    },\n    'tree' : {\n        'model' : DecisionTreeClassifier(),\n        'params' : {\n            'criterion': ['gini','entropy'],\n        }\n    }\n}","metadata":{"execution":{"iopub.status.busy":"2022-07-29T08:55:18.851128Z","iopub.execute_input":"2022-07-29T08:55:18.851564Z","iopub.status.idle":"2022-07-29T08:55:18.863437Z","shell.execute_reply.started":"2022-07-29T08:55:18.851530Z","shell.execute_reply":"2022-07-29T08:55:18.862070Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"scores = []\nfor model_name, mp in model_param.items():\n  clf = GridSearchCV(mp['model'],mp['params'],cv=3,return_train_score=False)\n  clf.fit(X_train,y_train)\n  scores.append({\n      'model' : model_name,\n      'best_score' : clf.best_score_,\n      'best_params' : clf.best_params_\n  })\nbest_sc = pd.DataFrame(scores,columns=['model','best_score','best_params'])\nbest_sc","metadata":{"execution":{"iopub.status.busy":"2022-07-29T08:55:18.864590Z","iopub.execute_input":"2022-07-29T08:55:18.865017Z","iopub.status.idle":"2022-07-29T08:55:20.577942Z","shell.execute_reply.started":"2022-07-29T08:55:18.864986Z","shell.execute_reply":"2022-07-29T08:55:20.576362Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"**So SVC model score is looking good. 81% accuracy with training samples**","metadata":{}},{"cell_type":"markdown","source":"**Lets try svc model with our test data**","metadata":{}},{"cell_type":"code","source":"model = SVC(C=1,kernel='rbf')\nmodel.fit(X_train,y_train)\nresult = model.predict(X_test)\nsubmission = pd.DataFrame({\n    \"PassengerId\" : test_data.PassengerId,\n    \"Survived\" : result\n})","metadata":{"execution":{"iopub.status.busy":"2022-07-29T08:55:20.579700Z","iopub.execute_input":"2022-07-29T08:55:20.580179Z","iopub.status.idle":"2022-07-29T08:55:20.615089Z","shell.execute_reply.started":"2022-07-29T08:55:20.580131Z","shell.execute_reply":"2022-07-29T08:55:20.614072Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"**convert submission to csv**","metadata":{}},{"cell_type":"code","source":"submission.to_csv('submission.csv', index=False)","metadata":{"execution":{"iopub.status.busy":"2022-07-29T08:55:20.616841Z","iopub.execute_input":"2022-07-29T08:55:20.617249Z","iopub.status.idle":"2022-07-29T08:55:20.626553Z","shell.execute_reply.started":"2022-07-29T08:55:20.617209Z","shell.execute_reply":"2022-07-29T08:55:20.625429Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"submission.head()","metadata":{"execution":{"iopub.status.busy":"2022-07-29T08:57:06.467967Z","iopub.execute_input":"2022-07-29T08:57:06.468606Z","iopub.status.idle":"2022-07-29T08:57:06.478161Z","shell.execute_reply.started":"2022-07-29T08:57:06.468572Z","shell.execute_reply":"2022-07-29T08:57:06.476923Z"},"trusted":true},"execution_count":null,"outputs":[]}]}