{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"pygments_lexer":"ipython3","nbconvert_exporter":"python","version":"3.6.4","file_extension":".py","codemirror_mode":{"name":"ipython","version":3},"name":"python","mimetype":"text/x-python"}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"## __Titanic - Machine Learning from Disaster__","metadata":{}},{"cell_type":"markdown","source":"I am not a expert in data science. This is my second try in this competition after lot of researches. There might be some issues. I hope you like this. These all written by myself and please leave a comment.. ","metadata":{}},{"cell_type":"markdown","source":"#### __Workflow Stages__","metadata":{}},{"cell_type":"markdown","source":"1. Problem definition.\n2. Acquire training and testing data.\n3. Wrangle, Prepare, Cleanse the data.\n4. Analyze, identify patterns and explore the data.\n5. Model, Predict and solve the problem.\n6. Visualize, report, and present the problem solving steps and final solution.\n7. Supply or submit the results.","metadata":{}},{"cell_type":"markdown","source":"The workflow indicates general sequence of how each stage may follow the other. However there are use cases with exceptions.","metadata":{}},{"cell_type":"markdown","source":"* We may combine multiple workflow stages. We may analyze by visualizing data.\n* Perform a stage earlier than indicated. We may analyze data before and after wrangling.\n* Perform a stage multiple times in our workflow. Visualize stage may be used multiple times.\n* We may drop a stage altogether","metadata":{}},{"cell_type":"markdown","source":"#### __Problem definition__","metadata":{}},{"cell_type":"markdown","source":"Use machine learning to create a model that predicts which passengers survived the Titanic shipwreck.","metadata":{}},{"cell_type":"markdown","source":"The sinking of the Titanic is one of the most infamous shipwrecks in history.\n\nOn April 15, 1912, during her maiden voyage, the widely considered “unsinkable” RMS Titanic sank after colliding with an iceberg. Unfortunately, there weren’t enough lifeboats for everyone onboard, resulting in the death of 1502 out of 2224 passengers and crew.\n\nWhile there was some element of luck involved in surviving, it seems some groups of people were more likely to survive than others.\n\nIn this challenge, we ask you to build a predictive model that answers the question: “what sorts of people were more likely to survive?” using passenger data (ie name, age, gender, socio-economic class, etc).","metadata":{}},{"cell_type":"markdown","source":"#### __Workflow goals__","metadata":{}},{"cell_type":"markdown","source":"The datascience solutions workflow solves for seven major goals.","metadata":{}},{"cell_type":"markdown","source":"__Classifing__. We may want to classify or categorize our samples. We may also want to undestand the implications or correlations of different classes with ourr solution goals\n\n__Correlation__. One can approach the problem based on available features within the training dataset. Which features within the dataset contribute significantly to our solution goal? Statistically speaking is there a correlation among a feature and solution goal? As the feature values change does the solution state change as well, and visa-versa? This can be tested both for numerical and categorical features in the given dataset. We may also want to determine correlation among features other than survival for subsequent goals and workflow stages. Correlating certain features may help in creating, completing, or correcting features.\n\n__Converting__. For modelling stage, one needs to prepare the data. Depending on the choice of model algorithm one may require all feature to be converted to numerical equivalent values. So for instance converting text categorical values to numerical values.\n\n__Completing__. Data preparation may also require us to estimate any missing values within a feature. Model algorithm may work best when there are no missing values.\n\n__Correcting__. We may also analyze the given training dataset for errors or possibly innacurate values within features and try to corrent these values or exclude the samples containing the errors. One way to do this is to detect any outliers among our samples or features. We may also completely discard a feature if it is not contributing to the analysis or may significantly skew the result.\n\n__Creating__. Can we create new feature based on an existing feature or a set of features, such that the new feature follows the correlation, conversion, completeness goals.\n\n__Charting__. How to slect the right visualization plots and charts depending on nature of the data and the solution goals.","metadata":{}},{"cell_type":"code","source":"# data analysis and wrangling\nimport pandas as pd\nimport numpy as np\nimport random as rnd\n\n# visualization\nimport seaborn as sns\nimport matplotlib.pyplot as plt\n%matplotlib inline\n\n# model training\nfrom sklearn.linear_model import LogisticRegression\nfrom sklearn.svm import SVC, LinearSVC\nfrom sklearn.ensemble import RandomForestClassifier\nfrom sklearn.neighbors import KNeighborsClassifier\nfrom sklearn.naive_bayes import GaussianNB\nfrom sklearn.linear_model import Perceptron\nfrom sklearn.linear_model import SGDClassifier\nfrom sklearn.tree import DecisionTreeClassifier\nfrom sklearn.model_selection import GridSearchCV\n","metadata":{"execution":{"iopub.status.busy":"2022-08-02T08:35:50.037308Z","iopub.execute_input":"2022-08-02T08:35:50.038118Z","iopub.status.idle":"2022-08-02T08:35:51.260012Z","shell.execute_reply.started":"2022-08-02T08:35:50.038031Z","shell.execute_reply":"2022-08-02T08:35:51.259060Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"#### __Acquire data__","metadata":{}},{"cell_type":"markdown","source":"Import train and test data sets using python pandas library and we create a variable called combine with combining train and test data sets together.","metadata":{}},{"cell_type":"code","source":"train_df = pd.read_csv('../input/titanic/train.csv')\ntest_df = pd.read_csv('../input/titanic/test.csv')\ncombine = [train_df,test_df]","metadata":{"execution":{"iopub.status.busy":"2022-08-02T08:35:51.261518Z","iopub.execute_input":"2022-08-02T08:35:51.261799Z","iopub.status.idle":"2022-08-02T08:35:51.284486Z","shell.execute_reply.started":"2022-08-02T08:35:51.261774Z","shell.execute_reply":"2022-08-02T08:35:51.283857Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"#### __Analize data by describing__","metadata":{}},{"cell_type":"markdown","source":"With pandas library we can use it to answer following questions.","metadata":{}},{"cell_type":"markdown","source":"__Which features are available in the dataset?__","metadata":{}},{"cell_type":"code","source":"print(train_df.columns.values)","metadata":{"execution":{"iopub.status.busy":"2022-08-02T08:35:51.285561Z","iopub.execute_input":"2022-08-02T08:35:51.286068Z","iopub.status.idle":"2022-08-02T08:35:51.291129Z","shell.execute_reply.started":"2022-08-02T08:35:51.286035Z","shell.execute_reply":"2022-08-02T08:35:51.290265Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"__Which features are categorical?__","metadata":{}},{"cell_type":"markdown","source":"These values classify the samples into sets of similar samples. (nominal, ordinal, ratio, interval based)","metadata":{}},{"cell_type":"markdown","source":"* Categorical\n    * Survived\n    * Sex\n    * Embarked\n* Ordinal\n    * Pclass","metadata":{}},{"cell_type":"markdown","source":"__Which features are numerical?__","metadata":{}},{"cell_type":"markdown","source":"These values change from sample to sample. (Discrete, Continuous, Timeseries based)","metadata":{}},{"cell_type":"markdown","source":"* Continuous\n    * Age\n    * Fare\n* Discrete\n    * SibSp\n    * Parch","metadata":{}},{"cell_type":"code","source":"train_df.head()","metadata":{"execution":{"iopub.status.busy":"2022-08-02T08:35:51.294163Z","iopub.execute_input":"2022-08-02T08:35:51.294989Z","iopub.status.idle":"2022-08-02T08:35:51.319433Z","shell.execute_reply.started":"2022-08-02T08:35:51.294953Z","shell.execute_reply":"2022-08-02T08:35:51.318591Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"__Which features are mixed types?__","metadata":{}},{"cell_type":"markdown","source":"Numerical, alphanumerical data whithin same feature.","metadata":{}},{"cell_type":"markdown","source":"* Numerical and alphanumerical\n    * Ticket\n* Aplhanumerical\n    * Cabin","metadata":{}},{"cell_type":"markdown","source":"__Which features may contain errors or typos?__","metadata":{}},{"cell_type":"markdown","source":"This can tell this by reviewing some samples.","metadata":{}},{"cell_type":"markdown","source":"* Name feature may contain errors or typos because there are several ways used in describing a name. (titles, round brackets, quots)","metadata":{}},{"cell_type":"code","source":"train_df.tail()","metadata":{"execution":{"iopub.status.busy":"2022-08-02T08:35:51.320665Z","iopub.execute_input":"2022-08-02T08:35:51.321375Z","iopub.status.idle":"2022-08-02T08:35:51.335902Z","shell.execute_reply.started":"2022-08-02T08:35:51.321341Z","shell.execute_reply":"2022-08-02T08:35:51.335278Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"__Which features contain blank, null or empty values?__","metadata":{}},{"cell_type":"markdown","source":"* In train data set,\n    * Age 177\n    * Cabin 687\n    * Embarked 2\n* In test data set,\n    * Age 86\n    * Cabin 327","metadata":{}},{"cell_type":"code","source":"train_df.isnull().sum()","metadata":{"execution":{"iopub.status.busy":"2022-08-02T08:35:51.336864Z","iopub.execute_input":"2022-08-02T08:35:51.337241Z","iopub.status.idle":"2022-08-02T08:35:51.348442Z","shell.execute_reply.started":"2022-08-02T08:35:51.337216Z","shell.execute_reply":"2022-08-02T08:35:51.347771Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"test_df.isnull().sum()","metadata":{"execution":{"iopub.status.busy":"2022-08-02T08:35:51.349272Z","iopub.execute_input":"2022-08-02T08:35:51.349556Z","iopub.status.idle":"2022-08-02T08:35:51.357638Z","shell.execute_reply.started":"2022-08-02T08:35:51.349533Z","shell.execute_reply":"2022-08-02T08:35:51.356885Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"__What are the data types for various features.?__","metadata":{}},{"cell_type":"markdown","source":"* 7 features are integer or float\n* 5 features are objects (strings)","metadata":{}},{"cell_type":"code","source":"train_df.info()","metadata":{"execution":{"iopub.status.busy":"2022-08-02T08:35:51.358646Z","iopub.execute_input":"2022-08-02T08:35:51.358953Z","iopub.status.idle":"2022-08-02T08:35:51.387497Z","shell.execute_reply.started":"2022-08-02T08:35:51.358920Z","shell.execute_reply":"2022-08-02T08:35:51.386481Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"__What is the distribution of numerical future values across the sample?__","metadata":{}},{"cell_type":"markdown","source":"* This dataset has 891 samples of data amoung actual number of passengers on the board (2,224)\n* Around 38% of samples represent survived\n* Most passengers > 75% did not travel with their parents or childrens.\n* Nearly 30% passengrs had siblingsand/or spouse aboard. \n* A very few elderly passengers withing age range 65-80","metadata":{}},{"cell_type":"code","source":"train_df.describe()","metadata":{"execution":{"iopub.status.busy":"2022-08-02T08:35:51.390545Z","iopub.execute_input":"2022-08-02T08:35:51.390865Z","iopub.status.idle":"2022-08-02T08:35:51.421206Z","shell.execute_reply.started":"2022-08-02T08:35:51.390839Z","shell.execute_reply":"2022-08-02T08:35:51.420308Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"__What is the distribution of categorical futures?__","metadata":{}},{"cell_type":"markdown","source":"* All names are unique\n* Sex variable have 2 possible values (Male:577, Female:314)\n* Ticket future has 681 unique values\n* Saveral passengers shared their cabins there are several duplicates\n* Embarked has 3 possible values and S is mostly used port (freq=644)","metadata":{}},{"cell_type":"code","source":"train_df.describe(include=['O'])","metadata":{"execution":{"iopub.status.busy":"2022-08-02T08:35:51.424145Z","iopub.execute_input":"2022-08-02T08:35:51.424920Z","iopub.status.idle":"2022-08-02T08:35:51.442559Z","shell.execute_reply.started":"2022-08-02T08:35:51.424893Z","shell.execute_reply":"2022-08-02T08:35:51.441689Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"#### __Assumtions based on data analysis__","metadata":{}},{"cell_type":"markdown","source":"We arrived following assumtions based on data analysis so far.","metadata":{}},{"cell_type":"markdown","source":"* __Correlation__:\n    We want to know how well does each feature correlate with Survival.","metadata":{}},{"cell_type":"markdown","source":"* __Completing__:\n    * We want to complete the age it is correlated with survival\n    * And Embarked feature also correlated with survival. So we should complete that","metadata":{}},{"cell_type":"markdown","source":"* __Creating__:\n    * We may want to create a new feature called family base on SibSp and Parch columns\n    * We may want to extract the title from name feature and create a new column\n    * We may want to create a new column named age brand. This turns a continuous numerical value to ordinal categotical feature\n    * We may want to create a fare rage feature","metadata":{}},{"cell_type":"markdown","source":"#### __Analyze by pivoting features__","metadata":{}},{"cell_type":"markdown","source":"We can analyze our future correlations by pivolating features against each other.We can only do this with non empty value features. So we can use Sex(categorical), Pclass(ordinal), SibSp(discrete) features.","metadata":{}},{"cell_type":"markdown","source":"* __Pclass__ We can observe significant correlation (>0.5) among Pclass=1 and Survived. So we decide to include this feature in our model\n\n* __Sex__ We can see female passangers had higher survival rate(74%). So Sex feature is important to our model\n\n* __SibSp and Parch__ These features have zero correlation for certain values. It may be best ot derive a feature or a set of features from these individual features.","metadata":{}},{"cell_type":"code","source":"train_df[['Pclass','Survived']].groupby(['Pclass'], as_index=False).mean().sort_values(by='Survived', ascending=False)","metadata":{"execution":{"iopub.status.busy":"2022-08-02T08:35:51.443726Z","iopub.execute_input":"2022-08-02T08:35:51.444402Z","iopub.status.idle":"2022-08-02T08:35:51.457994Z","shell.execute_reply.started":"2022-08-02T08:35:51.444363Z","shell.execute_reply":"2022-08-02T08:35:51.457094Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"train_df[['Sex','Survived']].groupby(['Sex'], as_index=False).mean().sort_values(by='Survived', ascending=False)","metadata":{"execution":{"iopub.status.busy":"2022-08-02T08:35:51.459124Z","iopub.execute_input":"2022-08-02T08:35:51.460951Z","iopub.status.idle":"2022-08-02T08:35:51.473599Z","shell.execute_reply.started":"2022-08-02T08:35:51.460917Z","shell.execute_reply":"2022-08-02T08:35:51.472904Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"train_df[['SibSp','Survived']].groupby(['SibSp'], as_index=False).mean().sort_values(by='SibSp', ascending=False)","metadata":{"execution":{"iopub.status.busy":"2022-08-02T08:35:51.474730Z","iopub.execute_input":"2022-08-02T08:35:51.475186Z","iopub.status.idle":"2022-08-02T08:35:51.487626Z","shell.execute_reply.started":"2022-08-02T08:35:51.475160Z","shell.execute_reply":"2022-08-02T08:35:51.486970Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"train_df[['Parch','Survived']].groupby(['Parch'], as_index=False).mean().sort_values(by='Parch', ascending=False)","metadata":{"execution":{"iopub.status.busy":"2022-08-02T08:35:51.488604Z","iopub.execute_input":"2022-08-02T08:35:51.489041Z","iopub.status.idle":"2022-08-02T08:35:51.500357Z","shell.execute_reply.started":"2022-08-02T08:35:51.489016Z","shell.execute_reply":"2022-08-02T08:35:51.499241Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"#### __Analyze by visualizing data__","metadata":{}},{"cell_type":"markdown","source":"We can continue our assumptions using visualizing data","metadata":{}},{"cell_type":"markdown","source":"__Correlation numerical features__","metadata":{}},{"cell_type":"markdown","source":"* __Observations__\n    * Age <= had higher survival rate\n    * Large number of 15-30 years olds did not survive\n    * Most passengers are in 15-35 range\n    * Oldest passengers around 80 years old survived","metadata":{}},{"cell_type":"markdown","source":"* __Decisions__\n    * Use age feature to model training.\n    * Complete the null values in age feaure\n    * Band age groups","metadata":{}},{"cell_type":"code","source":"g = sns.FacetGrid(train_df, col='Survived')\ng.map(plt.hist, 'Age', bins=20)","metadata":{"execution":{"iopub.status.busy":"2022-08-02T08:35:51.501472Z","iopub.execute_input":"2022-08-02T08:35:51.502284Z","iopub.status.idle":"2022-08-02T08:35:51.840937Z","shell.execute_reply.started":"2022-08-02T08:35:51.502257Z","shell.execute_reply":"2022-08-02T08:35:51.839952Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"__Correlation numerical and ordinal features__","metadata":{}},{"cell_type":"markdown","source":"We can combine multiple features for identify correlation using a single plot. This can be done with numerical and categorical features which have numeric values","metadata":{}},{"cell_type":"markdown","source":"* __Observation__\n    * In Pclass 1 most passengers survived\n    * In Pclass 2 around half of passengers survived.\n    * In Pckass 3 most passengers not survived and there were most passengers\n    * Pclass varies in term of Age distribution of passengers","metadata":{}},{"cell_type":"markdown","source":"* __Decisions__\n    * Consider Pclass for model training","metadata":{}},{"cell_type":"code","source":"grid = sns.FacetGrid(train_df, col='Survived', row='Pclass', size=2.2, aspect=1.6)\ngrid.map(plt.hist, 'Age', alpha=.5, bins=20)\ngrid.add_legend()","metadata":{"execution":{"iopub.status.busy":"2022-08-02T08:35:51.842141Z","iopub.execute_input":"2022-08-02T08:35:51.842691Z","iopub.status.idle":"2022-08-02T08:35:52.896627Z","shell.execute_reply.started":"2022-08-02T08:35:51.842665Z","shell.execute_reply":"2022-08-02T08:35:52.895714Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"__Correlation categorical features__","metadata":{}},{"cell_type":"markdown","source":"* __Observation__\n    * Female passengers had much better survival rate than male\n    * In Embarked C male had higher survival rate. \n    * Males had brtter survival rate in Pclass 3 when compare with Pclass 2 for C and Q ports\n    * Posts of Embarked have varying survival rate for Pclass 3 and among male passengers.","metadata":{}},{"cell_type":"markdown","source":"* __Decisions__\n    * Add Sex feature fot model training.\n    * Complete and add Embarked feature for model training","metadata":{}},{"cell_type":"code","source":"grid = sns.FacetGrid(train_df, row='Embarked', size=2.2, aspect=1.6)\ngrid.map(sns.pointplot, 'Pclass', 'Survived', 'Sex', palette='deep')\ngrid.add_legend()","metadata":{"execution":{"iopub.status.busy":"2022-08-02T08:35:52.897864Z","iopub.execute_input":"2022-08-02T08:35:52.898644Z","iopub.status.idle":"2022-08-02T08:35:53.841789Z","shell.execute_reply.started":"2022-08-02T08:35:52.898609Z","shell.execute_reply":"2022-08-02T08:35:53.840768Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"__Correlation categorical and numerical features__","metadata":{}},{"cell_type":"markdown","source":"We may also want to correlate categorical features (with non-numeric values) and numerical features. We can consider correlating Embarked (categorical non-numeric), Sex(categorical non-numeric), Fare(Numeric continuous), Survived(Categorical numeric)","metadata":{}},{"cell_type":"markdown","source":"* __Observations__\n    * Higher fare paying passengers had higher survival rate\n    * Port of embark is correlated with survival rate","metadata":{}},{"cell_type":"markdown","source":"* __Decisions__\n    * consider Fare feature for our model training\n    * Create a Fare range feature","metadata":{}},{"cell_type":"code","source":"grid = sns.FacetGrid(train_df, row='Embarked', col='Survived', size=2.2, aspect=1.6)\ngrid.map(sns.barplot, 'Sex', 'Fare', alpha=.5, ci=None)\ngrid.add_legend()","metadata":{"execution":{"iopub.status.busy":"2022-08-02T08:35:53.843037Z","iopub.execute_input":"2022-08-02T08:35:53.844073Z","iopub.status.idle":"2022-08-02T08:35:54.775543Z","shell.execute_reply.started":"2022-08-02T08:35:53.844038Z","shell.execute_reply":"2022-08-02T08:35:54.774607Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"#### __Wrangle data__","metadata":{}},{"cell_type":"markdown","source":"We have collected several assumptions and decisions regarding our datasets and solution requrements. So far we did not change any value or any feature. Lets execute our decisions and assumptions for correcting, creating, and completing goals.","metadata":{}},{"cell_type":"markdown","source":"__Correcting by dropping features__","metadata":{}},{"cell_type":"markdown","source":"Based on our assumptions we want to drop Cabin and Ticket features. ","metadata":{}},{"cell_type":"code","source":"train_df = train_df.drop(['Ticket', 'Cabin'], axis=1)\ntest_df = test_df.drop(['Ticket', 'Cabin'], axis=1)\ncombine = [train_df, test_df]","metadata":{"execution":{"iopub.status.busy":"2022-08-02T08:35:54.776829Z","iopub.execute_input":"2022-08-02T08:35:54.777402Z","iopub.status.idle":"2022-08-02T08:35:54.784936Z","shell.execute_reply.started":"2022-08-02T08:35:54.777368Z","shell.execute_reply":"2022-08-02T08:35:54.784114Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"train_df.shape, test_df.shape","metadata":{"execution":{"iopub.status.busy":"2022-08-02T08:35:54.786028Z","iopub.execute_input":"2022-08-02T08:35:54.786995Z","iopub.status.idle":"2022-08-02T08:35:54.801993Z","shell.execute_reply.started":"2022-08-02T08:35:54.786962Z","shell.execute_reply":"2022-08-02T08:35:54.801303Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"__Creating new feature extracting from existing__","metadata":{}},{"cell_type":"markdown","source":"Before dropping Name feature, We want to create a new feature called Titile from analyzing the existing Name feature.","metadata":{}},{"cell_type":"markdown","source":"* __Observation__\n    * Survival among Title brands varies slightly\n    * Certain titles (Rare) has low survival rate\n    ","metadata":{}},{"cell_type":"markdown","source":"* __Decisions__\n    * Consider new Title feature for model training","metadata":{}},{"cell_type":"code","source":"for dataset in combine:\n    dataset['Title'] = dataset.Name.str.extract('([A-Za-z]+)\\.', expand=False)\n\npd.crosstab(train_df['Title'], train_df['Sex'])","metadata":{"execution":{"iopub.status.busy":"2022-08-02T08:35:54.802855Z","iopub.execute_input":"2022-08-02T08:35:54.803380Z","iopub.status.idle":"2022-08-02T08:35:54.834971Z","shell.execute_reply.started":"2022-08-02T08:35:54.803354Z","shell.execute_reply":"2022-08-02T08:35:54.833940Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"We can replace less common titles with common titles or mark them as Rare","metadata":{}},{"cell_type":"code","source":"for dataset in combine:\n    dataset['Title'] = dataset['Title'].replace(['Lady', 'Countess', 'Capt', 'Col', 'Don', 'Dr', 'Major', 'Rev', 'Sir', 'Jonkheer', 'Done'], 'Rare')\n    dataset['Title'] = dataset['Title'].replace('Mlle', 'Miss')\n    dataset['Title'] = dataset['Title'].replace('Ms', 'Miss')\n    dataset['Title'] = dataset['Title'].replace('Mme', 'Mrs')\n\ntrain_df[['Title', 'Survived']].groupby(['Title'], as_index=False).mean()","metadata":{"execution":{"iopub.status.busy":"2022-08-02T08:35:54.838122Z","iopub.execute_input":"2022-08-02T08:35:54.838462Z","iopub.status.idle":"2022-08-02T08:35:54.863499Z","shell.execute_reply.started":"2022-08-02T08:35:54.838431Z","shell.execute_reply":"2022-08-02T08:35:54.862787Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"We can convert the categorical titles to ordinal","metadata":{}},{"cell_type":"code","source":"title_mapping = {'Mr':1, 'Miss':2, 'Mrs':3, 'Master':4, 'Rare':5 }\nfor dataset in combine:\n    dataset['Title'] = dataset['Title'].map(title_mapping)\n    dataset['Title'] = dataset['Title'].fillna(0)\n\ntrain_df.head()","metadata":{"execution":{"iopub.status.busy":"2022-08-02T08:35:54.864487Z","iopub.execute_input":"2022-08-02T08:35:54.864938Z","iopub.status.idle":"2022-08-02T08:35:54.884059Z","shell.execute_reply.started":"2022-08-02T08:35:54.864909Z","shell.execute_reply":"2022-08-02T08:35:54.883170Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"test_df.head()","metadata":{"execution":{"iopub.status.busy":"2022-08-02T08:35:54.884992Z","iopub.execute_input":"2022-08-02T08:35:54.885242Z","iopub.status.idle":"2022-08-02T08:35:54.903091Z","shell.execute_reply.started":"2022-08-02T08:35:54.885219Z","shell.execute_reply":"2022-08-02T08:35:54.901800Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"It seems the Title feature in test_df is a float value. So lets conver it into int","metadata":{}},{"cell_type":"code","source":"test_df['Title'] = test_df['Title'].astype(int)\ntest_df.head()","metadata":{"execution":{"iopub.status.busy":"2022-08-02T08:35:54.905032Z","iopub.execute_input":"2022-08-02T08:35:54.905890Z","iopub.status.idle":"2022-08-02T08:35:54.924476Z","shell.execute_reply.started":"2022-08-02T08:35:54.905855Z","shell.execute_reply":"2022-08-02T08:35:54.923503Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Now we can safely drop the Name feature from training and testing datasets. We also do not need the PassangerId feature in training set","metadata":{}},{"cell_type":"code","source":"train_df = train_df.drop(['Name', 'PassengerId'], axis=1)\ntest_df =  test_df.drop(['Name'], axis=1)\ncombine = [train_df, test_df]\ntrain_df.shape, test_df.shape","metadata":{"execution":{"iopub.status.busy":"2022-08-02T08:35:54.927348Z","iopub.execute_input":"2022-08-02T08:35:54.927661Z","iopub.status.idle":"2022-08-02T08:35:54.936338Z","shell.execute_reply.started":"2022-08-02T08:35:54.927635Z","shell.execute_reply":"2022-08-02T08:35:54.935501Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"__Converting a categorical feature__","metadata":{}},{"cell_type":"markdown","source":"We can convert features which contain string, to numerical values. This is necessary step for model training.","metadata":{}},{"cell_type":"code","source":"for dataset in combine:\n    dataset['Sex'] = dataset['Sex'].map({'female':0, 'male':1}).astype(int)\n\ntrain_df.head()","metadata":{"execution":{"iopub.status.busy":"2022-08-02T08:35:54.937485Z","iopub.execute_input":"2022-08-02T08:35:54.938209Z","iopub.status.idle":"2022-08-02T08:35:54.954217Z","shell.execute_reply.started":"2022-08-02T08:35:54.938184Z","shell.execute_reply":"2022-08-02T08:35:54.953543Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"__Completing a numerical continuous feature__","metadata":{}},{"cell_type":"markdown","source":"Now we should fill null empty values. There are several ways doing that","metadata":{}},{"cell_type":"markdown","source":"1. Generate random value between mean and standard deviation.\n2. Guessing missing values by using other correlated features.(Gender, Pclass)\n3. Combining 1 and 2 methods and use random value between mean and std base on set of Pclass and Gender combinations","metadata":{}},{"cell_type":"markdown","source":"Method 1 and 3 will introduce random noise into our models. So results might vary. We will use method 2","metadata":{}},{"cell_type":"code","source":"grid = sns.FacetGrid(train_df, row='Pclass', col='Sex', size=2.2, aspect=1.6)\ngrid.map(plt.hist, 'Age', alpha=.5, bins=20)\ngrid.add_legend()","metadata":{"execution":{"iopub.status.busy":"2022-08-02T08:35:54.959284Z","iopub.execute_input":"2022-08-02T08:35:54.960005Z","iopub.status.idle":"2022-08-02T08:35:55.985401Z","shell.execute_reply.started":"2022-08-02T08:35:54.959979Z","shell.execute_reply":"2022-08-02T08:35:55.984295Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Lets make a empty array contain guessed age based on Pclass and Sex","metadata":{}},{"cell_type":"code","source":"guess_ages = np.zeros((2,3))\nguess_ages","metadata":{"execution":{"iopub.status.busy":"2022-08-02T08:35:55.986561Z","iopub.execute_input":"2022-08-02T08:35:55.986913Z","iopub.status.idle":"2022-08-02T08:35:55.993284Z","shell.execute_reply.started":"2022-08-02T08:35:55.986885Z","shell.execute_reply":"2022-08-02T08:35:55.992375Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Now iterate over Sex(0,1) and Pclass(1,2,3) to calculate guessed values of Age for 6 combinations","metadata":{}},{"cell_type":"code","source":"for dataset in combine:\n    for s in range(0,2):\n        for p in range(0,3):\n            guess_df = dataset[(dataset['Sex'] == s) & (dataset['Pclass'] == p+1)]['Age'].dropna()\n            age_guess = guess_df.median()\n            guess_ages[s,p] = int(age_guess/0.5+0.5)*0.5\n\n    for s in range(0,2):\n        for p in range(0,3):\n            dataset.loc[(dataset.Age.isnull()) & (dataset.Sex == s) & (dataset.Pclass == p+1), 'Age'] = guess_ages[s,p]\n\n    dataset['Age'] = dataset['Age'].astype(int)\ntrain_df.head()\n","metadata":{"execution":{"iopub.status.busy":"2022-08-02T08:35:55.994660Z","iopub.execute_input":"2022-08-02T08:35:55.995006Z","iopub.status.idle":"2022-08-02T08:35:56.037663Z","shell.execute_reply.started":"2022-08-02T08:35:55.994974Z","shell.execute_reply":"2022-08-02T08:35:56.036679Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Let's create a age brands and determine correlation with survival","metadata":{}},{"cell_type":"code","source":"train_df['AgeBand'] = pd.cut(train_df['Age'], 5)\ntrain_df[['AgeBand', 'Survived']].groupby(['AgeBand'], as_index=False).mean().sort_values(by='AgeBand', ascending=True)","metadata":{"execution":{"iopub.status.busy":"2022-08-02T08:35:56.039139Z","iopub.execute_input":"2022-08-02T08:35:56.039494Z","iopub.status.idle":"2022-08-02T08:35:56.062365Z","shell.execute_reply.started":"2022-08-02T08:35:56.039461Z","shell.execute_reply":"2022-08-02T08:35:56.061239Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Let us replace Age with ordinal based on these bands.","metadata":{}},{"cell_type":"code","source":"for dataset in combine:\n    dataset.loc[dataset['Age'] <= 16, 'Age'] = 0\n    dataset.loc[(dataset['Age'] > 16) & (dataset['Age'] <= 32), 'Age'] = 1\n    dataset.loc[(dataset['Age'] > 32) & (dataset['Age'] <= 48), 'Age'] = 2\n    dataset.loc[(dataset['Age'] > 48) & (dataset['Age'] <= 64), 'Age'] = 3\n    dataset.loc[dataset['Age'] > 64, 'Age'] = 4\ntrain_df.head()","metadata":{"execution":{"iopub.status.busy":"2022-08-02T08:35:56.065106Z","iopub.execute_input":"2022-08-02T08:35:56.065375Z","iopub.status.idle":"2022-08-02T08:35:56.089458Z","shell.execute_reply.started":"2022-08-02T08:35:56.065350Z","shell.execute_reply":"2022-08-02T08:35:56.088416Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"We can remove the AgeBand feature.","metadata":{}},{"cell_type":"code","source":"train_df = train_df.drop(['AgeBand'], axis=1)\ncombine = [train_df, test_df]\ntrain_df.head()","metadata":{"execution":{"iopub.status.busy":"2022-08-02T08:35:56.090529Z","iopub.execute_input":"2022-08-02T08:35:56.090817Z","iopub.status.idle":"2022-08-02T08:35:56.102900Z","shell.execute_reply.started":"2022-08-02T08:35:56.090792Z","shell.execute_reply":"2022-08-02T08:35:56.102074Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"__Create new feature combining existing features__","metadata":{}},{"cell_type":"markdown","source":"We can create a new feature called FamilySize by combining Parch and SibSp features","metadata":{}},{"cell_type":"code","source":"for dataset in combine:\n    dataset['FamilySize'] = dataset['SibSp'] + dataset['Parch'] +1\ntrain_df[['FamilySize', 'Survived']].groupby(['FamilySize'], as_index=False).mean().sort_values(by='Survived', ascending=False)","metadata":{"execution":{"iopub.status.busy":"2022-08-02T08:35:56.104004Z","iopub.execute_input":"2022-08-02T08:35:56.104610Z","iopub.status.idle":"2022-08-02T08:35:56.122609Z","shell.execute_reply.started":"2022-08-02T08:35:56.104583Z","shell.execute_reply":"2022-08-02T08:35:56.121623Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"We can create another feature called IsAlone","metadata":{}},{"cell_type":"code","source":"for dataset in combine:\n    dataset['IsAlone'] = 0\n    dataset.loc[dataset['FamilySize'] == 1, 'IsAlone'] = 1\n    \ntrain_df[['IsAlone', 'Survived']].groupby(['IsAlone'], as_index=False).mean()","metadata":{"execution":{"iopub.status.busy":"2022-08-02T08:35:56.123622Z","iopub.execute_input":"2022-08-02T08:35:56.124262Z","iopub.status.idle":"2022-08-02T08:35:56.137916Z","shell.execute_reply.started":"2022-08-02T08:35:56.124236Z","shell.execute_reply":"2022-08-02T08:35:56.136452Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Lets drop Parch, SibSp and FamilySize features in favor of IsAlone","metadata":{}},{"cell_type":"code","source":"train_df = train_df.drop(['Parch', 'SibSp', 'FamilySize'], axis=1)\ntest_df = test_df.drop(['Parch', 'SibSp', 'FamilySize'], axis=1)\ncombine = [train_df, test_df]\ntrain_df.head()","metadata":{"execution":{"iopub.status.busy":"2022-08-02T08:35:56.138926Z","iopub.execute_input":"2022-08-02T08:35:56.139400Z","iopub.status.idle":"2022-08-02T08:35:56.151948Z","shell.execute_reply.started":"2022-08-02T08:35:56.139375Z","shell.execute_reply":"2022-08-02T08:35:56.150817Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"We can also create an artificial feature combining Pclass and Age.","metadata":{}},{"cell_type":"code","source":"for dataset in combine:\n    dataset['Age*Class'] = dataset.Age * dataset.Pclass\n\ntrain_df.loc[:, ['Age*Class', 'Age', 'Pclass']].head(10)","metadata":{"execution":{"iopub.status.busy":"2022-08-02T08:35:56.153533Z","iopub.execute_input":"2022-08-02T08:35:56.153922Z","iopub.status.idle":"2022-08-02T08:35:56.165806Z","shell.execute_reply.started":"2022-08-02T08:35:56.153895Z","shell.execute_reply":"2022-08-02T08:35:56.164969Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"__Completing a categorical feature__","metadata":{}},{"cell_type":"markdown","source":"Our training dataset has two missing values.We simply fill these with the most common occurance","metadata":{}},{"cell_type":"code","source":"freq_port = train_df.Embarked.dropna().mode()[0]\nfreq_port","metadata":{"execution":{"iopub.status.busy":"2022-08-02T08:35:56.166821Z","iopub.execute_input":"2022-08-02T08:35:56.167610Z","iopub.status.idle":"2022-08-02T08:35:56.174531Z","shell.execute_reply.started":"2022-08-02T08:35:56.167578Z","shell.execute_reply":"2022-08-02T08:35:56.173322Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"for dataset in combine:\n    dataset['Embarked'] = dataset['Embarked'].fillna(freq_port)\n\ntrain_df[['Embarked', 'Survived']].groupby(['Embarked'], as_index=False).mean().sort_values(by='Survived', ascending=False)","metadata":{"execution":{"iopub.status.busy":"2022-08-02T08:35:56.175517Z","iopub.execute_input":"2022-08-02T08:35:56.175957Z","iopub.status.idle":"2022-08-02T08:35:56.190346Z","shell.execute_reply.started":"2022-08-02T08:35:56.175926Z","shell.execute_reply":"2022-08-02T08:35:56.189772Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"__Convert categorical features to numeric__","metadata":{}},{"cell_type":"markdown","source":"We can convert categorical values in Embarked feature to numeric","metadata":{}},{"cell_type":"code","source":"for dataset in combine:\n    dataset['Embarked'] = dataset['Embarked'].map({'S':0, 'C':1, 'Q':2}).astype(int)\n\ntrain_df.head()","metadata":{"execution":{"iopub.status.busy":"2022-08-02T08:35:56.191208Z","iopub.execute_input":"2022-08-02T08:35:56.192961Z","iopub.status.idle":"2022-08-02T08:35:56.213667Z","shell.execute_reply.started":"2022-08-02T08:35:56.192928Z","shell.execute_reply":"2022-08-02T08:35:56.212794Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"__Completing and converting a numerica feature__","metadata":{}},{"cell_type":"markdown","source":"We can fill missing values in fare feature using mode to get the value that occurs most frequently for this feature","metadata":{}},{"cell_type":"code","source":"test_df['Fare'].fillna(test_df['Fare'].dropna().median(), inplace=True)\ntest_df.isnull().sum()","metadata":{"execution":{"iopub.status.busy":"2022-08-02T08:35:56.215055Z","iopub.execute_input":"2022-08-02T08:35:56.215781Z","iopub.status.idle":"2022-08-02T08:35:56.229125Z","shell.execute_reply.started":"2022-08-02T08:35:56.215723Z","shell.execute_reply":"2022-08-02T08:35:56.228113Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Lets create a fare band","metadata":{}},{"cell_type":"code","source":"train_df['FareBand'] = pd.qcut(train_df['Fare'], 4)\ntrain_df[['FareBand', 'Survived']].groupby(['FareBand'], as_index=False).mean().sort_values(by='FareBand', ascending=True)","metadata":{"execution":{"iopub.status.busy":"2022-08-02T08:35:56.230112Z","iopub.execute_input":"2022-08-02T08:35:56.230951Z","iopub.status.idle":"2022-08-02T08:35:56.246965Z","shell.execute_reply.started":"2022-08-02T08:35:56.230918Z","shell.execute_reply":"2022-08-02T08:35:56.246273Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Convert fare feature to ordinal values based on FareBand","metadata":{}},{"cell_type":"code","source":"for dataset in combine:\n    dataset.loc[dataset['Fare'] <= 7.91, 'Fare'] = 0\n    dataset.loc[(dataset['Fare'] > 7.91) & (dataset['Fare'] <= 14.454), 'Fare'] = 1\n    dataset.loc[(dataset['Fare'] > 14.454) & (dataset['Fare'] <= 31), 'Fare'] = 2\n    dataset.loc[dataset['Fare'] > 31, 'Fare'] = 3\n    dataset['Fare'] = dataset['Fare'].astype(int)\n\ntrain_df = train_df.drop(['FareBand'], axis=1)\ncombine = [train_df, test_df]\n\ntrain_df.head(10)","metadata":{"execution":{"iopub.status.busy":"2022-08-02T08:35:56.248138Z","iopub.execute_input":"2022-08-02T08:35:56.248560Z","iopub.status.idle":"2022-08-02T08:35:56.268689Z","shell.execute_reply.started":"2022-08-02T08:35:56.248535Z","shell.execute_reply":"2022-08-02T08:35:56.267914Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"test_df.head(10)","metadata":{"execution":{"iopub.status.busy":"2022-08-02T08:35:56.271236Z","iopub.execute_input":"2022-08-02T08:35:56.272268Z","iopub.status.idle":"2022-08-02T08:35:56.284595Z","shell.execute_reply.started":"2022-08-02T08:35:56.272195Z","shell.execute_reply":"2022-08-02T08:35:56.283630Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"#### __Model train, predict and slove__","metadata":{}},{"cell_type":"markdown","source":"Now our dataset is looking good and we have to train the model. Then we can use the model to slove the problem solution. There are many model algorithms to use. But our problem is classification and reggression problem in supervised learning. So we can use these models.","metadata":{}},{"cell_type":"markdown","source":"* Logistic Reggression\n* KNN\n* Support vector mask\n* Naive bayes classifier\n* Decision tree\n* Random forest\n* Preception\n* Artificial neural network\n* RVM","metadata":{}},{"cell_type":"markdown","source":"Now we should categorize our data into train and test data","metadata":{}},{"cell_type":"code","source":"X_train = train_df.drop('Survived', axis=1)\ny_train = train_df['Survived']\nX_test = test_df.drop('PassengerId', axis=1)\nX_train.shape, y_train.shape, X_test.shape","metadata":{"execution":{"iopub.status.busy":"2022-08-02T08:35:56.285717Z","iopub.execute_input":"2022-08-02T08:35:56.286247Z","iopub.status.idle":"2022-08-02T08:35:56.298165Z","shell.execute_reply.started":"2022-08-02T08:35:56.286220Z","shell.execute_reply":"2022-08-02T08:35:56.297111Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Lets we check each model algorithm one by one. And check the score","metadata":{}},{"cell_type":"code","source":"#  Lodistic regression\nlogreg = LogisticRegression()\nlogreg.fit(X_train, y_train)\nprint(round(logreg.score(X_train, y_train)*100,2),'%')","metadata":{"execution":{"iopub.status.busy":"2022-08-02T08:35:56.299447Z","iopub.execute_input":"2022-08-02T08:35:56.299962Z","iopub.status.idle":"2022-08-02T08:35:56.320247Z","shell.execute_reply.started":"2022-08-02T08:35:56.299929Z","shell.execute_reply":"2022-08-02T08:35:56.319370Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# K Nearest Neighbors\nknn = KNeighborsClassifier(n_neighbors=7)\nknn.fit(X_train, y_train)\nprint(round(knn.score(X_train, y_train)*100,2),'%')","metadata":{"execution":{"iopub.status.busy":"2022-08-02T08:35:56.321315Z","iopub.execute_input":"2022-08-02T08:35:56.321567Z","iopub.status.idle":"2022-08-02T08:35:56.353584Z","shell.execute_reply.started":"2022-08-02T08:35:56.321542Z","shell.execute_reply":"2022-08-02T08:35:56.352695Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Support vector machines\nsvc = SVC()\nsvc.fit(X_train, y_train)\nprint(round(svc.score(X_train, y_train)*100,2),'%')","metadata":{"execution":{"iopub.status.busy":"2022-08-02T08:35:56.355391Z","iopub.execute_input":"2022-08-02T08:35:56.356161Z","iopub.status.idle":"2022-08-02T08:35:56.395607Z","shell.execute_reply.started":"2022-08-02T08:35:56.356135Z","shell.execute_reply":"2022-08-02T08:35:56.395023Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Gaussian Naive bayes\ngaussian = GaussianNB()\ngaussian.fit(X_train, y_train)\nprint(round(gaussian.score(X_train, y_train)*100,2),'%')","metadata":{"execution":{"iopub.status.busy":"2022-08-02T08:35:56.396624Z","iopub.execute_input":"2022-08-02T08:35:56.397044Z","iopub.status.idle":"2022-08-02T08:35:56.406164Z","shell.execute_reply.started":"2022-08-02T08:35:56.397019Z","shell.execute_reply":"2022-08-02T08:35:56.405196Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Perceptron\nperceptron = Perceptron()\nperceptron.fit(X_train, y_train)\nprint(round(perceptron.score(X_train, y_train)*100,2),'%')","metadata":{"execution":{"iopub.status.busy":"2022-08-02T08:35:56.407132Z","iopub.execute_input":"2022-08-02T08:35:56.407835Z","iopub.status.idle":"2022-08-02T08:35:56.418061Z","shell.execute_reply.started":"2022-08-02T08:35:56.407803Z","shell.execute_reply":"2022-08-02T08:35:56.416809Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Linear SVC\nlinearsvc = LinearSVC()\nlinearsvc.fit(X_train, y_train)\nprint(round(linearsvc.score(X_train, y_train)*100,2),'%')","metadata":{"execution":{"iopub.status.busy":"2022-08-02T08:35:56.419247Z","iopub.execute_input":"2022-08-02T08:35:56.420112Z","iopub.status.idle":"2022-08-02T08:35:56.477655Z","shell.execute_reply.started":"2022-08-02T08:35:56.420081Z","shell.execute_reply":"2022-08-02T08:35:56.477008Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Stochastic Gradient Descent\nsgd = SGDClassifier()\nsgd.fit(X_train, y_train)\nprint(round(sgd.score(X_train, y_train)*100,2),'%')","metadata":{"execution":{"iopub.status.busy":"2022-08-02T08:35:56.478709Z","iopub.execute_input":"2022-08-02T08:35:56.479146Z","iopub.status.idle":"2022-08-02T08:35:56.490539Z","shell.execute_reply.started":"2022-08-02T08:35:56.479121Z","shell.execute_reply":"2022-08-02T08:35:56.489724Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Decision Tree\ndecisiontr = DecisionTreeClassifier()\ndecisiontr.fit(X_train, y_train)\nprint(round(decisiontr.score(X_train, y_train)*100,2),'%')","metadata":{"execution":{"iopub.status.busy":"2022-08-02T08:35:56.491465Z","iopub.execute_input":"2022-08-02T08:35:56.492971Z","iopub.status.idle":"2022-08-02T08:35:56.502860Z","shell.execute_reply.started":"2022-08-02T08:35:56.492938Z","shell.execute_reply":"2022-08-02T08:35:56.501972Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Random forest\nrndforest = RandomForestClassifier(n_estimators=100)\nrndforest.fit(X_train, y_train)\nprint(round(rndforest.score(X_train, y_train)*100,2),'%')","metadata":{"execution":{"iopub.status.busy":"2022-08-02T08:35:56.504056Z","iopub.execute_input":"2022-08-02T08:35:56.504590Z","iopub.status.idle":"2022-08-02T08:35:56.665251Z","shell.execute_reply.started":"2022-08-02T08:35:56.504565Z","shell.execute_reply":"2022-08-02T08:35:56.664343Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"#### __Model tuning__","metadata":{}},{"cell_type":"markdown","source":"Lets try these different models with different parameters and find the best algorithm for our solution","metadata":{}},{"cell_type":"code","source":"# define a dictionary for models and their parameters\nmodel_param = {\n    'svc': {\n        'model': SVC(gamma='auto'),\n        'params' : {\n            'C': [1,10,20],\n            'kernel': ['rbf','linear']\n        }  \n    },\n    'random_forest': {\n        'model': RandomForestClassifier(),\n        'params' : {\n            'n_estimators': [1,5,10]\n        }\n    },\n    'logistic_regression' : {\n        'model': LogisticRegression(solver='liblinear',multi_class='auto'),\n        'params': {\n            'C': [1,5,10]\n        }\n    },\n    'gaussian' :{\n        'model' : GaussianNB(),\n        'params' : {\n            \n        }\n    },\n    'knn' : {\n        'model' : KNeighborsClassifier(),\n        'params' : {\n            'n_neighbors' : [1,3,5,7,9]\n        }\n    },\n    'tree' : {\n        'model' : DecisionTreeClassifier(),\n        'params' : {\n            'criterion': ['gini','entropy'],\n        }\n    },\n    'perceptron' : {\n        'model' : Perceptron(),\n        'params' : {\n            'penalty' : ['l2','l1','elasticnet']\n        }\n    },\n    'linearsvc' : {\n        'model' : LinearSVC(),\n        'params' : {\n                 \n        }\n    },\n    'sgd' : {\n        'model' : SGDClassifier(),\n        'params' : {\n\n        }\n    }\n}","metadata":{"execution":{"iopub.status.busy":"2022-08-02T08:35:56.666259Z","iopub.execute_input":"2022-08-02T08:35:56.666514Z","iopub.status.idle":"2022-08-02T08:35:56.674289Z","shell.execute_reply.started":"2022-08-02T08:35:56.666491Z","shell.execute_reply":"2022-08-02T08:35:56.673462Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"#find the best model\nscores = []\nfor model_name, mp in model_param.items():\n  clf = GridSearchCV(mp['model'],mp['params'],cv=3,return_train_score=False)\n  clf.fit(X_train, y_train)\n  scores.append({\n      'model' : model_name,\n      'best_score' : clf.best_score_,\n      'best_params' : clf.best_params_\n  })\ndf = pd.DataFrame(scores,columns=['model','best_score','best_params'])\ndf","metadata":{"execution":{"iopub.status.busy":"2022-08-02T08:35:56.675234Z","iopub.execute_input":"2022-08-02T08:35:56.675472Z","iopub.status.idle":"2022-08-02T08:35:57.573095Z","shell.execute_reply.started":"2022-08-02T08:35:56.675449Z","shell.execute_reply":"2022-08-02T08:35:57.572253Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"#### __Model training and predict__","metadata":{}},{"cell_type":"markdown","source":"So according to above chart we can see SVC is the best scored model with these parameters. So let's use that to train our model.","metadata":{}},{"cell_type":"code","source":"model = SVC(C=20, kernel='rbf')\nmodel.fit(X_train, y_train)\ny_pred = model.predict(X_test)","metadata":{"execution":{"iopub.status.busy":"2022-08-02T08:35:57.574170Z","iopub.execute_input":"2022-08-02T08:35:57.574409Z","iopub.status.idle":"2022-08-02T08:35:57.606461Z","shell.execute_reply.started":"2022-08-02T08:35:57.574386Z","shell.execute_reply":"2022-08-02T08:35:57.605800Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"__Submission__","metadata":{}},{"cell_type":"markdown","source":"Lets do our submission","metadata":{}},{"cell_type":"code","source":"submission = pd.DataFrame({\n    'PassengerId':test_df['PassengerId'],\n    'Survived':y_pred\n})\nsubmission.to_csv('submission.csv', index=False)","metadata":{"execution":{"iopub.status.busy":"2022-08-02T08:35:57.607452Z","iopub.execute_input":"2022-08-02T08:35:57.608258Z","iopub.status.idle":"2022-08-02T08:35:57.615023Z","shell.execute_reply.started":"2022-08-02T08:35:57.608231Z","shell.execute_reply":"2022-08-02T08:35:57.614302Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"submission.head()","metadata":{"execution":{"iopub.status.busy":"2022-08-02T08:35:57.616071Z","iopub.execute_input":"2022-08-02T08:35:57.616961Z","iopub.status.idle":"2022-08-02T08:35:57.626557Z","shell.execute_reply.started":"2022-08-02T08:35:57.616926Z","shell.execute_reply":"2022-08-02T08:35:57.625707Z"},"trusted":true},"execution_count":null,"outputs":[]}]}