{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"pygments_lexer":"ipython3","nbconvert_exporter":"python","version":"3.6.4","file_extension":".py","codemirror_mode":{"name":"ipython","version":3},"name":"python","mimetype":"text/x-python"}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"code","source":"import numpy as np # linear algebra\nimport pandas as pd # data processing, CSV file I/O (e.g. pd.read_csv)\n\nimport os\nfor dirname, _, filenames in os.walk('/kaggle/input'):\n    for filename in filenames:\n        print(os.path.join(dirname, filename))","metadata":{"_uuid":"8f2839f25d086af736a60e9eeb907d3b93b6e0e5","_cell_guid":"b1076dfc-b9ad-4769-8c92-a6c4dae69d19","execution":{"iopub.status.busy":"2022-08-04T11:01:02.977173Z","iopub.execute_input":"2022-08-04T11:01:02.977930Z","iopub.status.idle":"2022-08-04T11:01:02.987861Z","shell.execute_reply.started":"2022-08-04T11:01:02.977894Z","shell.execute_reply":"2022-08-04T11:01:02.986404Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## ****Titanic EDA With Python & Applying Logistic Regression****\n\nWe will be working with the Titanic Dataset from Kaggle. This is a very famous dataset and often is a student's first step in machine learning.\n\nWe'll be trying to predict a classification - survival or deceased. Let's begin our understanding of implementing logistic regression in Python for classification.","metadata":{}},{"cell_type":"markdown","source":"## Import Libraries\n\nLet's import some libraries to get started!","metadata":{}},{"cell_type":"code","source":"import matplotlib.pyplot as plt\nimport seaborn as sns\n%matplotlib inline","metadata":{"execution":{"iopub.status.busy":"2022-08-04T11:01:02.991526Z","iopub.execute_input":"2022-08-04T11:01:02.993846Z","iopub.status.idle":"2022-08-04T11:01:03.000250Z","shell.execute_reply.started":"2022-08-04T11:01:02.993810Z","shell.execute_reply":"2022-08-04T11:01:02.999025Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## The Data\n\nLet's start by reading the titanic_train.csv into a pandas dataframe","metadata":{}},{"cell_type":"code","source":"train = pd.read_csv(\"/kaggle/input/titanic-dataset-semicleaned/titanic_train.csv\")","metadata":{"execution":{"iopub.status.busy":"2022-08-04T11:01:03.002406Z","iopub.execute_input":"2022-08-04T11:01:03.002777Z","iopub.status.idle":"2022-08-04T11:01:03.016911Z","shell.execute_reply.started":"2022-08-04T11:01:03.002746Z","shell.execute_reply":"2022-08-04T11:01:03.015800Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"train.head()","metadata":{"execution":{"iopub.status.busy":"2022-08-04T11:01:03.055588Z","iopub.execute_input":"2022-08-04T11:01:03.056033Z","iopub.status.idle":"2022-08-04T11:01:03.073147Z","shell.execute_reply.started":"2022-08-04T11:01:03.055997Z","shell.execute_reply":"2022-08-04T11:01:03.072165Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## Exploratory Data Analysis\n\nLet's begin some exploratory data analysis! We'll start by checking out missing data!","metadata":{}},{"cell_type":"markdown","source":"### Missing Data\n\nWe can use seaborn to create a simple heatmap to see where we are missing data","metadata":{}},{"cell_type":"code","source":"train.isnull()","metadata":{"execution":{"iopub.status.busy":"2022-08-04T11:01:03.108797Z","iopub.execute_input":"2022-08-04T11:01:03.109172Z","iopub.status.idle":"2022-08-04T11:01:03.135696Z","shell.execute_reply.started":"2022-08-04T11:01:03.109141Z","shell.execute_reply":"2022-08-04T11:01:03.134313Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"sns.heatmap(train.isnull(), yticklabels=False, cbar=False)","metadata":{"execution":{"iopub.status.busy":"2022-08-04T11:01:03.137949Z","iopub.execute_input":"2022-08-04T11:01:03.138406Z","iopub.status.idle":"2022-08-04T11:01:03.382430Z","shell.execute_reply.started":"2022-08-04T11:01:03.138361Z","shell.execute_reply":"2022-08-04T11:01:03.381659Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Roughly 20% of the Age data is missing. The proportion of Age data missing is likely small enough for reasonable replacement with some form of imputation. Looking at the Cabin column, it looks like we are missing too much of that data to do something useful with at a basic level. We'll probably drop this later, or change it to another feature, like \"Cabin Known: 1 or 0\" \n\nLet's continue by visualizing some more of the data.","metadata":{}},{"cell_type":"code","source":"sns.set_style('whitegrid')\nsns.countplot(x='Survived', data=train)","metadata":{"execution":{"iopub.status.busy":"2022-08-04T11:01:03.383936Z","iopub.execute_input":"2022-08-04T11:01:03.384248Z","iopub.status.idle":"2022-08-04T11:01:03.567385Z","shell.execute_reply.started":"2022-08-04T11:01:03.384218Z","shell.execute_reply":"2022-08-04T11:01:03.566362Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"sns.set_style('whitegrid')\nsns.countplot(x=\"Survived\", hue='Sex', data=train, palette='RdBu_r')","metadata":{"execution":{"iopub.status.busy":"2022-08-04T11:01:03.570421Z","iopub.execute_input":"2022-08-04T11:01:03.570897Z","iopub.status.idle":"2022-08-04T11:01:03.777369Z","shell.execute_reply.started":"2022-08-04T11:01:03.570855Z","shell.execute_reply":"2022-08-04T11:01:03.776377Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"sns.set_style('whitegrid')\nsns.countplot(x='Survived', hue='Pclass', data=train, palette='rainbow')","metadata":{"execution":{"iopub.status.busy":"2022-08-04T11:01:03.778799Z","iopub.execute_input":"2022-08-04T11:01:03.779145Z","iopub.status.idle":"2022-08-04T11:01:04.005291Z","shell.execute_reply.started":"2022-08-04T11:01:03.779113Z","shell.execute_reply":"2022-08-04T11:01:04.004051Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"sns.histplot(train['Age'].dropna(), color='darkred', bins=40)","metadata":{"execution":{"iopub.status.busy":"2022-08-04T11:01:04.006846Z","iopub.execute_input":"2022-08-04T11:01:04.007181Z","iopub.status.idle":"2022-08-04T11:01:04.339943Z","shell.execute_reply.started":"2022-08-04T11:01:04.007151Z","shell.execute_reply":"2022-08-04T11:01:04.338847Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"sns.countplot(x='SibSp', data=train)","metadata":{"execution":{"iopub.status.busy":"2022-08-04T11:01:04.343455Z","iopub.execute_input":"2022-08-04T11:01:04.343819Z","iopub.status.idle":"2022-08-04T11:01:04.574924Z","shell.execute_reply.started":"2022-08-04T11:01:04.343786Z","shell.execute_reply":"2022-08-04T11:01:04.573779Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"sns.set_style('whitegrid')\nsns.histplot(x='Fare', data=train, color='green', bins=40)\ntrain['Fare'].hist(color='green', bins=40, figsize=(8,4))","metadata":{"execution":{"iopub.status.busy":"2022-08-04T11:01:04.576700Z","iopub.execute_input":"2022-08-04T11:01:04.577382Z","iopub.status.idle":"2022-08-04T11:01:04.965354Z","shell.execute_reply.started":"2022-08-04T11:01:04.577337Z","shell.execute_reply":"2022-08-04T11:01:04.964278Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## Data Cleaning\n\nWe want to fill in missing age data instead of just dropping the missing age data rows. One way to do this is by filling the mean age of all the passengers(imputation). However, we can be smarter about this and check the average age by passenger class. For example:","metadata":{}},{"cell_type":"code","source":"plt.figure(figsize=(12, 7))\nsns.boxplot(x='Pclass', y='Age', data=train, palette='winter')","metadata":{"execution":{"iopub.status.busy":"2022-08-04T11:01:04.970280Z","iopub.execute_input":"2022-08-04T11:01:04.970633Z","iopub.status.idle":"2022-08-04T11:01:05.219350Z","shell.execute_reply.started":"2022-08-04T11:01:04.970584Z","shell.execute_reply":"2022-08-04T11:01:05.218260Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"We can see that the wealthier passengers in the higher classes tend to be older, which makes sense. We'll use these average values to impute based on Pclass for Age.","metadata":{}},{"cell_type":"code","source":" def impute_age(cols):\n        Age = cols[0]\n        Pclass = cols[1]\n        \n        if pd.isnull(Age):\n            \n            if Pclass == 1:\n                return 37\n                \n            elif Pclass == 2:\n                return 29\n                \n            elif Pclass == 3:\n                return 24\n        \n        else:\n            return Age","metadata":{"execution":{"iopub.status.busy":"2022-08-04T11:01:05.221179Z","iopub.execute_input":"2022-08-04T11:01:05.221648Z","iopub.status.idle":"2022-08-04T11:01:05.228856Z","shell.execute_reply.started":"2022-08-04T11:01:05.221575Z","shell.execute_reply":"2022-08-04T11:01:05.227443Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"train['Age'] = train[['Age', 'Pclass']].apply(impute_age, axis=1)","metadata":{"execution":{"iopub.status.busy":"2022-08-04T11:01:05.230848Z","iopub.execute_input":"2022-08-04T11:01:05.231249Z","iopub.status.idle":"2022-08-04T11:01:05.251380Z","shell.execute_reply.started":"2022-08-04T11:01:05.231218Z","shell.execute_reply":"2022-08-04T11:01:05.250632Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Now, let's check the heatmap again!","metadata":{}},{"cell_type":"code","source":"sns.heatmap(data=train.isnull(), cmap='viridis', cbar=False, yticklabels=False)","metadata":{"execution":{"iopub.status.busy":"2022-08-04T11:01:05.253721Z","iopub.execute_input":"2022-08-04T11:01:05.254060Z","iopub.status.idle":"2022-08-04T11:01:05.497638Z","shell.execute_reply.started":"2022-08-04T11:01:05.254029Z","shell.execute_reply":"2022-08-04T11:01:05.496670Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Great! Let's go ahead and drop the Cabin column and the row in Embarked that is NaN.","metadata":{}},{"cell_type":"code","source":"train.drop('Cabin', axis=1, inplace=True)","metadata":{"execution":{"iopub.status.busy":"2022-08-04T11:01:05.499244Z","iopub.execute_input":"2022-08-04T11:01:05.499699Z","iopub.status.idle":"2022-08-04T11:01:05.506443Z","shell.execute_reply.started":"2022-08-04T11:01:05.499655Z","shell.execute_reply":"2022-08-04T11:01:05.505160Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"train.head()","metadata":{"execution":{"iopub.status.busy":"2022-08-04T11:01:05.507989Z","iopub.execute_input":"2022-08-04T11:01:05.508415Z","iopub.status.idle":"2022-08-04T11:01:05.527736Z","shell.execute_reply.started":"2022-08-04T11:01:05.508365Z","shell.execute_reply":"2022-08-04T11:01:05.526847Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"train.dropna(inplace=True)","metadata":{"execution":{"iopub.status.busy":"2022-08-04T11:01:05.529018Z","iopub.execute_input":"2022-08-04T11:01:05.529408Z","iopub.status.idle":"2022-08-04T11:01:05.539486Z","shell.execute_reply.started":"2022-08-04T11:01:05.529327Z","shell.execute_reply":"2022-08-04T11:01:05.538453Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## Converting Categorical Features\n\nWe'll need to convert categorical features to dummy variables using pandas! Otherwise our machine learning algorithm won't be able to directly take in those features as input.","metadata":{}},{"cell_type":"code","source":"train.info()","metadata":{"execution":{"iopub.status.busy":"2022-08-04T11:01:05.540877Z","iopub.execute_input":"2022-08-04T11:01:05.541319Z","iopub.status.idle":"2022-08-04T11:01:05.558434Z","shell.execute_reply.started":"2022-08-04T11:01:05.541273Z","shell.execute_reply":"2022-08-04T11:01:05.557680Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"pd.get_dummies(train['Embarked'], drop_first=True).head()","metadata":{"execution":{"iopub.status.busy":"2022-08-04T11:01:05.559329Z","iopub.execute_input":"2022-08-04T11:01:05.559650Z","iopub.status.idle":"2022-08-04T11:01:05.574228Z","shell.execute_reply.started":"2022-08-04T11:01:05.559595Z","shell.execute_reply":"2022-08-04T11:01:05.573064Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"sex = pd.get_dummies(train['Sex'], drop_first=True)\nembark = pd.get_dummies(train['Embarked'], drop_first=True)","metadata":{"execution":{"iopub.status.busy":"2022-08-04T11:01:05.575416Z","iopub.execute_input":"2022-08-04T11:01:05.575857Z","iopub.status.idle":"2022-08-04T11:01:05.586565Z","shell.execute_reply.started":"2022-08-04T11:01:05.575813Z","shell.execute_reply":"2022-08-04T11:01:05.585457Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"train.drop(['Sex', 'Embarked', 'Name', 'Ticket'], axis=1, inplace=True)","metadata":{"execution":{"iopub.status.busy":"2022-08-04T11:01:05.588368Z","iopub.execute_input":"2022-08-04T11:01:05.588722Z","iopub.status.idle":"2022-08-04T11:01:05.595459Z","shell.execute_reply.started":"2022-08-04T11:01:05.588692Z","shell.execute_reply":"2022-08-04T11:01:05.594682Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"train.head()","metadata":{"execution":{"iopub.status.busy":"2022-08-04T11:01:05.597077Z","iopub.execute_input":"2022-08-04T11:01:05.597789Z","iopub.status.idle":"2022-08-04T11:01:05.614102Z","shell.execute_reply.started":"2022-08-04T11:01:05.597754Z","shell.execute_reply":"2022-08-04T11:01:05.613050Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"train = pd.concat([train, sex, embark], axis=1)","metadata":{"execution":{"iopub.status.busy":"2022-08-04T11:01:05.615412Z","iopub.execute_input":"2022-08-04T11:01:05.615969Z","iopub.status.idle":"2022-08-04T11:01:05.954093Z","shell.execute_reply.started":"2022-08-04T11:01:05.615938Z","shell.execute_reply":"2022-08-04T11:01:05.953057Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"train.head()","metadata":{"execution":{"iopub.status.busy":"2022-08-04T11:01:05.955399Z","iopub.execute_input":"2022-08-04T11:01:05.956603Z","iopub.status.idle":"2022-08-04T11:01:05.973480Z","shell.execute_reply.started":"2022-08-04T11:01:05.956556Z","shell.execute_reply":"2022-08-04T11:01:05.972405Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Great! Our data is ready for our model.","metadata":{}},{"cell_type":"markdown","source":"## Building Our Logistic Regression Model\n\nLet's start by splitting our data into a training set and test set \n\n### Train Test Split","metadata":{}},{"cell_type":"code","source":"train.drop('Survived', axis=1).head()","metadata":{"execution":{"iopub.status.busy":"2022-08-04T11:01:05.974832Z","iopub.execute_input":"2022-08-04T11:01:05.975152Z","iopub.status.idle":"2022-08-04T11:01:05.993725Z","shell.execute_reply.started":"2022-08-04T11:01:05.975124Z","shell.execute_reply":"2022-08-04T11:01:05.992686Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"train['Survived'].head()","metadata":{"execution":{"iopub.status.busy":"2022-08-04T11:01:05.997552Z","iopub.execute_input":"2022-08-04T11:01:05.997896Z","iopub.status.idle":"2022-08-04T11:01:06.005435Z","shell.execute_reply.started":"2022-08-04T11:01:05.997866Z","shell.execute_reply":"2022-08-04T11:01:06.004109Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"from sklearn.model_selection import train_test_split\n\nX_train, X_test, y_train, y_test = train_test_split(train.drop('Survived', axis=1),\n                                                   train['Survived'], test_size=0.30,\n                                                   random_state=101)","metadata":{"execution":{"iopub.status.busy":"2022-08-04T11:01:06.007439Z","iopub.execute_input":"2022-08-04T11:01:06.008255Z","iopub.status.idle":"2022-08-04T11:01:06.018156Z","shell.execute_reply.started":"2022-08-04T11:01:06.008210Z","shell.execute_reply":"2022-08-04T11:01:06.017172Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### Training and Predicting","metadata":{}},{"cell_type":"code","source":"from sklearn.linear_model import LogisticRegression","metadata":{"execution":{"iopub.status.busy":"2022-08-04T11:01:06.020562Z","iopub.execute_input":"2022-08-04T11:01:06.020952Z","iopub.status.idle":"2022-08-04T11:01:06.029017Z","shell.execute_reply.started":"2022-08-04T11:01:06.020914Z","shell.execute_reply":"2022-08-04T11:01:06.027955Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"logmodel = LogisticRegression(max_iter=600)\nlogmodel.fit(X_train, y_train)","metadata":{"execution":{"iopub.status.busy":"2022-08-04T11:03:26.250069Z","iopub.execute_input":"2022-08-04T11:03:26.250907Z","iopub.status.idle":"2022-08-04T11:03:26.403544Z","shell.execute_reply.started":"2022-08-04T11:03:26.250867Z","shell.execute_reply":"2022-08-04T11:03:26.402352Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"predictions = logmodel.predict(X_test)","metadata":{"execution":{"iopub.status.busy":"2022-08-04T11:03:28.154229Z","iopub.execute_input":"2022-08-04T11:03:28.155196Z","iopub.status.idle":"2022-08-04T11:03:28.161580Z","shell.execute_reply.started":"2022-08-04T11:03:28.155155Z","shell.execute_reply":"2022-08-04T11:03:28.160722Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"from sklearn.metrics import confusion_matrix, accuracy_score\n\naccuracy = confusion_matrix(y_test, predictions)\naccuracy","metadata":{"execution":{"iopub.status.busy":"2022-08-04T11:03:57.029892Z","iopub.execute_input":"2022-08-04T11:03:57.030318Z","iopub.status.idle":"2022-08-04T11:03:57.039448Z","shell.execute_reply.started":"2022-08-04T11:03:57.030282Z","shell.execute_reply":"2022-08-04T11:03:57.037885Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"accuracy_perc = accuracy_score(y_test, predictions)\nprint(accuracy_perc)","metadata":{"execution":{"iopub.status.busy":"2022-08-04T11:04:32.553566Z","iopub.execute_input":"2022-08-04T11:04:32.554014Z","iopub.status.idle":"2022-08-04T11:04:32.560503Z","shell.execute_reply.started":"2022-08-04T11:04:32.553977Z","shell.execute_reply":"2022-08-04T11:04:32.559660Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"predictions","metadata":{"execution":{"iopub.status.busy":"2022-08-04T11:04:56.596606Z","iopub.execute_input":"2022-08-04T11:04:56.596983Z","iopub.status.idle":"2022-08-04T11:04:56.604014Z","shell.execute_reply.started":"2022-08-04T11:04:56.596954Z","shell.execute_reply":"2022-08-04T11:04:56.603205Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"","metadata":{},"execution_count":null,"outputs":[]}]}