{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"pygments_lexer":"ipython3","nbconvert_exporter":"python","version":"3.6.4","file_extension":".py","codemirror_mode":{"name":"ipython","version":3},"name":"python","mimetype":"text/x-python"}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"### Goal:\n\nThe goal is for understanding machine learning and the CRISP-DM process\n1. Business Understanding\n2. Data Understanding / COllection\n3. Data preparation\n4. Modelling\n5. Evaluation\n6. Deployment","metadata":{}},{"cell_type":"markdown","source":"### Business Understanding:\n- What are we doing: Predict the deaths / survived\n- What we have: A structured tabular data for learning and exploring machine learning. ","metadata":{}},{"cell_type":"markdown","source":"### Data Understanding / Collection:","metadata":{}},{"cell_type":"code","source":"# importing libraries\nimport pandas as pd\nimport numpy as np\nimport matplotlib.pyplot as plt\nimport seaborn as sns\nfrom sklearn.preprocessing import LabelEncoder\nfrom sklearn.model_selection import train_test_split\nfrom sklearn.neighbors import KNeighborsClassifier\nfrom sklearn.metrics import accuracy_score","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# we have to use the train.csv file for exploring and analyzing our model\ntrain_df = pd.read_csv('/kaggle/input/titanic/train.csv')","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# reading top few rows\ntrain_df.head(5)","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"train_df.info()","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"1. We have 11 columns. Both numeric and character.\n2. ID columns will not be useful so we can remove it in the training data or we can make them index column. \n3. We can remove the name, ticket columns too. The fare column might be useful to get the class like first, second and third class but we already have Pclass so we can remove fare column. \n3. Survived column is going to be our target feature that we are going to predict\n4. SubSp, Parch. Well, number of siblings and parents will not be useful here as we have dataset of whole. We will just keep for exploring purposes. \n5. Embarked --> might be useful for analysis but not prediction. If we use embarked then it's causation, not correlation. (Please check the difference between causation and correlation here - https://www.abs.gov.au/websitedbs/D3310114.nsf/home/statistical+language+-+correlation+and+causation","metadata":{}},{"cell_type":"code","source":"# removing the unwanted columns\ncolumns_remove = ['Name', 'Ticket', 'Cabin', 'Fare']\ntrain_df = train_df.drop(labels = columns_remove, axis = 1) # axis = 0 means remove rows, axis = 1 means remove columns\ntrain_df.head(5)","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Now we have few non-numerical data. \n- Check for missing values first and if missing values exists - REMOVE THEM !!\n- we need to convert the non-numerical data into numerical data. (Pclass and Embarked). \n    - before that we will do some data analysis\n    - we have 2 options here. we can use label encoder / one hot encoding. Please research about the label encoder / one hot encoding. \n- The data Age is continuous. Some machine learning algorithms need it in discrete values. We might need to use the concept of binning. ","metadata":{}},{"cell_type":"code","source":"train_df.isnull().sum()","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Age have 177 missing values. We can either ignore them or fill them but mostly we will remove them in real world.","metadata":{}},{"cell_type":"code","source":"train_df = train_df.dropna(axis = 0)\ntrain_df.isnull().sum()","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"print('\\nMinimum age', train_df['Age'].min())\nprint('Maximum age', train_df['Age'].max())\nprint('\\n')\nsns.boxplot(x = train_df['Survived'], y = train_df['Age'])","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# let's plot the Age and see which Age group affected the most.\n# reference: https://seaborn.pydata.org/generated/seaborn.boxplot.html\n# Analyzing gender\nprint('Number of Men and Women:')\nprint(train_df['Sex'].value_counts())\nprint('\\n')\nsns.boxplot(x = train_df['Age'], y = train_df['Sex'])","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"There are some outliers. In real world application we will remove this outliers. We won't remove to reduce complexity here now. ","metadata":{}},{"cell_type":"code","source":"# Analyzing embarked now. \n# for embarked it makes sense to use the bar chart. \nsns.barplot(x = train_df['Embarked'], y = train_df['Survived'])","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### Data Preparation\n1. Converting all the categorical data into numerical data. ","metadata":{}},{"cell_type":"code","source":"# removing the passenger ID column\ntrain_df = train_df.drop('PassengerId', axis = 1)","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# converting the Sex, Embarked into numerical values. \n# We will use label encoding here to reduce the complexity. In real world, we will usually use one-hot encoder\nle = LabelEncoder()\ntrain_df['Sex'] = le.fit_transform(train_df['Sex'])\ntrain_df['Embarked'] = le.fit_transform(train_df['Embarked'])\ntrain_df.head(5)","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# taking the input and output features\ntarget_feature = train_df[['Survived']]\ninput_features = train_df.drop(target_feature, axis = 1)","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Now taking 80% of training data for model training and 20% of training data for model evaluation\n# this is called hold out method\n# there are other methods - K-Fold Cross Validation etc. \nX_train, X_test, y_train, y_test = train_test_split(input_features, target_feature, test_size = 0.2, random_state = 42)\nprint(X_train.shape)\nprint(X_test.shape)\nprint(y_train.shape)\nprint(y_test.shape)","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### Modelling\nWe can use a simple KNN classification model for this. Its easy to understand and get to know machine learning algorithms. (But generally in this kind of discrete dataset, we will generally use logistic regression model which uses a logit function)","metadata":{}},{"cell_type":"code","source":"classifier = KNeighborsClassifier(n_neighbors = 5, metric = 'minkowski')\n# Teaching the model to learn the input and output data.\nclassifier.fit(X_train, y_train.values.ravel())","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### Evaluation","metadata":{}},{"cell_type":"code","source":"# predicting the 20% test data for accuracy\npredictions = classifier.predict(X_test)\npredictions","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# As this is a classification problem, we can check the accuracy, confusion matrix, f1-score, precision, recall statistics\n# for now, we will check the accuracy\naccuracy_score(y_true = y_test, y_pred = predictions)","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"71 % accuracy. Not a bad score for first time ! Good luck. ","metadata":{}},{"cell_type":"markdown","source":"### Deployment\nDeploy it to Kaggle (Optional for now) (omit)","metadata":{}}]}