{"cells":[{"metadata":{"trusted":true,"_uuid":"177d858cca6f1c6d7815077205d493b1f846dec0"},"cell_type":"markdown","source":"Hi, this is my first kernel in kaggle. I wanted to make a walkthrough on how I approach a machine learning problem. I try to make this as clear as possible by explaining the reasoning behind each step. Note however that I'm not an expert in machine learning modelling and I made a lot of assumptions in my analysis based on my current understanding. I appreciate any corrections or suggestions for this kernel."},{"metadata":{"trusted":true,"_uuid":"721c63c5e7a7aacb488cb280a1fb870dfd671f40"},"cell_type":"code","source":"import warnings\nwarnings.filterwarnings('ignore')\n\nimport pandas as pd\nimport numpy as np\n\nimport matplotlib.pyplot as plt\nimport seaborn as sns\n\nfrom sklearn.model_selection import StratifiedKFold\nimport lightgbm as lgb","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"f6dad1a9a1f3d467c553ed945498a2d8481e6c67"},"cell_type":"markdown","source":"# Introduction\nThe no free lunch theorem states that there is no single best method to solve all problem in machine learning. To get the best accuracy in a problem will require a unique solution. Trying to use every techniques to solve a problem will lead to overfitting and unoptimal accuracy. Therefore, it is very important to approach a machine learning problem systematically. One needs to start with a basic model, the baseline, then slowly applying more techniques to the model while monitoring the evaluation metric. By doing so, we can evaluate which methods work and which doesnt by looking at their impact to the baseline accuracy. \n\nIn this kernel, I will demonstrate how I approach the Titanic classification problem by improving from the baseline. In general, there are many techniques that can be applied in the machine learning, this range from feature engineering, model selection, hyperparameter tuning, etc. For now, I will only use feature engineering to try improving the baseline accuracy. For the model I will use lightgbm."},{"metadata":{"_uuid":"098424d1373d76ada525fa751189b2de3323d776"},"cell_type":"markdown","source":"# Data loading and a quick EDA"},{"metadata":{"_uuid":"9a0488884f00667dbb569dc6df1aeb7f74d01b0d"},"cell_type":"markdown","source":"The idea of this kernel is to show how to increase the validation score systematically and I am not going to submit the test result in the end so I will only need the train data. Let's read it first"},{"metadata":{"trusted":true,"_uuid":"b771a2eb4e3f1590f866c68f8fbd88e4df2740c2"},"cell_type":"code","source":"df_train = pd.read_csv('../input/train.csv')\ndf_train.head()","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"a4af6c11299c47724890e2bdb4c0e1b12bd50a16"},"cell_type":"markdown","source":"Our first objective is to build a baseline model, to do that we are going to do a quick EDA. The aim for this EDA is to answer following basic questions:\n1. How big is the data set?\n2. How many missing values in the data set?\n3. How many \"easy\" features do we have?\n\nWhat I mean by \"easy\" feature is the features that do not need too much data preprocessing, for example: numerical features or categorical features with low cardinality.\n\nTo get the basic idea about the dataset we can use dataframe info"},{"metadata":{"_kg_hide-input":false,"trusted":true,"scrolled":true,"_uuid":"5ff4c2216e54b66112a1e2f23ed3a3d243565c3b"},"cell_type":"code","source":"df_train.info()","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"aacf2f90cf444ba39f47ab25b8edd5645d7be807"},"cell_type":"markdown","source":"We can see here that we have 891 entries with number and object data types. We can also see that some features, namely \"Age\", \"Cabin\", and \"Embarked\" have missing values since they have less than 891 non null data. We will take all numerical features with no missing values as our baseline features. \n\nFeatures with \"object\" data types are usually categorical type, we need to determine the cardinality of these features. We do this by looking at the number of unique values in these features."},{"metadata":{"trusted":true,"_uuid":"09c7cdac69ca30d84e5313fd748afbd4acaf2664"},"cell_type":"code","source":"for c in df_train.select_dtypes('object').columns:\n    print('='*75)\n    print('column \"{}\" has {} unique values and {} missing values'.format(c,len(df_train[c].drop_duplicates()),\n                                                          df_train[c].isnull().sum()))\n    print(df_train[c].value_counts().head(5))","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"ea36fa1ca34ba1c2736eec7b2a4302f6f3060b32"},"cell_type":"markdown","source":"As we can see, \"Name\", \"Ticket\" and \"Cabin\" has a lot of unique values so they are \"hard\" features.\n\nAccording to the data description, \"Pclass\" is the passenger class which is a categorical data type. Let's make sure that it is not a \"hard\" feature:"},{"metadata":{"trusted":true,"_uuid":"d7b4991335d55c9f166dcbd4af05b16e3bcf9b41"},"cell_type":"code","source":"df_train['Pclass'].value_counts()","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"c69dfe4b24861f67dfb331d7beadf7f9c71e2680"},"cell_type":"markdown","source":"It has 3 unique values so it has low cardinality.\n\nTo conclude, we will take the following \"easy\" features as our baseline features:\n\n- Pclass\n- Sex\n- SibSp\n- Parch\n- Fare\n- Embarked\n"},{"metadata":{"_uuid":"43182e61aae6d86a3acc303789e2b0e2b67f7674"},"cell_type":"markdown","source":"# A model function"},{"metadata":{"_uuid":"ff38c678c114c2014b4aa057d337079f9c408a2e"},"cell_type":"markdown","source":"We are going to train the model several times so it is helpful to create a function to make the code shorter. The function is given in the following lines. It is quite long so I provide comments to explain how it works. Basically the function will take a dataframe as an input and train an lgbm model with a 3-fold cross validation. The output is the average training  and validation binary errors with plots of feature importance of the features."},{"metadata":{"trusted":true,"_uuid":"a8918f1c4467ec60619fb057e642e5022cbdbba9"},"cell_type":"code","source":"def model(df):\n    # get the columns for independent variables (X) and target (y)\n    X = df.drop(['PassengerId','Survived'],axis=1).values\n    y = df['Survived'].values\n\n    # this variable store the column names that will be used later to plot the feature importance\n    fcol = df.drop(['PassengerId','Survived'],axis=1).columns\n\n    # Create Stratified 3 fold split, make sure the random seed is set to make the result reproducible\n    skf = StratifiedKFold(n_splits=3,shuffle=True,random_state=666)\n\n    # This is parameters for LGBM, use binary error as a metric and set the random seed\n    param = {\n             'objective':'binary',\n             \"metric\": 'binary_error',\n             \"random_state\": 4590\n    }\n\n    # These variables is used to store the scores and feature importance\n    evals = {}\n    avg_tr_score = 0\n    avg_val_score = 0\n    df_feats = pd.DataFrame()\n\n    for fold, (trn_idx, val_idx) in enumerate(skf.split(X,y)):\n        print('='*75)\n\n        # load the data as lgb dataset\n        trn_data = lgb.Dataset(X[trn_idx], label=y[trn_idx])\n        val_data = lgb.Dataset(X[val_idx], label=y[val_idx])\n\n        # train the model, I set early_stopping_rounds = 500 to make sure we get the optimal score, it's probably an overkill though\n        num_round = 10000\n        clf = lgb.train(param, trn_data, num_round, valid_sets = [trn_data, val_data], \n                        verbose_eval=100, early_stopping_rounds = 500,evals_result = evals)\n\n        # add the scores for this fold to the final score\n        avg_tr_score += evals['training']['binary_error'][clf.best_iteration]/skf.n_splits\n        avg_val_score += evals['valid_1']['binary_error'][clf.best_iteration]/skf.n_splits\n\n        # add the feature importance data for this fold to the final feature importance dataframe\n        # Note: I get both the feature importance by 'split' and 'information gain'\n        df_feats_fold = pd.DataFrame()\n        df_feats_fold['feature'] = fcol\n        df_feats_fold['importance'] = clf.feature_importance(importance_type='split')\n        df_feats_fold['type'] = 'split'\n        df_feats_fold['fold'] = fold + 1\n        df_feats = pd.concat([df_feats_fold, df_feats], axis=0)\n\n        df_feats_fold = pd.DataFrame()\n        df_feats_fold['feature'] = fcol\n        df_feats_fold['importance'] = clf.feature_importance(importance_type='gain')\n        df_feats_fold['type'] = 'gain'\n        df_feats_fold['fold'] = fold + 1\n        df_feats = pd.concat([df_feats_fold, df_feats], axis=0)\n\n    print('='*75)    \n    print('Average train score: {:.3f} Average validation score: {:.3f}'.format(avg_tr_score,avg_val_score))\n    \n    # Plot the feature importances\n    fig, (ax1,ax2) = plt.subplots(1,2,figsize=(15,10))\n    sns.barplot(x='importance',y='feature',\n                data=df_feats[df_feats['type']=='split'].sort_values('importance',ascending=False),ax=ax1)\n    sns.barplot(x='importance',y='feature',\n                data=df_feats[df_feats['type']=='gain'].sort_values('importance',ascending=False),ax=ax2)\n\n    ax1.set_title('Feature importance by split')\n    ax2.set_title('Feature importance by gain')\n\n    plt.tight_layout()\n    plt.show()","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"4be19ece5771b964f37e2112310d8d80729a7d2d"},"cell_type":"markdown","source":"# Baseline Modeling"},{"metadata":{"_uuid":"abf607c88d7cfaf897556b384a1cdc2791a78273"},"cell_type":"markdown","source":"We can now start training our baseline model. First we define the clean dataset:"},{"metadata":{"trusted":true,"_uuid":"73168d2a8b3fccf72429b3495f2694d3b809eda2"},"cell_type":"code","source":"df_train_clean = df_train[['PassengerId','Survived','Pclass','Sex','SibSp','Parch','Fare','Embarked']]\ndf_train_clean.head()","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"22cba83dfd864fcd383433c4084fb916a190043d"},"cell_type":"markdown","source":"Next, we label encode all the categorical features. Label encoding works fine for tree based model like lgbm, you might need to use one hot encoding if you use parametric models. We also fill the missing values in Embarked with its mode ('S')."},{"metadata":{"trusted":true,"_uuid":"73095a6d16241b7de6c947149e6f4a3e6f4f3b13"},"cell_type":"code","source":"df_train_clean['Sex'] = np.where(df_train_clean['Sex']=='male',1,0).astype('int64')\n\nm = {'S':1,'C':2,'Q':3}\ndf_train_clean['Embarked'] = df_train_clean['Embarked'].map(m)\ndf_train_clean['Embarked'] = df_train_clean['Embarked'].fillna(1).astype('int64')\ndf_train_clean.head()","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"c5679698b603a94d976158fb78049cd542812a45"},"cell_type":"markdown","source":"With all the features converted to numbers we can now start training our model."},{"metadata":{"trusted":true,"_uuid":"174982d989bfe08d5c12efbbca1874362819dc9d"},"cell_type":"code","source":"clf = model(df_train_clean)","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"627cfb5fd1f73bff55a1197cc303392f1f7aaa22"},"cell_type":"markdown","source":"Our 3-fold validation gives average training score = 0.142 and validation score = 0.180 (the lower the better). We will use the validation score as the main metric to evaluate our model. \n\nThe plots are the feature importance plots by split (left) and information gain (right). The feature importance by split shows the features ranked by the total number of splits in the trained model. The feature importance by information gain shows the features ranked by the total information gained in the trained model.\n\nLet us try to analyze the plots to gain some insights. First, the split of \"Fare\" is very high compared to the rest of the features. This means that the model has a lot of split points using feature \"Fare\". One of the reason is because \"Fare\" is a continuous numerical feature. We can study this further by looking at the distribution of \"Fare\"."},{"metadata":{"trusted":true,"_uuid":"ca29970a3ccb2c08d619a7f32f662bd03442bd37"},"cell_type":"code","source":"sns.distplot(df_train_clean['Fare'])","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"7f4b269d9240d3960cfe138411f8affe3b95c1ae"},"cell_type":"markdown","source":"As we can see, the distribution is far from normal and it has a lot of bumps which can means multiple segments. It also has a very long tail (outlier) which can create problems to our model. So to gain information from this feature, the model has to split it up a lot of times.\n\nNext, we look at the feature importance by information gain. We see here, the top three features that give information gain to our model is \"Sex\", \"Fare\" and \"Pclass\". \"Sex\" and \"Pclass\" features can gain a lot information with low splits so they are good features. \"Fare\", however, needs a lot of split to gain half the information of \"Sex\". Maybe if we reduce the splits in \"Feature\" we can make our model generalize better to reduce overfitting and improve our score. We will do that in the next section."},{"metadata":{"_uuid":"bc2e60c17090f7f13203be2f2a6c9f76d8ac7334"},"cell_type":"markdown","source":"# Fare binning"},{"metadata":{"_uuid":"6bd118b753112cc59de598076c48be3af1617e07"},"cell_type":"markdown","source":"As we have seen before, the distribution of \"Fare\" is somewhat complicated. In this part we will try to smooth it out and see whether this will effect our score. We do this by putting the values of \"Fare\" into small bins, this will reduce the noise in the distribution and make it smoother. We demonstate below what happens to the distribution when we bin it into 100 evenly spaced bins."},{"metadata":{"trusted":true,"_uuid":"986750f43ecf33dd9c5770c5878994c8807676d8"},"cell_type":"code","source":"plt.figure(figsize=(15,5))\nsns.countplot(pd.cut(df_train['Fare'],100,labels=False))","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"100dff82e011d549afba2626390bd257c60507f7"},"cell_type":"markdown","source":"After binning, the distribution is converted into histogram.\n\nLet's restart the data preprocessing and see what happens to our model if we bin the \"Fare\" feature."},{"metadata":{"trusted":true,"_uuid":"19a553bd2fe468a0e86ff1f62fcbc04910c44034"},"cell_type":"code","source":"df_train_clean = df_train[['PassengerId','Survived','Pclass','Sex','SibSp','Parch','Fare','Embarked']]\ndf_train_clean.head()","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"c07d79699f37566dd8560866f0d2e988abde7f22"},"cell_type":"code","source":"df_train_clean['Sex'] = np.where(df_train_clean['Sex']=='male',1,0).astype('int64')\n\nm = {'S':1,'C':2,'Q':3}\ndf_train_clean['Embarked'] = df_train_clean['Embarked'].map(m)\ndf_train_clean['Embarked'] = df_train_clean['Embarked'].fillna(1).astype('int64')\ndf_train_clean.head()","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"ee255b91a3775b6fbd6a18612a500f5ca8975744"},"cell_type":"markdown","source":"This is part where we bin the \"Fare\". We further assume that bin > 21 are outliers so we collect them in one bin labelled 21 to make the distribution even simpler."},{"metadata":{"trusted":true,"_uuid":"b65ba028f89f990846714da4accda6574516fa98"},"cell_type":"code","source":"df_train_clean['Fare'] = pd.cut(df_train_clean['Fare'],100,labels=False)\ndf_train_clean['Fare'] = np.where(df_train_clean['Fare']>20,21,df_train_clean['Fare'])\ndf_train_clean.head()","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"94634c94275014a574fffc6ac90a9a016692778e"},"cell_type":"markdown","source":"This is what the \"Fare\" distribution looks like after the previous process."},{"metadata":{"trusted":true,"_uuid":"7447cc18b79a80c78e6c0e31d38540980c10b116"},"cell_type":"code","source":"plt.figure(figsize=(15,5))\nsns.countplot(df_train_clean['Fare'])","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"db872570c7ac3e64c958756a3e6a559b0f5e403c"},"cell_type":"markdown","source":"Notice the outliers are now grouped in bin 21. Now, let's train the model again."},{"metadata":{"trusted":true,"_uuid":"43bf87f22aa43d27336afe88e7d77bb23f266875"},"cell_type":"code","source":"model(df_train_clean)","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"891edccb07f75e15d34e2c9b38a0f372c7e53a60"},"cell_type":"markdown","source":"Unfortunately the validation score gets worse by 0.015 so smoothing the feature \"Fare\" is a bad idea. Before reverting back to the baseline and try a new method, let's analyze this and try to understand what happens.\n\nLet's look at the plots first. Previously, our aim was to reduce the splits in \"Fare\", we see from the plots that binning actually accomplished this. The left plot shows that the number of split in \"Fare\" has been reduced significantly, it is now almost on par with the other features in term of total split count. However, the information gain also significantly dropped, the feature \"Fare\" now gain less information than feature \"Pclass\". With less information gain, it becomes harder for the model to fit to the dataset.\n\nWhat actually happened to our base model is a case of underfitting. Therefore, our previous attempt to reduce overfitting makes the model perform worse. There are two ways to combat underfitting, by including more data or more features. We will do the latter in the next section."},{"metadata":{"_uuid":"1d90e6cc0d922a7efa24b5a6facdf87e4b3c07b9"},"cell_type":"markdown","source":"# Adding Age"},{"metadata":{"_uuid":"95780a0cd8e54ef27125e1c55a38e7d7b4ea955f"},"cell_type":"markdown","source":"Let's add one more feature to the model to reduce underfitting. We will add feature \"Age\", this is a numerical feature. This feature has a slight issue with missing values so we have to take care of it in the data preprocessing part.\n\nWe start from beginning as usual."},{"metadata":{"trusted":true,"_uuid":"910c7867af069b84b50bd619f876ee5dcce9d772"},"cell_type":"code","source":"df_train_clean = df_train[['PassengerId','Survived','Pclass','Sex','SibSp','Parch','Fare','Embarked','Age']]\ndf_train_clean.head()","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"e38c9f20b234f0e3df41b2eb286860d9fc5ca2e5"},"cell_type":"markdown","source":"We fill the missing values in \"Age\" with -1. This is basically tells the model that these values are something else. this technique works tree based model, for parametric model, usually we can fill the missing values with mean or median."},{"metadata":{"trusted":true,"_uuid":"3adaa111b76a95e99fe7665908383cf0c84efde6"},"cell_type":"code","source":"df_train_clean['Sex'] = np.where(df_train_clean['Sex']=='male',1,0).astype('int64')\n\nm = {'S':1,'C':2,'Q':3}\ndf_train_clean['Embarked'] = df_train_clean['Embarked'].map(m)\ndf_train_clean['Embarked'] = df_train_clean['Embarked'].fillna(1).astype('int64')\ndf_train_clean['Age'] = df_train_clean['Age'].fillna(-1)\ndf_train_clean.head()","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"30990441ecaa9e9cdf5766eda5fe7f90524eb0c3"},"cell_type":"markdown","source":"Now, train the model."},{"metadata":{"trusted":true,"_uuid":"ef27792af70234c354ff4d1d05594cf752377e6b"},"cell_type":"code","source":"model(df_train_clean)","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"9b43a71d3db37ae6b2b19708418fa08f53204145"},"cell_type":"markdown","source":"The validation score has improved by 0.006 from the base model. We see that \"Age\" is a good feature with information gain more than \"Pclass\". Therefore, adding \"Age\" works so we can keep it. Notice too, that age has a lot of splits because it is a continuous numerical feature. This might also imply that the \"Age\" has a complex distribution just like \"Fare\". Let's check it out."},{"metadata":{"trusted":true,"_uuid":"bee980cf619b726f216fbfd6e1a899ccee9ddd10"},"cell_type":"code","source":"sns.distplot(df_train_clean['Age'])","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"d95f3aaec919fe12aa81036a27ee46e021ce2006"},"cell_type":"markdown","source":"Not as complex as \"Fare\" but still far from a normal distribution."},{"metadata":{"_uuid":"a02755e9228ebdc5d2b2cb37279a52b579c37fd5"},"cell_type":"markdown","source":"# Adding feature interaction"},{"metadata":{"_uuid":"dd0c5e9fdd5da0c07286f6f80608512a0027f34b"},"cell_type":"markdown","source":"We can try adding more features to improve our score but this time we will try adding feature interaction. From the previous analysis we can see that the three worst performing features are \"SibSp\", \"Parch\" and \"Embark\". Maybe if we combine some of these worst performing features we can create an interaction feature that performs better that the total of individual features."},{"metadata":{"trusted":true,"_uuid":"3ffc330b92ba290bc65123fca3d00fb9af50d04c"},"cell_type":"code","source":"df_train_clean = df_train[['PassengerId','Survived','Pclass','Sex','SibSp','Parch','Fare','Embarked','Age']]\ndf_train_clean.head()","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"981436a8a957aec1a5ff9cbe9e1fa651250e081a"},"cell_type":"markdown","source":"From the dataset description, \"SibSp\" is the number of sibling/spouse and \"Parch\" is parent/children. We can add these two to make a new feature \"Family\"."},{"metadata":{"trusted":true,"_uuid":"3842227d8627aa32fef347634910826ee5ec93f5"},"cell_type":"code","source":"df_train_clean['Sex'] = np.where(df_train_clean['Sex']=='male',1,0).astype('int64')\ndf_train_clean['Family'] = df_train_clean['Parch'] + df_train_clean['SibSp']\n\nm = {'S':1,'C':2,'Q':3}\ndf_train_clean['Embarked'] = df_train_clean['Embarked'].map(m)\ndf_train_clean['Embarked'] = df_train_clean['Embarked'].fillna(1).astype('int64')\ndf_train_clean['Age'] = df_train_clean['Age'].fillna(-1)\ndf_train_clean.head()","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"95024d256a405d9fcb8954818b8d453c034bc26d"},"cell_type":"markdown","source":"Let's see what happens now to our model."},{"metadata":{"trusted":true,"_uuid":"b92bfc41aa9ccd632772e9e80e5c4c941694fbc1"},"cell_type":"code","source":"model(df_train_clean)","execution_count":null,"outputs":[]},{"metadata":{"trusted":false,"_uuid":"e37981f6d1393803e9a0fe42eb8dda23da491d53"},"cell_type":"markdown","source":"Adding \"Family\" has improved our validation score even further by 0.006 so this is a good feature to keep. We can also see that the feature \"Family\" has higher information gain than the total information gain from the \"SibSp\" and \"Parch\""},{"metadata":{"_uuid":"1a9b78127ef12a9e7833a9217d1178f961183d20"},"cell_type":"markdown","source":"# Conclusion\nTo further improve the score we can add more features but the rest of the features that we have are high cardinal features. Processing this type of feature will require more advanced techniques that I will not discuss here. For now, you should get the basic idea on how to improve the score from the baseline. Once you are satisfied with the features, you can proceed to the next step with hyperparameter tuning and model selection/stacking."},{"metadata":{"trusted":true,"_uuid":"2c1ad2605ab59b7a5178e5c10453201929f73b24"},"cell_type":"code","source":"","execution_count":null,"outputs":[]}],"metadata":{"kernelspec":{"display_name":"Python 3","language":"python","name":"python3"},"language_info":{"name":"python","version":"3.6.6","mimetype":"text/x-python","codemirror_mode":{"name":"ipython","version":3},"pygments_lexer":"ipython3","nbconvert_exporter":"python","file_extension":".py"}},"nbformat":4,"nbformat_minor":1}