{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"pygments_lexer":"ipython3","nbconvert_exporter":"python","version":"3.6.4","file_extension":".py","codemirror_mode":{"name":"ipython","version":3},"name":"python","mimetype":"text/x-python"}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"## Learn Overfitting and Underfitting with Titanic Dataset","metadata":{}},{"cell_type":"markdown","source":"Hello kagglers ..\n\nIn this notebook we will explain the terms **Overfitting** and **Underfitting**, Then we will apply what we learned on the Titanic Dataset and see the difference in scores.\n\nI hope this notebook will be useful to you ..  \nhappy learning :)\n\n**Note:** You can download the slides below and more learning resources from my LinkedIn account: [Here](https://www.linkedin.com/in/oday-mourad-97b126217)\n\n\n\n\n","metadata":{}},{"cell_type":"markdown","source":"#### Before beginning let's explain what Overfitting and Underfitting are.","metadata":{}},{"cell_type":"markdown","source":"\n<table><tr>\n<td> <img src=\"https://user-images.githubusercontent.com/50684002/179947170-61d217c8-9996-471e-8037-44673bbed61f.png?raw=true\" alt=\"Drawing\" style=\"width: 600px;\"/> </td>\n<td> <img src=\"https://user-images.githubusercontent.com/50684002/179947117-37e98d6e-717b-4025-8ca3-50005e2b6590.png?raw=true\" alt=\"Drawing\" style=\"width: 600px;\"/> </td>\n</tr></table>","metadata":{}},{"cell_type":"markdown","source":"##### from the image above we can see that:  \n\n\n##### **Underfitting:**    \nThe Model is too simple and can not fit the training data. In this case **training** and **cross-validation** errors (costs) will be **high**.  \n\n##### **Overfitting:**  \nThe Model is too complex and **can not generalize** to new examples. In this case **training** cost will be low and **cross-validation** cost will be **high**.  \n\n##### **Sweet Spot:**  \nThis is our target, The model have the ability to **learns** from the trainig exampes and **generalizes** to new examples. ","metadata":{}},{"cell_type":"markdown","source":"<div>\n<img src=\"https://user-images.githubusercontent.com/50684002/179947152-778c2fcd-ff9b-48ba-8779-98ea1353a08f.png?raw=true\" width=800/>\n</div>","metadata":{}},{"cell_type":"markdown","source":"##### from the image above:  \nLet's suppose that complexity of our model is the degree of polynomial. Our target is to find the point that has **minimum cross-validation error**.","metadata":{}},{"cell_type":"markdown","source":"#### Let's begin:","metadata":{}},{"cell_type":"code","source":"import numpy as np # linear algebra\nimport pandas as pd # data processing, CSV file I/O (e.g. pd.read_csv)\nimport matplotlib.pyplot as plt # data visualization\nimport seaborn as sns # data visualization\nimport os \n\nfor dirname, _, filenames in os.walk('/kaggle/input'):\n    for filename in filenames:\n        print(os.path.join(dirname, filename))","metadata":{"_uuid":"8f2839f25d086af736a60e9eeb907d3b93b6e0e5","_cell_guid":"b1076dfc-b9ad-4769-8c92-a6c4dae69d19","execution":{"iopub.status.busy":"2022-07-20T08:22:00.927624Z","iopub.execute_input":"2022-07-20T08:22:00.928099Z","iopub.status.idle":"2022-07-20T08:22:00.942890Z","shell.execute_reply.started":"2022-07-20T08:22:00.928062Z","shell.execute_reply":"2022-07-20T08:22:00.941183Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### Acquire data:","metadata":{}},{"cell_type":"code","source":"def read_data():\n    train_data = pd.read_csv(\"/kaggle/input/titanic/train.csv\")\n    print(\"Train data imported successfully!!\")\n    print(\"-\"*50)\n    test_data = pd.read_csv(\"/kaggle/input/titanic/test.csv\")\n    print(\"Test data imported successfully!!\")\n    return train_data , test_data","metadata":{"execution":{"iopub.status.busy":"2022-07-20T08:22:01.270906Z","iopub.execute_input":"2022-07-20T08:22:01.271760Z","iopub.status.idle":"2022-07-20T08:22:01.280043Z","shell.execute_reply.started":"2022-07-20T08:22:01.271701Z","shell.execute_reply":"2022-07-20T08:22:01.278694Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"train_data , test_data = read_data()","metadata":{"execution":{"iopub.status.busy":"2022-07-20T08:22:01.572158Z","iopub.execute_input":"2022-07-20T08:22:01.572918Z","iopub.status.idle":"2022-07-20T08:22:01.591819Z","shell.execute_reply.started":"2022-07-20T08:22:01.572874Z","shell.execute_reply":"2022-07-20T08:22:01.590565Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"#### We are going to get the **true solution** for testing on unseen data.","metadata":{}},{"cell_type":"markdown","source":"**Important Note:** We will not depend on it, Just for learning purposes.","metadata":{}},{"cell_type":"code","source":"# You can skip this cell \n\nimport re\nimport warnings\nimport io\nimport requests\n\nwarnings.filterwarnings(\"ignore\")\n\nurl=\"https://github.com/thisisjasonjafari/my-datascientise-handcode/raw/master/005-datavisualization/titanic.csv\"\ns=requests.get(url).content\nc=pd.read_csv(io.StringIO(s.decode('utf-8')))\ntest_data_with_labels = c\nfor i, name in enumerate(test_data_with_labels['name']):\n    if '\"' in name:\n        test_data_with_labels['name'][i] = re.sub('\"', '', name)\n        \nfor i, name in enumerate(test_data['Name']):\n    if '\"' in name:\n        test_data['Name'][i] = re.sub('\"', '', name)\n        \nsurvived = []\n\nfor name in test_data['Name']:\n    survived.append(int(test_data_with_labels.loc[test_data_with_labels['name'] == name]['survived'].values[-1]))\n    \ntrue_solution = pd.read_csv('../input/titanic/gender_submission.csv')\ntrue_solution['Survived'] = survived\ntrue_solution.head()","metadata":{"execution":{"iopub.status.busy":"2022-07-20T08:22:01.893243Z","iopub.execute_input":"2022-07-20T08:22:01.893749Z","iopub.status.idle":"2022-07-20T08:22:02.986178Z","shell.execute_reply.started":"2022-07-20T08:22:01.893709Z","shell.execute_reply":"2022-07-20T08:22:02.984958Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"##### In this notebook we will not focus on data analysis, **Just modeling.**","metadata":{}},{"cell_type":"code","source":"train_data.head()","metadata":{"execution":{"iopub.status.busy":"2022-07-20T08:22:02.988504Z","iopub.execute_input":"2022-07-20T08:22:02.988849Z","iopub.status.idle":"2022-07-20T08:22:03.007053Z","shell.execute_reply.started":"2022-07-20T08:22:02.988816Z","shell.execute_reply":"2022-07-20T08:22:03.005852Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"#### Filling Blanks and Missed Data:","metadata":{"execution":{"iopub.status.busy":"2022-07-20T06:39:05.363521Z","iopub.execute_input":"2022-07-20T06:39:05.363944Z","iopub.status.idle":"2022-07-20T06:39:05.369594Z","shell.execute_reply.started":"2022-07-20T06:39:05.363911Z","shell.execute_reply":"2022-07-20T06:39:05.368235Z"}}},{"cell_type":"code","source":"train_data.isnull().sum()","metadata":{"execution":{"iopub.status.busy":"2022-07-20T08:22:03.008658Z","iopub.execute_input":"2022-07-20T08:22:03.009059Z","iopub.status.idle":"2022-07-20T08:22:03.025990Z","shell.execute_reply.started":"2022-07-20T08:22:03.009027Z","shell.execute_reply":"2022-07-20T08:22:03.024846Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"test_data.isnull().sum()","metadata":{"execution":{"iopub.status.busy":"2022-07-20T08:22:03.028581Z","iopub.execute_input":"2022-07-20T08:22:03.029360Z","iopub.status.idle":"2022-07-20T08:22:03.038611Z","shell.execute_reply.started":"2022-07-20T08:22:03.029305Z","shell.execute_reply":"2022-07-20T08:22:03.037485Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"##### It's important to fill Age, Embarked, Fare features: ","metadata":{}},{"cell_type":"code","source":"train_data['Embarked'] = train_data.Embarked.fillna(train_data.Embarked.dropna().max())\ntest_data['Fare'] = test_data.Fare.fillna(test_data.Fare.dropna().mean())","metadata":{"execution":{"iopub.status.busy":"2022-07-20T08:22:03.040006Z","iopub.execute_input":"2022-07-20T08:22:03.040368Z","iopub.status.idle":"2022-07-20T08:22:03.053148Z","shell.execute_reply.started":"2022-07-20T08:22:03.040328Z","shell.execute_reply":"2022-07-20T08:22:03.051961Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"##### we will guess the age from Pclass and Sex:","metadata":{}},{"cell_type":"code","source":"guess_ages = np.zeros((2,3))\nguess_ages","metadata":{"execution":{"iopub.status.busy":"2022-07-20T08:22:03.054530Z","iopub.execute_input":"2022-07-20T08:22:03.054829Z","iopub.status.idle":"2022-07-20T08:22:03.068778Z","shell.execute_reply.started":"2022-07-20T08:22:03.054802Z","shell.execute_reply":"2022-07-20T08:22:03.067766Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Now we iterate over Sex (0 or 1) and Pclass (1, 2, 3) to calculate guessed values of Age for the six combinations.","metadata":{}},{"cell_type":"code","source":"combine = [train_data , test_data]\n\n# Converting Sex categories (male and female) to 0 and 1:\nfor dataset in combine:\n    dataset['Sex'] = dataset['Sex'].map( {'female': 1, 'male': 0} ).astype(int)\n\n# Filling missed age feature:\n\nfor dataset in combine:\n    for i in range(0, 2):\n        for j in range(0, 3):\n            guess_df = dataset[(dataset['Sex'] == i) & \\\n                                  (dataset['Pclass'] == j+1)]['Age'].dropna()\n            age_guess = guess_df.median()\n\n            # Convert random age float to nearest .5 age\n            guess_ages[i,j] = int( age_guess/0.5 + 0.5 ) * 0.5\n            \n    for i in range(0, 2):\n        for j in range(0, 3):\n            dataset.loc[ (dataset.Age.isnull()) & (dataset.Sex == i) & (dataset.Pclass == j+1),\\\n                    'Age'] = guess_ages[i,j]\n\n    dataset['Age'] = dataset['Age'].astype(int)\n\ntrain_data.head()","metadata":{"execution":{"iopub.status.busy":"2022-07-20T08:22:03.192224Z","iopub.execute_input":"2022-07-20T08:22:03.192663Z","iopub.status.idle":"2022-07-20T08:22:03.251717Z","shell.execute_reply.started":"2022-07-20T08:22:03.192629Z","shell.execute_reply":"2022-07-20T08:22:03.250883Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"train_data.isnull().sum()","metadata":{"execution":{"iopub.status.busy":"2022-07-20T08:22:03.390398Z","iopub.execute_input":"2022-07-20T08:22:03.390794Z","iopub.status.idle":"2022-07-20T08:22:03.400914Z","shell.execute_reply.started":"2022-07-20T08:22:03.390761Z","shell.execute_reply":"2022-07-20T08:22:03.399587Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"test_data.isnull().sum()","metadata":{"execution":{"iopub.status.busy":"2022-07-20T08:22:03.579081Z","iopub.execute_input":"2022-07-20T08:22:03.580088Z","iopub.status.idle":"2022-07-20T08:22:03.588611Z","shell.execute_reply.started":"2022-07-20T08:22:03.580049Z","shell.execute_reply":"2022-07-20T08:22:03.587867Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"##### There are no important missed data anymore!!!","metadata":{}},{"cell_type":"markdown","source":"### Modeling:","metadata":{"execution":{"iopub.status.busy":"2022-07-20T07:00:02.080654Z","iopub.execute_input":"2022-07-20T07:00:02.081752Z","iopub.status.idle":"2022-07-20T07:00:02.087938Z","shell.execute_reply.started":"2022-07-20T07:00:02.081692Z","shell.execute_reply":"2022-07-20T07:00:02.086151Z"}}},{"cell_type":"code","source":"# ========================================================= #\nfrom sklearn.metrics import accuracy_score\nfrom sklearn.model_selection import cross_val_score\n# ========================================================= #\nfrom sklearn.tree import DecisionTreeClassifier\n# ========================================================= #\nfrom colorama import Fore","metadata":{"execution":{"iopub.status.busy":"2022-07-20T08:22:03.738863Z","iopub.execute_input":"2022-07-20T08:22:03.739793Z","iopub.status.idle":"2022-07-20T08:22:03.745947Z","shell.execute_reply.started":"2022-07-20T08:22:03.739737Z","shell.execute_reply":"2022-07-20T08:22:03.745126Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# First we will select the features that we will use:\n\nfeatures = [\"Pclass\" , \"Sex\" , \"Age\" , \"SibSp\" , \"Parch\" , \"Fare\" , \"Embarked\"]\n\n# Categorical to indicator variables:\nX_train = pd.get_dummies(train_data[features])\nY_train = train_data[\"Survived\"]\nX_test = pd.get_dummies(test_data[features])\n\nX_train.head()","metadata":{"execution":{"iopub.status.busy":"2022-07-20T08:22:04.014476Z","iopub.execute_input":"2022-07-20T08:22:04.015129Z","iopub.status.idle":"2022-07-20T08:22:04.038746Z","shell.execute_reply.started":"2022-07-20T08:22:04.015091Z","shell.execute_reply":"2022-07-20T08:22:04.037907Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"def print_scores(model ,X_train , Y_train,predictions , cv_splites=10):\n    print(Fore.BLUE , \"The mean accuracy score of the train data is %.5f\" % model.score(X_train, Y_train))\n    CV_scores = cross_val_score(model, X_train, Y_train, cv=cv_splites)\n    print(Fore.BLACK ,\"The individual cross-validation scores are: \\n\",CV_scores)\n    print(Fore.BLACK ,\"The minimum cross-validation score is %.3f\" % min(CV_scores))\n    print(Fore.BLACK ,\"The maximum cross-validation score is %.3f\" % max(CV_scores))\n    print(Fore.YELLOW ,\"The mean  cross-validation   score is %.5f ± %0.2f\" % (CV_scores.mean(), CV_scores.std() * 2))\n    print(Fore.RED ,\"The test (i.e. leaderboard)  score is %.5f (this score is unknown)\" % accuracy_score(true_solution[\"Survived\"],predictions))","metadata":{"execution":{"iopub.status.busy":"2022-07-20T08:22:04.161586Z","iopub.execute_input":"2022-07-20T08:22:04.162446Z","iopub.status.idle":"2022-07-20T08:22:04.170203Z","shell.execute_reply.started":"2022-07-20T08:22:04.162409Z","shell.execute_reply":"2022-07-20T08:22:04.168934Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### 1 ) Underfitting:","metadata":{}},{"cell_type":"markdown","source":"Let's try to use **too simple** model:","metadata":{}},{"cell_type":"code","source":"model = DecisionTreeClassifier(max_depth=1 , max_features=2 ,random_state=7)\nmodel.fit(X_train, Y_train)\npredictions = model.predict(X_test)\nprint_scores(model, X_train, Y_train, predictions)","metadata":{"execution":{"iopub.status.busy":"2022-07-20T08:22:04.224812Z","iopub.execute_input":"2022-07-20T08:22:04.225191Z","iopub.status.idle":"2022-07-20T08:22:04.300665Z","shell.execute_reply.started":"2022-07-20T08:22:04.225159Z","shell.execute_reply":"2022-07-20T08:22:04.299389Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"- It's obvious that we have **underfitting** state. the training score (66.7%) and the cross-validation score (66.7%). so we need to use more complex model.\n- The low score refers to high bias (Underfitting = High Bias). \n- This state leads to bad leaderboard result (60.2%).","metadata":{}},{"cell_type":"markdown","source":"### 2) Overfitting:","metadata":{}},{"cell_type":"markdown","source":"Let's try to use **too complex** model:","metadata":{}},{"cell_type":"code","source":"model = DecisionTreeClassifier(random_state=7)\nmodel.fit(X_train, Y_train)\npredictions = model.predict(X_test)\nprint_scores(model, X_train, Y_train, predictions)","metadata":{"execution":{"iopub.status.busy":"2022-07-20T08:22:04.435390Z","iopub.execute_input":"2022-07-20T08:22:04.436596Z","iopub.status.idle":"2022-07-20T08:22:04.525895Z","shell.execute_reply.started":"2022-07-20T08:22:04.436549Z","shell.execute_reply":"2022-07-20T08:22:04.524662Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"- It's obvious that we have **overfitting** state. the training score (97.9%) is much bigger than cross-validation score (77.8%). so we need to use simpler model.\n- The individual cross-validation scores have high variance (Overfitting = High variane) ranging from 71.9% to 83.1%\n- This state leads to bad leaderboard result (71.7%).","metadata":{}},{"cell_type":"markdown","source":"### 3 ) Sweet spot: ","metadata":{}},{"cell_type":"markdown","source":"let's try to use a model,**Not too simple nor too complex.**","metadata":{}},{"cell_type":"code","source":"model = DecisionTreeClassifier(max_depth=3 , max_features=4 ,random_state=7)\nmodel.fit(X_train, Y_train)\npredictions = model.predict(X_test)\nprint_scores(model, X_train, Y_train, predictions)","metadata":{"execution":{"iopub.status.busy":"2022-07-20T08:22:04.761546Z","iopub.execute_input":"2022-07-20T08:22:04.761954Z","iopub.status.idle":"2022-07-20T08:22:04.835879Z","shell.execute_reply.started":"2022-07-20T08:22:04.761923Z","shell.execute_reply":"2022-07-20T08:22:04.834662Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"- This score is much better because the training score (82.3%) is not far from cross-validation score (82.1%)\n- The individual cross-validation scores variance became less (from 77.5% to 85.4%).\n- This state leads to higher leaderboard result (78.9%). But we don't know it yet:(, So we want to find best cross-validation score without Overfitting nor Underfitting.","metadata":{}},{"cell_type":"markdown","source":"Let's try to use more effecient mode","metadata":{}},{"cell_type":"code","source":"from sklearn.ensemble import RandomForestClassifier\n\nmodel = RandomForestClassifier(n_estimators= 80 ,max_depth=5 , max_features=8 ,min_samples_split=3 ,random_state=7)\nmodel.fit(X_train, Y_train)\npredictions = model.predict(X_test)\nprint_scores(model, X_train, Y_train, predictions)","metadata":{"execution":{"iopub.status.busy":"2022-07-20T08:52:32.893536Z","iopub.execute_input":"2022-07-20T08:52:32.893933Z","iopub.status.idle":"2022-07-20T08:52:35.016521Z","shell.execute_reply.started":"2022-07-20T08:52:32.893902Z","shell.execute_reply":"2022-07-20T08:52:35.015053Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"- This score is much better because the training score (85.8%) is not far from cross-validation score (82.9%)\n- The individual cross-validation scores variance is not low enough (from 75.3% to 91.0%), and that's normal because Titanic Dataset hasn't enough examples.\n- This state leads to higher leaderboard result (79.4%) :)","metadata":{}},{"cell_type":"markdown","source":"#### Conclusions:","metadata":{}},{"cell_type":"markdown","source":"- Titanic Dataset hasn't enough training examples, And this leads to Overfitting.\n\n- For other problems with big datasets you can control Overfitting and Underfitting by:  \n\n    -1- tradeoff between simple and complex models to find the sweet spot.  \n    -2- cross validation score very important to measure Overfitting and Underfitting  \n    -3- use grid search to find the best hyperparameters.  ","metadata":{}}]}