{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"pygments_lexer":"ipython3","nbconvert_exporter":"python","version":"3.6.4","file_extension":".py","codemirror_mode":{"name":"ipython","version":3},"name":"python","mimetype":"text/x-python"}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"## Feature Selection & Data Engineering Tutorial (Step by Step For New Comers):\n\nHello kagglers.\n\nIn this notebook we will learn:\n - 1 ) What is feature selection?    \n - 2 ) Why is feature selection very important?  \n - 3 ) Dimensionality reduction and feature selection.  \n - 4 ) Filter Methods.  \n\nThen we will apply what we learned on the Titanic Dataset.\n\nI hope this notebook will be useful to you ..  \nNow let's start, happy learning :)","metadata":{}},{"cell_type":"code","source":"import numpy as np # linear algebra\nimport pandas as pd # data processing, CSV file I/O (e.g. pd.read_csv)\nimport matplotlib.pyplot as plt # data visualization\nimport seaborn as sns # data visualization\nimport os \nimport warnings\nwarnings.filterwarnings(\"ignore\")\n\nfor dirname, _, filenames in os.walk('/kaggle/input'):\n    for filename in filenames:\n        print(os.path.join(dirname, filename))","metadata":{"execution":{"iopub.status.busy":"2022-07-21T09:43:32.624720Z","iopub.execute_input":"2022-07-21T09:43:32.625691Z","iopub.status.idle":"2022-07-21T09:43:32.639230Z","shell.execute_reply.started":"2022-07-21T09:43:32.625650Z","shell.execute_reply":"2022-07-21T09:43:32.638364Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"def read_data():\n    train_data = pd.read_csv(\"/kaggle/input/titanic/train.csv\")\n    print(\"Train data imported successfully!!\")\n    print(\"-\"*50)\n    test_data = pd.read_csv(\"/kaggle/input/titanic/test.csv\")\n    print(\"Test data imported successfully!!\")\n    return train_data , test_data","metadata":{"execution":{"iopub.status.busy":"2022-07-21T09:43:35.113324Z","iopub.execute_input":"2022-07-21T09:43:35.114153Z","iopub.status.idle":"2022-07-21T09:43:35.118849Z","shell.execute_reply.started":"2022-07-21T09:43:35.114113Z","shell.execute_reply":"2022-07-21T09:43:35.118090Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"train_data , test_data = read_data()","metadata":{"execution":{"iopub.status.busy":"2022-07-21T09:43:37.372226Z","iopub.execute_input":"2022-07-21T09:43:37.372958Z","iopub.status.idle":"2022-07-21T09:43:37.388594Z","shell.execute_reply.started":"2022-07-21T09:43:37.372920Z","shell.execute_reply":"2022-07-21T09:43:37.387467Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### 1 ) What is feature selection?","metadata":{}},{"cell_type":"markdown","source":"Feature Selection is the process of selecting the **most significant** features from a given dataset. In many cases, Feature Selection can also **enhance** the performance of a machine learning model.\n\nThe problem of having unnecessary features:\n\n- Unnecessary resource allocation for these features.\n- These features act as a noise for which the machine learning model can perform terribly poorly.\n- The machine model takes more time to get trained.\n\nFeature selection is also known as **Variable selection** or **Attribute selection**.\n\n\n","metadata":{}},{"cell_type":"markdown","source":"### 2 ) Why is feature selection very important?","metadata":{}},{"cell_type":"markdown","source":"The importance of feature selection can best be recognized when you are dealing with a dataset that contains a vast number of features. This type of dataset is often referred to as a high dimensional dataset. Now, with this high dimensionality, comes a lot of problems such as - this high dimensionality will significantly increase the training time of your machine learning model, it can make your model very complicated which in turn may lead to **Overfitting**.  \n\n**The objective of variable selection is three-fold:**  \n1 - improving the prediction performance of the predictors.  \n2 - providing faster and more cost-effective predictors.  \n3 - providing a better understanding of the underlying process that generated the data.\n\n","metadata":{}},{"cell_type":"markdown","source":"### 3 ) Dimensionality reduction Vs Feature selection:","metadata":{}},{"cell_type":"markdown","source":"Sometimes, feature selection is mistaken with dimensionality reduction. But they are different. Feature selection is different from dimensionality reduction. Both methods tend to reduce the number of attributes in the dataset, but:\n- **Dimensionality reduction** method does so by creating new combinations of attributes (sometimes known as feature transformation). whereas  \n- **Feature selection** methods include and exclude attributes present in the data without changing them.","metadata":{}},{"cell_type":"markdown","source":"### 4 ) Filter methods:","metadata":{}},{"cell_type":"markdown","source":"The following image best describes filter-based feature selection methods:","metadata":{}},{"cell_type":"markdown","source":"<img src=\"https://res.cloudinary.com/dyd911kmh/image/upload/f_auto,q_auto:best/v1537552825/Image3_fqsh79.png\">\n","metadata":{}},{"cell_type":"markdown","source":"<center>(Image Source: Analytics Vidhya)<\\center>","metadata":{}},{"cell_type":"markdown","source":"The filter method selects subsets of features based on their relationship with the target by:\n- Statistical Methods  \n- Feature Importance Methods   \n\nIt is common to use correlation type statistical measures between input and output variables as the basis for filter feature selection.\n\nAs such, the choice of statistical measures is **highly dependent** upon the variable **data types**.\n\nCommon data types include numerical (such as height) and categorical (such as a label), although each may be further subdivided such as integer and floating point for numerical variables, and boolean, ordinal, or nominal for categorical variables.\n","metadata":{}},{"cell_type":"markdown","source":"<img src=\"https://machinelearningmastery.com/wp-content/uploads/2020/06/Overview-of-Data-Variable-Types2.png\">\n\n","metadata":{}},{"cell_type":"markdown","source":"In this section, we will consider two broad categories of variable types: numerical and categorical; also, the two main groups of variables to consider: input and output.  \n**Numerical Output:** Regression predictive modeling problem.  \n**Categorical Output:** Classification predictive modeling problem. (This is our state)  \nThe statistical measures used in filter-based feature selection are generally calculated one input variable at a time with the target variable. As such, they are referred to as **univariate** statistical measures. This may mean that any interaction between input variables is **not considered** in the filtering process.","metadata":{}},{"cell_type":"markdown","source":"\n<img src=\"https://machinelearningmastery.com/wp-content/uploads/2019/11/How-to-Choose-Feature-Selection-Methods-For-Machine-Learning.png\">\n\n","metadata":{}},{"cell_type":"markdown","source":"#### Let's see our dataset: ","metadata":{}},{"cell_type":"code","source":"train_data.head(3)","metadata":{"_uuid":"8f2839f25d086af736a60e9eeb907d3b93b6e0e5","_cell_guid":"b1076dfc-b9ad-4769-8c92-a6c4dae69d19","execution":{"iopub.status.busy":"2022-07-21T09:43:43.883828Z","iopub.execute_input":"2022-07-21T09:43:43.884580Z","iopub.status.idle":"2022-07-21T09:43:43.903848Z","shell.execute_reply.started":"2022-07-21T09:43:43.884538Z","shell.execute_reply":"2022-07-21T09:43:43.902692Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# ==============================================================================\n#  Data Types:\n# ==============================================================================\n\ntrain_data.info()","metadata":{"execution":{"iopub.status.busy":"2022-07-21T09:43:46.600098Z","iopub.execute_input":"2022-07-21T09:43:46.600738Z","iopub.status.idle":"2022-07-21T09:43:46.614790Z","shell.execute_reply.started":"2022-07-21T09:43:46.600691Z","shell.execute_reply":"2022-07-21T09:43:46.613369Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"**Numerical Data:**\n\n - PassengerId (int)\n - SibSp (int)\n - Parch (int)\n - Age (float)\n - Fare (float)  \n\n**Categorical Data:**\n - Pclass (Ordinal)\n - Name (Nominal)\n - Ticket (Nominal)\n - Cabin (Nominal)\n - Embarked (Nominal)\n","metadata":{}},{"cell_type":"code","source":"# ==============================================================================\n#  Missed Values\n# ==============================================================================\n\ntrain_data.isnull().sum()","metadata":{"execution":{"iopub.status.busy":"2022-07-21T09:43:49.900277Z","iopub.execute_input":"2022-07-21T09:43:49.900976Z","iopub.status.idle":"2022-07-21T09:43:49.910187Z","shell.execute_reply.started":"2022-07-21T09:43:49.900942Z","shell.execute_reply":"2022-07-21T09:43:49.908612Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# ==============================================================================\n# Data Cleaning:\n# ==============================================================================\n\n# Dropping Unuseful feature because it has too many missed values:\n\ntrain_data.drop(columns = [\"Cabin\"] , inplace = True)\ntrain_data.drop(columns = [\"Ticket\"] , inplace = True)\n\n\n# ==============================================================================\n#  Fill missed embarked values:  \n\ntrain_data.Embarked = train_data.Embarked.fillna(train_data.Embarked.dropna().max())\n\n# ==============================================================================\n#  Fill missed age values:  \n\ntrain_data['Sex'] = train_data['Sex'].map( {'female': 1, 'male': 0} ).astype(int)\nguess_ages = np.zeros((2,3))\n\nfor i in range(0, 2):\n    for j in range(0, 3):\n        guess_df = train_data[(train_data['Sex'] == i) & \\\n                              (train_data['Pclass'] == j+1)]['Age'].dropna()\n        age_guess = guess_df.median()\n\n        # Convert random age float to nearest .5 age\n        guess_ages[i,j] = int( age_guess/0.5 + 0.5 ) * 0.5\n\nfor i in range(0, 2):\n    for j in range(0, 3):\n        train_data.loc[ (train_data.Age.isnull()) & (train_data.Sex == i) & (train_data.Pclass == j+1),\\\n                'Age'] = guess_ages[i,j]\n\ntrain_data['Age'] = train_data['Age'].astype(int)\ntrain_data.head()","metadata":{"execution":{"iopub.status.busy":"2022-07-21T09:43:52.373047Z","iopub.execute_input":"2022-07-21T09:43:52.374076Z","iopub.status.idle":"2022-07-21T09:43:52.410950Z","shell.execute_reply.started":"2022-07-21T09:43:52.374039Z","shell.execute_reply":"2022-07-21T09:43:52.410231Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"train_data.isnull().sum()","metadata":{"execution":{"iopub.status.busy":"2022-07-21T09:43:55.956791Z","iopub.execute_input":"2022-07-21T09:43:55.957889Z","iopub.status.idle":"2022-07-21T09:43:55.967686Z","shell.execute_reply.started":"2022-07-21T09:43:55.957843Z","shell.execute_reply":"2022-07-21T09:43:55.966525Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"**No More Missed Values !!!**","metadata":{}},{"cell_type":"markdown","source":"#### Now Let's see the correlation between the features and our target.","metadata":{}},{"cell_type":"code","source":"train_data.corr()[\"Survived\"].sort_values(ascending=False)","metadata":{"execution":{"iopub.status.busy":"2022-07-21T09:43:58.943672Z","iopub.execute_input":"2022-07-21T09:43:58.944816Z","iopub.status.idle":"2022-07-21T09:43:58.954329Z","shell.execute_reply.started":"2022-07-21T09:43:58.944776Z","shell.execute_reply":"2022-07-21T09:43:58.953485Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"From the correlation above:\n- Sex and Fare has strong positive correlation with the target. \n- PassengerId  has very weak correlation with the target. \n- Pclass has strong negative correlation with the target. ","metadata":{}},{"cell_type":"markdown","source":"It's obviouse that **Data Engineering** is very important for this problem.","metadata":{}},{"cell_type":"markdown","source":"**Note : This approach to feature selection will likely fail if there are important interactions between attributes where only one of the attributes is significant**","metadata":{}},{"cell_type":"markdown","source":"Let's see the correlation between all our features:","metadata":{}},{"cell_type":"code","source":"sns.set(rc = {'figure.figsize':(10,6)})\nsns.heatmap(train_data.corr(), annot = True, fmt='.2g',cmap= 'YlGnBu')","metadata":{"execution":{"iopub.status.busy":"2022-07-21T09:44:02.054150Z","iopub.execute_input":"2022-07-21T09:44:02.054570Z","iopub.status.idle":"2022-07-21T09:44:02.574770Z","shell.execute_reply.started":"2022-07-21T09:44:02.054527Z","shell.execute_reply":"2022-07-21T09:44:02.573896Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"train_data.head()","metadata":{"execution":{"iopub.status.busy":"2022-07-21T09:44:06.040214Z","iopub.execute_input":"2022-07-21T09:44:06.040897Z","iopub.status.idle":"2022-07-21T09:44:06.054602Z","shell.execute_reply.started":"2022-07-21T09:44:06.040850Z","shell.execute_reply":"2022-07-21T09:44:06.053147Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# ==========================================================================================\n# Data Engineering:\n# ==========================================================================================\n\n# Family Size:\n\ntrain_data['Family_Size'] = train_data[\"Parch\"] + train_data[\"SibSp\"] + 1\n\n# ==========================================================================================\n# Is Alone:\n\ntrain_data['IsAlone'] = 0\ntrain_data.loc[train_data['Family_Size'] == 1, 'IsAlone'] = 1\n\n# ==========================================================================================\n# Age Band:\n\ntrain_data.loc[ train_data['Age'] <= 16, 'Age'] = 0\ntrain_data.loc[(train_data['Age'] > 16) & (train_data['Age'] <= 32), 'Age'] = 1\ntrain_data.loc[(train_data['Age'] > 32) & (train_data['Age'] <= 48), 'Age'] = 2\ntrain_data.loc[(train_data['Age'] > 48) & (train_data['Age'] <= 64), 'Age'] = 3\ntrain_data.loc[ train_data['Age'] > 64, 'Age']\n\n# ==========================================================================================\n# Fare Band:\n\ntrain_data.loc[ train_data['Fare'] <= 130, 'Fare'] = 0\ntrain_data.loc[(train_data['Fare'] > 130) & (train_data['Fare'] <= 256), 'Fare'] = 1\ntrain_data.loc[(train_data['Fare'] > 256) & (train_data['Fare'] <= 384), 'Fare'] = 2\ntrain_data.loc[ train_data['Fare'] > 384, 'Fare'] = 3\ntrain_data['Fare'] = train_data['Fare'].astype(int)\n\n# ==========================================================================================\n# Name Title:\n\ntrain_data['Title'] = train_data.Name.str.extract(' ([A-Za-z]+)\\.', expand=False)\ntrain_data['Title'] = train_data['Title'].replace(['Lady', 'Countess','Capt', 'Col',\\\n'Don', 'Dr', 'Major', 'Rev', 'Sir', 'Jonkheer', 'Dona'], 'Rare')\ntrain_data['Title'] = train_data['Title'].replace('Mlle', 'Miss')\ntrain_data['Title'] = train_data['Title'].replace('Ms', 'Miss')\ntrain_data['Title'] = train_data['Title'].replace('Mme', 'Mrs')\n\ntitle_mapping = {\"Mr\": 1, \"Miss\": 2, \"Mrs\": 3, \"Master\": 4, \"Rare\": 5}\ntrain_data['Title'] = train_data['Title'].map(title_mapping)\ntrain_data['Title'] = train_data['Title'].fillna(0)\n\ntrain_data.drop(columns = [\"Name\"] , inplace = True)\n\n\n# ==========================================================================================\n# Embarked:\ntrain_data['Embarked'] = train_data['Embarked'].map( {'S': 0, 'C': 1, 'Q': 2} ).astype(int)\n\n# ==========================================================================================\n# Passenger Id:\ntrain_data.drop(columns = [\"PassengerId\"] , inplace = True)\n\n# ==========================================================================================\n# ==========================================================================================\n\nprint(\"Data Engineering Finished !!!\")","metadata":{"execution":{"iopub.status.busy":"2022-07-21T09:44:09.739481Z","iopub.execute_input":"2022-07-21T09:44:09.740197Z","iopub.status.idle":"2022-07-21T09:44:09.773020Z","shell.execute_reply.started":"2022-07-21T09:44:09.740160Z","shell.execute_reply":"2022-07-21T09:44:09.772254Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"train_data.head()","metadata":{"execution":{"iopub.status.busy":"2022-07-21T09:44:15.725245Z","iopub.execute_input":"2022-07-21T09:44:15.725767Z","iopub.status.idle":"2022-07-21T09:44:15.741347Z","shell.execute_reply.started":"2022-07-21T09:44:15.725717Z","shell.execute_reply":"2022-07-21T09:44:15.740450Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"train_data.corr()[\"Survived\"].sort_values(ascending=False)","metadata":{"execution":{"iopub.status.busy":"2022-07-21T09:44:18.193061Z","iopub.execute_input":"2022-07-21T09:44:18.193520Z","iopub.status.idle":"2022-07-21T09:44:18.205713Z","shell.execute_reply.started":"2022-07-21T09:44:18.193477Z","shell.execute_reply":"2022-07-21T09:44:18.204315Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"sns.set(rc = {'figure.figsize':(10,6)})\nsns.heatmap(train_data.corr(), annot = True, fmt='.2g',cmap= 'YlGnBu')","metadata":{"execution":{"iopub.status.busy":"2022-07-21T09:44:20.543943Z","iopub.execute_input":"2022-07-21T09:44:20.544789Z","iopub.status.idle":"2022-07-21T09:44:21.274240Z","shell.execute_reply.started":"2022-07-21T09:44:20.544736Z","shell.execute_reply":"2022-07-21T09:44:21.273176Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# ==========================================================================================\n# Feature Selection\n# We will select best 8 features\n# ==========================================================================================\n\n\nfrom sklearn.feature_selection import SelectKBest\nfrom sklearn.feature_selection import f_classif\n\nfs = SelectKBest(score_func=f_classif, k=8)\n\nprint(\"Data shape before feature selection:\")\nprint(train_data.shape)\n\n# apply feature selection\nSelected_train_data = fs.fit_transform(train_data.iloc[:,1:], train_data[\"Survived\"])\nprint(\"Data shape After feature selection:\")\nprint(Selected_train_data.shape)\n","metadata":{"execution":{"iopub.status.busy":"2022-07-21T09:45:07.793729Z","iopub.execute_input":"2022-07-21T09:45:07.794102Z","iopub.status.idle":"2022-07-21T09:45:07.807211Z","shell.execute_reply.started":"2022-07-21T09:45:07.794075Z","shell.execute_reply":"2022-07-21T09:45:07.805996Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"#### Feature Selection Finished !!","metadata":{}},{"cell_type":"markdown","source":"### References:\n\n- [Data Camp Python Feature Selection Tutorial: A Beginner's Guide](https://www.datacamp.com/tutorial/feature-selection-python)\n- [Machine Learning Mastery (How to Choose a Feature Selection Method For Machine Learning Article)](https://machinelearningmastery.com/feature-selection-with-real-and-categorical-data/)\n\n","metadata":{}}]}