{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"pygments_lexer":"ipython3","nbconvert_exporter":"python","version":"3.6.4","file_extension":".py","codemirror_mode":{"name":"ipython","version":3},"name":"python","mimetype":"text/x-python"}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"### **Mutual Information and What it Measures**","metadata":{}},{"cell_type":"markdown","source":"In this notebook, we will explore the use of **Mutual Information (MI)** as a metric for feature selection for our further model.\n\nMutual information describes relationships in terms of uncertainty. The MI between two quantities is a measure of the extent to which knowledge of one quantity reduces uncertainty about the other. If you knew the value of a feature, how much more confident would you be about the target?\n\nMI is similar to correlation, but with the advantage of being able to detect *any* kind of relationship, not just linear ones. Moreover, it is easy to use and interpret, computationally efficient, and theoretically well-founded, and also resistant to overfitting.\n\nSo, we will apply MI to our dataset and demonstrate how it can be used to identify the most informative predictors and eliminate the redundant ones, which helps to reduce overfitting and improve the generalization performance of the model. Overall, throughout this notebook, you will gain a better understanding of how MI can be used as a powerful metric for feature selection and its potential to improve the accuracy of our predictions.","metadata":{}},{"cell_type":"code","source":"import pandas as pd\nimport numpy as np\n\nimport glob\n\nimport matplotlib.pyplot as plt\nimport seaborn as sns\n\nfrom sklearn.cluster import KMeans\nfrom sklearn.feature_selection import mutual_info_classif","metadata":{"execution":{"iopub.status.busy":"2023-05-07T10:41:54.174440Z","iopub.execute_input":"2023-05-07T10:41:54.174841Z","iopub.status.idle":"2023-05-07T10:41:54.180834Z","shell.execute_reply.started":"2023-05-07T10:41:54.174812Z","shell.execute_reply":"2023-05-07T10:41:54.179759Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### Loading the data","metadata":{}},{"cell_type":"code","source":"data_dir = '/kaggle/input/tlvmc-parkinsons-freezing-gait-prediction/'\n# \ndaily_metadata_file = f'{data_dir}daily_metadata.csv'\ndefog_metadata_file = f'{data_dir}defog_metadata.csv'\ntdcsfog_metadata_file = f'{data_dir}tdcsfog_metadata.csv'\nevents_data_file = f'{data_dir}events.csv'\nsubjects_data_file = f'{data_dir}subjects.csv'\ntasks_data_file = f'{data_dir}tasks.csv'\n\n# Read the meta data\ndaily_metadata = pd.read_csv(daily_metadata_file)\ndefog_metadata = pd.read_csv(defog_metadata_file)\ntdcsfog_metadata = pd.read_csv(tdcsfog_metadata_file)\nfull_metadata = pd.concat([tdcsfog_metadata, defog_metadata])\n\nevents_data = pd.read_csv(events_data_file)\nsubjects_data = pd.read_csv(subjects_data_file)\ntasks_data = pd.read_csv(tasks_data_file)","metadata":{"execution":{"iopub.status.busy":"2023-05-07T10:11:40.480151Z","iopub.execute_input":"2023-05-07T10:11:40.480480Z","iopub.status.idle":"2023-05-07T10:11:40.559941Z","shell.execute_reply.started":"2023-05-07T10:11:40.480437Z","shell.execute_reply":"2023-05-07T10:11:40.559187Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Read the data and append the Id\ndef read_data(path):\n    df = pd.read_csv(path)\n    df['Id'] = path.split(\"/\")[-1].split(\".\")[0]\n    \n    return df\n\n\n# Read and concatenate all files of the train data from specified dataset\ndef create_full_data(dataset_name):\n    paths = glob.glob(data_dir + f'train/{dataset_name}/*')\n    final_df = pd.concat([read_data(p) for p in paths])\n    final_df['dataset'] = dataset_name\n    \n    return final_df","metadata":{"execution":{"iopub.status.busy":"2023-05-07T10:42:08.859525Z","iopub.execute_input":"2023-05-07T10:42:08.859999Z","iopub.status.idle":"2023-05-07T10:42:08.868311Z","shell.execute_reply.started":"2023-05-07T10:42:08.859964Z","shell.execute_reply":"2023-05-07T10:42:08.866629Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Create a dataframe with all tdcsfog data\ntdcsfog_data = create_full_data('tdcsfog')","metadata":{"execution":{"iopub.status.busy":"2023-05-07T10:42:12.950217Z","iopub.execute_input":"2023-05-07T10:42:12.950644Z","iopub.status.idle":"2023-05-07T10:42:31.431882Z","shell.execute_reply.started":"2023-05-07T10:42:12.950613Z","shell.execute_reply":"2023-05-07T10:42:31.430728Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"tdcsfog_data.head(3)","metadata":{"execution":{"iopub.status.busy":"2023-05-07T10:12:02.660673Z","iopub.execute_input":"2023-05-07T10:12:02.661107Z","iopub.status.idle":"2023-05-07T10:12:02.689520Z","shell.execute_reply.started":"2023-05-07T10:12:02.661064Z","shell.execute_reply":"2023-05-07T10:12:02.688406Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Create a dataframe with all defog data\ndefog_data = create_full_data('defog')","metadata":{"execution":{"iopub.status.busy":"2023-05-07T10:12:02.690590Z","iopub.execute_input":"2023-05-07T10:12:02.690876Z","iopub.status.idle":"2023-05-07T10:12:30.606862Z","shell.execute_reply.started":"2023-05-07T10:12:02.690852Z","shell.execute_reply":"2023-05-07T10:12:30.605700Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"defog_data.head(3)","metadata":{"execution":{"iopub.status.busy":"2023-05-07T10:12:30.608317Z","iopub.execute_input":"2023-05-07T10:12:30.608770Z","iopub.status.idle":"2023-05-07T10:12:30.625611Z","shell.execute_reply.started":"2023-05-07T10:12:30.608733Z","shell.execute_reply":"2023-05-07T10:12:30.624087Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"defog_data_valid = defog_data.loc[ \\\n                        (defog_data.Valid == True) & (defog_data.Task == True)].copy()\n\ndefog_data_valid.reset_index(drop=True, inplace=True)\ndefog_data_valid.drop(['Valid', 'Task'], axis=1, inplace=True)","metadata":{"execution":{"iopub.status.busy":"2023-05-07T10:12:30.627571Z","iopub.execute_input":"2023-05-07T10:12:30.628009Z","iopub.status.idle":"2023-05-07T10:12:32.813088Z","shell.execute_reply.started":"2023-05-07T10:12:30.627970Z","shell.execute_reply":"2023-05-07T10:12:32.811687Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"full_data = pd.concat([tdcsfog_data, defog_data_valid])","metadata":{"execution":{"iopub.status.busy":"2023-05-07T10:12:32.814746Z","iopub.execute_input":"2023-05-07T10:12:32.815139Z","iopub.status.idle":"2023-05-07T10:12:33.798803Z","shell.execute_reply.started":"2023-05-07T10:12:32.815108Z","shell.execute_reply":"2023-05-07T10:12:33.797739Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## Feature engineering","metadata":{}},{"cell_type":"markdown","source":"Let's perform some clustering operations on two datasets `tasks_data` and `subjects_data` to extract new features.\n\nFor `tasks_data`, we will calculate the duration of each task by subtracting the `Begin` time from the `End` time, and then obtain the sum of durations for each task and each unique `Id` value. Then we will apply a k-means clustering algorithm with 10 clusters on the numeric columns.\n\nFor `subjects_data`, we will group the data by `Subject` column and computing the median values for the selected columns. And then we will also apply a k-means clustering algorithm with 10 clusters on the numeric columns.","metadata":{}},{"cell_type":"code","source":"tasks_data['Duration'] = tasks_data['End'] - tasks_data['Begin']\n\ntasks_data = pd.pivot_table(tasks_data, values=['Duration'],\n                                index=['Id'], columns=['Task'],\n                                aggfunc='sum', fill_value=0)\n\ntasks_data.columns = [c[-1] for c in tasks_data.columns]\ntasks_data = tasks_data.reset_index()\n\ntasks_data['t_kmeans'] = KMeans(n_clusters=10, random_state=42, n_init=10) \\\n                    .fit_predict(tasks_data[tasks_data.columns[1:]])\n\n\nsubjects_data = subjects_data.fillna(0).groupby('Subject') \\\n    [['Visit', 'Age', 'YearsSinceDx', 'UPDRSIII_On', 'UPDRSIII_Off', 'NFOGQ']].median()\nsubjects_data = subjects_data.reset_index()\n\nsubjects_data['s_kmeans'] = KMeans(n_clusters=10, random_state=42, n_init=10) \\\n                    .fit_predict(subjects_data[subjects_data.columns[1:]])","metadata":{"execution":{"iopub.status.busy":"2023-05-07T10:12:33.802004Z","iopub.execute_input":"2023-05-07T10:12:33.802339Z","iopub.status.idle":"2023-05-07T10:12:33.966297Z","shell.execute_reply.started":"2023-05-07T10:12:33.802312Z","shell.execute_reply":"2023-05-07T10:12:33.965453Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"full_data = full_data.merge(full_metadata, on='Id', how='inner') \\\n                        .merge(tasks_data[['Id','t_kmeans']], how='left', on='Id').fillna(-1) \\\n                        .merge(subjects_data.drop('Visit', axis=1),\n                                       on='Subject', how='left').fillna(-1)\nfull_data.head()","metadata":{"execution":{"iopub.status.busy":"2023-05-07T10:12:33.967766Z","iopub.execute_input":"2023-05-07T10:12:33.968352Z","iopub.status.idle":"2023-05-07T10:13:07.030257Z","shell.execute_reply.started":"2023-05-07T10:12:33.968321Z","shell.execute_reply":"2023-05-07T10:13:07.029186Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"full_data[full_data.dataset == 'defog'].index[0]","metadata":{"execution":{"iopub.status.busy":"2023-05-07T10:13:07.031336Z","iopub.execute_input":"2023-05-07T10:13:07.031815Z","iopub.status.idle":"2023-05-07T10:13:09.727764Z","shell.execute_reply.started":"2023-05-07T10:13:07.031786Z","shell.execute_reply":"2023-05-07T10:13:09.726649Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Let's take the first one million rows of the `full_data` dataframe that will correcpond to `tdscfog` dataset rows, and another one million rows (starting from index 7,062,672 to 8,062,672) will be from the `defog` dataset. This will give us a balanced final dataset with equal number of samples from both datasets.","metadata":{}},{"cell_type":"code","source":"full_data_sample = pd.concat([full_data[:1000_000], full_data[7_062_672:8_062_672]])","metadata":{"execution":{"iopub.status.busy":"2023-05-07T10:13:09.729075Z","iopub.execute_input":"2023-05-07T10:13:09.729461Z","iopub.status.idle":"2023-05-07T10:13:10.078209Z","shell.execute_reply.started":"2023-05-07T10:13:09.729434Z","shell.execute_reply":"2023-05-07T10:13:10.076794Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"full_data_sample.shape","metadata":{"execution":{"iopub.status.busy":"2023-05-07T10:13:10.079626Z","iopub.execute_input":"2023-05-07T10:13:10.079954Z","iopub.status.idle":"2023-05-07T10:13:10.085736Z","shell.execute_reply.started":"2023-05-07T10:13:10.079926Z","shell.execute_reply":"2023-05-07T10:13:10.084893Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Create multiclass event type column\nfull_data_sample['target'] = 'Normal'\ntargets = ['StartHesitation', 'Turn', 'Walking']\nfor target in targets:\n    full_data_sample.loc[full_data_sample[target] == 1, 'target'] = target\n    \nfull_data_sample.Medication = np.where(full_data_sample.Medication == 'off', 0, 1)","metadata":{"execution":{"iopub.status.busy":"2023-05-07T10:13:10.086917Z","iopub.execute_input":"2023-05-07T10:13:10.087856Z","iopub.status.idle":"2023-05-07T10:13:10.283811Z","shell.execute_reply.started":"2023-05-07T10:13:10.087826Z","shell.execute_reply":"2023-05-07T10:13:10.282866Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Below we are creating the function that takes a dataframe as input and generates new features based on the AccV, AccML, and AccAP columns. The function returns the original dataframe with the new features appended.","metadata":{}},{"cell_type":"code","source":"# Create new features based on accelerometer data columns and returns the updated dataframe\ndef create_features(df):\n    acc_cols = ['AccV', 'AccML', 'AccAP']\n\n    for acc in acc_cols:\n        df[f'{acc}_lag_2'] = df.groupby('Id')[acc].shift(2).fillna(method=\"backfill\")\n        df[f'{acc}_lag_3'] = df.groupby('Id')[acc].shift(3).fillna(method=\"backfill\")\n        df[f'{acc}_lag_4'] = df.groupby('Id')[acc].shift(4).fillna(method=\"backfill\")\n        df[f'{acc}_lag_5'] = df.groupby('Id')[acc].shift(5).fillna(method=\"backfill\")\n\n        df[f'{acc}_cumsum'] = (df[acc]).groupby(df['Id']).cumsum()\n\n        df[f'{acc}_first_value'] = df.groupby('Id')[acc].transform('first')\n        df[f'{acc}_last_value'] = df.groupby('Id')[acc].transform('last')\n\n        df[f'{acc}_mean'] = df.groupby('Id')[acc].transform('mean')\n        df[f'{acc}_median'] = df.groupby('Id')[acc].transform('median')\n        df[f'{acc}_std'] = df.groupby('Id')[acc].transform('std')\n\n        df[f'{acc}_min'] = df.groupby('Id')[acc].transform('min')\n        df[f'{acc}_max'] = df.groupby('Id')[acc].transform('max')\n\n        df[f'{acc}_delta'] = df[f'{acc}_max'] - df[f'{acc}_min']\n        for lag in [1,2,3]:\n            df[f'ma_{lag}_{acc}'] = df[acc].rolling(lag).mean().fillna(method=\"backfill\")\n            df[f'ma{lag}_{acc}'] = df[acc].rolling(lag).mean().fillna(method=\"backfill\")\n            df[f'ma_{lag}_{acc}'] = df[acc].rolling(lag).mean().fillna(method=\"backfill\")\n        \n    return df","metadata":{"execution":{"iopub.status.busy":"2023-05-07T10:13:10.285085Z","iopub.execute_input":"2023-05-07T10:13:10.285429Z","iopub.status.idle":"2023-05-07T10:13:10.297838Z","shell.execute_reply.started":"2023-05-07T10:13:10.285386Z","shell.execute_reply":"2023-05-07T10:13:10.296661Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Apply the above function to our dataset\nfull_data_sample = create_features(full_data_sample)","metadata":{"execution":{"iopub.status.busy":"2023-05-07T10:13:10.299617Z","iopub.execute_input":"2023-05-07T10:13:10.300657Z","iopub.status.idle":"2023-05-07T10:13:18.447116Z","shell.execute_reply.started":"2023-05-07T10:13:10.300601Z","shell.execute_reply":"2023-05-07T10:13:18.445923Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Assign variables with predictors and target\nX = full_data_sample.drop(['Id', 'dataset', 'Test', 'Subject', 'UPDRSIII_Off',\n                               'StartHesitation', 'Turn', 'Walking'], axis=1).copy()\ny = X.pop('target')","metadata":{"execution":{"iopub.status.busy":"2023-05-07T10:13:18.448545Z","iopub.execute_input":"2023-05-07T10:13:18.449479Z","iopub.status.idle":"2023-05-07T10:13:21.367182Z","shell.execute_reply.started":"2023-05-07T10:13:18.449435Z","shell.execute_reply":"2023-05-07T10:13:21.365669Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Check the missing values in our X dataset\nX.isna().sum()[X.isna().sum() > 0]","metadata":{"execution":{"iopub.status.busy":"2023-05-07T10:13:21.369391Z","iopub.execute_input":"2023-05-07T10:13:21.369845Z","iopub.status.idle":"2023-05-07T10:13:21.806340Z","shell.execute_reply.started":"2023-05-07T10:13:21.369809Z","shell.execute_reply":"2023-05-07T10:13:21.804879Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Check the categorical variables in our X dataset\nX.dtypes[X.dtypes == object]","metadata":{"execution":{"iopub.status.busy":"2023-05-07T10:15:45.995960Z","iopub.execute_input":"2023-05-07T10:15:45.996374Z","iopub.status.idle":"2023-05-07T10:15:46.011904Z","shell.execute_reply.started":"2023-05-07T10:15:45.996343Z","shell.execute_reply":"2023-05-07T10:15:46.010392Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### Calculating Mutual information scores","metadata":{}},{"cell_type":"code","source":"discrete_features = X.dtypes == int\n\n# Calculate MI scores for the features in X with respect to the target variable y\ndef make_mi_scores(X, y, discrete_features):\n    mi_scores = mutual_info_classif(X, y, discrete_features=discrete_features)\n    mi_scores = pd.Series(mi_scores, name=\"MI Scores\", index=X.columns)\n    mi_scores = mi_scores.sort_values(ascending=False)\n    return mi_scores\n\n# Apply this function to our data\nmi_scores = make_mi_scores(X, y, discrete_features)","metadata":{"execution":{"iopub.status.busy":"2023-05-07T10:16:06.340164Z","iopub.execute_input":"2023-05-07T10:16:06.340612Z","iopub.status.idle":"2023-05-07T10:36:29.790586Z","shell.execute_reply.started":"2023-05-07T10:16:06.340580Z","shell.execute_reply":"2023-05-07T10:36:29.788617Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Show 10 features with the highest MI score\nmi_scores.to_frame().head(10)","metadata":{"execution":{"iopub.status.busy":"2023-05-07T10:40:39.491965Z","iopub.execute_input":"2023-05-07T10:40:39.492870Z","iopub.status.idle":"2023-05-07T10:40:39.506345Z","shell.execute_reply.started":"2023-05-07T10:40:39.492811Z","shell.execute_reply":"2023-05-07T10:40:39.505137Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Let's plot a bar plot to make comparisions easier.","metadata":{}},{"cell_type":"code","source":"def plot_mi_scores(scores):\n    scores = scores.sort_values(ascending=False)\n    fig, ax = plt.subplots(figsize=(13, 11))\n    ax = sns.barplot(x=scores.values, y=scores.index, palette=\"coolwarm\", orient='h')\n    ax.set_title(\"Mutual Information Scores\")\n    ax.set_xlabel(\"MI Scores\")\n    ax.set_ylabel(\"Features\")\n    plt.show()\n\nplot_mi_scores(mi_scores)\nplt.show()","metadata":{"execution":{"iopub.status.busy":"2023-05-07T10:40:50.035942Z","iopub.execute_input":"2023-05-07T10:40:50.036372Z","iopub.status.idle":"2023-05-07T10:40:51.467859Z","shell.execute_reply.started":"2023-05-07T10:40:50.036341Z","shell.execute_reply":"2023-05-07T10:40:51.465494Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"So, as we can see from the above plot the most relevant features for our further model are:\n- the difference between the maximum and minimum values of all accelerometer data columns: `AccV_delta`, `AccAP_delta`, `AccML_delta`,\n- the minimum and maximum values of the accelerometer columns, especially `AccML_min`, `AccAP_min`, `AccML_max`, \n- the mean, median, and standard deviation of the columns, especially `AccV_median`, `AccV_mean`, `AccV_std`,\n- the first and last values of the columns, especially `AccV_last_value`,\n- the cumulative sum of the accelerometer data columns.\n \nOther highly informative features include `Age`, `UPDRSIII_On`, and `NFOGQ` which respectively represent the age of patients, the motor score of patients when taking medication, and the presence of freezing of gait.\n\nOn the other hand, the most uninformative features are the `Medication` status column, the moving average features using the rolling method such as `ma_2_AccML`, and the lagged versions of accelerometer data, such as `AccML_lag_3`. These features do not seem to have a strong relationship with the target variable and may not be useful in predicting it.\n\nOverall, these results can guide us in selecting the most relevant features for building our predictive model.","metadata":{}}]}