{"metadata":{"kaggle":{"accelerator":"none","dataSources":[{"sourceId":81933,"databundleVersionId":9643020,"sourceType":"competition"}],"dockerImageVersionId":30761,"isInternetEnabled":true,"language":"python","sourceType":"notebook","isGpuEnabled":false},"kernelspec":{"display_name":"Python 3 (ipykernel)","language":"python","name":"python3"},"language_info":{"codemirror_mode":{"name":"ipython","version":3},"file_extension":".py","mimetype":"text/x-python","name":"python","nbconvert_exporter":"python","pygments_lexer":"ipython3","version":"3.11.8"}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"# CMI - Problematic Internet Usage (EDA)\n\nThe Child Mind Institute (CMI) data challenge aims to understand the various metrics that influence problematic internet usage among people in the specific age group of 5 to 22 years (i.e. children to young adults). This notebook contains an initial exploratory data analysis (EDA) on only the ```csv``` data. Here, we mostly focus on finding whether or not correlations exist among variables. Moreover, since the dataset is only partially labeled, we aim to discover the existence of clusters in the unlabeled portion of the dataset.","metadata":{}},{"cell_type":"code","source":"import numpy as np\nimport pandas as pd\nimport re\nimport matplotlib.pyplot as plt\n%matplotlib inline\nimport seaborn as sns\n\n#plt.rc('text', usetex=True)\nplt.rc('font', family='serif')\nplt.rc('xtick',labelsize=12)\nplt.rc('ytick',labelsize=12)\n\nsns.set_theme()\n\nimport warnings\nwarnings.filterwarnings('ignore')\n\nimport plotly.express as px\n\nfrom sklearn.impute import SimpleImputer\nfrom sklearn.compose import ColumnTransformer\nfrom sklearn.pipeline import Pipeline\nfrom sklearn.preprocessing import OneHotEncoder, StandardScaler\nfrom sklearn.cluster import KMeans\nfrom sklearn.decomposition import PCA ","metadata":{"_cell_guid":"b1076dfc-b9ad-4769-8c92-a6c4dae69d19","_uuid":"8f2839f25d086af736a60e9eeb907d3b93b6e0e5"},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"We define the main folder path and import the ```.csv``` files here. The training dataset consists of $3960$ IDs with $82$ different attributes. The descriptions of these attributes and the values that they take are included in the ```data_dictionary.csv``` file.","metadata":{}},{"cell_type":"code","source":"folder_path = '/kaggle/input/child-mind-institute-problematic-internet-use/'\n\ntrain_df = pd.read_csv(folder_path + 'train.csv')\ndata_dict = pd.read_csv(folder_path + 'data_dictionary.csv')","metadata":{},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"train_df.head()","metadata":{},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"data_dict.head()","metadata":{},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"train_df.describe()","metadata":{},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"These are the different attributes contained the training data. The metadata file associated with these attributes is ```data_dictionary.csv``` which describes what these attributes are and what kinds of values they take. The feature columns can be broken down into feature categories depending on their prefixes. This can be done by observing that each category prefix is followed by a hyphen '-'.","metadata":{}},{"cell_type":"code","source":"features = train_df.columns.to_list()\n\nfeatures","metadata":{},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"The different feature categories are extracted using the ```feature_cat()``` function, which takes in the features column and applies regex to extract features before return a sorted list of unique categories.","metadata":{}},{"cell_type":"code","source":"# Function for making feature categories\n\ndef feature_cat(features):\n    ''' \n    This function takes in the feature column and applies regex iteratively to obtain the different unique\n    feature categories. It returns a sorted list of the different categories.\n    '''\n    types = []\n\n    for n, feature in enumerate(features):\n        match = re.findall('.+-', feature) # using regex to extract subfeature\n\n        if len(match) == 1:\n            match = match[0][:-1]\n            types.append(match)\n\n    # Reduces the list to only contain unique categories and sorts it\n    unique_list = list(set(types))\n    sorted_list = sorted(unique_list)\n\n    return sorted_list\n\nfeature_categories = feature_cat(features) \n\nprint(feature_categories)","metadata":{},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"The subfeatures within each feature category can also be extracted using the ```subfeatures()``` function. The following function extracts the subcategories, expect the 'Season' categories.","metadata":{}},{"cell_type":"code","source":"# Function for making subfeatures of each category\n\ndef subfeatures(features, feature_cat, simplify=True):\n    ''' \n    This functions takes in the feature column and a particular feature category as inputs and returns a sorted \n    list of subcategories.\n\n    The optional 'simplify' argument, if set to False, will prepend the feature category to the subcategories. \n    '''\n    subfeatures = []\n\n    for n, feature in enumerate(features):\n        match = re.findall(feature_cat, feature)\n\n        if not 'Enroll_Season' or not 'Season' in feature:\n            \n            if len(match) != 0:\n                if simplify==True:\n                    subfeature = feature.split('-')[-1]\n                    subfeatures.append(subfeature)\n                else:\n                    subfeature = feature.split('-')[-1]\n                    subfeatures.append(feature)\n    \n    return sorted(subfeatures)\n\nprint(subfeatures(features, 'Basic_Demos'))\nprint(subfeatures(features, 'Basic_Demos', simplify=False))","metadata":{},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"We can confirm that there are a large numbe of missing entries for a lot of these attributes. Even the target feature ```sii``` contains missing values, hinting at the fact that the dataset will require a semi-supervised learning approach. The highest proportion of missing data arises from the 'Physical Activity Questionnaire' ```PAQ``` with both FitnessGram data following suite (```Fitness_Endurance``` and ```FGC```). ","metadata":{}},{"cell_type":"code","source":"def plot_na(data):\n    if data.isna().sum().sum() != 0: # checks if there are NaN values\n        na_data = (data.isna().sum().sort_values(ascending=False))/len(data) *100\n        missing_data = pd.DataFrame({'Feature': na_data.index, 'Percent missing': na_data.values})\n\n        fig, ax = plt.subplots(figsize=(8, 15))\n        sns.barplot(missing_data, y='Feature', x='Percent missing', ax=ax)\n    else:\n        print('No missing data found.')\n\nplot_na(train_df)","metadata":{},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"The plot of missing data can be filtered on the subset of the data for which the training label has been provided. We observe a very strong correlation between ```sii``` and the different categories of ```PCIAT```, which is the 'Parent-Child Internet Addiction Test'.","metadata":{}},{"cell_type":"code","source":"# Missing values in the subset where the target label is provided\n\nplot_na(train_df[train_df['sii'].notna()])","metadata":{},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Is there any correlation between ```PCIAT_Total``` and ```sii```? There is certainly a correlation which proceedes as a collection of step functions. As the ```data_dictionary.csv``` file suggests, the 'Severity Impairment Index' is given by\n* None: $0-30$\n* Mild: $31-49$\n* Moderate: $50-79$\n* Severe: $80-100$\n\nBased on this, the ```sii``` mapping is formed.","metadata":{}},{"cell_type":"code","source":"sns.scatterplot(train_df, x='PCIAT-PCIAT_Total', y='sii')\nplt.show()","metadata":{},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## EDA\nHere we perform exploratory data analysis (EDA) on the different feature categories to study the data spread, finding outliers and correlations.","metadata":{}},{"cell_type":"markdown","source":"### Demographic information\n\nThe demographic category contains information on two subcategories -\n* Sex: labeled as '0' (male) and '1' (female)\n* Age\n\nWe see that male participants are overly representative of the data. The median age of the participants is $10.5\\:\\rm{years}$.","metadata":{}},{"cell_type":"code","source":"sns.countplot(train_df, x='Basic_Demos-Sex')\nplt.show()","metadata":{},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"fig, (ax_box, ax_hist) = plt.subplots(2, sharex=True, \n                                    gridspec_kw={\"height_ratios\": (.15, .85)})\n\nsns.boxplot(train_df, x='Basic_Demos-Age', ax=ax_box) \nsns.histplot(train_df, x='Basic_Demos-Age', ax=ax_hist, hue='Basic_Demos-Sex', bins=15, kde=True, edgecolor=\"k\")\n\nplt.show()","metadata":{},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### Physical\nThe 'Physical' feature contains seven numerical features\n1. Body-Mass Index (BMI): a ratio of the mass in kilograms and height in meters\n2. Height in meters\n3. Weight in kilograms\n4. Waist cirumference in meters\n5. Systolic blood pressure in mmHg\n6. Diastolic blood pressure in mmHg\n7. Heart rate in BPM: It is not mentioned if it is resting heart rate or not. It is likely not, since there is a wide spread of values.\n\nFrom here, we also see that the sex labeled with '0' displays greater heights and weights, from which we can conclude that '0' and '1' refer to 'male' and 'female'. The histograms for these subcategories are shown where the dashed, black lines indicate the $25\\%$ quantile, median and $75\\%$ quantile respectively.","metadata":{}},{"cell_type":"code","source":"PHYSICAL_TYPES = subfeatures(features, 'Physical', simplify=False)\n\nplt.figure(figsize=(15, 12))\n\nstats = train_df.describe()\nfor n, TYPE in enumerate(PHYSICAL_TYPES):\n    ax_hist = plt.subplot(3, 3, n+1)\n\n    # calculating stats\n    median = train_df[TYPE].median()\n    lq, uq = train_df[TYPE].quantile([0.25, 0.75])\n\n    sns.histplot(train_df, x=TYPE, bins=20, hue='Basic_Demos-Sex', ax=ax_hist, kde=True, edgecolor='k')\n    ax_hist.axvline(lq, linestyle='dashed', linewidth=1, color='black')\n    ax_hist.axvline(median, linestyle='dashed', linewidth=1, color='black')\n    ax_hist.axvline(uq, linestyle='dashed', linewidth=1, color='black')\n\nplt.show()","metadata":{},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"The ```Physical``` attributes display significant outlier behavior.","metadata":{}},{"cell_type":"code","source":"sns.boxplot(\n    pd.melt(train_df[PHYSICAL_TYPES]),   \n    y='variable',\n    x='value',\n    hue='variable',\n    whis=[1,99]\n)\nplt.show()","metadata":{},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"We observe that there are some discrepancies in the blood pressure readings. The systolic BP is always higher than diastolic. However, we find a systolic BP readings few that are equal to or lower diastollic. ","metadata":{}},{"cell_type":"code","source":"sns.scatterplot(train_df, x='Physical-Diastolic_BP', y='Physical-Systolic_BP', hue='Basic_Demos-Sex')\nplt.plot([0, 200], [0, 200], linewidth=0.5, color='black')\nplt.show()","metadata":{},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Moreover, there are very high and low BPs and it will be interesting to check if there is any correlation with heart rate. This might indicate if the blood pressures were recorded during activity or rest.","metadata":{}},{"cell_type":"code","source":"fig = plt.figure(figsize=(10, 10))\nax = fig.add_subplot(projection='3d')\n\nax.scatter(\n    train_df['Physical-Diastolic_BP'],\n    train_df['Physical-Systolic_BP'],\n    train_df['Physical-HeartRate'],\n    s=50,\n    edgecolor='w'\n)\nax.set_xlabel('Diastolic BP')\nax.set_ylabel('Systolic BP')\nax.set_zlabel('Heart Rate')\n\nax.set_box_aspect(None, zoom=0.85)\nplt.show()","metadata":{},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"We observe that there is a positive relation between heart rate and systolic/diastolic BPs, as can be seen from the correlation heatmap. Other subfeatures in 'Physical' show some correlation among each other, as expected. We should also expect to see some correlation with those from the 'Fitness' category as well.\n\nAs for the improbable diastolic-systolic points, they can be change by adopting a simple relation\n$$ \\text{systolic BP} \\approx 1.5 \\times \\text{diastolic BP} $$\nThis comes from the fact that the normal blood pressure (at least in adults) is $120\\:\\text{mmHg}$ for systolic and $80\\:\\text{mmHg}$ for diastolic, which is a ratio of $1.5$. ","metadata":{}},{"cell_type":"code","source":"xticklabels = subfeatures(features, 'Physical')\nyticklabels = subfeatures(features, 'Physical')\n\nsns.heatmap(train_df[PHYSICAL_TYPES].corr(), annot=True, xticklabels=xticklabels, yticklabels=yticklabels)\nplt.show()","metadata":{},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"A few entries also have zero BMI, which is erroneous.","metadata":{}},{"cell_type":"code","source":"fig, ax = plt.subplots(1, 2, figsize=(12, 6))\n\nsns.scatterplot(train_df, x='Physical-BMI', y='Physical-Diastolic_BP', ax=ax[0], hue='Basic_Demos-Sex')\nsns.scatterplot(train_df, x='Physical-BMI', y='Physical-Systolic_BP', ax=ax[1], hue='Basic_Demos-Sex')\n\nplt.show()","metadata":{},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"FGC_TYPES = subfeatures(features, 'FGC', simplify=False)\n\nplt.figure(figsize=(18, 18))\n\n#stats = train_df.describe()\nfor n, TYPE in enumerate(FGC_TYPES):\n    ax = plt.subplot(4, 4, n+1)\n    if TYPE.split('_')[-1] != 'Zone':\n        sns.histplot(train_df, x=TYPE, ax=ax, bins=20, hue='Basic_Demos-Sex', kde=True, edgecolor='k')\n        #calculating stats\n        #median = train_df[TYPE].median()\n        #lq, uq = train_df[TYPE].quantile([0.25, 0.75])\n\n        #ax.axvline(lq, linestyle='dashed', linewidth=1, color='black')\n        #ax.axvline(median, linestyle='dashed', linewidth=1, color='black')\n        #ax.axvline(uq, linestyle='dashed', linewidth=1, color='black')\n\n    else:\n        sns.countplot(train_df, x=TYPE, hue='Basic_Demos-Sex')\n   \n\nplt.show()","metadata":{},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"sns.boxplot(\n    pd.melt(train_df[FGC_TYPES]),   \n    y='variable',\n    x='value',\n    hue='variable',\n    whis=[1,99]\n)\nplt.show()","metadata":{},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"xticklabels = subfeatures(features, 'Physical') + subfeatures(features, 'FGC')\nyticklabels = subfeatures(features, 'Physical') + subfeatures(features, 'FGC')\n\nplt.figure(figsize=(15, 15))\nsns.heatmap(train_df[PHYSICAL_TYPES + FGC_TYPES].corr(), annot=True, xticklabels=xticklabels, yticklabels=yticklabels)\n\nplt.show()","metadata":{},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### Bio-electric Impedance Analysis","metadata":{}},{"cell_type":"markdown","source":"It appears that there are significant discrepencies in the various BIA subcategories. Although the values of these subcategories are clustered around a reasonable range, they usually include a rather range or small entry.","metadata":{}},{"cell_type":"code","source":"BIA_TYPES = subfeatures(features, 'BIA', simplify=False)\n\nplt.figure(figsize=(18, 18))\n\nfor n, TYPE in enumerate(BIA_TYPES):\n    ax = plt.subplot(4, 4, n+1)\n    if TYPE.split('_')[-1] != 'num':\n        sns.histplot(train_df, x=TYPE, ax=ax, hue='Basic_Demos-Sex', kde=True)\n\n    else:\n        sns.countplot(train_df, x=TYPE, hue='Basic_Demos-Sex')\n   \n#plt.savefig('BMI_plots.png', dpi=300, bbox_inches='tight')\nplt.show()","metadata":{},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"train_df[BIA_TYPES].describe().T","metadata":{},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"The same histograms are now plotted with the $x$-axes constrained within the $1^\\text{st}$ and $99^\\text{th}$ quantiles.","metadata":{}},{"cell_type":"code","source":"plt.figure(figsize=(18, 18))\n\nstats = train_df.describe()\nfor n, TYPE in enumerate(BIA_TYPES):\n    ax = plt.subplot(4, 4, n+1)\n    if TYPE.split('_')[-1] != 'num':\n        sns.histplot(train_df, x=TYPE, ax=ax, hue='Basic_Demos-Sex', edgecolor='k')\n        #calculating stats\n        #median = train_df[TYPE].median()\n        low, high = train_df[TYPE].quantile([0.01, 0.99])\n        plt.xlim(low, high)\n\n        #ax.axvline(lq, linestyle='dashed', linewidth=1, color='black')\n        #ax.axvline(median, linestyle='dashed', linewidth=1, color='black')\n        #ax.axvline(uq, linestyle='dashed', linewidth=1, color='black')\n\n    else:\n        sns.countplot(train_df, x=TYPE, hue='Basic_Demos-Sex')\n   \n#plt.savefig('BMI_plots.png', dpi=300, bbox_inches='tight')\nplt.show()","metadata":{},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"There are two sets of BMI measurements for some participants (```Physical-BMI``` and ```BIA-BIA_BMI```). The BMI measurements are fairly consistent, displaying some variation.","metadata":{}},{"cell_type":"code","source":"def BMI_compare(data):\n    idx1 = data[data['Physical-BMI'].notna()].index # indices of Physical-BMI values\n    idx2 = data[data['BIA-BIA_BMI'].notna()].index # indices of BIA-BIA_BMI values\n\n    matches = idx1.intersection(idx2)\n\n    data_matches = data.loc[matches, :]\n\n    sns.scatterplot(data_matches, x='Physical-BMI', y='BIA-BIA_BMI', hue='Basic_Demos-Sex')\n    plt.show()\n\nBMI_compare(train_df)","metadata":{},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### Parent-Child Internet Addiction Test (PCIAT)\nThe target label ```sii``` is based on the scores from the PCIAT, which consist of a series of $20$ questions related to childrens' internet usage, followed by an aggregate of the score.","metadata":{}},{"cell_type":"code","source":"PCIAT_TYPES = subfeatures(features, 'PCIAT', simplify=False)\n\nplt.figure(figsize=(18, 25))\n\nfor n, TYPE in enumerate(PCIAT_TYPES):\n    ax = plt.subplot(7, 3, n+1)\n    if TYPE.split('_')[-1] == 'Total':\n        sns.histplot(train_df, x=TYPE, ax=ax, hue='Basic_Demos-Sex', kde=True)\n\n    else:\n        sns.countplot(train_df, x=TYPE, hue='Basic_Demos-Sex')\n   \nplt.show()","metadata":{},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"SDS_TYPES = subfeatures(features, 'SDS', simplify=False)\n\nprint(SDS_TYPES)","metadata":{},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"sns.heatmap(train_df.loc[:, ['PCIAT-PCIAT_Total', 'SDS-SDS_Total_Raw', 'SDS-SDS_Total_T', 'sii']].corr(), annot=True)\nplt.show()","metadata":{},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### Which features correlate most strongly with ```PCIAT_Total``` and ```sii```?\nIt will be interesting to check which features correlate most strongly (most positive and negative) with ```PCIAT_Total``` and ```sii``` values, even though there are a lot of missing entries in the dataset.\n\nThe function ```corr_with_sii()``` calculates the correlation between ```sii``` and other numeric features and sorts them in a descending order. The does not seem to show overly large correlations with ```sii``` with any feature other than the various ```PCIAT``` columns. Nevertheless, there are *relatively* strong positively correlations with height, age, weight, BMI and SDS total.","metadata":{}},{"cell_type":"code","source":"def corr_with_sii(data):\n    data_numeric = data.select_dtypes(['int64', 'float64'])\n    corr_matrix = data_numeric.corr()\n\n    corr_with_sii = corr_matrix.loc['sii'].sort_values(ascending=False)\n    return pd.DataFrame(\n        {\n            'Features': corr_with_sii.index.to_list()[1:],\n            'correlation': corr_with_sii.to_list()[1:]\n        }\n    )\n\ncorr_with_sii = corr_with_sii(train_df)\n\nfig, ax = plt.subplots(figsize=(5, 15))\nsns.barplot(corr_with_sii, x='correlation', y='Features', ax=ax)\nplt.show()","metadata":{},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## Finding clusters in the data\n\nThe dataset contains $82$ features. Hence we are presented with the so-called *curse of dimensionality* in that there is no simple way to visualize such a high dimensional dataset. However, we can use *dimensionality reduction* techniques to see if the dataset displays localized clusters when mapped to a lower dimension. Here, we will use principal component analysis (PCA) on the subset of the data with ```sii``` values present. \n\nAs a baseline, we will simply impute the missing values using some preferred statistic (mean or median) and perform the clustering.","metadata":{}},{"cell_type":"code","source":"# Pre-processing the date before applying unsupervised learning techniques\n\ndef preprocess(data, data_dictionary, label_missing=False):\n    ''' \n    This function preprocesses the data. The inputs are the dataframe and data_dictionary with an optional\n    label_missing argument set to 'False' by default. When 'False', the subset of the dataframe with labels present\n    will be chosen. Else, the opposite will be true.\n    \n    The input dataframe is first split into numerical and categorical dataframes.\n    The missing entries in the numerical and categorical features are imputed using the mean and mode respectively.\n    Then, the numerical column is standardized.\n\n    Except for the sex column, the other categorical features are not one-hot encoded and we assume an ordinal ordering.\n    '''\n    data_preprocessing = data.copy()\n\n    ''' \n    The dataframe is first reduced to contain only the subset for which the sii values exist.\n    '''\n    if label_missing==False:\n        data_reduced = data_preprocessing[data_preprocessing.loc[:,'sii'].notna()]\n        data_reduced = data_reduced.reset_index() \n\n        X = data_reduced.iloc[:, 0:-1]\n        y = data_reduced.loc[:, 'sii']\n    else:\n        data_reduced = data_preprocessing[data_preprocessing.loc[:,'sii'].isna()]\n        data_reduced = data_reduced.reset_index() \n\n        X = data_reduced.iloc[:, 0:-1]\n\n    ''' \n    The numeric separated out. Among them, we must also separate out that ones are the defined as 'categorial int' since\n    they are also categorial features that have been encoded as integers.\n    '''\n    \n    categorical_cols = data_dictionary[\n        data_dictionary['Type'] == 'categorical int'\n    ].loc[:, 'Field'].to_list()\n\n    numerical_cols = data_dictionary[\n        data_dictionary['Type'] == 'float'\n    ].loc[:, 'Field'].to_list()\n\n    X_categorcial = X.loc[:, categorical_cols]\n    X_numerical = X.loc[:, numerical_cols]\n\n    ct = ColumnTransformer(\n        transformers=[\n            ('ohe', OneHotEncoder(sparse_output=False), [categorical_cols[0]]),\n            ('imp_cat', SimpleImputer(strategy='most_frequent'), categorical_cols[1:])\n        ]\n    )\n\n    ct.set_output(transform='pandas')\n    X_categorcial = ct.fit_transform(X_categorcial)\n\n    pipeline = Pipeline(\n        [\n            ('imp_num', SimpleImputer(strategy='median')),\n            ('ss', StandardScaler())\n        ]\n    )\n    pipeline.set_output(transform='pandas')\n    X_numerical = pipeline.fit_transform(X_numerical)\n\n    #X = X_categorcial.join(X_numerical)\n    if label_missing==False:\n        return X_categorcial, X_numerical, y\n    else:\n        return X_categorical, X_numerical\n\nX_categorical, X_numerical, y = preprocess(train_df, data_dict)","metadata":{},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### Principal Component Analysis (PCA)\nPCA is a powerful technique for handling the complexity of higher dimensional data by reducing it into a small number of dimensions called *principal components* which explains the greatest variance in the data. Hence, a dataset with a very large number of features can be compressed to one with only two or three dimensions. This, however, comes at a cost of some information loss.\n\nPCA is designed to work on continuous data. However, a lot of the features here are ordinal. Here we first perform PCA only on the numerical features and then on both numerical and ordinal features (excluding sex) in the hopes that the ordinal features follow some hypothetical linear mapping.","metadata":{}},{"cell_type":"markdown","source":"We first check how many PCA components are needed to explain a goot chunk of the variance in the data. This is done through a *scree plot* which plots the number of PCA components with the *explained variance*. The scree plot on the scaled numerical data shows that over $70\\%$ of the variance is explained by the first three PCA components. This is not great, but still something.","metadata":{}},{"cell_type":"code","source":"def scree_plot(X):\n    pca = PCA()\n\n    pca.fit_transform(X)\n    #cov_matrix = pca.get_covariance()\n\n    explained_variance = pca.explained_variance_ratio_.flatten()\n    PCA_components = np.arange(1, len(explained_variance)+1)\n\n    plt.scatter(PCA_components, explained_variance.cumsum(), c='red', edgecolors='black')\n    plt.plot(PCA_components, explained_variance.cumsum(), c='red')\n    plt.bar(PCA_components, explained_variance)\n    plt.xlabel('PCA components', fontsize=15)\n    plt.ylabel('Explained variance ratio', fontsize=15)\n    plt.show()","metadata":{},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"scree_plot(X_numerical)","metadata":{},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Now, we perform PCA with ```n_components=3``` on the numerical data. We plot each PCA component against each other. However, for ```X_numerical```, we find no clusters appear. Will it be better with the inclusion of ```X_categorical```?","metadata":{}},{"cell_type":"code","source":"def PCA_pairplot(X, y, n_components, labeled=True):\n    pca = PCA(n_components=n_components)\n    X_PCA = pca.fit_transform(X)\n    X_PCA = pd.DataFrame(X_PCA, columns=['PCA_{}'.format(n) for n in range(1, n_components+1)])\n    \n    if labeled==True:\n        df = X_PCA.join(y)\n\n        sns.pairplot(df, hue='sii', diag_kind='None')\n        plt.show()\n    else:\n        df = X_PCA\n\n        sns.pairplot(df, diag_kind='None')\n        plt.show()\n\n    return df\n\n_ = PCA_pairplot(X_numerical, y, 3)","metadata":{},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"We now consider a joined list of ```X_categorial``` and ```X_numerical``` with the consideration that the categorial features are ordinal (which they are for the most part). In the ```X_joined``` set, we ignore the sex information since it is definitely a categorical variable with no intrinsic ordering. \n\nWith the joint dataframe, we see that the features display very well structured clusters for ```sii=0,1,2,3``` with the greatest variations occuring along the $1^{\\text{st}}$ and $3^{\\text{rd}}$ PCA axes. However, we note that the data is not actually spatially separated in PCA space and they only appear separated when the color labels are applied.","metadata":{}},{"cell_type":"code","source":"X_joined = X_categorical.iloc[:,2:].join(X_numerical)\n\nPCA_df = PCA_pairplot(X_joined, y, 3)","metadata":{},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"fig = px.scatter_3d(\n    PCA_df,\n    x='PCA_1',\n    y='PCA_2',\n    z='PCA_3',\n    color='sii', \n    title='PCA on labeled data'\n)\nfig.update_traces(\n    marker=dict(size=5,line=dict(width=2, color='DarkSlateGrey'))\n)\nfig.update_layout(\n    font_family=\"Times New Roman\",\n    title_font_family=\"Times New Roman\",\n)\nfig.show()","metadata":{},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Let us now apply PCA on the unlabeled portion of the dataset to check for patterns. Interestingly, PCA on the unlabeled set shows more variation on the $\\text{PCA}_2$ axis while retaining the cluster on the $\\text{PCA}_1-\\text{PCA}_3$ plane.","metadata":{}},{"cell_type":"code","source":"X_numerical_2, X_categorical_2 = preprocess(train_df, data_dict, label_missing=True)\n\nX_joined_2 = X_categorical_2.iloc[:, 2:].join(X_numerical_2)\n\nPCA_df_2 = PCA_pairplot(X_joined_2, y, 3, labeled=False)","metadata":{},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"We can now apply $k$-means clustering to the unlabeled dataset. Since we are aware that the dataset should contain four ```sii``` labels, we initialized ```KMeans()``` with ```n_cluster=4```. The generated labels are then joined to the processed, unlabeled data, which is then plotted. The $k$-means algorithm is able to identify four clusters, albeit with greater variance along the $\\text{PCA}_2$ axis.","metadata":{}},{"cell_type":"code","source":"kmeans = KMeans(n_clusters=4, init='k-means++', random_state=10)\nkmeans.fit(X_joined_2)\n\nknn_labels = kmeans.labels_","metadata":{},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"knn_labels = pd.DataFrame(knn_labels, columns=['knn_labels'])\n\nPCA_df_2_knn = PCA_df_2.join(knn_labels)\n\nfig = px.scatter_3d(\n    PCA_df_2_knn,\n    x='PCA_1',\n    y='PCA_2',\n    z='PCA_3',\n    color='knn_labels',\n    title='PCA on unlabeled data with k-means'\n)\nfig.update_traces(\n    marker=dict(size=5,line=dict(width=2, color='DarkSlateGrey'))\n)\nfig.update_layout(\n    margin=dict(l=20, r=20, t=20, b=20),\n    font_family=\"Times New Roman\",\n    title_font_family=\"Times New Roman\",\n)\nfig.show()","metadata":{},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## Conclusions\nHere we performed a detailed EDA on the ```train.csv``` data. We find that the data contains substantial outliers in all the subcategories. Particular examples of these outliers are in the ```Physical``` and ```BIA``` categories. In the former category, along with erroneous blood pressure entries, we observe some very high and low heart rates and weights. In the latter category, we also see some points that give rise to nonsensical data. These outliers and erroneous entries need to be fixed before a training algorithm is created.\n\nMoving on to the missing entries in the various features, in this notebook we used a very simple imutation technique where the numerical and categorical features are imputed using median and mode respectively. Evidently, this strategy is very simple and one must devise a more robust way to filling in these values using patterns present in the data. This notebook has not taken into account the actigraphy data and it can be assumed that some of these missing values can be filled in with the actigraphy data taken into account.\n\nFinally, we perform PCA on the labeled and unlabeled sets. PCA on the labeled set reveals clear clusters of four ```sii``` labels. Performing $k$-means clustering on the PCA result of the unlabeled set similarly reveals the existence of similar clusters.","metadata":{}},{"cell_type":"markdown","source":"","metadata":{}}]}