{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"name":"python","version":"3.11.11","mimetype":"text/x-python","codemirror_mode":{"name":"ipython","version":3},"pygments_lexer":"ipython3","nbconvert_exporter":"python","file_extension":".py"},"kaggle":{"accelerator":"none","dataSources":[{"sourceId":81933,"databundleVersionId":9643020,"sourceType":"competition"},{"sourceId":12375679,"sourceType":"datasetVersion","datasetId":7803330},{"sourceId":198686398,"sourceType":"kernelVersion"},{"sourceId":203489307,"sourceType":"kernelVersion"}],"dockerImageVersionId":31040,"isInternetEnabled":false,"language":"python","sourceType":"notebook","isGpuEnabled":false}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"# **Child Mind Institute: Problematic Internet Use**\n\nView on [GitHub](https://github.com/tomragus/CMI-PIU-Model)\n\n### ‣ 🧑‍💻 The Problem at Hand *(taken from the comptetition homepage 👉  [read here](https://www.kaggle.com/competitions/child-mind-institute-problematic-internet-use/overview))*\n\n\"In today’s digital age, problematic internet use among children and adolescents is a growing concern. Better understanding this issue is crucial for addressing mental health problems such as depression and anxiety.\n\nCurrent methods for measuring problematic internet use in children and adolescents are often complex and require professional assessments. This creates access, cultural, and linguistic barriers for many families. Due to these limitations, problematic internet use is often not measured directly, but is instead associated with issues such as depression and anxiety in youth.\n\nConversely, physical & fitness measures are extremely accessible and widely available with minimal intervention or clinical expertise. Changes in physical habits, such as poorer posture, irregular diet, and reduced physical activity, are common in excessive technology users. We propose using these easily obtainable physical fitness indicators as proxies for identifying problematic internet use, especially in contexts lacking clinical expertise or suitable assessment tools.\"\n\n**What does this mean?** The Child Mind Institute has tasked the public with building predictive machine learning models that will determine a participant's Severity Impairment Index (SII) - a metric measuring the level of problematic internet use among children and adolescents - based on physical activity, health, and other factors. The aim is to improve our ability to identify signs of problematic internet use early so preventative intervention can take place.\n\n### ‣ 📊 The Data at Hand *(available to download from competition page 👉 [here](https://www.kaggle.com/competitions/child-mind-institute-problematic-internet-use/data)*\n\nFrom the institute's *Healthy Brain Network*, we have been provided with a roughly 4,000-entry dataset comprising measurements from various instruments, assessments, and questionairres - in particular, an assessment called the \"Parent-Child Internet Addiction Test\" (PCIAT), which is used to calculate the SII of each participant. We have also been given a collection of time-series data, collected via a wrist accelerometer given to roughly 1,000 participants to wear for up to 30 days continually while at home and going about their daily lives. The competition is largely concerned with the first dataset, but we will touch on it a little bit at the end of the notebook.\n\nWe have been provided with a **train set** on which we will train our models, and a **test set** on which we will evaluate their performance and generate our submission. The train set is the full dataset with **over 70 features**, including SII (**target variable**) and the PCIAT results used to calculate it. The test set is a much smaller collection of data with about 20 rows, and missing these features. Our objective then is to train the models to accurately predict SII values for each entry in the test set. \n\n### ‣ 📝 Competition Evaluation\n\nThe result will be evaluated based on thequadratic weighted kappa, a metric measuring the agreement between two outcomes. Our submission will consist of two rows, one for id and one for SII, with an entry for each participant in the test set. An example submission has been given to us to view.","metadata":{}},{"cell_type":"code","source":"# load pandas\nimport pandas as pd\n\n# load sample submission\nsample = pd.read_csv(\"/kaggle/input/child-mind-institute-problematic-internet-use/sample_submission.csv\")\n\n# display sample submission\nprint(\"Sample submission\")\nprint(f\"Submission shape: {sample.shape}\")\nsample","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-11-17T05:15:45.999655Z","iopub.execute_input":"2025-11-17T05:15:45.999918Z","iopub.status.idle":"2025-11-17T05:15:48.508753Z","shell.execute_reply.started":"2025-11-17T05:15:45.999895Z","shell.execute_reply":"2025-11-17T05:15:48.507595Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"### ‣ 📚 Credit\n\n*Parts of this notebook's EDA section were adapted from [Antonina Dolgorukova](https://datadelic.dev/)'s feature EDA notebook for this competition. I highly encourage checking out her work - it is extremely in-depth and well written!* 👉   [read it here](https://www.kaggle.com/code/antoninadolgorukova/cmi-piu-features-eda/notebook)","metadata":{}},{"cell_type":"markdown","source":"***\n\n## 🎬 ***Getting Started***\n\nLet's start by taking a peek into our data and observing the features of the set.","metadata":{}},{"cell_type":"code","source":"import warnings\n\n# load train set\ntrain = pd.read_csv(\"/kaggle/input/child-mind-institute-problematic-internet-use/train.csv\")\n\n# display first 5 rows of train set\nprint(\"\"\"Train set: where the 'features' live\"\"\")\nprint(f\"Train shape: {train.shape}\")\nwith warnings.catch_warnings():\n    warnings.simplefilter(\"ignore\", category=RuntimeWarning)\n    display(train.head())\n","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-11-17T05:15:48.511479Z","iopub.execute_input":"2025-11-17T05:15:48.511769Z","iopub.status.idle":"2025-11-17T05:15:48.614413Z","shell.execute_reply.started":"2025-11-17T05:15:48.511747Z","shell.execute_reply":"2025-11-17T05:15:48.613447Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"While this is just a snippet, we can see the SII and the PCIAT scores on the right side of the set. The SII scores range from 0 to 3, with 0 representing no impairment and 3 representing severe impairment. So, we can think of the problem as training our models to **classify** each id in the test set into one of the 4 SII classes (0, 1, 2 or 3).","metadata":{}},{"cell_type":"code","source":"# load test set\ntest = pd.read_csv(\"/kaggle/input/child-mind-institute-problematic-internet-use/test.csv\")\n\n# display first 5 rows of test set\nprint(\"\"\"Test set: what we will evaluate our models on\"\"\")\nprint(f\"Test shape: {test.shape}\")\nwith warnings.catch_warnings():\n    warnings.simplefilter(\"ignore\", category=RuntimeWarning)\n    display(test.head())","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-11-17T05:15:48.615323Z","iopub.execute_input":"2025-11-17T05:15:48.615639Z","iopub.status.idle":"2025-11-17T05:15:48.657544Z","shell.execute_reply.started":"2025-11-17T05:15:48.615619Z","shell.execute_reply":"2025-11-17T05:15:48.655097Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"The Child Mind Institute was kind enough to include a **data dictionary** for this competition, which gives some extra information on each feature. Here is a little preview, but [you can view the full file on the Kaggle page.](https://www.kaggle.com/competitions/child-mind-institute-problematic-internet-use/data?select=data_dictionary.csv)","metadata":{}},{"cell_type":"code","source":"# load data dictionary\ndata_dict = pd.read_csv(\"/kaggle/input/child-mind-institute-problematic-internet-use/data_dictionary.csv\")\n\n# display first 5 rows of data dictionary\nprint(\"\"\"Data Dictionary: what each feature means\"\"\")\nprint(f\"Data Dictionary shape: {data_dict.shape}\")\ndisplay(data_dict.head())","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-11-17T05:15:48.658238Z","iopub.execute_input":"2025-11-17T05:15:48.658518Z","iopub.status.idle":"2025-11-17T05:15:48.705968Z","shell.execute_reply.started":"2025-11-17T05:15:48.658493Z","shell.execute_reply":"2025-11-17T05:15:48.704883Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"# load all other libraries\nimport numpy as np\nimport matplotlib\nimport matplotlib.pyplot as plt\nimport seaborn as sns\nimport os\nimport xgboost as xgb\nimport lightgbm as lgb\nfrom pandas.api.types import is_numeric_dtype, is_object_dtype, is_categorical_dtype, CategoricalDtype\nfrom scipy import stats\nfrom scipy.stats import pearsonr, spearmanr\nfrom sklearn.model_selection import train_test_split, StratifiedKFold, cross_val_score, KFold, GridSearchCV\nfrom sklearn.preprocessing import StandardScaler, RobustScaler, LabelEncoder\nfrom sklearn.linear_model import Ridge\nfrom sklearn.metrics import mean_squared_error, mean_absolute_error\nfrom lightgbm import LGBMRegressor, early_stopping, log_evaluation\nwarnings.filterwarnings(\"ignore\", category=RuntimeWarning)","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-11-17T05:15:48.709504Z","iopub.execute_input":"2025-11-17T05:15:48.711534Z","iopub.status.idle":"2025-11-17T05:15:57.130937Z","shell.execute_reply.started":"2025-11-17T05:15:48.711500Z","shell.execute_reply":"2025-11-17T05:15:57.129778Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"***\n\n## 🧐 ***Exploratory Data Analysis (EDA)***\n\nLet's begin by taking a closer look at the PCIAT features, or in other words, all the features that in the train set but not in the test set.","metadata":{}},{"cell_type":"code","source":"# isolating train-only features\ntrain_cols = set(train.columns)\ntest_cols = set(test.columns)\ncolumns_not_in_test = sorted(list(train_cols - test_cols))\n\n# addind additional information using data dictionary\ndata_dict[data_dict['Field'].isin(columns_not_in_test)]","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-11-17T05:15:57.131932Z","iopub.execute_input":"2025-11-17T05:15:57.132568Z","iopub.status.idle":"2025-11-17T05:15:57.156754Z","shell.execute_reply.started":"2025-11-17T05:15:57.132542Z","shell.execute_reply":"2025-11-17T05:15:57.155369Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"Here we see that each item in the Parent-Child Internet Addiction Test (PCIAT) is a question assessing a different aspect of a child's behavior related to internet use, with responses given on a scale from 0 to 5 - the total score, PCIAT_Total, indicates the severity of internet addiction. PCIAT_Total should directly align with SII.","metadata":{}},{"cell_type":"code","source":"# calculate max and min\npciat_min_max = train.groupby('sii', observed=True)['PCIAT-PCIAT_Total'].agg(['min', 'max'])\npciat_min_max = pciat_min_max.rename(columns={'min': 'Minimum PCIAT total Score', 'max': 'Maximum total PCIAT Score'})\ndisplay(pciat_min_max)\n\n# print range for each level of severity\nprint(data_dict[data_dict['Field'] == 'PCIAT-PCIAT_Total']['Value Labels'].iloc[0])","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-11-17T05:15:57.160611Z","iopub.execute_input":"2025-11-17T05:15:57.160962Z","iopub.status.idle":"2025-11-17T05:15:57.199154Z","shell.execute_reply.started":"2025-11-17T05:15:57.160940Z","shell.execute_reply":"2025-11-17T05:15:57.197808Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"So essentially, the answers to the PCIAT questions are totaled, and each participant is given an SII of 0-3 based on the range of their total score.\n\nBut wait, what about missing values? Say, for instance, a participant sees a question that *really hits close to home*, and they skip it. Maybe they skip many questions. How does the assessment account for this?","metadata":{}},{"cell_type":"code","source":"# display section with missing values highlighted\ntrain_with_sii = train[train['sii'].notna()][columns_not_in_test]\ntrain_with_sii[train_with_sii.isna().any(axis=1)].head().style.map(lambda x: 'background-color: #FFC0CB' if pd.isna(x) else '')","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-11-17T05:15:57.199966Z","iopub.execute_input":"2025-11-17T05:15:57.200237Z","iopub.status.idle":"2025-11-17T05:15:57.290095Z","shell.execute_reply.started":"2025-11-17T05:15:57.200215Z","shell.execute_reply":"2025-11-17T05:15:57.288962Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"**The problem:** the SII score is calculated by adding together all of the *non-NA* (non-empty) values in the test, meaning a participant can ignore every question (like the second participant) and still get a score of 0. Entries like this are obviously not reliable, and will skew the results. We can assume then that the SII score can sometimes be incorrect.\n\nWe can address this though, by recalculating the SII scores ourselves based on the maximum possible PCIAT_Total score *IF* missing values were answered. That way, participants who skip the survey will not be labeled as having no severity, but will instead be given a \"missing\" label.","metadata":{}},{"cell_type":"code","source":"# isolate PCIAT columns\nPCIAT_cols = [f'PCIAT-PCIAT_{i+1:02d}' for i in range(20)]\n\n# define function to recalculate SII accounting for missing PCIAT values\ndef recalculate_sii(row):\n    if pd.isna(row['PCIAT-PCIAT_Total']):\n        return np.nan\n    max_possible = row['PCIAT-PCIAT_Total'] + row[PCIAT_cols].isna().sum() * 5\n    if row['PCIAT-PCIAT_Total'] <= 30 and max_possible <= 30:\n        return 0\n    elif 31 <= row['PCIAT-PCIAT_Total'] <= 49 and max_possible <= 49:\n        return 1\n    elif 50 <= row['PCIAT-PCIAT_Total'] <= 79 and max_possible <= 79:\n        return 2\n    elif row['PCIAT-PCIAT_Total'] >= 80 and max_possible >= 80:\n        return 3\n    return np.nan\ntrain['recalc_sii'] = train.apply(recalculate_sii, axis=1)\n\n# overwriting SII with recalc_SII and adding labels 'missing', 'none', 'mild', 'moderate', 'severe'\ntrain['sii'] = train['recalc_sii']\ntrain['complete_resp_total'] = train['PCIAT-PCIAT_Total'].where(train[PCIAT_cols].notna().all(axis=1), np.nan)\nsii_map = {0: '0 (None)', 1: '1 (Mild)', 2: '2 (Moderate)', 3: '3 (Severe)'}\ntrain['sii'] = train['sii'].map(sii_map).fillna('Missing')\nsii_order = ['Missing', '0 (None)', '1 (Mild)', '2 (Moderate)', '3 (Severe)']\ntrain['sii'] = pd.Categorical(train['sii'], categories=sii_order, ordered=True)\ntrain.drop(columns='recalc_sii', inplace=True)","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-11-17T05:15:57.291196Z","iopub.execute_input":"2025-11-17T05:15:57.291552Z","iopub.status.idle":"2025-11-17T05:15:58.507621Z","shell.execute_reply.started":"2025-11-17T05:15:57.291527Z","shell.execute_reply":"2025-11-17T05:15:58.506602Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"Now with our SII values recalculated and updated, we can remove all duplicate rows, and all rows with \"missing\" in the SII column.","metadata":{}},{"cell_type":"code","source":"# remove rows with no SII\ninitial_rows = len(train)\ntrain = train[train['sii'] != 'Missing']\ntrain['sii'] = train['sii'].cat.remove_unused_categories()\nremoved_rows = initial_rows - len(train)\nprint(f\"Removed {removed_rows} rows with 'Missing' SII values.\")\nprint(f\"Train shape: {train.shape}\")\n\n# check for/ remove duplicate rows (if present)\nduplicate_count = train.duplicated().sum()\nprint(f\"Duplicate rows: {duplicate_count}\")","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-11-17T05:15:58.508466Z","iopub.execute_input":"2025-11-17T05:15:58.508779Z","iopub.status.idle":"2025-11-17T05:15:58.547831Z","shell.execute_reply.started":"2025-11-17T05:15:58.508752Z","shell.execute_reply":"2025-11-17T05:15:58.546803Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"# define helper function for gathering information from given column(s)\ndef calculate_stats(data, columns):\n    if isinstance(columns, str):\n        columns = [columns]\n    stats = []\n    for col in columns:\n        if data[col].dtype in ['object', 'category']:\n            counts = data[col].value_counts(dropna=False, sort=False)\n            percents = data[col].value_counts(normalize=True, dropna=False, sort=False) * 100\n            formatted = counts.astype(str) + ' (' + percents.round(2).astype(str) + '%)'\n            stats_col = pd.DataFrame({'count (%)': formatted})\n            stats.append(stats_col)\n        else:\n            stats_col = data[col].describe().to_frame().transpose()\n            stats_col['missing'] = data[col].isnull().sum()\n            stats_col.index.name = col\n            stats.append(stats_col)\n    return pd.concat(stats, axis=0)","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-11-17T05:15:58.548827Z","iopub.execute_input":"2025-11-17T05:15:58.549245Z","iopub.status.idle":"2025-11-17T05:15:58.556616Z","shell.execute_reply.started":"2025-11-17T05:15:58.549219Z","shell.execute_reply":"2025-11-17T05:15:58.555606Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"🖼️ ***With that, it's time to start visualizing!***\n\nFirst, let's view the distributions for the target variable SII, as well as for the corresponding PCIAT scores.","metadata":{}},{"cell_type":"code","source":"warnings.filterwarnings(\"ignore\", message=\".*use_inf_as_na option is deprecated.*\")\n\n# helper variables for EDA\nsii_counts = train['sii'].value_counts().reset_index()\nsii_counts.columns = ['SII', 'Count']\ntotal = sii_counts['Count'].sum()\nsii_counts['percentage'] = (sii_counts['Count'] / total) * 100\nfig, axes = plt.subplots(1, 2, figsize=(14, 5))\n\n# generate custom color pallete: bright yellow, lavender, mint green, light pink, sky blue\ncustom_palette = ['#FFFB46', '#E2A0FF', '#B7FFD8', '#FFC1CF', '#C4F5FC']\n\n# create pie chart: distribution of participant SII\naxes[0].pie(sii_counts['Count'], labels=sii_counts['SII'], autopct='%1.1f%%', colors=custom_palette, startangle=140, wedgeprops={'edgecolor': 'black'})\naxes[0].set_title('Distribution of Severity Impairment Index (SII)', fontsize=14)\naxes[0].axis('equal')\n\n# create histogram: distribution of participant PCIAT_Total\nsns.histplot(train['complete_resp_total'].dropna(), bins=20, color='#E2A0FF', ax=axes[1])\naxes[1].set_title('Distribution of PCIAT_Total', fontsize=14)\naxes[1].set_xlabel('PCIAT_Total')\naxes[1].set_ylabel('Participants')\nplt.tight_layout()\nplt.show()","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-11-17T05:15:58.557606Z","iopub.execute_input":"2025-11-17T05:15:58.557850Z","iopub.status.idle":"2025-11-17T05:15:59.061465Z","shell.execute_reply.started":"2025-11-17T05:15:58.557831Z","shell.execute_reply":"2025-11-17T05:15:59.060452Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"A couple things here - first, we can see that about 60% of participants have an SII score of 0, which we can basically interpret as, 60% of participants reporting that they are not impaired by internet use whatsoever. *Yeah, doubtful.* Following that, 26% are mildly impaired, and only a little over 1% are severely impaired. PCIAT scores seem to mostly follow a normal curve, with obvious exception being the giant spike at 0. With reasonable assurance, we can assume that the self-reported format of the assessment has led to some responses that are... *disingenuous,* to say the least.\n\nBut, let us in this moment give humanity the benefit of the doubt and proceed with our analysis. Next we'll use our handy **calculate_stats** funtion to take a look at some basic participant demographics like age and sex, and see how SII distributes over these metrics.","metadata":{}},{"cell_type":"code","source":"assert train['Basic_Demos-Age'].isna().sum() == 0\nassert train['Basic_Demos-Sex'].isna().sum() == 0\n\n# create table: distribution of participant Age Group\ntrain['Age Group'] = pd.cut(train['Basic_Demos-Age'], bins=[4, 12, 18, 22], labels=['Children (5-12)', 'Adolescents (13-18)', 'Adults (19-22)'])\ncalculate_stats(train, 'Age Group')","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-11-17T05:15:59.062385Z","iopub.execute_input":"2025-11-17T05:15:59.062785Z","iopub.status.idle":"2025-11-17T05:15:59.084982Z","shell.execute_reply.started":"2025-11-17T05:15:59.062764Z","shell.execute_reply":"2025-11-17T05:15:59.083965Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"# create table: distribution of participant Sex\nsex_map = {0: 'Male', 1: 'Female'}\ntrain['Basic_Demos-Sex'] = train['Basic_Demos-Sex'].map(sex_map)\ncalculate_stats(train, 'Basic_Demos-Sex')","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-11-17T05:15:59.085951Z","iopub.execute_input":"2025-11-17T05:15:59.086216Z","iopub.status.idle":"2025-11-17T05:15:59.100837Z","shell.execute_reply.started":"2025-11-17T05:15:59.086197Z","shell.execute_reply":"2025-11-17T05:15:59.099530Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"# create table: distribution of SII by Age Group\nstats_age = train.groupby(['Age Group', 'sii'], observed=False).size().unstack(fill_value=0)\nstats_age_prop = stats_age.div(stats_age.sum(axis=1), axis=0) * 100\nstats_age = stats_age.astype(str) +' (' + stats_age_prop.round(1).astype(str) + '%)'\nstats_age","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-11-17T05:15:59.102295Z","iopub.execute_input":"2025-11-17T05:15:59.102721Z","iopub.status.idle":"2025-11-17T05:15:59.140688Z","shell.execute_reply.started":"2025-11-17T05:15:59.102690Z","shell.execute_reply":"2025-11-17T05:15:59.139815Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"# create table: distribution of SII by Sex\nstats_sex = train.groupby(['Basic_Demos-Sex', 'sii'], observed=False).size().unstack(fill_value=0)\nstats_sex_prop = stats_sex.div(stats_sex.sum(axis=1), axis=0) * 100\nstats_sex = stats_sex.astype(str) +' (' + stats_sex_prop.round(1).astype(str) + '%)'\nstats_sex","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-11-17T05:15:59.141742Z","iopub.execute_input":"2025-11-17T05:15:59.142077Z","iopub.status.idle":"2025-11-17T05:15:59.161772Z","shell.execute_reply.started":"2025-11-17T05:15:59.142050Z","shell.execute_reply":"2025-11-17T05:15:59.160915Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"The tables show us that the majority of the participants are children, with fewer adolescents and very few adults. The participant pool is also interestingly two-thirds male, **which is significant - this means the model will be essentially gender-biased.** It is plausible that this was intentional by the institute.\n\nThe distribution of SII scores for the \"children\" and \"adults\" groups is skewed towards lower values, while scores for adolescents are relatively balanced across the categories, following something of a reverse-U shape. We also might notice that there are significantly more severe cases for males, but considering the sample is heavily skewed male, the differences between SII distribution in males and females are likely not significant.","metadata":{}},{"cell_type":"markdown","source":"📲 ***Internet Use:***\n\nYou might have noticed in our data samples a variable called \"PreInt_EduHx-computerinternet_hoursday\". This measures the number of hours participants spend on the internet per day - naturally, since we are investigating internet addiction, this variable warrants some deeper analysis. Let us first take a look at how long different age groups spend online.","metadata":{}},{"cell_type":"code","source":"warnings.filterwarnings(\"ignore\", message=\".*observed=False is deprecated.*\")\nfig, axes = plt.subplots(1, 3, figsize=(24, 6)) \n\n# create bar chart: distribution of Hours of Internet Use\nax1 = sns.countplot(x='PreInt_EduHx-computerinternet_hoursday', data=train, palette=custom_palette[:4], ax=axes[0], edgecolor='black', linewidth=0.8)\naxes[0].set_title('Distribution of Hours of Internet Use')\naxes[0].set_xlabel('Hours per Day Group')\naxes[0].set_ylabel('Count')\naxes[0].tick_params(axis='x', which='both', bottom=False, labelbottom=True)\n\n# create boxplot: Hours of Internet Use vs Age\nsns.boxplot(y=train['Basic_Demos-Age'], x=train['PreInt_EduHx-computerinternet_hoursday'], palette=custom_palette[:4], ax=axes[1])\naxes[1].set_title('Hours of Internet Use by Age')\naxes[1].set_ylabel('Age')\naxes[1].set_xlabel('Hours per Day Group')\n\n# create boxplot: Hours of Internet Use vs Age Group\nsns.boxplot(y='PreInt_EduHx-computerinternet_hoursday', x='Age Group', data=train, palette=custom_palette[:3], ax=axes[2])\naxes[2].set_title('Internet Use by Age Group')\naxes[2].set_ylabel('Hours per Day')\naxes[2].set_xlabel('Age Group');","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-11-17T05:15:59.162786Z","iopub.execute_input":"2025-11-17T05:15:59.163160Z","iopub.status.idle":"2025-11-17T05:15:59.841272Z","shell.execute_reply.started":"2025-11-17T05:15:59.163128Z","shell.execute_reply":"2025-11-17T05:15:59.840253Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"Following from the tables we analyzed previously, we can see that higher internet usage is linked to older age, hinting towards a correlation between hours of internet use and SII scores.","metadata":{}},{"cell_type":"code","source":"warnings.filterwarnings(\"ignore\", message=\".*observed=False is deprecated.*\")\nfig = plt.figure(figsize=(12, 10))\ngs = fig.add_gridspec(2, 2, height_ratios=[1, 1.5])\n\n# create boxplot: SII vs Hours of Internet Use\nax1 = fig.add_subplot(gs[0, 0])\nsns.boxplot(x='sii', y='PreInt_EduHx-computerinternet_hoursday', data=train, ax=ax1, palette=custom_palette[:4])\nax1.set_title('SII by Internet Use')\nax1.set_ylabel('Hours per Day')\nax1.set_xlabel('SII')\n\n# create boxplot: PCIAT_Total vs Hours of Internet Use\nax2 = fig.add_subplot(gs[0, 1])\nsns.boxplot(x='PreInt_EduHx-computerinternet_hoursday', y='complete_resp_total', data=train, palette=custom_palette[:4], ax=ax2)\nax2.set_title('PCIAT_Total by Internet Use')\nax2.set_ylabel('PCIAT_Total')\nax2.set_xlabel('Hours per Day');","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-11-17T05:15:59.842250Z","iopub.execute_input":"2025-11-17T05:15:59.842558Z","iopub.status.idle":"2025-11-17T05:16:00.250339Z","shell.execute_reply.started":"2025-11-17T05:15:59.842537Z","shell.execute_reply":"2025-11-17T05:16:00.249033Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"Simply tracing your finger across the plots, a slight upward trend is visible between SII/ PCIAT and hours of internet use, signaling their clear correlation. We can discern from our analysis so far that *while adults spend the most time using the internet, adolescents are the most likely to report a high SII score.*\n\nIt might also be notable that there are plenty of participants across all age groups that report less than an hour of internet use AND a high SII score, showing that for some participants, only an hour of internet use can be enough to negatively impact their mental health. ","metadata":{}},{"cell_type":"markdown","source":"🫀 ***Physcial Activity:***\n\nWhat about the participants' physical attributes? A significant portion of the dataset is dedicated to metrics like height, weight, BMI, heart rate, etc. The relationship between internet use and SII makes sense, but analysis of these metrics might yield some interesting and unexpected results.\n\nLooking at our data, there just so happen to be 7 primary physical metrics: BMI, Height, Weight, Waist Circumference, Heart Rate, and Diastolic and Systolic BP (Blood Pressure). This is the perfect number of features for building a **correlation matrix**, where each feature is given a score from -1 to 1, based on its correlation with the target variable (positive values indicate positive correlation, negative values indicate negative correlation, and 0 indicates no correlation).","metadata":{}},{"cell_type":"code","source":"# covert SII values to float type\ntrain['sii'] = train['sii'].str[0].astype(float)\n\n# isolate physical features\nphysical_cols = ['Physical-BMI', 'Physical-Height', 'Physical-Weight', 'Physical-Waist_Circumference', 'Physical-Diastolic_BP', 'Physical-HeartRate', 'Physical-Systolic_BP']\ndata_subset = train[physical_cols + ['sii']]\n\n# generate correlation heatmap matrix for physical features\ncorr_matrix = data_subset.corr()\nplt.figure(figsize=(10, 8))\nsns.heatmap(corr_matrix, annot=True, cmap='coolwarm', fmt='.2f', vmin=-1, vmax=1)\nplt.title('Correlation Heatmap')\nplt.show()","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-11-17T05:16:00.251500Z","iopub.execute_input":"2025-11-17T05:16:00.251770Z","iopub.status.idle":"2025-11-17T05:16:00.726026Z","shell.execute_reply.started":"2025-11-17T05:16:00.251751Z","shell.execute_reply":"2025-11-17T05:16:00.725084Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"Each square reads the correlation coefficeint between the two features intersecting at that square, with positive correlations marked red and negative correlations marked blue. Right away we notice strong positive correlations in the top right corner - this is where variables like height, weight, BMI and waist circumference all intersect with one another.\n\nLooking at the bottom row, we see that height, weight and waist circumference all have positive correlations with SII, while dialostic blood pressure, systolic blood pressure and heart rate are weakly correlated. This would suggest that taller and heavier participants are more likely to have higher SII scores, while cardiovascular health is not a determining factor for SII.\n\n*It should be noted:* since height and weight increase with age, some of this correlation is likely our previously observed age correlation in disguise.","metadata":{}},{"cell_type":"code","source":"plt.figure(figsize=(18, 5))\n\n# create scatterplot: Height vs Age\nplt.subplot(1, 3, 1)\nsns.scatterplot(x='Basic_Demos-Age', y='Physical-Height', data=train)\nplt.title('Physical-Height by Age')\nplt.xlabel('Age')\nplt.ylabel('Height (cm)')\n\n# create scatterplot: Weight vs Age\nplt.subplot(1, 3, 2)\nsns.scatterplot(x='Basic_Demos-Age', y='Physical-Weight', data=train)\nplt.title('Physical-Weight by Age')\nplt.xlabel('Age')\nplt.ylabel('Weight (kg)')\nplt.tight_layout()\nplt.show()","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-11-17T05:16:00.726994Z","iopub.execute_input":"2025-11-17T05:16:00.727319Z","iopub.status.idle":"2025-11-17T05:16:01.189364Z","shell.execute_reply.started":"2025-11-17T05:16:00.727292Z","shell.execute_reply":"2025-11-17T05:16:01.188300Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"We can see a number of outliers in the weight measurements, possibly suggesting some input error. This might have an effect on the correlation between weight and SII.","metadata":{}},{"cell_type":"code","source":"fig, axes = plt.subplots(1, 3, figsize=(18, 5))\n\n# create boxplot: SII vs BMI\nsns.boxplot(x='sii', y='Physical-BMI', data=train, ax=axes[0], palette=custom_palette[:4])\naxes[0].set_title('SII vs BMI')\naxes[0].set_xlabel('SII Category')\naxes[0].set_ylabel('BMI')\n\n# create boxplot: SII vs Height\nsns.boxplot(x='sii', y='Physical-Height', data=train, ax=axes[1], palette=custom_palette[:4])\naxes[1].set_title('SII vs Height')\naxes[1].set_xlabel('SII Category')\naxes[1].set_ylabel('Height (cm)')\n\n# create boxplot: SII vs Weight\nsns.boxplot(x='sii', y='Physical-Weight', data=train, ax=axes[2], palette=custom_palette[:4])\naxes[2].set_title('SII vs Weight')\naxes[2].set_xlabel('SII Category')\naxes[2].set_ylabel('Weight (kg)')\nplt.tight_layout()\nplt.show()","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-11-17T05:16:01.190404Z","iopub.execute_input":"2025-11-17T05:16:01.190747Z","iopub.status.idle":"2025-11-17T05:16:01.968325Z","shell.execute_reply.started":"2025-11-17T05:16:01.190725Z","shell.execute_reply":"2025-11-17T05:16:01.967406Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"Here we can see a pretty sharp jump in BMI and weight when SII is 3, suggesting that **weight** may play a significant factor in severe internet addiction, which is plausible. A positive correlation also exists with height, but the increase is gradual and less dramatic.","metadata":{}},{"cell_type":"markdown","source":"🏠 ***Bringing it Home:***\n\nAt this point, we have a decent understanding of the data at hand. To conclude our analysis, let us calculate the **Pearson correlation coefficient** between each feature and SII, like we did with the physical measures.","metadata":{}},{"cell_type":"code","source":"# isolate non-PCIAT features\ntrain_subset = train.drop(train.columns[55:76].tolist() + ['complete_resp_total'], axis=1)\nnumerical_features = train_subset.select_dtypes(include=[np.number]).columns.tolist()\n\n# calculate Pearson correlations with 'sii'\ntarget_correlations = []\nfor feature in numerical_features:\n    if feature != 'sii':\n        corr = train_subset[[feature, 'sii']].corr().iloc[0, 1]\n        if not np.isnan(corr):\n            target_correlations.append((feature, abs(corr)))\n\n# sort and select top 20\ntop_corr = sorted(target_correlations, key=lambda x: x[1], reverse=True)[:20]\nfeatures, corrs = zip(*top_corr)\n\n# create bar chart of features by correlation with SII\nplt.figure(figsize=(12, 8))\nplt.barh(range(len(features)), corrs, color='#E2A0FF', edgecolor='black')\nplt.yticks(range(len(features)), features)\nplt.xlabel('Absolute Correlation with SII')\nplt.title('Top 20 Features by Correlation with SII')\nplt.gca().invert_yaxis()  # Optional: highest at top\nplt.tight_layout()\nplt.show()","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-11-17T05:16:01.972663Z","iopub.execute_input":"2025-11-17T05:16:01.972958Z","iopub.status.idle":"2025-11-17T05:16:02.380445Z","shell.execute_reply.started":"2025-11-17T05:16:01.972938Z","shell.execute_reply":"2025-11-17T05:16:02.379423Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"This confirms that features like height, weight, age, and internet use, have a strong positive correlation with SII score. In interpreting these results, we must remember the time-old logical fallacy: *\"correlation does not imply causation\"*. Pearson correlation (same correlation as used in the matrix) doesn't say that any of these features cause SII, only that when one is higher, the other also tends to be higher.","metadata":{}},{"cell_type":"markdown","source":"🥡 ***Key Takeaways:***\n\n‣ Roughly 85% of participants report and SII of either 0 or 1.\n\n‣ The largest share of high SII scores are found in adolescents, with children and adults following.\n\n‣ Physical attribures (BMI, height, weight, waist circumference) are among the features with the strongest positive correlations with SII.\n\n‣ SII increases with internet use, and internet use increases with age.\n\n‣ The dataset overall skews male.","metadata":{}},{"cell_type":"markdown","source":"***\n\n## 🏋️ ***Model Training!***\n\nWe will train two models: a Ridge Regression model and a LightGBM (Gradient Boosting Machine) model. \n\nBefore we begin, we want to ensure our dataset only contains features we want to train on, so let us remove all trivial features/ columns like PCIAT scores, and arbitrary identifiers like \"id\". ","metadata":{}},{"cell_type":"code","source":"# permanentlyremove all PCIAT columns\npciat_cols = [col for col in train.columns if 'PCIAT' in col]\ntrain = train.drop(columns=pciat_cols)\n\n# remove \"complete_resp\" and \"age group\" columns\ntrain = train.drop(\"complete_resp_total\", axis=1)\ntrain = train.drop(\"Age Group\", axis=1) \n\n# remove id column\ntrain = train.drop(\"id\", axis=1)","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-11-17T05:16:02.381546Z","iopub.execute_input":"2025-11-17T05:16:02.381840Z","iopub.status.idle":"2025-11-17T05:16:02.394114Z","shell.execute_reply.started":"2025-11-17T05:16:02.381819Z","shell.execute_reply":"2025-11-17T05:16:02.392876Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"Here we separate our data into categorical/ numerical columns and re-calculate the absolute Pearson correlations with the target variable (SII) for the numerical features.","metadata":{}},{"cell_type":"code","source":"# separate SII from features\nX = train.drop('sii', axis=1)\ny = train['sii']\n\n# create sets for categorical and numerical features\ncategorical_cols = []\nnumerical_cols = []\nfor col in X.columns:\n    if X[col].dtype == 'object':\n        categorical_cols.append(col)\n    else:\n        numerical_cols.append(col)\n\n# re-calculate Pearson correlation coefficients for numerical columns (corr. with SII)\ndef calculate_correlations(X, y, numerical_cols):\n    correlations = {}\n    for col in numerical_cols:\n        mask = ~(X[col].isnull() | y.isnull())\n        if mask.sum() > 1:\n            corr, p_value = pearsonr(X[col][mask], y[mask])\n            correlations[col] = {'correlation': corr, 'p_value': p_value}\n    return correlations\ncorrelations = calculate_correlations(X, y, numerical_cols)\n\n# sort features by correlation\nsorted_correlations = sorted(correlations.items(), key=lambda x: abs(x[1]['correlation']), reverse=True)\nprint(\"\\nTop 20 correlations with SII (absolute correlation):\")\nfor i, (col, stats) in enumerate(sorted_correlations[:20]):\n    print(f\"{i+1:2d}. {col:<40} | Corr: {stats['correlation']:6.3f} | p-value: {stats['p_value']:.4f}\")","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-11-17T05:16:02.395226Z","iopub.execute_input":"2025-11-17T05:16:02.395579Z","iopub.status.idle":"2025-11-17T05:16:02.465704Z","shell.execute_reply.started":"2025-11-17T05:16:02.395552Z","shell.execute_reply":"2025-11-17T05:16:02.464124Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"⚙️ ***Feature Engineering:***\n\nWe will engineer the following features:","metadata":{}},{"cell_type":"markdown","source":"***Custom features:***\n\n‣ Difference between Physical BMI and Bio-electric Impedance Analysis (BIA) BMI (different assessments - BIA incorporates internal composition, while Physical BMI is a crude ratio)\n\n‣ Duration of Total Exercise (in seconds)\n\n‣ Ratio of Fat to FFM (body fat percentage to fat-free mass)\n\n‣ Z-score normalation of cardiovascular health statistics (heart rate, systolic and diastolic blood pressure)\n\n‣ Age Squared and Age Group (self-explanatory)","metadata":{}},{"cell_type":"code","source":"# define function to engineer new features\ndef create_engineered_features(df, correlations=None):\n    df_eng = df.copy()\n    \n    # BMI-related features\n    if 'Physical-BMI' in df_eng.columns and 'BIA-BIA_BMI' in df_eng.columns:\n        df_eng['BMI_difference'] = df_eng['Physical-BMI'] - df_eng['BIA-BIA_BMI']\n    \n    # fitness ratios\n    if 'Fitness_Endurance-Time_Mins' in df_eng.columns and 'Fitness_Endurance-Time_Sec' in df_eng.columns:\n        df_eng['Total_Fitness_Time'] = df_eng['Fitness_Endurance-Time_Mins'] * 60 + df_eng['Fitness_Endurance-Time_Sec']\n    \n    # body composition ratios\n    if 'BIA-BIA_Fat' in df_eng.columns and 'BIA-BIA_FFM' in df_eng.columns:\n        df_eng['Fat_to_FFM_ratio'] = df_eng['BIA-BIA_Fat'] / (df_eng['BIA-BIA_FFM'] + 1e-8)\n    \n    # physical health composite\n    if all(col in df_eng.columns for col in ['Physical-HeartRate', 'Physical-Systolic_BP', 'Physical-Diastolic_BP']):\n        health_cols = ['Physical-HeartRate', 'Physical-Systolic_BP', 'Physical-Diastolic_BP']\n        for col in health_cols:\n            if df_eng[col].notna().sum() > 0:\n                mean_val = df_eng[col].mean()\n                std_val = df_eng[col].std()\n                df_eng[f'{col}_normalized'] = (df_eng[col] - mean_val) / (std_val + 1e-8)\n    \n    # age-related\n    if 'Basic_Demos-Age' in df_eng.columns:\n        df_eng['Age_squared'] = df_eng['Basic_Demos-Age'] ** 2\n        df_eng['Age_group'] = pd.cut(df_eng['Basic_Demos-Age'], \n                                   bins=[0, 8, 12, 16, 22], \n                                   labels=['child', 'preteen', 'teen', 'young_adult'])\n    \n    # features based on correlation\n    if correlations is not None:\n        top_corr_features = [col for col, stats in sorted(correlations.items(), key=lambda x: abs(x[1]['correlation']), reverse=True)[:10] if col in df_eng.columns]\n        for i, col1 in enumerate(top_corr_features[:5]):\n            for col2 in top_corr_features[i+1:5]:\n                if col1 in df_eng.columns and col2 in df_eng.columns:\n                    if df_eng[col1].dtype in ['int64', 'float64'] and df_eng[col2].dtype in ['int64', 'float64']:\n                        df_eng[f'{col1}_x_{col2}'] = df_eng[col1] * df_eng[col2]\n    return df_eng\n\n# apply feature engineering function to train set\nX_engineered = create_engineered_features(X, correlations)\nprint(\"New features created:\")\nnew_features = set(X_engineered.columns) - set(X.columns)\nfor feat in new_features:\n    print(f\"  - {feat}\")","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-11-17T05:16:02.466704Z","iopub.execute_input":"2025-11-17T05:16:02.467027Z","iopub.status.idle":"2025-11-17T05:16:02.496266Z","shell.execute_reply.started":"2025-11-17T05:16:02.466999Z","shell.execute_reply":"2025-11-17T05:16:02.494794Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"Here, we can see all of the features created through correlation marked with an \"_x_\", with our custom features sprinkled in throughout. Since many of these features were calculated using correlation, many of them will likely be more highly correlated with SII than the original features.","metadata":{}},{"cell_type":"code","source":"# re-create categorical and numerical sets to include engineered features\ncategorical_cols_eng = []\nnumerical_cols_eng = []\nfor col in X_engineered.columns:\n    if is_object_dtype(X_engineered[col]) or isinstance(X_engineered[col].dtype, CategoricalDtype):\n        categorical_cols_eng.append(col)\n    elif is_numeric_dtype(X_engineered[col]):\n        numerical_cols_eng.append(col)\nprint(f\"\\nFinal categorical columns: {len(categorical_cols_eng)}\")\nprint(f\"Final numerical columns: {len(numerical_cols_eng)}\")\n\n# recalculate correlations with engineered features\ncorrelations_eng = calculate_correlations(X_engineered, y, numerical_cols_eng)\n\n# sort by absolute correlation\nsorted_correlations_eng = sorted(correlations_eng.items(), key=lambda x: abs(x[1]['correlation']), reverse=True)\nprint(\"\\nTop 15 correlations with SII (after feature engineering):\")\nfor i, (col, stats) in enumerate(sorted_correlations_eng[:15]):\n    print(f\"{i+1:2d}. {col:<40} | Corr: {stats['correlation']:6.3f} | p-value: {stats['p_value']:.4f}\")","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-11-17T05:16:02.497478Z","iopub.execute_input":"2025-11-17T05:16:02.497841Z","iopub.status.idle":"2025-11-17T05:16:02.576547Z","shell.execute_reply.started":"2025-11-17T05:16:02.497812Z","shell.execute_reply":"2025-11-17T05:16:02.575611Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"Important thing here is to split our train set into a training set and a **validation set** to experiment on. ","metadata":{}},{"cell_type":"code","source":"# split engineered data into train and validation sets, 80/ 20\nX_train, X_val, y_train, y_val = train_test_split(X_engineered, y, test_size=0.2, random_state=42, stratify=y)\nprint(f\"\\nTrain set shape: {X_train.shape}\")\nprint(f\"Validation set shape: {X_val.shape}\")","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-11-17T05:16:02.577552Z","iopub.execute_input":"2025-11-17T05:16:02.577813Z","iopub.status.idle":"2025-11-17T05:16:02.590274Z","shell.execute_reply.started":"2025-11-17T05:16:02.577793Z","shell.execute_reply":"2025-11-17T05:16:02.589104Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"⛰ ***Time to Train: Ridge Model***\n\nThe pre-processing pipeline will select which features to use (based on correlation), fill missing values, scale the features to have zero mean and unit variance, and  encode the categorical variables using one-hot encoding. After running through our engineered train and validation sets through this pipeline, we'll train the Ridge model with a few different correlation thresholds and evaluate each on the validation set to calculate RMSE between the model's predictions for SII and the actual SII values.\n\nWe'll try 4 different correlation thresholds: 0.01, 0.03, 0.05, and 0.1, as well as one model with no threshold.","metadata":{}},{"cell_type":"code","source":"# RIDGE REGRESSION PREPROCESSING PIPELINE\nprint(\"\\n\" + \"=\"*50)\nprint(\"RIDGE REGRESSION PREPROCESSING\")\nprint(\"=\"*50)\n\n# define function for feature selection based on threshold\ndef select_features_by_correlation(X, y, numerical_cols, threshold=0.05):\n    correlations = calculate_correlations(X, y, numerical_cols)\n    selected_features = []\n    for col, stats in correlations.items():\n        if abs(stats['correlation']) >= threshold and stats['p_value'] < 0.05:\n            selected_features.append(col)\n    return selected_features\n\n# define ridge model preprocessing function\ndef preprocess_for_ridge(X_train, X_val, categorical_cols, numerical_cols, correlation_threshold=0.05, use_correlation_filtering=True):\n    \n    # create copies of sets\n    X_train_processed = X_train.copy()\n    X_val_processed = X_val.copy()\n    \n    # missing values (fill numerical with median, categorical with mode)\n    for col in numerical_cols:\n        if col in X_train_processed.columns:\n            median_val = X_train_processed[col].median()\n            X_train_processed[col] = X_train_processed[col].fillna(median_val)\n            X_val_processed[col] = X_val_processed[col].fillna(median_val)\n    for col in categorical_cols:\n        if col in X_train_processed.columns:\n            mode_val = X_train_processed[col].mode()[0] if not X_train_processed[col].mode().empty else 'Unknown'\n            X_train_processed[col] = X_train_processed[col].fillna(mode_val)\n            X_val_processed[col] = X_val_processed[col].fillna(mode_val)\n    \n    # feature selection based on correlation\n    selected_numerical_features = numerical_cols\n    if use_correlation_filtering:\n        selected_numerical_features = select_features_by_correlation(X_train_processed, y_train, numerical_cols, correlation_threshold)\n        print(f\"Selected {len(selected_numerical_features)} numerical features out of {len(numerical_cols)} based on correlation > {correlation_threshold}\")\n        print(\"Selected features:\", selected_numerical_features[:10], \"...\" if len(selected_numerical_features) > 10 else \"\")\n        \n        # keep only selected numerical features\n        features_to_keep = selected_numerical_features + categorical_cols\n        X_train_processed = X_train_processed[features_to_keep]\n        X_val_processed = X_val_processed[features_to_keep]\n    \n    # one-hot encoding for categorical variables\n    X_train_encoded = pd.get_dummies(X_train_processed, columns=categorical_cols, drop_first=True)\n    X_val_encoded = pd.get_dummies(X_val_processed, columns=categorical_cols, drop_first=True)\n    \n    # validate continuity between sets\n    missing_cols_val = set(X_train_encoded.columns) - set(X_val_encoded.columns)\n    missing_cols_train = set(X_val_encoded.columns) - set(X_train_encoded.columns)\n    for col in missing_cols_val:\n        X_val_encoded[col] = 0\n    for col in missing_cols_train:\n        X_train_encoded[col] = 0\n    X_val_encoded = X_val_encoded.reindex(columns=X_train_encoded.columns, fill_value=0)\n    \n    # scale features\n    scaler = StandardScaler()\n    X_train_scaled = scaler.fit_transform(X_train_encoded)\n    X_val_scaled = scaler.transform(X_val_encoded)\n    encoded_cols = list(X_train_encoded.columns)\n    return X_train_scaled, X_val_scaled, scaler, encoded_cols, selected_numerical_features\n\n# train Ridge models with different correlation thresholds\ncorrelation_thresholds = [0.01, 0.03, 0.05, 0.1]\nridge_results = {}\nprint(\"\\nTesting different correlation thresholds for Ridge:\")\nfor threshold in correlation_thresholds:\n    print(f\"\\n--- Testing threshold: {threshold} ---\")\n    \n    # preprocessing\n    X_train_ridge, X_val_ridge, scaler, encoded_cols, selected_features = preprocess_for_ridge(X_train, X_val, categorical_cols_eng, numerical_cols_eng, correlation_threshold=threshold, use_correlation_filtering=True)\n    print(f\"Ridge - Train shape after preprocessing: {X_train_ridge.shape}\")\n    \n    # train\n    ridge_model = Ridge(alpha=1.0, random_state=42)\n    ridge_model.fit(X_train_ridge, y_train)\n    \n    # predictions and evaluation\n    y_train_pred_ridge = ridge_model.predict(X_train_ridge)\n    y_val_pred_ridge = ridge_model.predict(X_val_ridge)\n    ridge_train_rmse = np.sqrt(mean_squared_error(y_train, y_train_pred_ridge))\n    ridge_val_rmse = np.sqrt(mean_squared_error(y_val, y_val_pred_ridge))\n    ridge_results[threshold] = {'train_rmse': ridge_train_rmse, 'val_rmse': ridge_val_rmse, 'num_features': X_train_ridge.shape[1], 'selected_features': selected_features}\n    print(f\"Train RMSE: {ridge_train_rmse:.4f}\")\n    print(f\"Validation RMSE: {ridge_val_rmse:.4f}\")\n\n# test without correlation fitting\nprint(f\"\\n--- Testing without correlation filtering ---\")\nX_train_ridge_no_filter, X_val_ridge_no_filter, ridge_scaler_no_filter, ridge_columns_no_filter, _ = preprocess_for_ridge(X_train, X_val, categorical_cols_eng, numerical_cols_eng, use_correlation_filtering=False)\nprint(f\"Ridge - Train shape after preprocessing: {X_train_ridge_no_filter.shape}\")\nridge_model_no_filter = Ridge(alpha=1.0, random_state=42)\nridge_model_no_filter.fit(X_train_ridge_no_filter, y_train)\n\n# evaluate without correlation fitting\ny_train_pred_ridge_no_filter = ridge_model_no_filter.predict(X_train_ridge_no_filter)\ny_val_pred_ridge_no_filter = ridge_model_no_filter.predict(X_val_ridge_no_filter)\nridge_train_rmse_no_filter = np.sqrt(mean_squared_error(y_train, y_train_pred_ridge_no_filter))\nridge_val_rmse_no_filter = np.sqrt(mean_squared_error(y_val, y_val_pred_ridge_no_filter))\nridge_results['no_filter'] = {'train_rmse': ridge_train_rmse_no_filter, 'val_rmse': ridge_val_rmse_no_filter, 'num_features': X_train_ridge_no_filter.shape[1], 'selected_features': None}\nprint(f\"Train RMSE: {ridge_train_rmse_no_filter:.4f}\")\nprint(f\"Validation RMSE: {ridge_val_rmse_no_filter:.4f}\")\n\n# select best Ridge configuration for final model\nbest_ridge_config = min(ridge_results.items(), key=lambda x: x[1]['val_rmse'])\nprint(f\"\\nBest Ridge configuration: {best_ridge_config[0]} (Validation RMSE: {best_ridge_config[1]['val_rmse']:.4f})\")\nif best_ridge_config[0] == 'no_filter':\n    X_train_ridge_final, X_val_ridge_final = X_train_ridge_no_filter, X_val_ridge_no_filter\n    ridge_model_final = ridge_model_no_filter\n    final_ridge_train_rmse, final_ridge_val_rmse = ridge_train_rmse_no_filter, ridge_val_rmse_no_filter\nelse:\n    X_train_ridge_final, X_val_ridge_final, _, _, _ = preprocess_for_ridge(X_train, X_val, categorical_cols_eng, numerical_cols_eng, correlation_threshold=best_ridge_config[0], use_correlation_filtering=True)\n    ridge_model_final = Ridge(alpha=1.0, random_state=42)\n    ridge_model_final.fit(X_train_ridge_final, y_train)\n    y_train_pred_final = ridge_model_final.predict(X_train_ridge_final)\n    y_val_pred_final = ridge_model_final.predict(X_val_ridge_final)\n    final_ridge_train_rmse = np.sqrt(mean_squared_error(y_train, y_train_pred_final))\n    final_ridge_val_rmse = np.sqrt(mean_squared_error(y_val, y_val_pred_final))\nselected_features = best_ridge_config[1]['selected_features']","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-11-17T05:16:02.591324Z","iopub.execute_input":"2025-11-17T05:16:02.591670Z","iopub.status.idle":"2025-11-17T05:16:03.609869Z","shell.execute_reply.started":"2025-11-17T05:16:02.591638Z","shell.execute_reply":"2025-11-17T05:16:03.608778Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"We see our validation RMSEs mostly hovering around 0.68, with the clear winner being the model with threshold 0.1, with a validation RMSE of 0.6768. This means that on average, the model's SII predictions are off by 0.678 points. ","metadata":{}},{"cell_type":"markdown","source":"💡 ***Time to Train: LightGBM Model***\n\nUnlike Ridge, LightGBM doesn't require correlation filterning like Ridge does, and selects automatically features internally based on a calculated metric called \"importance\".\n\nIt also requires us to pre-define a set of **hyperparameters** for tuning the model.","metadata":{}},{"cell_type":"code","source":"# LIGHTGBM PREPROCESSING PIPELINE\nprint(\"\\n\" + \"=\"*50)\nprint(\"LIGHTGBM PREPROCESSING\")\nprint(\"=\"*50)\n\ndef preprocess_for_lightgbm(X_train, X_val, categorical_cols, numerical_cols):\n    \n    # create copies of sets\n    X_train_processed = X_train.copy()\n    X_val_processed = X_val.copy()\n    \n    # missing values for numerical columns\n    for col in numerical_cols:\n        if col in X_train_processed.columns:\n            median_val = X_train_processed[col].median()\n            X_train_processed[col] = X_train_processed[col].fillna(median_val)\n            X_val_processed[col] = X_val_processed[col].fillna(median_val)\n    \n    # label encoding for categorical variables\n    label_encoders = {}\n    for col in categorical_cols:\n        if col in X_train_processed.columns:\n            le = LabelEncoder()\n            \n            # fill missing values first\n            mode_val = X_train_processed[col].mode()[0] if not X_train_processed[col].mode().empty else 'Unknown'\n            X_train_processed[col] = X_train_processed[col].fillna(mode_val)\n            X_val_processed[col] = X_val_processed[col].fillna(mode_val)\n\n            # fit on train data\n            X_train_processed[col] = le.fit_transform(X_train_processed[col]).astype(int)\n\n            # handle unseen categories in validation\n            train_categories = set(le.classes_)\n            X_val_processed[col] = X_val_processed[col].map(lambda x: le.transform([x])[0] if x in train_categories else 0).astype(int)\n            label_encoders[col] = le\n    return X_train_processed, X_val_processed, label_encoders\n\n# preprocess for LightGBM\nX_train_lgb, X_val_lgb, lgb_encoders = preprocess_for_lightgbm(X_train, X_val, categorical_cols_eng, numerical_cols_eng)\ncategorical_feature_names = [col for col in categorical_cols_eng if col in X_train_lgb.columns]\nprint(f\"LightGBM - Train shape after preprocessing: {X_train_lgb.shape}\")\nprint(f\"LightGBM - Validation shape after preprocessing: {X_val_lgb.shape}\")\n\n# define LightGBM parameters\nlgb_params = {'objective': 'regression', 'metric': 'rmse', 'boosting_type': 'gbdt', 'num_leaves': 31, 'learning_rate': 0.05, 'colsample_bytree': 0.9, 'subsample': 0.8, 'bagging_freq': 5, 'verbose': 0, 'random_state': 42, 'force_col_wise': True}\n\n# create train and validation sets for LightGBM\ntrain_data = lgb.Dataset(X_train_lgb, label=y_train, categorical_feature=categorical_feature_names)\nval_data = lgb.Dataset(X_val_lgb, label=y_val, reference=train_data, categorical_feature=categorical_feature_names)\n\n# train LightGBM model\nlgb_model = lgb.train(lgb_params, train_data, valid_sets=[train_data, val_data], num_boost_round=1000, callbacks=[lgb.early_stopping(stopping_rounds=50), lgb.log_evaluation(0)])\n\n# evaluate trained model on train and validation sets for result\ny_train_pred_lgb = lgb_model.predict(X_train_lgb, num_iteration=lgb_model.best_iteration)\ny_val_pred_lgb = lgb_model.predict(X_val_lgb, num_iteration=lgb_model.best_iteration)\nlgb_train_rmse = np.sqrt(mean_squared_error(y_train, y_train_pred_lgb))\nlgb_val_rmse = np.sqrt(mean_squared_error(y_val, y_val_pred_lgb))\nprint(f\"\\nLightGBM Results:\")\nprint(f\"Train RMSE: {lgb_train_rmse:.4f}\")\nprint(f\"Validation RMSE: {lgb_val_rmse:.4f}\")","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-11-17T05:16:03.610935Z","iopub.execute_input":"2025-11-17T05:16:03.611273Z","iopub.status.idle":"2025-11-17T05:16:04.321258Z","shell.execute_reply.started":"2025-11-17T05:16:03.611245Z","shell.execute_reply":"2025-11-17T05:16:04.320559Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"We get a slightly lower validation RMSE of 0.6693, implying a better performance than our Ridge model. However, we also see a significant difference between the train and validation RMSE's, with a train RMSE of 0.5296, evidence of substantial overfitting.","metadata":{}},{"cell_type":"code","source":"# print overfitting analysis\nprint(\"\\nOverfitting Analysis:\")\nridge_overfit = final_ridge_val_rmse - final_ridge_train_rmse\nlgb_overfit = lgb_val_rmse - lgb_train_rmse\nprint(f\"   - Ridge overfitting: {ridge_overfit:.4f}\")\nprint(f\"   - LightGBM overfitting: {lgb_overfit:.4f}\")\nif ridge_overfit < lgb_overfit:\n    print(\"   - Ridge is more robust (less overfitting)\")\nelse:\n    print(\"   - LightGBM is more robust (less overfitting)\")","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-11-17T05:16:04.321864Z","iopub.execute_input":"2025-11-17T05:16:04.322097Z","iopub.status.idle":"2025-11-17T05:16:04.329208Z","shell.execute_reply.started":"2025-11-17T05:16:04.322076Z","shell.execute_reply":"2025-11-17T05:16:04.328318Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"🥡 ***Key Takeaways:***\n\n‣ Both models preformed fairly well on the validation set, with the LightGBM model performing slightly better.\n\n‣ Both models show signs of overfitting, but much moreso in the LightGBM model.","metadata":{}},{"cell_type":"markdown","source":"***\n\n## 🎛️ ***Tuning Our Models!***\n\nRidge Regression only has one parameter to tune *alpha*, which controls the strength of the penalty applied to the model's coefficients. To experiment with different alpha, we can use a **grid search cross-validation function**.","metadata":{}},{"cell_type":"code","source":"# determine best alpha for Ridge\nridge = Ridge(random_state=42)\nparam_grid = {'alpha': [0.01, 0.1, 1.0, 10.0, 100.0]}\ngrid_search = GridSearchCV(ridge, param_grid, cv=5, scoring='neg_root_mean_squared_error')\ngrid_search.fit(X_train_ridge_final, y_train)\nprint(f\"Best alpha: {grid_search.best_params_['alpha']}\")\n\n# retrain with best alpha from CV\nbest_alpha = grid_search.best_params_['alpha']\nridge_model_final = Ridge(alpha=best_alpha, random_state=42)\nridge_model_final.fit(X_train_ridge_final, y_train)\n\n# re-evaluate Ridge model on validation set with best alpha\ny_train_pred = ridge_model_final.predict(X_train_ridge_final)\ny_val_pred = ridge_model_final.predict(X_val_ridge_final)\n\n# calculate train and validation RMSE\ntrain_rmse = np.sqrt(mean_squared_error(y_train, y_train_pred))\nval_rmse = np.sqrt(mean_squared_error(y_val, y_val_pred))\nprint(f\"\\nFinal Ridge Model with alpha={best_alpha}:\")\nprint(f\"  Train RMSE: {train_rmse:.4f}\")\nprint(f\"  Validation RMSE: {val_rmse:.4f}\")","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-11-17T05:16:04.330403Z","iopub.execute_input":"2025-11-17T05:16:04.330864Z","iopub.status.idle":"2025-11-17T05:16:04.438891Z","shell.execute_reply.started":"2025-11-17T05:16:04.330837Z","shell.execute_reply":"2025-11-17T05:16:04.438049Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"Slight improvement!\n\nFor LightGBM there are many parameters to choose from, but we'll focus on a simple configuration focusing on learning speed, tree complexity and overfitting. With these parameters, grid search will go through 192 different configurations.","metadata":{}},{"cell_type":"code","source":"# define hyperparameters for tuning\nparam_grid = {'learning_rate': [0.01, 0.05], 'num_leaves': [15, 31], 'max_depth': [-1, 5], 'colsample_bytree': [0.8, 1.0], 'subsample': [0.8, 1.0], 'min_child_samples': [20, 50]}\nlgb_params['verbose'] = -1\n\n# define new model for parameter testing\nlgb_model = LGBMRegressor(objective='regression', random_state=42, n_estimators=1000, verbose=-1)\ngrid_search = GridSearchCV(estimator=lgb_model, param_grid=param_grid, scoring='neg_root_mean_squared_error', cv=3, verbose=0, n_jobs=-1)\n\n# fit new parameters on train lgb data\ngrid_search.fit(X_train_lgb, y_train)\nprint(f\"Best Parameters:\\n{grid_search.best_params_}\")\nprint(f\"Best CV RMSE: {-grid_search.best_score_:.4f}\")\nbest_lgb_params = grid_search.best_params_\n\n# retrain final model with tuned parameters\nfinal_lgb_model = LGBMRegressor(**best_lgb_params, objective='regression', random_state=42, n_estimators=1000)\n\n# fit final model on full training set\nfinal_lgb_model.fit(X_train_lgb, y_train,eval_set=[(X_val_lgb, y_val)], eval_metric='rmse', callbacks=[early_stopping(stopping_rounds=50), log_evaluation(0)])\n\n# evaluate tuned/ trained model on train and validation sets for final result\ny_train_pred = final_lgb_model.predict(X_train_lgb)\ny_val_pred = final_lgb_model.predict(X_val_lgb)\ntrain_rmse = np.sqrt(mean_squared_error(y_train, y_train_pred))\nval_rmse = np.sqrt(mean_squared_error(y_val, y_val_pred))\nprint(f\"\\nTuned LightGBM Results:\")\nprint(f\"  Train RMSE: {train_rmse:.4f}\")\nprint(f\"  Validation RMSE: {val_rmse:.4f}\")","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-11-17T05:16:04.439867Z","iopub.execute_input":"2025-11-17T05:16:04.440123Z","iopub.status.idle":"2025-11-17T05:23:54.339176Z","shell.execute_reply.started":"2025-11-17T05:16:04.440102Z","shell.execute_reply":"2025-11-17T05:23:54.337873Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"The tuned LightGBM model has a validation RMSE of 0.6657, a slight improvement over the previous 0.6693. But what is likely more significant is the increase in RMSE from 0.5296 to 0.5901 - combined with the decreased validation RMSE, this nearly cuts model overfitting in half, from 0.1397 to 0.0756.\n\nThe Ridge model is saved as *ridge_model_final* and the LightGBM model is saved as *final_lgb_model*.","metadata":{}},{"cell_type":"markdown","source":"***\n\n## 🧮 ***Evaluating the Test Set***\n\nTo evaluate on the test set and get predictions, we need to engineer the same features from the train/val sets onto the test set.","metadata":{}},{"cell_type":"code","source":"# convert sex binary to string (we changed this in the train set in the EDA section)\ntest['Basic_Demos-Sex'] = test['Basic_Demos-Sex'].map({0: 'Male', 1: 'Female'})\n\n# separate id from test set\ntest_ids = test[\"id\"].copy()\ntest = test.drop(\"id\", axis=1)\n\n# update test set with engineered features\nX_test_engineered = create_engineered_features(test, correlations)","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-11-17T05:23:54.340547Z","iopub.execute_input":"2025-11-17T05:23:54.340909Z","iopub.status.idle":"2025-11-17T05:23:54.364848Z","shell.execute_reply.started":"2025-11-17T05:23:54.340880Z","shell.execute_reply":"2025-11-17T05:23:54.363334Z"}},"outputs":[],"execution_count":null},{"cell_type":"raw","source":"For Ridge:","metadata":{}},{"cell_type":"code","source":"# define helper function to round predictions\ndef clip_predictions(predictions):\n    rounded = np.round(predictions)\n    clipped = np.clip(rounded, 0, 3)\n    return clipped\n\n# define Ridge preprocessing function for test set\ndef preprocess_test_for_ridge(X_test, categorical_cols, numerical_cols, selected_numerical_features, scaler, encoded_column_names):\n    X_test_processed = X_test.copy()\n    \n    # missing values (fill numerical with median, categorical with mode)\n    for col in numerical_cols:\n        if col in X_test_processed.columns:\n            median_val = X_test_processed[col].median()\n            X_test_processed[col] = X_test_processed[col].fillna(median_val)\n    for col in categorical_cols:\n        if col in X_test_processed.columns:\n            mode_val = X_test_processed[col].mode()[0] if not X_test_processed[col].mode().empty else 'Unknown'\n            X_test_processed[col] = X_test_processed[col].fillna(mode_val)\n\n    # keep selected features and one-hot encode\n    features_to_keep = selected_numerical_features + categorical_cols\n    X_test_processed = X_test_processed[features_to_keep]\n    X_test_encoded = pd.get_dummies(X_test_processed, columns=categorical_cols, drop_first=True)\n\n    # add absent columns, reorder columns and scale\n    for col in encoded_column_names:\n        if col not in X_test_encoded.columns:\n            X_test_encoded[col] = 0\n    X_test_encoded = X_test_encoded[encoded_column_names]\n    X_test_scaled = scaler.transform(X_test_encoded)\n    return X_test_scaled\n\n# apply Ridge preprocessing function to test set\nX_test_scaled = preprocess_test_for_ridge(X_test_engineered, categorical_cols = categorical_cols_eng, numerical_cols = numerical_cols_eng, selected_numerical_features = selected_features, scaler = scaler, encoded_column_names = encoded_cols)\n\n# make predictions\npredictions_ridge = ridge_model_final.predict(X_test_scaled)\ny_pred_ridge = clip_predictions(predictions_ridge)\n\n# generate Ridge predictions as dataframe\nsubmission_ridge = pd.DataFrame({'id': test_ids, 'SII': y_pred_ridge})\ndisplay(submission_ridge)","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-11-17T05:23:54.365689Z","iopub.execute_input":"2025-11-17T05:23:54.365999Z","iopub.status.idle":"2025-11-17T05:23:54.441997Z","shell.execute_reply.started":"2025-11-17T05:23:54.365977Z","shell.execute_reply":"2025-11-17T05:23:54.440889Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"If you recall in the EDA, about 85% of participants reported an SII of 0 or 1, leaving only 15% of participants with an SII of 2 or 3. It's fairly likely that one of these 2's and 3's just didn't slip into the test.\n\nNow same process for the LightGBM model:","metadata":{}},{"cell_type":"code","source":"# define LightGBM preprocessing function for test set\ndef preprocess_test_for_lightgbm(X_test, categorical_cols, numerical_cols, label_encoders):\n    X_test_processed = X_test.copy()\n    \n    # missing values for numerical columns\n    for col in numerical_cols:\n        if col in X_test_processed.columns:\n            median_val = X_test_processed[col].median()\n            X_test_processed[col] = X_test_processed[col].fillna(median_val)\n    \n    # label encode using training encoders\n    for col in categorical_cols:\n        if col in X_test_processed.columns:\n            le = label_encoders.get(col)\n            if le:\n                mode_val = X_test_processed[col].mode()[0] if not X_test_processed[col].mode().empty else 'Unknown'\n                X_test_processed[col] = X_test_processed[col].fillna(mode_val)\n                train_categories = set(le.classes_)\n                X_test_processed[col] = X_test_processed[col].map(lambda x: le.transform([x])[0] if x in train_categories else 0).astype(int)\n            else:\n                X_test_processed[col] = 0\n    return X_test_processed\n\n# apply LightGBM preprocessing function to test set\nX_test_lgb = preprocess_test_for_lightgbm(X_test_engineered, categorical_cols = categorical_cols_eng, numerical_cols = numerical_cols_eng, label_encoders = lgb_encoders)\n\n# make predictions\npredictions_lgb = final_lgb_model.predict(X_test_lgb)\ny_pred_lgb = clip_predictions(predictions_lgb)\n\n# generate LightGBM predictions as dataframe\nsubmission_gbm = pd.DataFrame({'id': test_ids, 'SII': y_pred_lgb})\ndisplay(submission_gbm)","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-11-17T05:23:54.442966Z","iopub.execute_input":"2025-11-17T05:23:54.443302Z","iopub.status.idle":"2025-11-17T05:23:54.508508Z","shell.execute_reply.started":"2025-11-17T05:23:54.443272Z","shell.execute_reply":"2025-11-17T05:23:54.507230Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"While the competition ended a while back, I'll still submit a prediction. We'll go with... Ridge model :)","metadata":{}},{"cell_type":"code","source":"# calcuate weighted average of predictions\nsubmission_ridge.to_csv('submission.csv', index=False)","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-11-17T05:23:54.509567Z","iopub.execute_input":"2025-11-17T05:23:54.509827Z","iopub.status.idle":"2025-11-17T05:23:54.521267Z","shell.execute_reply.started":"2025-11-17T05:23:54.509807Z","shell.execute_reply":"2025-11-17T05:23:54.520186Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"Thanks for checking out my first ML project! Hope it was interesting, informative, illuminating, and totally incredible! Cheers, Tom","metadata":{}}]}