{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"name":"python","version":"3.10.14","mimetype":"text/x-python","codemirror_mode":{"name":"ipython","version":3},"pygments_lexer":"ipython3","nbconvert_exporter":"python","file_extension":".py"},"kaggle":{"accelerator":"none","dataSources":[{"sourceId":81933,"databundleVersionId":9643020,"sourceType":"competition"}],"dockerImageVersionId":30775,"isInternetEnabled":false,"language":"python","sourceType":"notebook","isGpuEnabled":false}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"# Introduction\n\nThe rapid rise of internet usage among children and adolescents has brought a pressing concern into focus: **Problematic Internet Use (PIU)**. This phenomenon, characterized by excessive or unhealthy engagement with the digital world, can significantly disrupt physical health, emotional well-being, and social interactions. Left unchecked, PIU can lead to adverse outcomes, including poor mental health, disrupted daily routines, and diminished physical activity. Recognizing these behaviors early is crucial to enabling timely interventions that foster healthier digital habits and overall well-being.\n\nThis project seeks to predict the **Severity Impairment Index (SII)**, a critical measure of PIU, by analyzing data on adolescents' physical activity and fitness levels. By uncovering meaningful patterns and trends, we aim to provide actionable insights that can inform interventions and support balanced digital engagement.\n\nIn this notebook, I will:\n\n* **Analyze the dataset** to explore correlations between physical activity, fitness patterns, and internet use behaviors.\n* **Prepare and preprocess the data** to ensure it is clean, comprehensive, and ready for modeling.\n* **Develop and fine-tune machine learning models** to predict PIU with a focus on accuracy and precision.\n* **Evaluate the models** using the Quadratic Weighted Kappa (QWK) score, the primary evaluation metric for this competition.\n\n  \nThe overarching goal is to leverage data-driven insights to empower caregivers, educators, and mental health professionals. By identifying unhealthy internet usage patterns early, we can help adolescents maintain mental health, stay physically active, and develop healthier relationships with technology.","metadata":{}},{"cell_type":"code","source":"# import libraries \n\nimport numpy as np \nimport pandas as pd","metadata":{"execution":{"iopub.status.busy":"2024-12-12T22:56:28.243324Z","iopub.execute_input":"2024-12-12T22:56:28.243701Z","iopub.status.idle":"2024-12-12T22:56:28.664744Z","shell.execute_reply.started":"2024-12-12T22:56:28.243667Z","shell.execute_reply":"2024-12-12T22:56:28.663614Z"},"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"# 1. Dataset Overview\n\nThe Child Mind Institute’s **Healthy Brain Network (HBN)** provides a unique opportunity to explore the interplay between physical activity, behavioral patterns, and problematic internet use (PIU) in children and adolescents. The dataset combines **survey-based tabular data** with **time-series actigraphy data**, offering a comprehensive view of participants' physical activity, health metrics, and internet behaviors. The ultimate objective is to predict the **Severity Impairment Index (SII)**—a key measure of PIU derived from the Parent-Child Internet Addiction Test (PCIAT). The SII categorizes internet use into four levels: None, Mild, Moderate, and Severe.\n\n\n\nThis project seeks to leverage this rich dataset to uncover early signs of PIU, paving the way for targeted mental health interventions.\n\n\n1. **Tabular Data (CSV Files)**\n \nThe tabular dataset serves as a foundation for understanding participants' demographic, physical, and behavioral profiles. Key areas of focus include:\n\n* **Demographics**: Participant attributes like age, gender, and enrollment season.\n\n* **Physical Health Metrics**: Measures such as BMI, blood pressure, heart rate, weight, and waist circumference.\n\n* **Fitness Assessments**: Results from tests evaluating aerobic capacity, muscular strength, and endurance.\n\n* **Internet Usage Patterns**: Self-reported daily hours spent on computers and the internet.\n\n* **Parent-Child Internet Addiction Test (PCIAT)**: The PCIAT scores underpin the SII, which is the competition's target variable.\n\n* **Sleep Quality Surveys**: Insights into participants' sleep patterns and disturbances.\n\n\n\nThis structured data provides a static snapshot of participants’ behaviors and physical health.\n\n\n\n\n2. **Actigraphy Data (Parquet Files)**\n\nThe actigraphy dataset, derived from wrist-worn accelerometers, complements the tabular data by capturing dynamic physical activity patterns. It includes:\n\n* **Motion Data (X, Y, Z Axes)**: Tracks movement intensity across three axes.\n\n* **ENMO**: A measure of overall movement intensity.\n\n* **Non-wear Flag**: Highlights periods when the device was not worn, indicating inactivity.\n\n* **Light Exposure**: Captures environmental light levels during activity.\n\n* **Battery Voltage**: Ensures data reliability by tracking device status.\n\n* **Time of Day and Weekday**: Helps identify trends across different times and days.\n\n\nTogether, these data points offer a temporal perspective on physical activity and lifestyle.\n\n\n\n\n**Combining the Datasets**\n\nThe strength of this dataset lies in the integration of static and dynamic features. By merging survey-based behaviors with real-time activity data, we aim to identify the key factors contributing to PIU. This dual-layered approach not only enhances prediction accuracy but also enables a deeper understanding of how physical activity and internet use interact in adolescents’ daily lives.\n\n\n\n\n\n**Challenge and Evaluation Metric**\n\n\nThe competition’s evaluation metric, the **Quadratic Weighted Kappa (QWK)** score, measures the agreement between predicted and actual SII levels. This nuanced metric ensures that models accurately capture the distinctions between the severity levels of internet use, emphasizing precision in predictions.\n\n\n","metadata":{}},{"cell_type":"markdown","source":"# Loading Training Data \n","metadata":{}},{"cell_type":"code","source":"# loading the tabular data files\ntrain_data = pd.read_csv('/kaggle/input/child-mind-institute-problematic-internet-use/train.csv')\n\n# Displaying the first few rows of the dataframe to get an initial look at the data\nprint(train_data.head())\n","metadata":{"execution":{"iopub.status.busy":"2024-12-12T22:56:31.860783Z","iopub.execute_input":"2024-12-12T22:56:31.861268Z","iopub.status.idle":"2024-12-12T22:56:31.960621Z","shell.execute_reply.started":"2024-12-12T22:56:31.861235Z","shell.execute_reply":"2024-12-12T22:56:31.959316Z"},"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"code","source":"print(train_data.columns)","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-12-12T22:56:34.381362Z","iopub.execute_input":"2024-12-12T22:56:34.381753Z","iopub.status.idle":"2024-12-12T22:56:34.388489Z","shell.execute_reply.started":"2024-12-12T22:56:34.381720Z","shell.execute_reply":"2024-12-12T22:56:34.387216Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"# Loading Test Data \n","metadata":{}},{"cell_type":"code","source":"# loading the tabular data files\ntest_data = pd.read_csv('/kaggle/input/child-mind-institute-problematic-internet-use/test.csv')\n\n# Displaying the first few rows of the dataframe to get an initial look at the data\nprint(test_data.head())","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-12-12T22:56:36.492999Z","iopub.execute_input":"2024-12-12T22:56:36.493368Z","iopub.status.idle":"2024-12-12T22:56:36.516575Z","shell.execute_reply.started":"2024-12-12T22:56:36.493336Z","shell.execute_reply":"2024-12-12T22:56:36.515282Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"print(test_data.columns)","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-12-12T22:56:38.873755Z","iopub.execute_input":"2024-12-12T22:56:38.874159Z","iopub.status.idle":"2024-12-12T22:56:38.880513Z","shell.execute_reply.started":"2024-12-12T22:56:38.874122Z","shell.execute_reply":"2024-12-12T22:56:38.879270Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"common_features = list(set(train_data.columns) & set(test_data.columns))\nprint(\"Common features:\", common_features)\n","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-12-12T22:56:42.045607Z","iopub.execute_input":"2024-12-12T22:56:42.046028Z","iopub.status.idle":"2024-12-12T22:56:42.052281Z","shell.execute_reply.started":"2024-12-12T22:56:42.045990Z","shell.execute_reply":"2024-12-12T22:56:42.051051Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"# Train and Test csv comparison\n\nIt is evident that the train_data and test_data CSV files have differing sets of columns. Specifically:\n\n* The train dataset (train_data) includes additional columns like PCIAT-PCIAT_01 through PCIAT-PCIAT_20, PCIAT-PCIAT_Total, and sii which are absent in test_data.\n\n* The test dataset (test_data) has columns like Fitness_Endurance-Time_Mins, Fitness_Endurance-Time_Sec, PAQ_A-Season, and PAQ_A-PAQ_A_Total that are missing in the train dataset.","metadata":{}},{"cell_type":"markdown","source":"# Load Series Train Parquet file ","metadata":{}},{"cell_type":"code","source":"import os\nimport pandas as pd\nimport gc\n\n# Base path for Parquet files\nparquet_base_path = '/kaggle/input/child-mind-institute-problematic-internet-use/series_train.parquet'\n\n# List of participant IDs\nparticipant_ids = [\n    folder for folder in os.listdir(parquet_base_path) \n    if os.path.isdir(os.path.join(parquet_base_path, folder))\n]\n\naggregated_data = []\n\nfor participant_id in participant_ids:\n    try:\n        # Load and aggregate Parquet data\n        parquet_path = os.path.join(parquet_base_path, participant_id, 'part-0.parquet')\n        data = pd.read_parquet(parquet_path)\n\n        aggregated_features = data.agg({\n            'X': ['mean', 'std', 'min', 'max'],\n            'Y': ['mean', 'std', 'min', 'max'],\n            'Z': ['mean', 'std', 'min', 'max'],\n            'enmo': ['mean', 'std', 'min', 'max'],\n            'light': ['mean', 'std', 'min', 'max'],\n            'battery_voltage': ['mean', 'std', 'min', 'max']\n        }).stack().to_frame().T\n\n        aggregated_features['id'] = participant_id.split('=')[1]\n        aggregated_data.append(aggregated_features)\n\n        # Clear memory for the current participant\n        del data\n        gc.collect()\n\n    except Exception as e:\n        print(f\"Error processing participant {participant_id}: {e}\")\n\n# Combine aggregated data\naggregated_df = pd.concat(aggregated_data, ignore_index=True)\n\n# Downcast numerical columns\nfor col in aggregated_df.select_dtypes(include=['float64']).columns:\n    aggregated_df[col] = aggregated_df[col].astype('float32')\n\n# Flatten MultiIndex columns\naggregated_df.columns = [\n    f\"{stat}_{feature}\" if isinstance(stat, str) else feature\n    for stat, feature in aggregated_df.columns\n]\naggregated_df = aggregated_df.rename(columns={\"id_\": \"id\"})\n\n# Ensure compatible data types for merging\ntrain_data['id'] = train_data['id'].astype(str)\naggregated_df['id'] = aggregated_df['id'].astype(str)\n\n# Merge training CSV with Parquet features using a left join\ntry:\n    combined_data = pd.merge(train_data, aggregated_df, on='id', how='left')\n    print(\"Merged successfully!\")\n    print(combined_data.head())\nexcept Exception as e:\n    print(f\"Error during merging: {e}\")\n\n# Save combined data\ncombined_data.to_csv('combined_train_data.csv', index=False)\nprint(\"Aggregated features and metadata combined successfully!\")\n","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-12-12T22:56:44.331264Z","iopub.execute_input":"2024-12-12T22:56:44.331705Z","iopub.status.idle":"2024-12-12T22:59:45.080034Z","shell.execute_reply.started":"2024-12-12T22:56:44.331673Z","shell.execute_reply":"2024-12-12T22:59:45.078897Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"# Combined Training Data \nTo ensure the training dataset is enriched with meaningful features while managing memory efficiently, we processed the Parquet files corresponding to individual participants. For each file, key metrics like mean, standard deviation, minimum, and maximum were calculated for sensor data (e.g., X, Y, Z coordinates, light, and battery voltage). These aggregated features were then consolidated into a single DataFrame, ensuring compatibility with the training CSV by aligning participant IDs.\n\nTo optimize memory usage, numerical columns were downcast to smaller data types, and garbage collection was applied after processing each participant's file. Finally, the enriched features were merged with the training CSV to create a comprehensive training dataset, ready for model building. This approach balances feature engineering and resource management, ensuring scalability for larger datasets.","metadata":{}},{"cell_type":"code","source":"# Display the columns and shape of the combined training DataFrame\nprint(\"Columns in combined_data:\")\nprint(combined_data.columns)\n\nprint(\"\\nShape of combined_data:\")\nprint(combined_data.shape)\n","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-12-12T23:00:06.668996Z","iopub.execute_input":"2024-12-12T23:00:06.669794Z","iopub.status.idle":"2024-12-12T23:00:06.681452Z","shell.execute_reply.started":"2024-12-12T23:00:06.669741Z","shell.execute_reply":"2024-12-12T23:00:06.680292Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"# Handling Missing Values in Combined Training Data","metadata":{}},{"cell_type":"code","source":"# Handle missing values in training data\n# Identify numerical and categorical columns in combined_data only\nnumerical_columns = combined_data.select_dtypes(include=['number']).columns\ncategorical_columns = combined_data.select_dtypes(include=['object']).columns\n\n# Fill missing values for numerical columns using training data means\ncombined_data[numerical_columns] = combined_data[numerical_columns].fillna(combined_data[numerical_columns].mean())\n\n# Fill missing values for categorical columns using training data modes\nfor col in categorical_columns:\n    if combined_data[col].isnull().any():\n        combined_data[col] = combined_data[col].fillna(combined_data[col].mode()[0])\n\nprint(\"Missing values in training data handled successfully!\")\n","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-12-12T23:00:16.808493Z","iopub.execute_input":"2024-12-12T23:00:16.808906Z","iopub.status.idle":"2024-12-12T23:00:16.869308Z","shell.execute_reply.started":"2024-12-12T23:00:16.808871Z","shell.execute_reply":"2024-12-12T23:00:16.868159Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"# Training the Model (RandomForest) ","metadata":{}},{"cell_type":"code","source":"from imblearn.over_sampling import SMOTE\nfrom sklearn.preprocessing import OneHotEncoder\nfrom sklearn.ensemble import RandomForestClassifier\nfrom sklearn.pipeline import Pipeline\nimport pandas as pd\n\ntarget_column = 'PCIAT-PCIAT_Total'\nif target_column not in combined_data.columns:\n    raise ValueError(f\"Target column '{target_column}' is missing in the combined dataset.\")\n\n# Select features from training data only\nfeatures = [col for col in combined_data.columns if col not in [target_column, 'id']]\nX_train = combined_data[features]\ny_train = combined_data[target_column]\n\n# Convert the target variable into categorical bins (if classification is intended)\ny_train = pd.cut(y_train, bins=[-1, 50, 100, 150, 200], labels=[0, 1, 2, 3]).astype('int')\n\n# Identify numerical and categorical columns\nnumerical_columns = X_train.select_dtypes(include=['number']).columns\ncategorical_columns = X_train.select_dtypes(include=['object']).columns\n\n# Ensure categorical columns are all strings\ncombined_data[categorical_columns] = combined_data[categorical_columns].astype(str)\n\n# Encode categorical features\nencoder = OneHotEncoder(handle_unknown='ignore', sparse_output=False)\nX_train_encoded = pd.DataFrame(encoder.fit_transform(X_train[categorical_columns]))\nX_train_encoded.columns = encoder.get_feature_names_out(categorical_columns)\nX_train_encoded.index = X_train.index\n\n# Combine numerical and encoded features\nX_train_preprocessed = pd.concat([X_train[numerical_columns], X_train_encoded], axis=1)\n\n# Apply SMOTE to handle class imbalance\nsmote = SMOTE(random_state=42)\nX_train_balanced, y_train_balanced = smote.fit_resample(X_train_preprocessed, y_train)\n\n# Train the model\npipeline = Pipeline([\n    ('classifier', RandomForestClassifier(n_estimators=100, random_state=42))\n])\npipeline.fit(X_train_balanced, y_train_balanced)\n\nprint(\"Model training completed successfully!\")\n","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-12-12T23:00:19.540957Z","iopub.execute_input":"2024-12-12T23:00:19.541364Z","iopub.status.idle":"2024-12-12T23:00:24.014716Z","shell.execute_reply.started":"2024-12-12T23:00:19.541326Z","shell.execute_reply":"2024-12-12T23:00:24.013091Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"feature_names = X_train_balanced.columns.tolist()\nfeature_names","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-12-12T23:00:31.332048Z","iopub.execute_input":"2024-12-12T23:00:31.332616Z","iopub.status.idle":"2024-12-12T23:00:31.342224Z","shell.execute_reply.started":"2024-12-12T23:00:31.332579Z","shell.execute_reply":"2024-12-12T23:00:31.341130Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"train_numeric_means = X_train[numerical_columns].mean()\ntrain_categorical_modes = {col: X_train[col].mode()[0] for col in categorical_columns}\nprint(\"Stored training means and modes for use in test preprocessing.\")\n\n","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-12-12T23:00:40.844276Z","iopub.execute_input":"2024-12-12T23:00:40.844690Z","iopub.status.idle":"2024-12-12T23:00:40.876761Z","shell.execute_reply.started":"2024-12-12T23:00:40.844653Z","shell.execute_reply":"2024-12-12T23:00:40.875459Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"# Load Series Test Parquet file","metadata":{}},{"cell_type":"code","source":"import os\nimport pandas as pd\nimport gc\n\n# Paths for the test data\ntest_csv_path = '/kaggle/input/child-mind-institute-problematic-internet-use/test.csv'\ntest_parquet_base_path = '/kaggle/input/child-mind-institute-problematic-internet-use/series_test.parquet'\n\ntest_data = pd.read_csv(test_csv_path)\n\n# Aggregate the test parquet data\ntest_participant_ids = [\n    folder for folder in os.listdir(test_parquet_base_path)\n    if os.path.isdir(os.path.join(test_parquet_base_path, folder))\n]\n\ntest_aggregated_data = []\nfor participant_id in test_participant_ids:\n    parquet_path = os.path.join(test_parquet_base_path, participant_id, 'part-0.parquet')\n    data = pd.read_parquet(parquet_path)\n\n    aggregated_features = data.agg({\n        'X': ['mean', 'std', 'min', 'max'],\n        'Y': ['mean', 'std', 'min', 'max'],\n        'Z': ['mean', 'std', 'min', 'max'],\n        'enmo': ['mean', 'std', 'min', 'max'],\n        'light': ['mean', 'std', 'min', 'max'],\n        'battery_voltage': ['mean', 'std', 'min', 'max']\n    }).stack().to_frame().T\n\n    aggregated_features['id'] = participant_id.split('=')[1]\n    test_aggregated_data.append(aggregated_features)\n\n    del data\n    gc.collect()\n\ntest_aggregated_df = pd.concat(test_aggregated_data, ignore_index=True)\ntest_aggregated_df.columns = [\n    f\"{stat}_{feature}\" if isinstance(stat, str) else feature\n    for stat, feature in test_aggregated_df.columns\n]\ntest_aggregated_df = test_aggregated_df.rename(columns={\"id_\": \"id\"})\n\ntest_combined_data = pd.merge(test_data, test_aggregated_df, on='id', how='left')\nprint(\"Merged test data shape:\", test_combined_data.shape)\n\n\n# Ensure all features exist\nmissing_features = [col for col in features if col not in test_combined_data.columns]\nfor col in missing_features:\n    if col in categorical_columns:\n        test_combined_data.loc[:, col] = \"missing\"\n    else:\n        test_combined_data.loc[:, col] = 0.0\n\nX_test = test_combined_data[features].copy()\n\n# Impute missing values using training stats\nX_test.loc[:, numerical_columns] = X_test[numerical_columns].fillna(train_numeric_means)\nfor col in categorical_columns:\n    X_test.loc[:, col] = X_test[col].fillna(train_categorical_modes[col])\n\n# Ensure categorical columns are strings AFTER imputation and BEFORE encoding\nX_test.loc[:, categorical_columns] = X_test[categorical_columns].astype(str)\n\n# Use the encoder's recorded feature names to maintain consistent ordering\ntrain_categorical_order = encoder.feature_names_in_\n\nfor col in train_categorical_order:\n    if col not in X_test.columns:\n        X_test.loc[:, col] = \"missing\"\n\nfinal_order = list(numerical_columns) + list(train_categorical_order)\nX_test = X_test.reindex(columns=final_order)\n\n# Transform using the encoder (now that all categorical columns are strings)\nX_test_encoded = pd.DataFrame(encoder.transform(X_test[train_categorical_order]))\nX_test_encoded.columns = encoder.get_feature_names_out(train_categorical_order)\nX_test_encoded.index = X_test.index\n\nX_test_preprocessed = pd.concat([X_test[numerical_columns], X_test_encoded], axis=1)\nX_test_preprocessed = X_test_preprocessed.reindex(columns=feature_names, fill_value=0)\n\npredictions = pipeline.predict(X_test_preprocessed)\nprint(\"Predictions on test data:\")\nprint(predictions)","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-12-12T23:00:44.208640Z","iopub.execute_input":"2024-12-12T23:00:44.209054Z","iopub.status.idle":"2024-12-12T23:00:44.692267Z","shell.execute_reply.started":"2024-12-12T23:00:44.209015Z","shell.execute_reply":"2024-12-12T23:00:44.691204Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"from sklearn.model_selection import cross_val_score\n\nscores = cross_val_score(pipeline, X_train_preprocessed, y_train, cv=5, scoring='accuracy')\nprint(\"Cross-Validation Scores:\", scores)\nprint(\"Mean Cross-Validation Accuracy:\", scores.mean())\n","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-12-12T23:00:50.030232Z","iopub.execute_input":"2024-12-12T23:00:50.030662Z","iopub.status.idle":"2024-12-12T23:00:52.814680Z","shell.execute_reply.started":"2024-12-12T23:00:50.030624Z","shell.execute_reply":"2024-12-12T23:00:52.813534Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"# Create a submission DataFrame\n# 'id' to match the identifier in the test data\nsubmission = pd.DataFrame({\n    'id': test_combined_data['id'],  # 'id' column from test_combined_data\n    'sii': predictions  \n})\n\n# Save the submission file\nsubmission.to_csv('submission.csv', index=False)\nprint(\"Submission file created: submission.csv\")\n","metadata":{"execution":{"iopub.status.busy":"2024-12-12T23:08:38.107356Z","iopub.execute_input":"2024-12-12T23:08:38.107820Z","iopub.status.idle":"2024-12-12T23:08:38.117520Z","shell.execute_reply.started":"2024-12-12T23:08:38.107783Z","shell.execute_reply":"2024-12-12T23:08:38.116070Z"},"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"# 3. Exploratory Data Analysis \n\nIn this section, I will conduct an exploratory data analysis (EDA) to gain insights into the structure, distribution, and relationships between features. EDA is crucial for understanding the underlying patterns within the dataset, identifying any anomalies or outliers, and informing further modeling decisions. \n\n","metadata":{}},{"cell_type":"code","source":"import pandas as pd\nimport matplotlib.pyplot as plt\n\n# Load the time-series data for a specific participant ID from the Parquet file\nparticipant_id = '00f332d1'\nparquet_file_path = f'/kaggle/input/child-mind-institute-problematic-internet-use/series_train.parquet/id={participant_id}'\ndata = pd.read_parquet(parquet_file_path)\n\n# Inspect the first few values of 'time_of_day' to understand its format\nprint(data['time_of_day'].head())\n\n# Convert 'time_of_day' to seconds since midnight if it's a datetime.time object\nif data['time_of_day'].dtype == 'object':\n    # Convert the time strings to datetime.time objects, then to seconds since midnight\n    data['time_of_day'] = pd.to_datetime(data['time_of_day'], format='%H:%M:%S.%f').dt.hour * 3600 + \\\n                          pd.to_datetime(data['time_of_day'], format='%H:%M:%S.%f').dt.minute * 60 + \\\n                          pd.to_datetime(data['time_of_day'], format='%H:%M:%S.%f').dt.second\n\n# Plot X, Y, Z accelerometer data\nplt.figure(figsize=(12, 6))\nplt.plot(data['time_of_day'], data['X'], label='X-axis', alpha=0.7, color='blue')\nplt.plot(data['time_of_day'], data['Y'], label='Y-axis', alpha=0.7, color='orange')\nplt.plot(data['time_of_day'], data['Z'], label='Z-axis', alpha=0.7, color='green')\nplt.xlabel('Time of Day (Seconds since Midnight)')\nplt.ylabel('Acceleration (g)')\nplt.legend()\nplt.title(f'Accelerometer Data (X, Y, Z) for Participant ID: {participant_id}')\nplt.xticks(rotation=45)\nplt.show()","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-12-12T00:57:06.134557Z","iopub.execute_input":"2024-12-12T00:57:06.136976Z","iopub.status.idle":"2024-12-12T00:57:09.327556Z","shell.execute_reply.started":"2024-12-12T00:57:06.136841Z","shell.execute_reply":"2024-12-12T00:57:09.326218Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"# Accelerometer Data (X, Y, Z):\n\nThe X, Y, and Z axes display variations corresponding to movements. The Z-axis shows consistent dominance, suggesting the orientation of the device or specific motion patterns.","metadata":{}},{"cell_type":"code","source":"\n# Plot ENMO and Angle-Z\nplt.figure(figsize=(12, 6))\nplt.plot(data['time_of_day'], data['enmo'], label='ENMO', linestyle='--', alpha=0.7, color='brown')\nplt.plot(data['time_of_day'], data['anglez'], label='Angle-Z', linestyle='--', alpha=0.7, color='purple')\nplt.xlabel('Time of Day (Seconds since Midnight)')\nplt.ylabel('Measurement')\nplt.legend()\nplt.title(f'ENMO and Angle-Z for Participant ID: {participant_id}')\nplt.xticks(rotation=45)\nplt.show()","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-12-12T00:57:09.329474Z","iopub.execute_input":"2024-12-12T00:57:09.329858Z","iopub.status.idle":"2024-12-12T00:57:13.216348Z","shell.execute_reply.started":"2024-12-12T00:57:09.329824Z","shell.execute_reply":"2024-12-12T00:57:13.215203Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"# ENMO and Angle-Z:\n\nENMO values reflect movement intensity, with periods of activity (higher values) and rest (near-zero values).\nAngle-Z represents the orientation of the device along the Z-axis, showing periodic variations as the participant moves.","metadata":{}},{"cell_type":"code","source":"# Plot Light and Battery Voltage\nplt.figure(figsize=(12, 6))\nplt.plot(data['time_of_day'], data['light'], label='Light (lux)', linestyle=':', alpha=0.7, color='pink')\nplt.plot(data['time_of_day'], data['battery_voltage'], label='Battery Voltage', linestyle=':', alpha=0.7, color='red')\nplt.xlabel('Time of Day (Seconds since Midnight)')\nplt.ylabel('Measurement')\nplt.legend()\nplt.title(f'Light and Battery Voltage for Participant ID: {participant_id}')\nplt.xticks(rotation=45)\nplt.show()","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-12-12T00:57:13.831777Z","iopub.execute_input":"2024-12-12T00:57:13.832194Z","iopub.status.idle":"2024-12-12T00:57:15.615615Z","shell.execute_reply.started":"2024-12-12T00:57:13.832157Z","shell.execute_reply":"2024-12-12T00:57:15.614503Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"# Light and Battery Voltage:\n\n* The light (in lux) fluctuates, indicating exposure to varying ambient light levels, likely representing day-night cycles or changes in indoor and outdoor environments.\n\n* The battery voltage remains relatively stable, reflecting the consistent functioning of the device during data collection.","metadata":{}},{"cell_type":"markdown","source":"# 3.1 Exploring the Relationship between Weight, BMI, and Phone Usage\n\nI will first explore the relationship between physical weight, BMI, and daily phone/internet usage. These variables provide critical insights into participants' physical health and their internet usage patterns, which are relevant for understanding problematic behaviors. By visualizing these dimensions together in a 3D scatter plot, we can identify potential patterns or trends that might help in building predictive models. \n\nThe 3D scatter plot provides an interactive view of how weight, BMI, and internet usage relate to each other across participants.\n\n* X-axis (Physical Weight): As the weight increases, we see that BMI tends to increase as well, which is expected since BMI is calculated using weight and height.\n\n* Y-axis (BMI): Participants' BMI ranges widely, with some higher outliers indicating potential obesity or other health conditions.\n\n* Z-axis and Color (Phone Usage in Hours/Day): The plot shows that phone usage varies across individuals but is capped at 3 hours per day, indicating that the dataset's maximum recorded value for daily phone/internet usage is 3 hours.\n\n\nThe color gradient makes it easy to identify participants with higher phone usage (yellow) versus those with lower usage (purple). It appears that high phone usage is not exclusively tied to any specific weight or BMI cluster, suggesting that problematic phone use could occur across various body types and physical profiles.\n\nThis plot provides a good starting point for further analysis, such as checking if higher internet usage correlates with lower physical fitness or higher BMI, which could offer insights into potential health impacts of prolonged screen time.\n\n","metadata":{}},{"cell_type":"markdown","source":"\n","metadata":{}},{"cell_type":"code","source":"import plotly.express as px\n\n# Ensure your dataset has relevant columns\nfig = px.scatter_3d(train_data, \n                    x='Physical-Weight', \n                    y='Physical-BMI', \n                    z='PreInt_EduHx-computerinternet_hoursday',\n                    color='PreInt_EduHx-computerinternet_hoursday', \n                    title='3D Scatter Plot: Weight, BMI, and Phone Usage',\n                    labels={'PreInt_EduHx-computerinternet_hoursday': 'Phone Usage (Hours/Day)'})\n\nfig.show()","metadata":{"execution":{"iopub.status.busy":"2024-12-12T00:58:12.170839Z","iopub.execute_input":"2024-12-12T00:58:12.171321Z","iopub.status.idle":"2024-12-12T00:58:12.237013Z","shell.execute_reply.started":"2024-12-12T00:58:12.171282Z","shell.execute_reply":"2024-12-12T00:58:12.235687Z"},"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"# Hourly Trends in Accelerometer Data\n\nThis plot shows hourly averages of accelerometer data (X, Y, Z axes) for a participant, aggregated over the course of the recorded time. Here's what it does and why it's important:\n\nThe accelerometer data measures the participant's wrist motion in three directions (X, Y, Z). By calculating hourly averages, the plot highlights motion trends or patterns in activity over time. The data aggregation logic in the plot is based on grouping the time-series data by hourly intervals and calculating the mean for each hour. \n\nX-Axis: Represents time, divided into hours of the day (converted from seconds since midnight).\n\nY-Axis: Shows the average acceleration (in \"g\", the gravitational force unit) for each axis.\n\n* Helps detect daily activity patterns like exercise, rest, or sleep times.\n  \n* Useful for understanding behavior or physical routines based on motion data.\n  \n* Can be correlated with other variables like time of day, light levels, or events.","metadata":{}},{"cell_type":"code","source":"# Resample time-series data by hourly average\ndata['hour'] = data['time_of_day'] // 3600\naggregated_data = data.groupby('hour').mean()\n\n# Plot aggregated data\nplt.figure(figsize=(12, 6))\nplt.plot(aggregated_data.index, aggregated_data['X'], label='X-axis', alpha=0.7, color='blue')\nplt.plot(aggregated_data.index, aggregated_data['Y'], label='Y-axis', alpha=0.7, color='orange')\nplt.plot(aggregated_data.index, aggregated_data['Z'], label='Z-axis', alpha=0.7, color='green')\nplt.xlabel('Hour of Day')\nplt.ylabel('Average Acceleration (g)')\nplt.legend()\nplt.title('Hourly Average Accelerometer Data')\nplt.show()","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-12-12T00:58:23.997803Z","iopub.execute_input":"2024-12-12T00:58:23.998199Z","iopub.status.idle":"2024-12-12T00:58:24.709004Z","shell.execute_reply.started":"2024-12-12T00:58:23.998167Z","shell.execute_reply":"2024-12-12T00:58:24.707690Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"# Insights : \n\n* Spikes in the lines indicate periods of high activity.\n  \n* Flat or minimal variations suggest periods of rest or low motion.\n \n* Variations between axes (X, Y, Z) could indicate directional movement or specific wrist motions.","metadata":{}},{"cell_type":"markdown","source":"# Exploring Correlations Between Accelerometer Variables\n\nNext, I aim to explore the relationships between accelerometer variables, such as X, Y, Z, ENMO, Angle-Z, light, and battery_voltage. By calculating and visualizing the correlations between these variables using a heatmap, I can identify patterns or dependencies. This helps in understanding which features are strongly or weakly related, which is essential for feature engineering and model selection in the next steps.\n","metadata":{}},{"cell_type":"code","source":"import seaborn as sns\nimport matplotlib.pyplot as plt\n\n# Select only accelerometer-related variables for correlation\naccelerometer_columns = ['X', 'Y', 'Z', 'enmo', 'anglez', 'light', 'battery_voltage']\ncorrelation_data = data[accelerometer_columns].corr()\n\n# Plot the heatmap\nplt.figure(figsize=(10, 8))\nsns.heatmap(correlation_data, annot=True, cmap='coolwarm', fmt=\".2f\", cbar=True)\nplt.title('Correlation Heatmap of Accelerometer Variables')\nplt.show()\n","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-12-12T00:58:27.654055Z","iopub.execute_input":"2024-12-12T00:58:27.655407Z","iopub.status.idle":"2024-12-12T00:58:28.500111Z","shell.execute_reply.started":"2024-12-12T00:58:27.655349Z","shell.execute_reply":"2024-12-12T00:58:28.498836Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"# Key Insights \n\n* High Correlation Between Z and Angle-Z: A strong positive correlation (~0.99) indicates that Z and Angle-Z are closely related. This might imply redundancy, and we can explore whether one of these features can be removed to avoid multicollinearity.\n\n* Low Correlation Between ENMO and Other Variables: ENMO shows low correlation with most other variables, suggesting it may provide independent information useful for modeling.\n  \n* Minimal Relationship Between Light and Accelerometer Data: Variables like light and battery_voltage have low correlation with X, Y, Z, and other accelerometer measurements, indicating they measure distinct aspects.\n  \n* No Strong Relationships Between X, Y, and Other Variables: X and Y do not show significant correlations with other variables, suggesting they capture independent axes of motion.\n","metadata":{}},{"cell_type":"markdown","source":"# Correlation Matrix for Weight, BMI, and Phone Usage\n\nIn this step, I explored the relationships between participants' physical characteristics (Physical-Weight and Physical-BMI) and their self-reported daily phone usage (PreInt_EduHx-computerinternet_hoursday). The goal was to identify any significant linear relationships that could inform clustering or further feature engineering for predictive modeling.\n","metadata":{}},{"cell_type":"code","source":"import seaborn as sns\nimport matplotlib.pyplot as plt\n\n# Select relevant columns for correlation analysis\ncorrelation_data = train_data[['Physical-Weight', 'Physical-BMI', 'PreInt_EduHx-computerinternet_hoursday']]\n\n# Compute the correlation matrix\ncorr_matrix = correlation_data.corr()\n\n# Plot the heatmap\nplt.figure(figsize=(8, 6))\nsns.heatmap(corr_matrix, annot=True, fmt=\".2f\", cmap=\"coolwarm\", cbar=True)\nplt.title(\"Correlation Matrix for Weight, BMI, and Phone Usage\")\nplt.show()\n","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-12-12T00:58:32.012859Z","iopub.execute_input":"2024-12-12T00:58:32.013488Z","iopub.status.idle":"2024-12-12T00:58:32.350836Z","shell.execute_reply.started":"2024-12-12T00:58:32.013449Z","shell.execute_reply":"2024-12-12T00:58:32.349777Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"# Key Insights:\n\n1. Strong Positive Correlation Between Weight and BMI:\n\nAs expected, Physical-Weight and Physical-BMI show a strong correlation (0.85). This relationship aligns with the fact that BMI is derived from weight and height.\n\n\n2. Weak Positive Correlation Between Weight and Phone Usage:\n\nPhysical-Weight shows a moderate correlation (0.34) with Phone Usage. This might indicate that individuals with higher weight tend to report slightly more phone usage.\n\n\n3. Minimal Correlation Between BMI and Phone Usage:\n\nThe correlation between Physical-BMI and Phone Usage (0.26) is relatively weak, suggesting BMI has limited linear association with phone usage patterns.\n\n\n4. Practical Implication:\n\nWhile weight shows some relationship with phone usage, it is not strong enough to directly influence feature selection for predictive modeling. However, this insight may guide further clustering or interaction term exploration.\n","metadata":{}},{"cell_type":"markdown","source":"In this analysis, I tested k=2 and k=3 for clustering. While k=2 provided simpler clusters, I chose k=3 as it revealed more detailed patterns and offered better segmentation of the data for a richer interpretation of behavioral similarities.","metadata":{}},{"cell_type":"markdown","source":"#  K-Means Clustering for Weight, BMI, and Phone Usage\n\nK-Means Clustering is an unsupervised machine learning algorithm used to group data points into k clusters based on their similarity. By minimizing the distance between data points and their assigned cluster centroids, it identifies patterns in the dataset.","metadata":{}},{"cell_type":"code","source":"from sklearn.cluster import KMeans\nimport matplotlib.pyplot as plt\nimport pandas as pd\n\ncluster_data = train_data[['Physical-Weight', 'Physical-BMI', 'PreInt_EduHx-computerinternet_hoursday']].dropna()\n\n# Initialize K-Means\nkmeans = KMeans(n_clusters=3, random_state=42, n_init='auto') \n\n# Fit the K-Means model and predict clusters\ncluster_data['Cluster'] = kmeans.fit_predict(cluster_data[['Physical-Weight', 'Physical-BMI']])\n\n\ncluster_data = cluster_data.copy()  \ncluster_data.loc[:, 'Cluster'] = kmeans.fit_predict(cluster_data[['Physical-Weight', 'Physical-BMI']])\n\n# Plot the clusters\nplt.figure(figsize=(10, 6))\nfor cluster in cluster_data['Cluster'].unique():\n    clustered = cluster_data[cluster_data['Cluster'] == cluster]\n    plt.scatter(\n        clustered['Physical-Weight'],\n        clustered['Physical-BMI'],\n        label=f\"Cluster {cluster}\"\n    )\n\nplt.xlabel('Physical Weight')\nplt.ylabel('Physical BMI')\nplt.title('K-Means Clustering for Weight, BMI, and Phone Usage')\nplt.legend()\nplt.show()\n","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-12-12T00:58:54.519857Z","iopub.execute_input":"2024-12-12T00:58:54.520295Z","iopub.status.idle":"2024-12-12T00:58:54.927562Z","shell.execute_reply.started":"2024-12-12T00:58:54.520259Z","shell.execute_reply":"2024-12-12T00:58:54.926455Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"# Key Insights\n\nIn this analysis, clustering was performed on attributes like weight, BMI, and phone usage hours. For k=3, the algorithm segmented the dataset into three distinct groups:\n\n* Cluster 0: Individuals with lower weight and BMI, indicating a healthier range.\n* Cluster 1: Individuals with moderate weight and BMI, representing an average group.\n* Cluster 2: Individuals with higher weight and BMI, possibly suggesting at-risk health profiles.\n\n\nThese clusters reveal behavioral and physical trends in the population, allowing for targeted interventions or deeper analysis of patterns across groups.\n","metadata":{}}]}