{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"name":"python","version":"3.10.14","mimetype":"text/x-python","codemirror_mode":{"name":"ipython","version":3},"pygments_lexer":"ipython3","nbconvert_exporter":"python","file_extension":".py"},"kaggle":{"accelerator":"none","dataSources":[{"sourceId":81933,"databundleVersionId":9643020,"sourceType":"competition"},{"sourceId":196454,"sourceType":"modelInstanceVersion","modelInstanceId":167527,"modelId":189843}],"dockerImageVersionId":30786,"isInternetEnabled":false,"language":"python","sourceType":"notebook","isGpuEnabled":false}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":" # **Predicting Problematic Internet Usage Based on Physical Activity**","metadata":{}},{"cell_type":"markdown","source":"# **Introduction**","metadata":{}},{"cell_type":"markdown","source":"\nIn today's digital age, the internet has become an essential part of our lives, offering countless benefits for learning, entertainment, and communication. However, excessive use, especially among children and adolescents, has given rise to a growing concern: problematic internet usage. This can manifest in negative behaviors such as compulsivity, dependency, and escapism, often linked to mental health challenges like anxiety, depression, and social withdrawal.\n\nCurrent methods for identifying problematic internet use rely on professional assessments, which can be complex, costly, and inaccessible for many families. These barriers highlight the urgent need for innovative, scalable solutions that can address this issue effectively.\n\nThis project seeks to bridge this gap by using physical fitness and activity data—simple, widely available indicators—as a proxy to predict early signs of problematic internet use. By leveraging machine learning and data-driven insights, we aim to build a predictive model that identifies at-risk individuals, enabling timely interventions to encourage healthier digital habits and promote overall well-being.\n\nThis work isn’t just about building a model; it’s about making a meaningful impact on the mental health and quality of life for future generations.","metadata":{}},{"cell_type":"markdown","source":"# **Problem Statement**","metadata":{}},{"cell_type":"markdown","source":"Problematic internet use among children and adolescents is a growing issue, with far-reaching implications for their mental and physical well-being. Excessive screen time and technology dependency have been linked to a range of negative outcomes, including poor posture, irregular diets, reduced physical activity, and mental health challenges such as anxiety and depression.\n\nDespite its seriousness, identifying problematic internet use early remains a challenge. Current methods rely heavily on professional assessments, which are often complex, costly, and inaccessible to many families due to cultural, linguistic, and logistical barriers. As a result, early intervention opportunities are often missed, leaving children and their families without the support they need.\n\nOn the other hand, physical activity and fitness data, such as movement patterns, sleep quality, and cardiovascular health, are widely available and easy to collect. These metrics could provide valuable insights into behavioral changes associated with excessive internet use, offering a practical, non-invasive alternative for early detection.\n\nThis project seeks to address the gap by using physical activity and fitness indicators to develop a predictive model for identifying problematic internet use. By leveraging accessible data, we aim to enable proactive interventions, promoting healthier digital habits and supporting the mental and physical health of children and adolescents.","metadata":{}},{"cell_type":"markdown","source":"# **Dataset Overview**","metadata":{}},{"cell_type":"markdown","source":"\n* Healthy Brain Network (HBN) Dataset\n    * Participants: 5,000 children and adolescents (5-22 years old).\n    * Data Types:\n          * Tabular: Demographics, physical measures, fitness tests, internet usage behavior.\n          * Time-series: Accelerometer data from wrist-worn devices.\n    * Target: Severity Impairment Index (sii) - None (0), Mild (1), Moderate (2), Severe (3).\n* Key Features:\n    * ENMO (activity measure), sleep disturbances, BMI, internet usage hours.\n","metadata":{}},{"cell_type":"markdown","source":"# Load the Data","metadata":{}},{"cell_type":"code","source":"# Importing necessary libraries\nimport pandas as pd\nimport pyarrow.parquet as pq\n# Importing necessary libraries\nimport numpy as np\nimport pandas as pd\nimport matplotlib.pyplot as plt\nimport seaborn as sns\nimport os\n\n# For machine learning models\nfrom sklearn.model_selection import train_test_split, StratifiedKFold\nfrom sklearn.impute import SimpleImputer\nfrom sklearn.preprocessing import StandardScaler, OneHotEncoder\nfrom sklearn.ensemble import RandomForestClassifier\nfrom sklearn.metrics import cohen_kappa_score, confusion_matrix\n\nimport tensorflow as tf\nfrom tensorflow.keras.models import Sequential\nfrom tensorflow.keras.layers import Dense, LSTM, Dropout, Flatten\n\n# For time-series processing\nimport pyarrow.parquet as pq\nfrom tsfresh import extract_features\n\n# For evaluation metric (Quadratic Weighted Kappa)\nfrom sklearn.metrics import make_scorer\n\n# Load CSV files\ntrain_data = pd.read_csv('/kaggle/input/child-mind-institute-problematic-internet-use/train.csv')\ntest_data = pd.read_csv('/kaggle/input/child-mind-institute-problematic-internet-use/test.csv')\ndata_dictionary = pd.read_csv('/kaggle/input/child-mind-institute-problematic-internet-use/data_dictionary.csv')","metadata":{"execution":{"iopub.status.busy":"2024-12-13T00:14:11.554096Z","iopub.execute_input":"2024-12-13T00:14:11.554886Z","iopub.status.idle":"2024-12-13T00:14:29.445776Z","shell.execute_reply.started":"2024-12-13T00:14:11.554849Z","shell.execute_reply":"2024-12-13T00:14:29.444524Z"},"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"code","source":"# Display the head of each dataset\ntrain_data.head()","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-12-13T00:14:29.448367Z","iopub.execute_input":"2024-12-13T00:14:29.448803Z","iopub.status.idle":"2024-12-13T00:14:29.490560Z","shell.execute_reply.started":"2024-12-13T00:14:29.448762Z","shell.execute_reply":"2024-12-13T00:14:29.489495Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"test_data.head()","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-12-13T00:14:29.491682Z","iopub.execute_input":"2024-12-13T00:14:29.491996Z","iopub.status.idle":"2024-12-13T00:14:29.515606Z","shell.execute_reply.started":"2024-12-13T00:14:29.491966Z","shell.execute_reply":"2024-12-13T00:14:29.514513Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"data_dictionary.head()","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-12-13T00:14:29.517064Z","iopub.execute_input":"2024-12-13T00:14:29.517518Z","iopub.status.idle":"2024-12-13T00:14:29.534233Z","shell.execute_reply.started":"2024-12-13T00:14:29.517466Z","shell.execute_reply":"2024-12-13T00:14:29.533180Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"# Load Parquet files for time-series data\n# Example for a single file, assuming the directory structure is consistent\nimport glob\n\n# Directory paths for time-series data\ntrain_parquet_files = glob.glob('/kaggle/input/child-mind-institute-problematic-internet-use/series_train.parquet/*')\ntest_parquet_files = glob.glob('/kaggle/input/child-mind-institute-problematic-internet-use/series_test.parquet/*')\n\n# Load a sample Parquet file to understand its structure\nsample_parquet = pq.read_table(train_parquet_files[0]).to_pandas()\n\nprint(\"\\nSample Parquet Data:\")\nprint(sample_parquet.head())\n\n# Save all Parquet file paths for further processing\ntrain_parquet_dict = {file.split('=')[-1]: file for file in train_parquet_files}\ntest_parquet_dict = {file.split('=')[-1]: file for file in test_parquet_files}\n\n# Check number of Parquet files loaded\nprint(f\"\\nNumber of train Parquet files: {len(train_parquet_dict)}\")\nprint(f\"Number of test Parquet files: {len(test_parquet_dict)}\")","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-12-13T00:14:29.536898Z","iopub.execute_input":"2024-12-13T00:14:29.537296Z","iopub.status.idle":"2024-12-13T00:14:29.652272Z","shell.execute_reply.started":"2024-12-13T00:14:29.537237Z","shell.execute_reply":"2024-12-13T00:14:29.651043Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"## Data Preprocessing","metadata":{}},{"cell_type":"code","source":"pseudo_train = train_data.copy()\npseudo_train = pseudo_train.dropna(subset=['sii'])\npseudo_train.head()","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-12-13T00:14:29.653397Z","iopub.execute_input":"2024-12-13T00:14:29.653684Z","iopub.status.idle":"2024-12-13T00:14:29.684570Z","shell.execute_reply.started":"2024-12-13T00:14:29.653656Z","shell.execute_reply":"2024-12-13T00:14:29.683407Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"pseudo_train.info()","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-12-13T00:14:29.685726Z","iopub.execute_input":"2024-12-13T00:14:29.686071Z","iopub.status.idle":"2024-12-13T00:14:29.713249Z","shell.execute_reply.started":"2024-12-13T00:14:29.686041Z","shell.execute_reply":"2024-12-13T00:14:29.712248Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"criterio = 0.4 * len(train_data)\ncolumnas = pseudo_train.columns[pseudo_train.isnull().sum() < criterio]\n\npseudo_train = pseudo_train[columnas]\npseudo_train = pseudo_train.fillna(0)\n\npseudo_train.info()","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-12-13T00:14:29.714606Z","iopub.execute_input":"2024-12-13T00:14:29.714946Z","iopub.status.idle":"2024-12-13T00:14:29.740080Z","shell.execute_reply.started":"2024-12-13T00:14:29.714914Z","shell.execute_reply":"2024-12-13T00:14:29.738842Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"tabla_cuali = ['Basic_Demos-Enroll_Season', 'CGAS-Season', \n               'Physical-Season', 'Fitness_Endurance-Season', \n               'FGC-Season','BIA-Season',\n               'PAQ_C-Season', 'PCIAT-Season',\n               'SDS-Season', 'PreInt_EduHx-Season']\ncolumnas_num = ['Basic_Demos-Age', 'Basic_Demos-Sex', 'CGAS-CGAS_Score', \n                'Physical-BMI', 'Physical-Height', 'Physical-Weight',\n               'Physical-Diastolic_BP', 'Physical-HeartRate', 'Physical-Systolic_BP', \n                'PCIAT-PCIAT_01',\n                'PCIAT-PCIAT_02', 'PCIAT-PCIAT_03', 'PCIAT-PCIAT_04', 'PCIAT-PCIAT_05',\n                'PCIAT-PCIAT_06', 'PCIAT-PCIAT_07', 'PCIAT-PCIAT_08', 'PCIAT-PCIAT_09',\n                'PCIAT-PCIAT_10', 'PCIAT-PCIAT_11', 'PCIAT-PCIAT_12', 'PCIAT-PCIAT_13',\n                'PCIAT-PCIAT_14', 'PCIAT-PCIAT_15', 'PCIAT-PCIAT_16', 'PCIAT-PCIAT_17',\n                'PCIAT-PCIAT_18', 'PCIAT-PCIAT_19', 'PCIAT-PCIAT_20', 'PCIAT-PCIAT_Total',\n                'SDS-SDS_Total_Raw', 'SDS-SDS_Total_T',\n       'PreInt_EduHx-computerinternet_hoursday', 'sii', 'id']\n","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-12-13T00:14:29.741555Z","iopub.execute_input":"2024-12-13T00:14:29.741924Z","iopub.status.idle":"2024-12-13T00:14:29.747848Z","shell.execute_reply.started":"2024-12-13T00:14:29.741892Z","shell.execute_reply":"2024-12-13T00:14:29.746674Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"# Tabular Data Preprocessing","metadata":{}},{"cell_type":"code","source":"import pandas as pd\nfrom sklearn.impute import SimpleImputer\n\n# Identify numeric and categorical columns\nnumeric_cols = train_data.select_dtypes(include=['float64', 'int64']).columns\ncategorical_cols = train_data.select_dtypes(include=['object']).columns\n\n# Check for missing columns between train and test data\nmissing_cols = set(train_data.columns) - set(test_data.columns)\n\n# Handle missing columns in test data (add them with default values if necessary)\nfor col in missing_cols:\n    if col in numeric_cols:\n        test_data[col] = 0  # Filling numeric columns with 0\n    elif col in categorical_cols:\n        test_data[col] = \"Unknown\"  # Filling categorical columns with \"Unknown\"\n\n# Initialize imputer for numeric columns (using mean for imputation)\nnumeric_imputer = SimpleImputer(strategy='mean')\n\n# Impute missing values in numeric columns for both train and test data\ntrain_data[numeric_cols] = numeric_imputer.fit_transform(train_data[numeric_cols])\ntest_data[numeric_cols] = numeric_imputer.transform(test_data[numeric_cols])\n\n# Initialize imputer for categorical columns (using most frequent value for imputation)\ncategorical_imputer = SimpleImputer(strategy='most_frequent')\n\n# Impute missing values in categorical columns for both train and test data\ntrain_data[categorical_cols] = categorical_imputer.fit_transform(train_data[categorical_cols])\ntest_data[categorical_cols] = categorical_imputer.transform(test_data[categorical_cols])\n\n# Verify if all missing values are handled\nprint(f\"Missing values in train data after preprocessing: {train_data.isnull().sum()}\")\nprint(f\"Missing values in test data after preprocessing: {test_data.isnull().sum()}\")\n","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-12-13T00:14:29.749039Z","iopub.execute_input":"2024-12-13T00:14:29.749401Z","iopub.status.idle":"2024-12-13T00:14:29.841640Z","shell.execute_reply.started":"2024-12-13T00:14:29.749351Z","shell.execute_reply":"2024-12-13T00:14:29.840512Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"# Time-Series Data Preprocessing","metadata":{}},{"cell_type":"code","source":"# Step 3b: Time-Series Data Preprocessing\n\n# Aggregating features from time-series data\ndef process_parquet(file_path):\n    # Load Parquet file\n    df = pq.read_table(file_path).to_pandas()\n\n    # Handle missing values in time-series data\n    df.fillna(0, inplace=True)\n\n    # Generate features: Mean, Std, Min, Max for each numeric column\n    feature_dict = {}\n    feature_dict['id'] = file_path.split('=')[-1]  # Extract ID from file path\n    for col in ['X', 'Y', 'Z', 'enmo', 'anglez', 'light', 'battery_voltage']:\n        feature_dict[f'{col}_mean'] = df[col].mean()\n        feature_dict[f'{col}_std'] = df[col].std()\n        feature_dict[f'{col}_min'] = df[col].min()\n        feature_dict[f'{col}_max'] = df[col].max()\n\n    return feature_dict\n\n# Process all Parquet files\ntrain_time_features = pd.DataFrame([process_parquet(file) for file in train_parquet_files])\ntest_time_features = pd.DataFrame([process_parquet(file) for file in test_parquet_files])\n\nprint(\"\\nTime-series features from train data:\")\nprint(train_time_features.head())\n","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-12-13T00:14:29.842885Z","iopub.execute_input":"2024-12-13T00:14:29.843177Z","iopub.status.idle":"2024-12-13T00:16:21.884887Z","shell.execute_reply.started":"2024-12-13T00:14:29.843148Z","shell.execute_reply":"2024-12-13T00:16:21.882649Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"# Combining Tabular and Time-Series Data","metadata":{}},{"cell_type":"code","source":"# Step 3c: Merge Tabular and Time-Series Features\n\n# Ensure 'id' columns are treated as strings\ntrain_time_features['id'] = train_time_features['id'].astype(str)\ntest_time_features['id'] = test_time_features['id'].astype(str)\n\n# Merge time-series features with tabular data\nmerged_train_data = pd.merge(train_data, train_time_features, on='id', how='left')\nmerged_test_data = pd.merge(test_data, test_time_features, on='id', how='left')\n\nprint(\"\\nFinal Train Data after merging:\")\nprint(train_data.head())","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-12-13T00:16:21.887503Z","iopub.execute_input":"2024-12-13T00:16:21.888830Z","iopub.status.idle":"2024-12-13T00:16:21.940802Z","shell.execute_reply.started":"2024-12-13T00:16:21.888767Z","shell.execute_reply":"2024-12-13T00:16:21.939630Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"# Exploratory Data Analysis (EDA)","metadata":{}},{"cell_type":"markdown","source":"1. Overview of the Dataset","metadata":{}},{"cell_type":"code","source":"# Display first few rows of the training and test data\nprint(\"Train Data Overview:\")\nprint(merged_train_data.head())\n\nprint(\"Test Data Overview:\")\nprint(merged_test_data.head())\n\n# Summary statistics for numerical columns\nprint(\"Train Data Summary Statistics:\")\nprint(merged_train_data.describe())\n\nprint(\"Test Data Summary Statistics:\")\nprint(merged_test_data.describe())\n","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-12-13T00:16:21.942317Z","iopub.execute_input":"2024-12-13T00:16:21.942654Z","iopub.status.idle":"2024-12-13T00:16:22.449550Z","shell.execute_reply.started":"2024-12-13T00:16:21.942623Z","shell.execute_reply":"2024-12-13T00:16:22.446527Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"2. Check for Missing Values","metadata":{}},{"cell_type":"code","source":"# Missing values in the train and test datasets\nprint(\"Missing Values in Train Data:\")\nprint(merged_train_data.isnull().sum())\n\nprint(\"Missing Values in Test Data:\")\nprint(merged_test_data.isnull().sum())\n\n# Percentage of missing values\nprint(\"Missing Values Percentage in Train Data:\")\nprint(merged_train_data.isnull().mean() * 100)\n\nprint(\"Missing Values Percentage in Test Data:\")\nprint(merged_test_data.isnull().mean() * 100)\n","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-12-13T00:16:22.459312Z","iopub.execute_input":"2024-12-13T00:16:22.460466Z","iopub.status.idle":"2024-12-13T00:16:22.497823Z","shell.execute_reply.started":"2024-12-13T00:16:22.460411Z","shell.execute_reply":"2024-12-13T00:16:22.495420Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"3. Data Distribution (Numerical Features)","metadata":{}},{"cell_type":"code","source":"import matplotlib.pyplot as plt\nimport seaborn as sns\n\n# Histograms for numerical features\nmerged_train_data.hist(figsize=(20, 16), bins=30)\nplt.suptitle('Histogram of Numerical Features in Train Data')\nplt.tight_layout()\nplt.show()\n","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-12-13T00:16:22.502039Z","iopub.execute_input":"2024-12-13T00:16:22.503636Z","iopub.status.idle":"2024-12-13T00:16:41.566791Z","shell.execute_reply.started":"2024-12-13T00:16:22.503475Z","shell.execute_reply":"2024-12-13T00:16:41.565658Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"4. Correlation Analysis","metadata":{}},{"cell_type":"code","source":"# Select only numerical columns for correlation matrix\nnumerical_columns = merged_train_data.select_dtypes(include=['float64', 'int64'])\n\n# Compute the correlation matrix\ncorr = numerical_columns.corr()\n\n# Heatmap of correlation matrix\nplt.figure(figsize=(32, 32))\nsns.heatmap(corr, annot=True, fmt=\".2f\", cmap='coolwarm', linewidths=0.5)\nplt.title('Correlation Heatmap of Numerical Features in Train Data')\nplt.tight_layout()\nplt.show()\n","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-12-13T00:16:41.568141Z","iopub.execute_input":"2024-12-13T00:16:41.568465Z","iopub.status.idle":"2024-12-13T00:16:53.507853Z","shell.execute_reply.started":"2024-12-13T00:16:41.568435Z","shell.execute_reply":"2024-12-13T00:16:53.506512Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"5. Data Distribution (Categorical Features)","metadata":{}},{"cell_type":"code","source":"# Plot categorical features (example for 'Basic_Demos-Sex')\nplt.figure(figsize=(8, 6))\nsns.countplot(x='Basic_Demos-Sex', data=merged_train_data)\nplt.title('Count of Categories in Basic_Demos-Sex')\nplt.show()\n\n# Example for 'CGAS-Season'\nplt.figure(figsize=(8, 6))\nsns.countplot(x='CGAS-Season', data=merged_train_data)\nplt.title('Count of Categories in CGAS-Season')\nplt.show()\n","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-12-13T00:16:53.509365Z","iopub.execute_input":"2024-12-13T00:16:53.509727Z","iopub.status.idle":"2024-12-13T00:16:53.988781Z","shell.execute_reply.started":"2024-12-13T00:16:53.509693Z","shell.execute_reply":"2024-12-13T00:16:53.987546Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"6. Outliers Detection\n\n   \nIdentify potential outliers using boxplots.","metadata":{}},{"cell_type":"code","source":"# Boxplot to detect outliers in 'Basic_Demos-Age'\nplt.figure(figsize=(8, 6))\nsns.boxplot(x=merged_train_data['Basic_Demos-Age'])\nplt.title('Boxplot of Basic_Demos-Age to detect outliers')\nplt.show()\n","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-12-13T00:16:53.990198Z","iopub.execute_input":"2024-12-13T00:16:53.990581Z","iopub.status.idle":"2024-12-13T00:16:54.208231Z","shell.execute_reply.started":"2024-12-13T00:16:53.990549Z","shell.execute_reply":"2024-12-13T00:16:54.207087Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"7. Target Variable Analysis (sii)\n   \nAnalyze the distribution of the target variable (sii).","metadata":{}},{"cell_type":"code","source":"# Count plot for target variable 'sii'\nplt.figure(figsize=(8, 6))\nsns.countplot(x='sii', data=merged_train_data)\nplt.title('Distribution of Target Variable (sii)')\nplt.show()\n\n# Analyze the target variable distribution (e.g., 'sii' as categorical)\nprint(merged_train_data['sii'].value_counts())\n","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-12-13T00:16:54.209657Z","iopub.execute_input":"2024-12-13T00:16:54.209998Z","iopub.status.idle":"2024-12-13T00:16:54.502205Z","shell.execute_reply.started":"2024-12-13T00:16:54.209964Z","shell.execute_reply":"2024-12-13T00:16:54.500987Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"8. Check for Duplicates\n\nCheck if there are any duplicate entries in the data.","metadata":{}},{"cell_type":"code","source":"# Check for duplicate rows\nprint(\"Number of duplicate rows in train data:\", merged_train_data.duplicated().sum())\nprint(\"Number of duplicate rows in test data:\", merged_test_data.duplicated().sum())\n","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-12-13T00:16:54.503866Z","iopub.execute_input":"2024-12-13T00:16:54.504345Z","iopub.status.idle":"2024-12-13T00:16:54.567438Z","shell.execute_reply.started":"2024-12-13T00:16:54.504296Z","shell.execute_reply":"2024-12-13T00:16:54.566246Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"9. Feature Engineering (if applicable)\nCreate new features or transform existing ones (e.g., extract day or month from timestamps).","metadata":{}},{"cell_type":"code","source":"# Z-score method for outlier detection\nfrom scipy.stats import zscore\n\nz_scores = zscore(merged_train_data[['Physical-BMI', 'Physical-Height', 'Physical-Weight']])\nmerged_train_data = merged_train_data[(z_scores < 3).all(axis=1)]\n","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-12-13T00:16:54.569048Z","iopub.execute_input":"2024-12-13T00:16:54.569521Z","iopub.status.idle":"2024-12-13T00:16:54.588312Z","shell.execute_reply.started":"2024-12-13T00:16:54.569474Z","shell.execute_reply":"2024-12-13T00:16:54.586840Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"10. Visualizing Relationships\n    \nYou can visualize relationships between key variables using scatter plots, box plots, etc.","metadata":{}},{"cell_type":"code","source":"# Scatter plot between 'Physical-BMI' and 'Physical-Height'\nsns.scatterplot(x='Physical-BMI', y='Physical-Height', data=merged_train_data)\nplt.title('Scatter plot between Physical BMI and Height')\nplt.show()\n\n# Boxplot for 'Physical-BMI'\nsns.boxplot(x='Physical-BMI', data=merged_train_data)\nplt.title('Boxplot for Physical BMI')\nplt.show()\n","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-12-13T00:16:54.589924Z","iopub.execute_input":"2024-12-13T00:16:54.591021Z","iopub.status.idle":"2024-12-13T00:16:55.132024Z","shell.execute_reply.started":"2024-12-13T00:16:54.590957Z","shell.execute_reply":"2024-12-13T00:16:55.130624Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"10. Pairplot for Multivariate Analysis","metadata":{}},{"cell_type":"code","source":"# Pairplot for a subset of features\nsns.pairplot(merged_train_data[['Basic_Demos-Age', 'CGAS-Season', 'sii']])\nplt.title('Pairplot for Age, CGAS-Season, and sii')\nplt.tight_layout()\nplt.show()\n","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-12-13T00:16:55.133762Z","iopub.execute_input":"2024-12-13T00:16:55.134330Z","iopub.status.idle":"2024-12-13T00:16:56.504334Z","shell.execute_reply.started":"2024-12-13T00:16:55.134246Z","shell.execute_reply":"2024-12-13T00:16:56.502822Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"# Feature Engineering","metadata":{}},{"cell_type":"markdown","source":"Encoding Categorical Features: Use techniques like one-hot encoding or label encoding for categorical features in merged_train_data and merged_test_data.","metadata":{}},{"cell_type":"code","source":"ohe = OneHotEncoder(handle_unknown='ignore', sparse=False)\ncat_features = ['Basic_Demos-Sex', 'CGAS-Season']  # Add relevant categorical columns\nencoded_features = ohe.fit_transform(merged_train_data[cat_features])\n","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-12-13T00:16:56.506348Z","iopub.execute_input":"2024-12-13T00:16:56.506895Z","iopub.status.idle":"2024-12-13T00:16:56.524549Z","shell.execute_reply.started":"2024-12-13T00:16:56.506840Z","shell.execute_reply":"2024-12-13T00:16:56.523131Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"Scaling Numerical Features: Apply standard scaling (or MinMax scaling) to numerical features for consistent input magnitudes, especially if you plan to use neural networks or distance-based models.","metadata":{}},{"cell_type":"code","source":"scaler = StandardScaler()\nnumerical_cols = merged_train_data.select_dtypes(include=['float64', 'int64']).columns\nmerged_train_data[numerical_cols] = scaler.fit_transform(merged_train_data[numerical_cols])","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-12-13T00:16:56.525889Z","iopub.execute_input":"2024-12-13T00:16:56.526227Z","iopub.status.idle":"2024-12-13T00:16:56.550439Z","shell.execute_reply.started":"2024-12-13T00:16:56.526196Z","shell.execute_reply":"2024-12-13T00:16:56.549346Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"# Categorical Features  ","metadata":{}},{"cell_type":"code","source":"# One-hot encode categorical variables\ncategorical_cols = train_data.select_dtypes(include=['object']).columns\nencoder = OneHotEncoder(sparse=False, handle_unknown='ignore')\n\nencoded_train = pd.DataFrame(encoder.fit_transform(train_data[categorical_cols]))\nencoded_test = pd.DataFrame(encoder.transform(test_data[categorical_cols]))\n\n# Add encoded columns to train/test data\ntrain_data = pd.concat([train_data.drop(columns=categorical_cols), encoded_train], axis=1)\ntest_data = pd.concat([test_data.drop(columns=categorical_cols), encoded_test], axis=1)\n\nprint(\"After Encoding:\")\nprint(train_data.head())\n","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-12-13T00:16:56.551838Z","iopub.execute_input":"2024-12-13T00:16:56.552171Z","iopub.status.idle":"2024-12-13T00:16:56.965313Z","shell.execute_reply.started":"2024-12-13T00:16:56.552140Z","shell.execute_reply":"2024-12-13T00:16:56.964150Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"# Time-Series Features","metadata":{}},{"cell_type":"code","source":"# Example: Rolling statistics for time-series features\ndef generate_time_series_features(data):\n    for col in ['X', 'Y', 'Z', 'enmo', 'anglez']:\n        data[f'{col}_mean'] = data[col].mean()\n        data[f'{col}_std'] = data[col].std()\n        data[f'{col}_rolling_mean'] = data[col].rolling(window=5).mean()\n    return data","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-12-13T00:16:56.966528Z","iopub.execute_input":"2024-12-13T00:16:56.966868Z","iopub.status.idle":"2024-12-13T00:16:56.972729Z","shell.execute_reply.started":"2024-12-13T00:16:56.966835Z","shell.execute_reply":"2024-12-13T00:16:56.971476Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"# Model Preparation","metadata":{}},{"cell_type":"code","source":"# Drop unnecessary columns (adjust as needed)\nX = train_data.drop(columns=['sii'])\n\n# Target variable\ny = train_data['sii']\n\n# Perform stratified train-validation split\nfrom sklearn.model_selection import train_test_split\nX_train, X_val, y_train, y_val = train_test_split(\n    X, y, test_size=0.2, stratify=y, random_state=42)\n","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-12-13T00:16:56.974252Z","iopub.execute_input":"2024-12-13T00:16:56.974622Z","iopub.status.idle":"2024-12-13T00:16:57.160858Z","shell.execute_reply.started":"2024-12-13T00:16:56.974590Z","shell.execute_reply":"2024-12-13T00:16:57.159519Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"# Identify categorical columns\ncategorical_cols = X_train.select_dtypes(include=['object']).columns\n\n# One-hot encode categorical variables\nfrom sklearn.preprocessing import OneHotEncoder\n\nencoder = OneHotEncoder(sparse=False, handle_unknown='ignore')\n\n# Fit encoder on training data\nencoder.fit(X_train[categorical_cols])\n\n# Transform training and validation data\nX_train_encoded = encoder.transform(X_train[categorical_cols])\nX_val_encoded = encoder.transform(X_val[categorical_cols])\n\n# Convert encoded features to DataFrame\nencoded_cols = encoder.get_feature_names_out(categorical_cols)\nX_train_encoded_df = pd.DataFrame(X_train_encoded, columns=encoded_cols, index=X_train.index)\nX_val_encoded_df = pd.DataFrame(X_val_encoded, columns=encoded_cols, index=X_val.index)\n\n# Drop original categorical columns and concatenate encoded columns\nX_train = X_train.drop(columns=categorical_cols).join(X_train_encoded_df)\nX_val = X_val.drop(columns=categorical_cols).join(X_val_encoded_df)\n","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-12-13T00:16:57.162246Z","iopub.execute_input":"2024-12-13T00:16:57.162597Z","iopub.status.idle":"2024-12-13T00:16:57.421491Z","shell.execute_reply.started":"2024-12-13T00:16:57.162560Z","shell.execute_reply":"2024-12-13T00:16:57.420234Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"# Separate features and target\nX = train_data.drop(columns=['sii', 'id'], errors='ignore')\ny = train_data['sii']\ntest_features = test_data.drop(columns=['sii', 'id'], errors='ignore')","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-12-13T00:16:57.422803Z","iopub.execute_input":"2024-12-13T00:16:57.423130Z","iopub.status.idle":"2024-12-13T00:16:57.470825Z","shell.execute_reply.started":"2024-12-13T00:16:57.423100Z","shell.execute_reply":"2024-12-13T00:16:57.469759Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"from sklearn.model_selection import train_test_split\nfrom sklearn.preprocessing import StandardScaler\n\n# Ensure column names are strings\nX.columns = X.columns.astype(str)\ntest_features.columns = test_features.columns.astype(str)\n\n# Split the data into training and validation sets\nX_train, X_val, y_train, y_val = train_test_split(X, y, test_size=0.2, stratify=y, random_state=42)\n\n# Ensure column names are consistent across training and validation sets\nX_train.columns = X_train.columns.astype(str)\nX_val.columns = X_val.columns.astype(str)\n\n# Align test_features to match X_train's columns\ntest_features = test_features[X_train.columns]\n\n# Standardize features\nscaler = StandardScaler()\nX_train_scaled = scaler.fit_transform(X_train)\nX_val_scaled = scaler.transform(X_val)\ntest_data_scaled = scaler.transform(test_features)\n","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-12-13T00:16:57.472273Z","iopub.execute_input":"2024-12-13T00:16:57.472651Z","iopub.status.idle":"2024-12-13T00:16:58.059589Z","shell.execute_reply.started":"2024-12-13T00:16:57.472616Z","shell.execute_reply":"2024-12-13T00:16:58.058586Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"print(X.shape)\nprint(pseudo_train.shape)\n","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-12-13T00:16:58.060806Z","iopub.execute_input":"2024-12-13T00:16:58.061123Z","iopub.status.idle":"2024-12-13T00:16:58.066832Z","shell.execute_reply.started":"2024-12-13T00:16:58.061093Z","shell.execute_reply":"2024-12-13T00:16:58.065406Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"columnas_tt = pseudo_train.columns.intersection(test_data.columns)\n\n# Check if 'id' exists before dropping\nif 'id' in columnas_tt:\n    columnas_tt = columnas_tt.drop('id')  # Exclude 'id' from selected columns\n\nX = pseudo_train.loc[:, columnas_tt]\ny = pseudo_train.loc[X.index, 'sii']\n\n# Perform the train-test split\nX_train, X_test, y_train, y_test = train_test_split(X, y, test_size=0.3, random_state=1111)\n","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-12-13T00:16:58.068235Z","iopub.execute_input":"2024-12-13T00:16:58.068612Z","iopub.status.idle":"2024-12-13T00:16:58.085918Z","shell.execute_reply.started":"2024-12-13T00:16:58.068579Z","shell.execute_reply":"2024-12-13T00:16:58.084759Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"# Random Forest Baseline","metadata":{}},{"cell_type":"code","source":"model = RandomForestClassifier(random_state=1111)","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-12-13T00:16:58.087511Z","iopub.execute_input":"2024-12-13T00:16:58.087944Z","iopub.status.idle":"2024-12-13T00:16:58.093496Z","shell.execute_reply.started":"2024-12-13T00:16:58.087911Z","shell.execute_reply":"2024-12-13T00:16:58.092319Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"from sklearn.decomposition import PCA\npca = PCA(n_components=0.8)\nX_train_pca = pca.fit_transform(X_train)\nX_test_pca = pca.transform(X_test)","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-12-13T00:16:58.094834Z","iopub.execute_input":"2024-12-13T00:16:58.095304Z","iopub.status.idle":"2024-12-13T00:16:58.152846Z","shell.execute_reply.started":"2024-12-13T00:16:58.095236Z","shell.execute_reply":"2024-12-13T00:16:58.150935Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"# Assuming you have your training data X_train and y_train\nfrom sklearn.ensemble import RandomForestClassifier\n\n# Initialize the Random Forest Classifier\nmodel = RandomForestClassifier()\n\n# Fit the model with the training data\nmodel.fit(X_train, y_train)\n\n# Now you can generate predictions on the test set\npredictions = model.predict(X_test)","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-12-13T00:16:58.153985Z","iopub.execute_input":"2024-12-13T00:16:58.154375Z","iopub.status.idle":"2024-12-13T00:16:58.653523Z","shell.execute_reply.started":"2024-12-13T00:16:58.154335Z","shell.execute_reply":"2024-12-13T00:16:58.652240Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"from sklearn.preprocessing import StandardScaler\n\nscaler = StandardScaler()\nX_train_norm = scaler.fit_transform(X_train)\nX_test_norm = scaler.transform(X_test)\n\n# Aplicaremos PCA para reducir la dimensionalidad\npca = PCA(n_components=0.95)\nX_train_pca = pca.fit_transform(X_train_norm)\nX_test_pca = pca.transform(X_test_norm)","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-12-13T00:16:58.655059Z","iopub.execute_input":"2024-12-13T00:16:58.655536Z","iopub.status.idle":"2024-12-13T00:16:58.692564Z","shell.execute_reply.started":"2024-12-13T00:16:58.655489Z","shell.execute_reply":"2024-12-13T00:16:58.689250Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"print(y_train[:10])  # Display a sample of y_train\nprint(y_train.dtype)  # Check the data type","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-12-13T00:16:58.694309Z","iopub.execute_input":"2024-12-13T00:16:58.698021Z","iopub.status.idle":"2024-12-13T00:16:58.714607Z","shell.execute_reply.started":"2024-12-13T00:16:58.697966Z","shell.execute_reply":"2024-12-13T00:16:58.711769Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"import numpy as np\n\n# Example: Categorize continuous values into bins\nbins = [0, 50, 100, 150]  # Define bins\nlabels = ['low', 'medium', 'high']  # Define class labels\ny_train_binned = pd.cut(y_train, bins=bins, labels=labels)\n\nprint(y_train_binned[:10])  # Verify the transformation\n","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-12-13T00:16:58.718903Z","iopub.execute_input":"2024-12-13T00:16:58.719371Z","iopub.status.idle":"2024-12-13T00:16:58.771042Z","shell.execute_reply.started":"2024-12-13T00:16:58.719307Z","shell.execute_reply":"2024-12-13T00:16:58.764105Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"from sklearn.ensemble import RandomForestRegressor\n\nrf = RandomForestRegressor(n_estimators=100, random_state=1111,\n                           max_depth=8, min_samples_split=10,\n                           min_samples_leaf=4)\nrf.fit(X_train_pca, y_train)","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-12-13T00:16:58.776220Z","iopub.execute_input":"2024-12-13T00:16:58.776845Z","iopub.status.idle":"2024-12-13T00:17:01.119796Z","shell.execute_reply.started":"2024-12-13T00:16:58.776780Z","shell.execute_reply":"2024-12-13T00:17:01.118476Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"from sklearn.metrics import r2_score, mean_squared_error, confusion_matrix\nimport numpy as np\nimport pandas as pd\nfrom sklearn.preprocessing import KBinsDiscretizer\nfrom sklearn.metrics import cohen_kappa_score","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-12-13T00:17:01.128516Z","iopub.execute_input":"2024-12-13T00:17:01.128943Z","iopub.status.idle":"2024-12-13T00:17:01.134486Z","shell.execute_reply.started":"2024-12-13T00:17:01.128906Z","shell.execute_reply":"2024-12-13T00:17:01.133351Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"# Predictions\ny_pred = rf.predict(X_test_pca)\n\n# R² Score\nr2_rf = r2_score(y_test, y_pred)\nprint(f\"R² Score: {r2_rf:.4f}\")\n\n# Mean Squared Error\nmse_rf = mean_squared_error(y_test, y_pred)\nprint(f\"Mean Squared Error: {mse_rf:.4f}\")","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-12-13T00:17:01.136024Z","iopub.execute_input":"2024-12-13T00:17:01.136480Z","iopub.status.idle":"2024-12-13T00:17:01.166276Z","shell.execute_reply.started":"2024-12-13T00:17:01.136434Z","shell.execute_reply":"2024-12-13T00:17:01.164968Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"# Define bins for discretization\nbins = [0, 50, 100, 150]  # Customize bins based on your data\nlabels = ['low', 'medium', 'high']  # Customize based on bins\n\n# Initialize KBinsDiscretizer\ndiscretizer = KBinsDiscretizer(n_bins=len(bins) - 1, encode='ordinal', strategy='uniform')\n\n# Ensure y_test and y_pred are NumPy arrays and discretize\ny_test_binned = discretizer.fit_transform(y_test.values.reshape(-1, 1)).flatten()\ny_pred_binned = discretizer.transform(y_pred.reshape(-1, 1)).flatten()\n","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-12-13T00:17:01.168072Z","iopub.execute_input":"2024-12-13T00:17:01.169031Z","iopub.status.idle":"2024-12-13T00:17:01.177272Z","shell.execute_reply.started":"2024-12-13T00:17:01.168978Z","shell.execute_reply":"2024-12-13T00:17:01.176067Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"# Confusion Matrix\nconf_matrix = confusion_matrix(y_test_binned, y_pred_binned)\nprint(\"Confusion Matrix:\")\nprint(conf_matrix)\n\n# Quadratic Weighted Kappa\nqwk_rf = cohen_kappa_score(y_test_binned, y_pred_binned, weights='quadratic')\nprint(f\"Quadratic Weighted Kappa (QWK) Score: {qwk_rf:.4f}\")\n","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-12-13T00:17:01.179388Z","iopub.execute_input":"2024-12-13T00:17:01.179841Z","iopub.status.idle":"2024-12-13T00:17:01.197354Z","shell.execute_reply.started":"2024-12-13T00:17:01.179794Z","shell.execute_reply":"2024-12-13T00:17:01.196057Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"import seaborn as sns\nimport matplotlib.pyplot as plt\n\n# Plotting the confusion matrix\nplt.figure(figsize=(8, 6))\nsns.heatmap(conf_matrix, annot=True, fmt='d', cmap='Blues', xticklabels=labels, yticklabels=labels)\nplt.title(\"Confusion Matrix\")\nplt.xlabel(\"Predicted Labels\")\nplt.ylabel(\"True Labels\")\nplt.show()\n","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-12-13T00:17:01.199100Z","iopub.execute_input":"2024-12-13T00:17:01.199462Z","iopub.status.idle":"2024-12-13T00:17:01.503735Z","shell.execute_reply.started":"2024-12-13T00:17:01.199428Z","shell.execute_reply":"2024-12-13T00:17:01.502619Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"# Plotting Predicted vs Actual\nplt.figure(figsize=(8, 6))\nplt.scatter(y_test, y_pred, color='blue', alpha=0.5)\nplt.plot([y_test.min(), y_test.max()], [y_test.min(), y_test.max()], color='red', linestyle='--')  # Ideal line\nplt.title(\"Predicted vs Actual Values\")\nplt.xlabel(\"Actual Values\")\nplt.ylabel(\"Predicted Values\")\nplt.show()\n","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-12-13T00:17:01.505513Z","iopub.execute_input":"2024-12-13T00:17:01.505978Z","iopub.status.idle":"2024-12-13T00:17:01.725032Z","shell.execute_reply.started":"2024-12-13T00:17:01.505930Z","shell.execute_reply":"2024-12-13T00:17:01.723543Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"# Cumulative Gains Chart (Lift Chart):","metadata":{}},{"cell_type":"code","source":"import numpy as np\nimport matplotlib.pyplot as plt\nfrom sklearn.metrics import roc_curve, auc\n\n# Cumulative Gains (Lift) Chart\n# Calculate True Positive Rate (Recall) for each bin\nsorted_true = np.sort(y_test)\nsorted_pred = np.sort(y_pred)\n\n# Calculate cumulative true positive rate for the sorted predictions\ncumulative_true = np.cumsum(sorted_true) / np.sum(sorted_true)\ncumulative_random = np.arange(1, len(sorted_true) + 1) / len(sorted_true)\n\nplt.figure(figsize=(8, 6))\nplt.plot(cumulative_random, label=\"Random Model\", linestyle='--', color='gray')\nplt.plot(cumulative_true, label=\"Model\", color='blue')\nplt.title(\"Cumulative Gains (Lift) Chart\")\nplt.xlabel(\"Percentile\")\nplt.ylabel(\"Cumulative Gain\")\nplt.legend()\nplt.show()\n","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-12-13T00:17:01.726682Z","iopub.execute_input":"2024-12-13T00:17:01.727225Z","iopub.status.idle":"2024-12-13T00:17:01.962197Z","shell.execute_reply.started":"2024-12-13T00:17:01.727174Z","shell.execute_reply":"2024-12-13T00:17:01.960980Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"# Residuals Histogram:","metadata":{}},{"cell_type":"code","source":"# Residuals calculation\nresiduals = y_test - y_pred\n\n# Plotting residuals as a histogram\nplt.figure(figsize=(8, 6))\nplt.hist(residuals, bins=30, edgecolor='black', color='purple', alpha=0.7)\nplt.title(\"Residuals Histogram\")\nplt.xlabel(\"Residuals (Actual - Predicted)\")\nplt.ylabel(\"Frequency\")\nplt.show()\n","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-12-13T00:17:01.963400Z","iopub.execute_input":"2024-12-13T00:17:01.963699Z","iopub.status.idle":"2024-12-13T00:17:02.212619Z","shell.execute_reply.started":"2024-12-13T00:17:01.963672Z","shell.execute_reply":"2024-12-13T00:17:02.211392Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"# Precision-Recall Curve (for each class):","metadata":{}},{"cell_type":"code","source":"from sklearn.preprocessing import LabelBinarizer\nfrom sklearn.metrics import precision_recall_curve, auc\nimport matplotlib.pyplot as plt\n\n# Binarize the labels for multiclass problem\nlb = LabelBinarizer()\ny_test_binarized = lb.fit_transform(y_test_binned)\ny_pred_binarized = lb.transform(y_pred_binned)\n\n# Plot precision-recall curve for each class and print AUC\nplt.figure(figsize=(8, 6))\nfor i in range(y_test_binarized.shape[1]):\n    precision, recall, _ = precision_recall_curve(y_test_binarized[:, i], y_pred_binarized[:, i])\n    auc_score = auc(recall, precision)  # Calculate AUC for each class\n    plt.plot(recall, precision, label=f'Class {i} (AUC = {auc_score:.2f})')\n\nplt.title(\"Precision-Recall Curve for Each Class\")\nplt.xlabel(\"Recall\")\nplt.ylabel(\"Precision\")\nplt.legend()\nplt.show()\n","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-12-13T00:17:02.214073Z","iopub.execute_input":"2024-12-13T00:17:02.214417Z","iopub.status.idle":"2024-12-13T00:17:02.511531Z","shell.execute_reply.started":"2024-12-13T00:17:02.214385Z","shell.execute_reply":"2024-12-13T00:17:02.510357Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"# XGBoost","metadata":{}},{"cell_type":"code","source":"import xgboost as xgb\nfrom sklearn.metrics import r2_score, mean_squared_error, confusion_matrix\nfrom sklearn.preprocessing import KBinsDiscretizer\nimport numpy as np\nimport pandas as pd\nfrom sklearn.metrics import cohen_kappa_score","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-12-13T00:17:02.512850Z","iopub.execute_input":"2024-12-13T00:17:02.513182Z","iopub.status.idle":"2024-12-13T00:17:02.714884Z","shell.execute_reply.started":"2024-12-13T00:17:02.513151Z","shell.execute_reply":"2024-12-13T00:17:02.713913Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"# Initialize the XGBoost Regressor\nxgb_model = xgb.XGBRegressor(\n    n_estimators=100,       # Number of trees\n    max_depth=8,            # Maximum tree depth\n    learning_rate=0.1,      # Step size shrinkage\n    min_child_weight=10,    # Minimum sum of instance weight\n    subsample=0.8,          # Subsample ratio of the training instance\n    colsample_bytree=0.8,   # Subsample ratio of columns when constructing each tree\n    random_state=1111       # For reproducibility\n)\n\n# Fit the model to training data\nxgb_model.fit(X_train_pca, y_train)\n","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-12-13T00:17:02.716251Z","iopub.execute_input":"2024-12-13T00:17:02.716629Z","iopub.status.idle":"2024-12-13T00:17:03.081844Z","shell.execute_reply.started":"2024-12-13T00:17:02.716596Z","shell.execute_reply":"2024-12-13T00:17:03.081016Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"# Make predictions on the test data\ny_pred = xgb_model.predict(X_test_pca)\n\n# R² Score\nr2_xgb = r2_score(y_test, y_pred)\nprint(f\"R² Score: {r2_xgb:.4f}\")\n\n# Mean Squared Error\nmse_xgb = mean_squared_error(y_test, y_pred)\nprint(f\"Mean Squared Error: {mse_xgb:.4f}\")","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-12-13T00:17:03.082774Z","iopub.execute_input":"2024-12-13T00:17:03.083092Z","iopub.status.idle":"2024-12-13T00:17:03.100873Z","shell.execute_reply.started":"2024-12-13T00:17:03.083061Z","shell.execute_reply":"2024-12-13T00:17:03.098418Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"# Define bins for discretization\nbins = [0, 50, 100, 150]  # Customize bins based on your data\nlabels = ['low', 'medium', 'high']  # Customize based on bins\n\n# Discretize actual and predicted values\ndiscretizer = KBinsDiscretizer(n_bins=len(bins) - 1, encode='ordinal', strategy='uniform')\n\ny_test_binned = discretizer.fit_transform(y_test.values.reshape(-1, 1)).flatten()\ny_pred_binned = discretizer.transform(y_pred.reshape(-1, 1)).flatten()\n","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-12-13T00:17:03.104822Z","iopub.execute_input":"2024-12-13T00:17:03.105167Z","iopub.status.idle":"2024-12-13T00:17:03.113377Z","shell.execute_reply.started":"2024-12-13T00:17:03.105133Z","shell.execute_reply":"2024-12-13T00:17:03.112033Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"# Confusion Matrix\nconf_matrix = confusion_matrix(y_test_binned, y_pred_binned)\nprint(\"Confusion Matrix:\")\nprint(conf_matrix)\n\n# Quadratic Weighted Kappa (QWK) Score\nqwk_xgb = cohen_kappa_score(y_test_binned, y_pred_binned, weights='quadratic')\nprint(f\"Quadratic Weighted Kappa (QWK) Score: {qwk_xgb:.4f}\")\n","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-12-13T00:17:03.115212Z","iopub.execute_input":"2024-12-13T00:17:03.116435Z","iopub.status.idle":"2024-12-13T00:17:03.137108Z","shell.execute_reply.started":"2024-12-13T00:17:03.116398Z","shell.execute_reply":"2024-12-13T00:17:03.135790Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"import seaborn as sns\nimport matplotlib.pyplot as plt\n\n# Plot Confusion Matrix\nplt.figure(figsize=(8, 6))\nsns.heatmap(conf_matrix, annot=True, fmt='d', cmap='Blues', xticklabels=labels, yticklabels=labels)\nplt.title(\"Confusion Matrix\")\nplt.xlabel(\"Predicted\")\nplt.ylabel(\"Actual\")\nplt.show()","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-12-13T00:17:03.138628Z","iopub.execute_input":"2024-12-13T00:17:03.139090Z","iopub.status.idle":"2024-12-13T00:17:03.435755Z","shell.execute_reply.started":"2024-12-13T00:17:03.139042Z","shell.execute_reply":"2024-12-13T00:17:03.434566Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"xgb_model.fit(X_train_pca, y_train)\nprint(\"Model fitting complete.\")","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-12-13T00:17:03.437021Z","iopub.execute_input":"2024-12-13T00:17:03.437384Z","iopub.status.idle":"2024-12-13T00:17:03.760787Z","shell.execute_reply.started":"2024-12-13T00:17:03.437340Z","shell.execute_reply":"2024-12-13T00:17:03.758980Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"xgb.plot_importance(xgb_model, importance_type='weight', max_num_features=10)\nplt.title(\"Top 10 Feature Importances\")\nplt.show()","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-12-13T00:17:03.761832Z","iopub.execute_input":"2024-12-13T00:17:03.762147Z","iopub.status.idle":"2024-12-13T00:17:04.044874Z","shell.execute_reply.started":"2024-12-13T00:17:03.762117Z","shell.execute_reply":"2024-12-13T00:17:04.043730Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"# Residuals Plot\nresiduals = y_test - y_pred\nplt.figure(figsize=(8, 6))\nplt.scatter(y_pred, residuals, alpha=0.5)\nplt.axhline(0, color='red', linestyle='--')\nplt.title(\"Residuals vs Predicted\")\nplt.xlabel(\"Predicted Values\")\nplt.ylabel(\"Residuals\")\nplt.show()\n\n# Predicted vs Actual Plot\nplt.figure(figsize=(8, 6))\nplt.scatter(y_test, y_pred, alpha=0.5, color='green')\nplt.plot([y_test.min(), y_test.max()], [y_test.min(), y_test.max()], color='red', linestyle='--', linewidth=2)\nplt.title(\"Predicted vs Actual\")\nplt.xlabel(\"Actual Values\")\nplt.ylabel(\"Predicted Values\")\nplt.show()","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-12-13T00:17:04.046156Z","iopub.execute_input":"2024-12-13T00:17:04.046522Z","iopub.status.idle":"2024-12-13T00:17:04.635054Z","shell.execute_reply.started":"2024-12-13T00:17:04.046490Z","shell.execute_reply":"2024-12-13T00:17:04.633938Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"# Cumulative Gains Chart (Lift Chart)","metadata":{}},{"cell_type":"code","source":"from sklearn.metrics import roc_curve\n\n# Calculate predicted probabilities\ny_prob = xgb_model.predict(X_test_pca)\n\n# Sort actual values by predicted probabilities\nsorted_indices = np.argsort(y_prob)[::-1]\ny_sorted = y_test.iloc[sorted_indices]\n\n# Calculate cumulative gains\ncum_gains = np.cumsum(y_sorted) / np.sum(y_sorted)\n\n# Plot\nplt.figure(figsize=(8, 6))\nplt.plot(np.arange(len(cum_gains)), cum_gains, label='Cumulative Gains')\nplt.title(\"Cumulative Gains Chart\")\nplt.xlabel(\"Number of Samples\")\nplt.ylabel(\"Cumulative Gains\")\nplt.legend()\nplt.show()\n","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-12-13T00:17:04.636749Z","iopub.execute_input":"2024-12-13T00:17:04.637204Z","iopub.status.idle":"2024-12-13T00:17:04.912030Z","shell.execute_reply.started":"2024-12-13T00:17:04.637158Z","shell.execute_reply":"2024-12-13T00:17:04.910960Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"# Residuals Histogram","metadata":{}},{"cell_type":"code","source":"residuals = y_test - y_pred\n\nplt.figure(figsize=(8, 6))\nsns.histplot(residuals, kde=True, color='purple')\nplt.title(\"Residuals Distribution\")\nplt.xlabel(\"Residuals\")\nplt.ylabel(\"Frequency\")\nplt.axvline(0, color='red', linestyle='--', label='Zero Error')\nplt.legend()\nplt.show()\n","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-12-13T00:17:04.913505Z","iopub.execute_input":"2024-12-13T00:17:04.913956Z","iopub.status.idle":"2024-12-13T00:17:05.724851Z","shell.execute_reply.started":"2024-12-13T00:17:04.913907Z","shell.execute_reply":"2024-12-13T00:17:05.723780Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"# Precision-Recall Curve","metadata":{}},{"cell_type":"code","source":"from sklearn.preprocessing import LabelBinarizer\nfrom sklearn.metrics import precision_recall_curve, roc_auc_score\nimport matplotlib.pyplot as plt\n\n# Binarize the labels for multiclass problem\nlb = LabelBinarizer()\ny_test_binarized = lb.fit_transform(y_test_binned)\ny_pred_binarized = lb.transform(y_pred_binned)\n\n# Plot precision-recall curve for each class\nplt.figure(figsize=(8, 6))\nfor i in range(y_test_binarized.shape[1]):\n    precision, recall, _ = precision_recall_curve(y_test_binarized[:, i], y_pred_binarized[:, i])\n    auc_score = roc_auc_score(y_test_binarized[:, i], y_pred_binarized[:, i])\n    plt.plot(recall, precision, label=f'Class {i} (AUC = {auc_score:.2f})')\n\nplt.title(\"Precision-Recall Curve for Each Class\")\nplt.xlabel(\"Recall\")\nplt.ylabel(\"Precision\")\nplt.legend()\nplt.show()\n","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-12-13T00:17:05.726536Z","iopub.execute_input":"2024-12-13T00:17:05.726972Z","iopub.status.idle":"2024-12-13T00:17:06.027405Z","shell.execute_reply.started":"2024-12-13T00:17:05.726926Z","shell.execute_reply":"2024-12-13T00:17:06.026220Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"# comparison","metadata":{}},{"cell_type":"code","source":"import matplotlib.pyplot as plt\nimport numpy as np\n\n# R² Scores for both models\nr2_scores = [r2_xgb, r2_rf]\n\n# Mean Squared Errors for both models\nmse_scores = [mse_xgb, mse_rf]\n\n# Quadratic Weighted Kappa for both models\nqwk_scores = [qwk_xgb, qwk_rf]\n\n# Labels for the models\nmodels = ['XGBoost', 'Random Forest']\n\n# Set the position for the bars\nx = np.arange(len(models))\n\n# Width of the bars\nwidth = 0.25\n\n# Create a figure and axis for the plot\nfig, ax1 = plt.subplots(figsize=(10, 6))\n\n# Plot R² Score as bars with a bright blue color\nbar1 = ax1.bar(x - width, r2_scores, width, label='R² Score', color='#1f77b4')  # Dark blue\n\n# Plot MSE as bars with a vibrant orange color (on a secondary y-axis)\nax2 = ax1.twinx()\nbar2 = ax2.bar(x, mse_scores, width, label='MSE', color='#ff7f0e')  # Vibrant orange\n\n# Plot QWK as bars with a fresh green color (on a third y-axis)\nax3 = ax1.twinx()\nax3.spines['right'].set_position(('outward', 60))  # To avoid overlap with ax2\nbar3 = ax3.bar(x + width, qwk_scores, width, label='QWK', color='#2ca02c')  # Fresh green\n\n# Add labels and titles\nax1.set_xlabel('Models', fontsize=14)\nax1.set_ylabel('R² Score', color='#1f77b4', fontsize=12)\nax2.set_ylabel('Mean Squared Error', color='#ff7f0e', fontsize=12)\nax3.set_ylabel('Quadratic Weighted Kappa (QWK)', color='#2ca02c', fontsize=12)\n\nax1.set_title('Model Comparison: R² Score, MSE, and QWK', fontsize=16, weight='bold')\n\n# Add a legend\nax1.legend(loc='upper left', fontsize=12)\nax2.legend(loc='upper center', fontsize=12)\nax3.legend(loc='upper right', fontsize=12)\n\n# Show the plot\nplt.xticks(x, models, fontsize=12)\nplt.tight_layout()\nplt.show()\n","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-12-13T00:17:06.029083Z","iopub.execute_input":"2024-12-13T00:17:06.029875Z","iopub.status.idle":"2024-12-13T00:17:07.288496Z","shell.execute_reply.started":"2024-12-13T00:17:06.029822Z","shell.execute_reply":"2024-12-13T00:17:07.287164Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"import xgboost as xgb\nfrom sklearn.metrics import precision_recall_curve, roc_auc_score, r2_score, mean_squared_error, confusion_matrix\nfrom sklearn.preprocessing import KBinsDiscretizer, LabelBinarizer\nimport seaborn as sns\nimport matplotlib.pyplot as plt\nimport numpy as np\nfrom sklearn.ensemble import RandomForestRegressor\n\n# Predictions for both models\ny_pred_rf = rf.predict(X_test_pca)\ny_pred_xgb = xgb_model.predict(X_test_pca)\n\n# Binning the predicted values for classification evaluation\nbins = [0, 50, 100, 150]\nlabels = ['low', 'medium', 'high']\n\n# Initialize KBinsDiscretizer\ndiscretizer = KBinsDiscretizer(n_bins=len(bins) - 1, encode='ordinal', strategy='uniform')\n\ny_test_binned = discretizer.fit_transform(y_test.values.reshape(-1, 1)).flatten()\ny_pred_rf_binned = discretizer.transform(y_pred_rf.reshape(-1, 1)).flatten()\ny_pred_xgb_binned = discretizer.transform(y_pred_xgb.reshape(-1, 1)).flatten()\n\n# Precision-Recall Curve for RandomForest\nlb = LabelBinarizer()\ny_test_binarized = lb.fit_transform(y_test_binned)\ny_pred_rf_binarized = lb.transform(y_pred_rf_binned)\n\nplt.figure(figsize=(8, 6))\nfor i in range(y_test_binarized.shape[1]):\n    precision_rf, recall_rf, _ = precision_recall_curve(y_test_binarized[:, i], y_pred_rf_binarized[:, i])\n    auc_score_rf = roc_auc_score(y_test_binarized[:, i], y_pred_rf_binarized[:, i])\n    plt.plot(recall_rf, precision_rf, label=f'RandomForest Class {i} (AUC = {auc_score_rf:.2f})')\n\n# Precision-Recall Curve for XGBoost\ny_pred_xgb_binarized = lb.transform(y_pred_xgb_binned)\n\nfor i in range(y_test_binarized.shape[1]):\n    precision_xgb, recall_xgb, _ = precision_recall_curve(y_test_binarized[:, i], y_pred_xgb_binarized[:, i])\n    auc_score_xgb = roc_auc_score(y_test_binarized[:, i], y_pred_xgb_binarized[:, i])\n    plt.plot(recall_xgb, precision_xgb, label=f'XGBoost Class {i} (AUC = {auc_score_xgb:.2f})')\n\nplt.title(\"Precision-Recall Curve for RandomForest and XGBoost\")\nplt.xlabel(\"Recall\")\nplt.ylabel(\"Precision\")\nplt.legend()\nplt.show()\n","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-12-13T00:17:07.289916Z","iopub.execute_input":"2024-12-13T00:17:07.290289Z","iopub.status.idle":"2024-12-13T00:17:07.649717Z","shell.execute_reply.started":"2024-12-13T00:17:07.290236Z","shell.execute_reply":"2024-12-13T00:17:07.648470Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"Submission Preparation","metadata":{}},{"cell_type":"code","source":"test_data['id'] = test_data.index  # Creates a new 'id' column based on the row index","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-12-13T00:18:39.765426Z","iopub.execute_input":"2024-12-13T00:18:39.765866Z","iopub.status.idle":"2024-12-13T00:18:39.772597Z","shell.execute_reply.started":"2024-12-13T00:18:39.765834Z","shell.execute_reply":"2024-12-13T00:18:39.771349Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"# Ensure predictions are aligned with the test data\naligned_predictions = predictions[:len(test_data)]  # Truncate predictions if needed\n\n# Create the submission DataFrame\nsubmission = pd.DataFrame({\n    'id': test_data['id'],  # Now this should work\n    'sii': aligned_predictions\n})\n\n\n# Save the submission file\nsubmission.to_csv('submission.csv', index=False)\nprint(\"Submission file created: submission.csv\")\n","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-12-13T00:18:56.953342Z","iopub.execute_input":"2024-12-13T00:18:56.953868Z","iopub.status.idle":"2024-12-13T00:18:56.967234Z","shell.execute_reply.started":"2024-12-13T00:18:56.953828Z","shell.execute_reply":"2024-12-13T00:18:56.965952Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"# # If predictions were made for a larger dataset, filter them to match `test_data`\n# predictions = predictions[:len(test_data)]  # Truncate or adjust as necessary\n\n# # Ensure alignment before creating the submission file\n# submission = pd.DataFrame({\n#     'id': test_data['id'],  # Ensure this matches test_data\n#     'sii': predictions      # Ensure this matches the adjusted predictions\n# })\n\n# # Save the submission file\n# submission.to_csv('submission.csv', index=False)\n\n# print(\"Submission file created: submission.csv\")\n","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-12-13T00:17:08.945502Z","iopub.status.idle":"2024-12-13T00:17:08.946010Z","shell.execute_reply.started":"2024-12-13T00:17:08.945801Z","shell.execute_reply":"2024-12-13T00:17:08.945824Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"# **Conclusion**","metadata":{}},{"cell_type":"markdown","source":"In an era where the internet plays an integral role in our lives, it is crucial to address the growing concern of problematic internet use among children and adolescents. This project demonstrates how physical activity and fitness data can serve as accessible and reliable indicators for identifying early signs of internet addiction. By leveraging machine learning models, we have shown that it is possible to predict problematic internet use, paving the way for timely interventions.\n\nOur work emphasizes the importance of proactive measures to promote healthier digital habits and reduce the negative impacts of excessive screen time on mental and physical well-being. Beyond the scope of this study, these findings open doors to creating practical tools that parents, educators, and healthcare professionals can use to support children in navigating the digital landscape responsibly.\n\nUltimately, this project contributes to building a healthier and more balanced future for the next generation—one where technology is used as a tool for growth and connection rather than a source of dependency. Together, we can empower children and adolescents to lead healthier, more fulfilling lives in this ever-evolving digital world.","metadata":{}}]}