{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"name":"python","version":"3.10.14","mimetype":"text/x-python","codemirror_mode":{"name":"ipython","version":3},"pygments_lexer":"ipython3","nbconvert_exporter":"python","file_extension":".py"},"kaggle":{"accelerator":"none","dataSources":[{"sourceId":81933,"databundleVersionId":9643020,"sourceType":"competition"}],"dockerImageVersionId":30804,"isInternetEnabled":true,"language":"python","sourceType":"notebook","isGpuEnabled":false}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"This code is a comprehensive pipeline for a machine learning project that involves data preprocessing, feature engineering, building and training a CNN model, evaluating several machine learning models, and generating a submission file. Below is a detailed explanation of each part:\nImport Libraries\n\n    General Libraries:\n        os, warnings: For file system operations and suppressing warnings.\n        numpy (np), pandas (pd): Core libraries for numerical computations and data manipulation.\n        matplotlib.pyplot (plt), seaborn (sns): Visualization tools.\n\n    Machine Learning Libraries:\n        scikit-learn: Provides models and utilities like train_test_split, metrics, scaling, encoding, and ML models.\n        xgboost, lightgbm: Popular libraries for gradient-boosted trees.\n        optuna: Used for hyperparameter optimization.\n\n    Deep Learning Libraries:\n        tensorflow.keras: Used for creating and training the CNN model.\n        tqdm: Displays a progress bar for iterable operations.\n        concurrent.futures.ThreadPoolExecutor: Manages concurrent tasks.\n\nGlobal Settings\n\n    Set Random Seed: Ensures reproducibility of results.\n\n    SEED = 42\n    np.random.seed(SEED)\n\nData Loading and Preprocessing\n1. Processing Time-Series Data\n\n    Function process_file:\n        Loads a .parquet file for a given filename and extracts statistical summaries (e.g., mean, std).\n        Drops unnecessary columns like step.\n\n    Function load_time_series:\n        Iterates over all files in a directory using ThreadPoolExecutor for parallelism.\n        Computes descriptive statistics for each file and combines them into a DataFrame.\n\n2. Merge Time-Series with Main Dataset\n\n    Loads CSV files for training and test datasets.\n    Merges processed time-series features into the main DataFrame using a common column id.\n\n3. Preprocessing\n\n    Drops irrelevant columns (id) and rows with missing target values (sii).\n    Encodes categorical features using OrdinalEncoder, handling unknown categories with a default value of -1.\n\nFeature Scaling\n\n    Identifies numeric columns shared between train and test datasets.\n    Fills missing values in numeric columns with the median of the respective column.\n    Scales data using MinMaxScaler to normalize values between 0 and 1.\n\nCNN Model for Feature Engineering\n\n    Model Definition:\n        Input: 1D numeric features reshaped to include a channel dimension.\n        Layers:\n            Conv1D: Extracts local patterns in time-series data.\n            MaxPooling1D: Down-samples features.\n            Flatten: Converts 2D data to 1D for fully connected layers.\n            Dense: Two layers for high-level representation and regression output.\n        Loss: Mean Squared Error (MSE), suitable for regression tasks.\n\n    Training:\n        CNN is trained for 50 epochs on the scaled training data with a batch size of 256.\n        Validation split: 20% of training data.\n        Training loss is plotted for monitoring convergence.\n\nComparing Machine Learning Models\n\n    Model Selection:\n        Random Forest Regressor\n        XGBoost Regressor\n        Logistic Regression\n\n    Evaluation:\n        Trains each model on a split of training data.\n        Predicts on validation data.\n        Computes RMSE (Root Mean Squared Error) for each model to compare performance.\n\nQuadratic Weighted Kappa (QWK)\n\n    A custom metric used to evaluate the agreement between predicted and true labels.\n    The function computes:\n        A weighted confusion matrix.\n        A kappa score that accounts for random agreement.\n\nEvaluation and Metrics\n\n    Confusion Matrix:\n        Visualizes the performance of the best model (assumed to be Random Forest).\n        Compares true labels with predicted values on validation data.\n\n    Save Submission:\n        Generates a CSV submission file with predictions on the test dataset.\n        Uses sample_submission to align the format with the competition requirements.\n\nOutput\n\n    Training Results:\n        Loss plots for CNN training.\n        RMSE for Random Forest, XGBoost, and Logistic Regression.\n\n    Evaluation Results:\n        Confusion matrix for the best model.\n        QWK score for Random Forest.\n\n    Submission:\n        A CSV file ready for upload to the competition platform.\n\nKey Algorithmic Logic:\n\n    Parallelism: ThreadPoolExecutor speeds up processing of time-series files.\n    Hybrid Models: Combines CNN for feature extraction and traditional models for regression.\n    Comprehensive Evaluation: Multiple metrics (RMSE, QWK) provide robust evaluation.\n    Reproducibility: Seed ensures consistent results.","metadata":{}},{"cell_type":"markdown","source":"# Graphical Plots","metadata":{}},{"cell_type":"markdown","source":"The accompanying image(s) consists of two plots:\n\n1. Confusion Matrix ():\n\nThe confusion matrix is a performance evaluation metric used to visualize how well a classification model is performing. Here's a detailed breakdown of this matrix:\n\n    Structure:\n        The rows represent the True Labels (actual class of the data points).\n        The columns represent the Predicted Labels (class predicted by the model).\n\n    Key Observations:\n        Diagonal elements: These represent correct predictions (i.e., when the predicted label matches the true label). For example:\n            Class 0 has 202 correct predictions.\n            Class 1 has 87 correct predictions.\n            Class 2 has 2 correct predictions.\n            Class 3 has 3 correct predictions.\n        Off-diagonal elements: These represent misclassifications (errors made by the model). For instance:\n            131 instances of class 0 were misclassified as class 1.\n            41 instances of class 1 were misclassified as class 0.\n            60 instances of class 2 were misclassified as class 1.\n\n    Overall Insights:\n        Class 0 is predicted relatively well with 202 correct predictions out of a total of 202+131+3+0=336202+131+3+0=336 instances, giving a rough accuracy for this class of around 60.1%.\n        Classes 2 and 3 have very few correct predictions, which suggests the model struggles significantly with these categories.\n\n    Color Intensity:\n        The intensity of the color in each cell represents the number of predictions, with darker blue shades indicating higher counts. For example, the cell corresponding to true label 0 and predicted label 0 (202 correct predictions) is the darkest.\n\n2. CNN Training Loss ():\n\nThis plot shows how the loss of a Convolutional Neural Network (CNN) decreases during the training process over multiple epochs. Here's a detailed explanation:\n\n    Axes:\n        The x-axis represents the Epochs (training iterations).\n        The y-axis represents the Loss (a measure of model error; lower is better).\n\n    Key Observations:\n        The loss starts at a relatively high value (~0.65) at the beginning of training (epoch 0).\n        It decreases rapidly during the first few epochs (steep drop between epochs 0–10), indicating that the model is learning quickly in the early stages.\n        After about 20 epochs, the loss plateaus, indicating diminishing returns from further training. The curve flattens out around 0.40, suggesting that the model is approaching convergence.\n\n    Interpretation:\n        A steady decline in loss over epochs suggests that the CNN is learning effectively.\n        The flattening of the curve could indicate that the model has reached its optimal performance on the training set and further training may not lead to significant improvement.\n\nGeneral Analysis and Recommendations:\n\n    Confusion Matrix:\n        The model performs well for class 0 but struggles with classes 2 and 3. Consider:\n            Data Augmentation: Increase training data for underperforming classes.\n            Class Balancing: Use techniques like oversampling for minority classes or weighted loss functions to handle class imbalance.\n            Model Refinement: Adjust the architecture or hyperparameters of the CNN to improve its ability to distinguish among all classes.\n\n    Training Loss:\n        While the loss reduction trend is promising, analyze the validation loss alongside training loss to ensure the model is not overfitting. If the gap between training and validation loss is large, regularization methods like dropout or early stopping might help.\n        Experiment with learning rate adjustments or advanced optimizers to achieve even faster convergence.","metadata":{}},{"cell_type":"code","source":"# Import Libraries\nimport os\nimport numpy as np\nimport pandas as pd\nimport matplotlib.pyplot as plt\nimport seaborn as sns\nfrom sklearn.model_selection import train_test_split\nfrom sklearn.metrics import mean_squared_error, confusion_matrix, cohen_kappa_score\nfrom sklearn.preprocessing import MinMaxScaler, OrdinalEncoder\nfrom sklearn.ensemble import RandomForestRegressor\nfrom sklearn.linear_model import LogisticRegression\nfrom sklearn.svm import SVR\nimport tensorflow as tf\nfrom tensorflow.keras.models import Model\nfrom tensorflow.keras.layers import Input, Conv1D, MaxPooling1D, Flatten, Dense\nfrom tensorflow.keras.optimizers import Adam\nfrom tensorflow.keras.callbacks import EarlyStopping\nfrom tqdm import tqdm\nfrom concurrent.futures import ThreadPoolExecutor\nimport warnings\nimport xgboost as xgb\nimport optuna\nimport lightgbm as lgb\n\n# Suppress warnings\nwarnings.filterwarnings('ignore')\n\n# Global Parameters\nSEED = 42\nnp.random.seed(SEED)\n\n# ------------------------------\n# Data Loading and Preprocessing\n# ------------------------------\ndef process_file(filename, dirname):\n    df = pd.read_parquet(os.path.join(dirname, filename, 'part-0.parquet'))\n    df.drop('step', axis=1, inplace=True)\n    return df.describe().values.reshape(-1), filename.split('=')[1]\n\ndef load_time_series(dirname):\n    ids = os.listdir(dirname)\n    with ThreadPoolExecutor() as executor:\n        results = list(tqdm(executor.map(lambda fname: process_file(fname, dirname), ids), total=len(ids)))\n    stats, indexes = zip(*results)\n    df = pd.DataFrame(stats, columns=[f\"Stat_{i}\" for i in range(len(stats[0]))])\n    df['id'] = indexes\n    return df\n\n# Load datasets\ntrain = pd.read_csv('/kaggle/input/child-mind-institute-problematic-internet-use/train.csv')\ntest = pd.read_csv('/kaggle/input/child-mind-institute-problematic-internet-use/test.csv')\nsample = pd.read_csv('/kaggle/input/child-mind-institute-problematic-internet-use/sample_submission.csv')\ntrain_ts = load_time_series(\"/kaggle/input/child-mind-institute-problematic-internet-use/series_train.parquet\")\ntest_ts = load_time_series(\"/kaggle/input/child-mind-institute-problematic-internet-use/series_test.parquet\")\n\n# Merge time-series data\ntrain = pd.merge(train, train_ts, how=\"left\", on='id')\ntest = pd.merge(test, test_ts, how=\"left\", on='id')\n\n# Drop irrelevant columns\ntrain = train.drop(columns=['id'])\ntest = test.drop(columns=['id'])\n\n# Handle missing target values\ntrain = train.dropna(subset=['sii'])\n\n# Categorical Columns\ncat_cols = ['Basic_Demos-Enroll_Season', 'CGAS-Season', 'Physical-Season']\nencoder = OrdinalEncoder(handle_unknown='use_encoded_value', unknown_value=-1)\ntrain[cat_cols] = encoder.fit_transform(train[cat_cols])\ntest[cat_cols] = encoder.transform(test[cat_cols])\n\n# ------------------------------\n# Feature Scaling\n# ------------------------------\n# Select common numeric columns\nnumeric_cols = train.select_dtypes(include=['number']).columns\ncommon_numeric_cols = list(set(numeric_cols) & set(test.columns))\n\n# Check and fill missing values only for numeric columns\nfor col in common_numeric_cols:\n    train[col] = train[col].fillna(train[col].median())\n    test[col] = test[col].fillna(test[col].median())\n\n# Scale data using MinMaxScaler\nscaler = MinMaxScaler()\ntrain_scaled = scaler.fit_transform(train[common_numeric_cols])\ntest_scaled = scaler.transform(test[common_numeric_cols])\n\n# ------------------------------\n# CNN Model for Feature Engineering\n# ------------------------------\ndef build_cnn(input_dim):\n    input_layer = Input(shape=(input_dim, 1))  # Assume 1D data\n    x = Conv1D(64, kernel_size=3, activation='relu')(input_layer)\n    x = MaxPooling1D(pool_size=2)(x)\n    x = Flatten()(x)\n    x = Dense(64, activation='relu')(x)\n    output_layer = Dense(1, activation='linear')(x)\n    model = Model(inputs=input_layer, outputs=output_layer)\n    model.compile(optimizer=Adam(learning_rate=0.001), loss='mse')\n    return model\n\n# Reshape data for CNN input (1D convolution)\ntrain_scaled_cnn = np.expand_dims(train_scaled, axis=-1)\ntest_scaled_cnn = np.expand_dims(test_scaled, axis=-1)\n\n# Build and train CNN\ncnn_model = build_cnn(train_scaled_cnn.shape[1])\nhistory = cnn_model.fit(train_scaled_cnn, train['sii'], epochs=50, batch_size=256, verbose=1, validation_split=0.2)\n\n# Plot training loss\nplt.figure(figsize=(10, 6))\nplt.plot(history.history['loss'], label='Loss')\nplt.xlabel('Epoch')\nplt.ylabel('Loss')\nplt.legend()\nplt.title('CNN Training Loss')\nplt.show()\n\n# ------------------------------\n# Comparing Machine Learning Models\n# ------------------------------\n\n# Prepare train and validation sets\nX_train, X_val, y_train, y_val = train_test_split(train_scaled, train['sii'], test_size=0.2, random_state=SEED)\n\n# RandomForest Model\nrf_model = RandomForestRegressor(random_state=SEED)\nrf_model.fit(X_train, y_train)\nrf_preds = rf_model.predict(X_val)\n\n# XGBoost Model\nxgb_model = xgb.XGBRegressor(random_state=SEED)\nxgb_model.fit(X_train, y_train)\nxgb_preds = xgb_model.predict(X_val)\n\n# Logistic Regression Model\nlr_model = LogisticRegression()\nlr_model.fit(X_train, y_train)\nlr_preds = lr_model.predict(X_val)\n\n# Calculate RMSE for each model\nrf_rmse = np.sqrt(mean_squared_error(y_val, rf_preds))\nxgb_rmse = np.sqrt(mean_squared_error(y_val, xgb_preds))\nlr_rmse = np.sqrt(mean_squared_error(y_val, lr_preds))\n\nprint(f\"Random Forest RMSE: {rf_rmse}\")\nprint(f\"XGBoost RMSE: {xgb_rmse}\")\nprint(f\"Logistic Regression RMSE: {lr_rmse}\")\n\n# ------------------------------\n# Optimized Quadratic Weighted Kappa (QWK) Score\n# ------------------------------\ndef quadratic_weighted_kappa(y_true, y_pred):\n    # Normalize predictions to integers\n    y_pred = np.round(y_pred).astype(int)\n    y_true = np.round(y_true).astype(int)\n\n    # Ensure values are within the range [0, n_classes-1]\n    n_classes = len(np.unique(y_true))\n    y_true = np.clip(y_true, 0, n_classes - 1)\n    y_pred = np.clip(y_pred, 0, n_classes - 1)\n\n    # Calculate the weighted confusion matrix\n    conf_matrix = confusion_matrix(y_true, y_pred)\n    \n    # Calculate weights\n    weights = np.zeros((n_classes, n_classes))\n    for i in range(n_classes):\n        for j in range(n_classes):\n            weights[i, j] = ((i - j) ** 2) / (n_classes - 1) ** 2\n\n    # Calculate QWK\n    weighted_conf_matrix = np.multiply(conf_matrix, weights)\n    numerator = np.sum(weighted_conf_matrix)\n    denominator = np.sum(conf_matrix) * np.sum(weights)\n    qwk = 1 - (numerator / denominator)\n    \n    return qwk\n\n# Compute QWK for RandomForest on validation data\nqwk_score = quadratic_weighted_kappa(y_val, rf_preds)\nprint(f\"Optimized Quadratic Weighted Kappa (QWK) Score: {qwk_score}\")\n\n# ------------------------------\n# Evaluation and Metrics\n# ------------------------------\n# Compare the best model's predictions with true values\nbest_model = rf_model  # Assume RF performed best\ntest_preds = best_model.predict(test_scaled)\n\n# Instead of using test['sii'], use y_val (true labels from validation set)\ncm = confusion_matrix(y_val, np.round(rf_preds).astype(int))  # rf_preds are predictions on validation set\nplt.figure(figsize=(10, 7))\nsns.heatmap(cm, annot=True, fmt='d', cmap='Blues')\nplt.xlabel('Predicted')\nplt.ylabel('True')\nplt.title('Confusion Matrix')\nplt.show()\n\n# ------------------------------\n# Save Submission\n# ------------------------------\nsubmission = pd.DataFrame({\n    'id': sample['id'],  # Use 'id' from the sample submission file\n    'sii': test_preds  # Use predictions on test data\n})\nsubmission.to_csv('/kaggle/working/submission.csv', index=False)\n\n# Display the first few rows of the saved submission file\nprint(\"Preview of the submission file:\")\nprint(submission.head())","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-12-19T22:26:12.435974Z","iopub.execute_input":"2024-12-19T22:26:12.437271Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"# Previously....","metadata":{}},{"cell_type":"markdown","source":"Code Explanation\n\nGlobal Setup\n\n    Imports and Dependencies:\n        The code imports key libraries like pandas, numpy, and sklearn for data manipulation and modeling.\n        Machine learning models like LightGBM, XGBoost, and CatBoost are used for their efficiency in tabular data.\n        Utilities like tqdm provide progress bars, and ThreadPoolExecutor allows parallel execution.\n\n    Global Parameters:\n        Constants like SEED for reproducibility and N_SPLITS for cross-validation setup.\n\nKey Functions\n\n    Quadratic Weighted Kappa Metric:\n        cohen_kappa_score is used to evaluate the agreement between predicted and true labels.\n\n    Threshold-Based Rounding:\n        Custom logic maps raw regression predictions to discrete class labels using thresholds.\n\n    Optimization Logic:\n        scipy.optimize.minimize tunes thresholds to maximize the quadratic weighted kappa.\n\n    Time-Series Processing:\n        process_file extracts statistical features (mean, std, etc.) from time-series data in .parquet files.\n\n    Categorical Encoding:\n        OneHotEncoder transforms categorical columns into binary dummy variables.\n\n    Model Training:\n        The train_ml function uses stratified K-Fold cross-validation for robust evaluation.\n        Voting ensemble (combination of LightGBM, XGBoost, CatBoost) aggregates predictions.\n\n    Submission Preparation:\n        Final predictions are saved in the required CSV format for competition submissions.\n\nDataset and Feature Processing\n\n    Loads and preprocesses train/test datasets.\n    Extracts time-series features and aligns train/test feature sets.\n\nModel Training and Submission\n\n    Constructs a VotingRegressor with three models.\n    Trains the model using cross-validation, optimizes thresholds, and generates a submission.","metadata":{}},{"cell_type":"code","source":"import pandas as pd\nimport numpy as np\nimport os\nfrom sklearn.base import clone\nfrom sklearn.metrics import cohen_kappa_score\nfrom sklearn.model_selection import StratifiedKFold\nfrom sklearn.impute import SimpleImputer\nfrom sklearn.preprocessing import OneHotEncoder\nfrom sklearn.pipeline import Pipeline\nfrom sklearn.ensemble import VotingRegressor\nfrom lightgbm import LGBMRegressor\nfrom xgboost import XGBRegressor\nfrom catboost import CatBoostRegressor\nfrom tqdm import tqdm\nfrom scipy.optimize import minimize\nfrom concurrent.futures import ThreadPoolExecutor\n\n# Global Parameters\nSEED = 42\nN_SPLITS = 5\n\n# Quadratic Weighted Kappa Metric\ndef quadratic_weighted_kappa(y_true, y_pred):\n    return cohen_kappa_score(y_true, y_pred, weights='quadratic')\n\n# Threshold-Based Prediction Rounding\ndef threshold_rounder(oof_non_rounded, thresholds):\n    return np.where(oof_non_rounded < thresholds[0], 0,\n                    np.where(oof_non_rounded < thresholds[1], 1,\n                             np.where(oof_non_rounded < thresholds[2], 2, 3)))\n\n# Evaluation Function for Optimization\ndef evaluate_predictions(thresholds, y_true, oof_non_rounded):\n    rounded_predictions = threshold_rounder(oof_non_rounded, thresholds)\n    return -quadratic_weighted_kappa(y_true, rounded_predictions)\n\n# Process Parquet Files to Extract Time Series Features\ndef process_file(filename, dirname):\n    df = pd.read_parquet(os.path.join(dirname, filename, 'part-0.parquet'))\n    df.drop('step', axis=1, inplace=True)\n    return df.describe().values.reshape(-1), filename.split('=')[1]\n\ndef load_time_series(dirname) -> pd.DataFrame:\n    ids = os.listdir(dirname)\n    with ThreadPoolExecutor() as executor:\n        results = list(tqdm(executor.map(lambda fname: process_file(fname, dirname), ids), total=len(ids)))\n    stats, indexes = zip(*results)\n    df = pd.DataFrame(stats, columns=[f\"stat_{i}\" for i in range(len(stats[0]))])\n    df['id'] = indexes\n    return df\n\n# Encode Categorical Columns\ndef encode_categorical_columns(df, categorical_columns):\n    existing_columns = [col for col in categorical_columns if col in df.columns]\n    if not existing_columns:\n        return df\n    encoder = OneHotEncoder(handle_unknown='ignore', sparse_output=False)  # Use sparse_output=False\n    encoded = encoder.fit_transform(df[existing_columns])\n    encoded_df = pd.DataFrame(encoded, columns=encoder.get_feature_names_out(existing_columns))\n    df = df.drop(columns=existing_columns).reset_index(drop=True)\n    return pd.concat([df, encoded_df], axis=1)\n\n# Train and Evaluate Models\ndef train_ml(model_class, train_data, test_data, sample_submission):\n    # Handle Missing Values in Target Column\n    imputer = SimpleImputer(strategy=\"most_frequent\")\n    train_data['sii'] = imputer.fit_transform(train_data[['sii']])\n\n    X = train_data.drop(['sii'], axis=1)\n    y = train_data['sii']\n\n    # Stratified K-Fold\n    skf = StratifiedKFold(n_splits=N_SPLITS, shuffle=True, random_state=SEED)\n    oof_non_rounded = np.zeros(len(y), dtype=float)\n    test_preds = np.zeros((len(test_data), N_SPLITS))\n\n    for fold, (train_idx, val_idx) in enumerate(tqdm(skf.split(X, y), total=N_SPLITS)):\n        X_train, X_val = X.iloc[train_idx], X.iloc[val_idx]\n        y_train, y_val = y.iloc[train_idx], y.iloc[val_idx]\n\n        # Train the Model\n        model = clone(model_class)\n        model.fit(X_train, y_train)\n        val_preds = model.predict(X_val)\n\n        # Store Predictions\n        oof_non_rounded[val_idx] = val_preds\n        test_preds[:, fold] = model.predict(test_data)\n\n        # Metrics\n        train_kappa = quadratic_weighted_kappa(y_train, model.predict(X_train).round())\n        val_kappa = quadratic_weighted_kappa(y_val, val_preds.round())\n        print(f\"Fold {fold + 1} - Train QWK: {train_kappa:.4f}, Validation QWK: {val_kappa:.4f}\")\n\n    # Optimize Thresholds\n    kappa_optimizer = minimize(evaluate_predictions, x0=[0.5, 1.5, 2.5], args=(y, oof_non_rounded), method='Nelder-Mead')\n    assert kappa_optimizer.success, \"Threshold optimization failed\"\n    thresholds = kappa_optimizer.x\n\n    # Final Predictions\n    test_preds_mean = test_preds.mean(axis=1)\n    final_predictions = threshold_rounder(test_preds_mean, thresholds)\n\n    # Create Submission\n    submission = pd.DataFrame({'id': sample_submission['id'], 'sii': final_predictions})\n    return submission\n\n# Load Dataset\ntrain = pd.read_csv('/kaggle/input/child-mind-institute-problematic-internet-use/train.csv')\ntest = pd.read_csv('/kaggle/input/child-mind-institute-problematic-internet-use/test.csv')\nsample_submission = pd.read_csv('/kaggle/input/child-mind-institute-problematic-internet-use/sample_submission.csv')\n\n# Process Time Series Data\ntrain_ts = load_time_series('/kaggle/input/child-mind-institute-problematic-internet-use/series_train.parquet')\ntest_ts = load_time_series('/kaggle/input/child-mind-institute-problematic-internet-use/series_test.parquet')\n\n# Merge Data\ntrain = pd.merge(train, train_ts, on='id', how='left').drop('id', axis=1)\ntest = pd.merge(test, test_ts, on='id', how='left').drop('id', axis=1)\n\n# Encode Categorical Columns\ncategorical_columns = [\n    'Basic_Demos-Enroll_Season', 'CGAS-Season', 'Physical-Season', 'Fitness_Endurance-Season',\n    'FGC-Season', 'BIA-Season', 'PAQ_A-Season', 'PAQ_C-Season', 'PCIAT-Season',\n    'SDS-Season', 'PreInt_EduHx-Season'\n]\ntrain = encode_categorical_columns(train, categorical_columns)\ntest = encode_categorical_columns(test, categorical_columns)\n\n# Align Features\ncommon_features = list(set(train.columns) & set(test.columns))\ntrain = train[common_features + ['sii']]\ntest = test[common_features]\n\n# Model Configuration\nlgb_params = {'learning_rate': 0.05, 'max_depth': 6, 'num_leaves': 31, 'n_estimators': 100, 'random_state': SEED}\nxgb_params = {'learning_rate': 0.05, 'max_depth': 6, 'n_estimators': 100, 'random_state': SEED}\ncat_params = {'learning_rate': 0.05, 'depth': 6, 'iterations': 100, 'random_seed': SEED}\n\n# Models\nlgb_model = LGBMRegressor(**lgb_params)\nxgb_model = XGBRegressor(**xgb_params)\ncat_model = CatBoostRegressor(**cat_params, silent=True)\n\n# Voting Regressor\nvoting_model = VotingRegressor([('lgb', lgb_model), ('xgb', xgb_model), ('cat', cat_model)])\n\n# Train and Generate Submission\nsubmission = train_ml(voting_model, train, test, sample_submission)\nsubmission.to_csv('submission.csv', index=False)\nprint(submission.head())","metadata":{"_uuid":"8f2839f25d086af736a60e9eeb907d3b93b6e0e5","_cell_guid":"b1076dfc-b9ad-4769-8c92-a6c4dae69d19","trusted":true},"outputs":[],"execution_count":null},{"cell_type":"code","source":"import pandas as pd\nimport numpy as np\nimport os\nfrom sklearn.base import clone\nfrom sklearn.metrics import cohen_kappa_score, confusion_matrix, ConfusionMatrixDisplay\nfrom sklearn.model_selection import StratifiedKFold\nfrom sklearn.impute import SimpleImputer\nfrom sklearn.preprocessing import OneHotEncoder\nfrom sklearn.pipeline import Pipeline\nfrom sklearn.ensemble import VotingRegressor\nfrom lightgbm import LGBMRegressor\nfrom xgboost import XGBRegressor\nfrom catboost import CatBoostRegressor\nfrom tqdm import tqdm\nfrom scipy.optimize import minimize\nfrom concurrent.futures import ThreadPoolExecutor\nimport matplotlib.pyplot as plt\nimport seaborn as sns\n\n# Global Parameters\nSEED = 42\nN_SPLITS = 5\n\n# Quadratic Weighted Kappa Metric\ndef quadratic_weighted_kappa(y_true, y_pred):\n    return cohen_kappa_score(y_true, y_pred, weights='quadratic')\n\n# Threshold-Based Prediction Rounding\ndef threshold_rounder(oof_non_rounded, thresholds):\n    return np.where(oof_non_rounded < thresholds[0], 0,\n                    np.where(oof_non_rounded < thresholds[1], 1,\n                             np.where(oof_non_rounded < thresholds[2], 2, 3)))\n\n# Evaluation Function for Optimization\ndef evaluate_predictions(thresholds, y_true, oof_non_rounded):\n    rounded_predictions = threshold_rounder(oof_non_rounded, thresholds)\n    return -quadratic_weighted_kappa(y_true, rounded_predictions)\n\n# Process Parquet Files to Extract Time Series Features\ndef process_file(filename, dirname):\n    df = pd.read_parquet(os.path.join(dirname, filename, 'part-0.parquet'))\n    df.drop('step', axis=1, inplace=True)\n    return df.describe().values.reshape(-1), filename.split('=')[1]\n\ndef load_time_series(dirname) -> pd.DataFrame:\n    ids = os.listdir(dirname)\n    with ThreadPoolExecutor() as executor:\n        results = list(tqdm(executor.map(lambda fname: process_file(fname, dirname), ids), total=len(ids)))\n    stats, indexes = zip(*results)\n    df = pd.DataFrame(stats, columns=[f\"stat_{i}\" for i in range(len(stats[0]))])\n    df['id'] = indexes\n    return df\n\n# Encode Categorical Columns\ndef encode_categorical_columns(df, categorical_columns):\n    existing_columns = [col for col in categorical_columns if col in df.columns]\n    if not existing_columns:\n        return df\n    encoder = OneHotEncoder(handle_unknown='ignore', sparse_output=False)\n    encoded = encoder.fit_transform(df[existing_columns])\n    encoded_df = pd.DataFrame(encoded, columns=encoder.get_feature_names_out(existing_columns))\n    df = df.drop(columns=existing_columns).reset_index(drop=True)\n    return pd.concat([df, encoded_df], axis=1)\n\n# Train and Evaluate Models\ndef train_ml(model_class, train_data, test_data, sample_submission):\n    # Handle Missing Values in Target Column\n    imputer = SimpleImputer(strategy=\"most_frequent\")\n    train_data['sii'] = imputer.fit_transform(train_data[['sii']])\n\n    X = train_data.drop(['sii'], axis=1)\n    y = train_data['sii']\n\n    # Stratified K-Fold\n    skf = StratifiedKFold(n_splits=N_SPLITS, shuffle=True, random_state=SEED)\n    oof_non_rounded = np.zeros(len(y), dtype=float)\n    test_preds = np.zeros((len(test_data), N_SPLITS))\n    metrics = []\n\n    for fold, (train_idx, val_idx) in enumerate(tqdm(skf.split(X, y), total=N_SPLITS)):\n        X_train, X_val = X.iloc[train_idx], X.iloc[val_idx]\n        y_train, y_val = y.iloc[train_idx], y.iloc[val_idx]\n\n        # Train the Model\n        model = clone(model_class)\n        model.fit(X_train, y_train)\n        val_preds = model.predict(X_val)\n\n        # Store Predictions\n        oof_non_rounded[val_idx] = val_preds\n        test_preds[:, fold] = model.predict(test_data)\n\n        # Metrics\n        train_kappa = quadratic_weighted_kappa(y_train, model.predict(X_train).round())\n        val_kappa = quadratic_weighted_kappa(y_val, val_preds.round())\n        metrics.append((train_kappa, val_kappa))\n        print(f\"Fold {fold + 1} - Train QWK: {train_kappa:.4f}, Validation QWK: {val_kappa:.4f}\")\n\n    # Optimize Thresholds\n    kappa_optimizer = minimize(evaluate_predictions, x0=[0.5, 1.5, 2.5], args=(y, oof_non_rounded), method='Nelder-Mead')\n    assert kappa_optimizer.success, \"Threshold optimization failed\"\n    thresholds = kappa_optimizer.x\n\n    # Final Predictions\n    test_preds_mean = test_preds.mean(axis=1)\n    final_predictions = threshold_rounder(test_preds_mean, thresholds)\n\n    # Create Submission\n    submission = pd.DataFrame({'id': sample_submission['id'], 'sii': final_predictions})\n\n    # Visualization Logic\n    plot_metrics(metrics)\n    plot_confusion_matrix(y, oof_non_rounded, thresholds)\n    return submission\n\n# Plot Metrics\ndef plot_metrics(metrics):\n    train_kappas, val_kappas = zip(*metrics)\n    plt.figure(figsize=(10, 6))\n    plt.plot(range(1, len(metrics) + 1), train_kappas, label=\"Train QWK\")\n    plt.plot(range(1, len(metrics) + 1), val_kappas, label=\"Validation QWK\")\n    plt.title(\"Quadratic Weighted Kappa Across Folds\")\n    plt.xlabel(\"Fold\")\n    plt.ylabel(\"QWK\")\n    plt.legend()\n    plt.grid(True)\n    plt.show()\n\n# Plot Confusion Matrix\ndef plot_confusion_matrix(y_true, oof_preds, thresholds):\n    final_preds = threshold_rounder(oof_preds, thresholds)\n    cm = confusion_matrix(y_true, final_preds)\n    disp = ConfusionMatrixDisplay(confusion_matrix=cm)\n    disp.plot(cmap=plt.cm.Blues)\n    plt.title(\"Confusion Matrix\")\n    plt.show()\n\n# Load Dataset\ntrain = pd.read_csv('/kaggle/input/child-mind-institute-problematic-internet-use/train.csv')\ntest = pd.read_csv('/kaggle/input/child-mind-institute-problematic-internet-use/test.csv')\nsample_submission = pd.read_csv('/kaggle/input/child-mind-institute-problematic-internet-use/sample_submission.csv')\n\n# Process Time Series Data\ntrain_ts = load_time_series('/kaggle/input/child-mind-institute-problematic-internet-use/series_train.parquet')\ntest_ts = load_time_series('/kaggle/input/child-mind-institute-problematic-internet-use/series_test.parquet')\n\n# Merge Data\ntrain = pd.merge(train, train_ts, on='id', how='left').drop('id', axis=1)\ntest = pd.merge(test, test_ts, on='id', how='left').drop('id', axis=1)\n\n# Encode Categorical Columns\ncategorical_columns = [\n    'Basic_Demos-Enroll_Season', 'CGAS-Season', 'Physical-Season', 'Fitness_Endurance-Season',\n    'FGC-Season', 'BIA-Season', 'PAQ_A-Season', 'PAQ_C-Season', 'PCIAT-Season',\n    'SDS-Season', 'PreInt_EduHx-Season'\n]\ntrain = encode_categorical_columns(train, categorical_columns)\ntest = encode_categorical_columns(test, categorical_columns)\n\n# Align Features\ncommon_features = list(set(train.columns) & set(test.columns))\ntrain = train[common_features + ['sii']]\ntest = test[common_features]\n\n# Model Configuration\nlgb_params = {'learning_rate': 0.05, 'max_depth': 6, 'num_leaves': 31, 'n_estimators': 100, 'random_state': SEED}\nxgb_params = {'learning_rate': 0.05, 'max_depth': 6, 'n_estimators': 100, 'random_state': SEED}\ncat_params = {'learning_rate': 0.05, 'depth': 6, 'iterations': 100, 'random_seed': SEED}\n\n# Models\nlgb_model = LGBMRegressor(**lgb_params)\nxgb_model = XGBRegressor(**xgb_params)\ncat_model = CatBoostRegressor(**cat_params, silent=True)\n\n# Voting Regressor\nvoting_model = VotingRegressor([('lgb', lgb_model), ('xgb', xgb_model), ('cat', cat_model)])\n\n# Train and Generate Submission\nsubmission = train_ml(voting_model, train, test, sample_submission)\nsubmission.to_csv('submission.csv', index=False)\nprint(submission.head())","metadata":{"trusted":true},"outputs":[],"execution_count":null}]}