{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"name":"python","version":"3.10.14","mimetype":"text/x-python","codemirror_mode":{"name":"ipython","version":3},"pygments_lexer":"ipython3","nbconvert_exporter":"python","file_extension":".py"},"kaggle":{"accelerator":"none","dataSources":[{"sourceId":81933,"databundleVersionId":9643020,"sourceType":"competition"}],"dockerImageVersionId":30761,"isInternetEnabled":false,"language":"python","sourceType":"notebook","isGpuEnabled":false}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"code","source":"# This Python 3 environment comes with many helpful analytics libraries installed\n# It is defined by the kaggle/python Docker image: https://github.com/kaggle/docker-python\n# For example, here's several helpful packages to load\n\nimport numpy as np # linear algebra\nimport pandas as pd # data processing, CSV file I/O (e.g. pd.read_csv)\n\n# Input data files are available in the read-only \"../input/\" directory\n# For example, running this (by clicking run or pressing Shift+Enter) will list all files under the input directory\n\nimport os\n\n# You can write up to 20GB to the current directory (/kaggle/working/) that gets preserved as output when you create a version using \"Save & Run All\" \n# You can also write temporary files to /kaggle/temp/, but they won't be saved outside of the current session","metadata":{"_uuid":"8f2839f25d086af736a60e9eeb907d3b93b6e0e5","_cell_guid":"b1076dfc-b9ad-4769-8c92-a6c4dae69d19","execution":{"iopub.status.busy":"2024-09-20T19:46:06.925310Z","iopub.execute_input":"2024-09-20T19:46:06.925833Z","iopub.status.idle":"2024-09-20T19:46:06.932714Z","shell.execute_reply.started":"2024-09-20T19:46:06.925786Z","shell.execute_reply":"2024-09-20T19:46:06.931377Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Multi-Target Regression with Neural Networks and Optimized Rounding\n\nIn this notebook, we will perform multi-target regression on tabular data using a neural network. We will use PyTorch for modeling and Optuna for optimizing the rounding thresholds to maximize the Quadratic Weighted Kappa (QWK) score. The workflow includes data loading, preprocessing, model training with cross-validation, and making predictions on the test set.","metadata":{}},{"cell_type":"markdown","source":"# Import Libraries\n\nWe start by importing all the necessary libraries:\n\n- warnings: To handle warning messages.\n- functools.partial: For creating partial functions.\n- pathlib.Path: For handling file paths.\n- numpy: For numerical computations.\n- optuna: For hyperparameter optimization.\n- pandas: For data manipulation.\n- polars: A fast DataFrame library for data loading and manipulation.\n- torch: PyTorch library for building neural networks.\n- sklearn: For preprocessing and evaluation metrics.\n\nWe also suppress specific warnings to keep the output clean.","metadata":{}},{"cell_type":"code","source":"import warnings\nfrom functools import partial\nfrom pathlib import Path\n\nimport numpy as np\nimport optuna\nimport pandas as pd\nimport polars as pl\nimport polars.selectors as cs\nimport torch\nimport torch.nn as nn\nimport torch.utils.data\nfrom numpy.typing import NDArray\nfrom sklearn.metrics import cohen_kappa_score\nfrom sklearn.model_selection import StratifiedKFold\nfrom sklearn.preprocessing import LabelEncoder, StandardScaler\n\nwarnings.filterwarnings(\"ignore\", message=\"Failed to optimize method\")\n\n\ndevice = torch.device('cuda' if torch.cuda.is_available() else 'cpu')","metadata":{"execution":{"iopub.status.busy":"2024-09-20T19:46:06.935439Z","iopub.execute_input":"2024-09-20T19:46:06.936022Z","iopub.status.idle":"2024-09-20T19:46:06.953558Z","shell.execute_reply.started":"2024-09-20T19:46:06.935965Z","shell.execute_reply":"2024-09-20T19:46:06.952238Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Define Constants\n\nWe define:\n\n* **DATA_DIR**: The directory where the data files are located.\n* **TARGET_COLS**: The list of target columns we aim to predict.\n* **FEATURE_COLS**: The list of feature columns we will use for modeling.","metadata":{}},{"cell_type":"code","source":"DATA_DIR = Path(\"../input/child-mind-institute-problematic-internet-use\")\n\nTARGET_COLS = [\"PCIAT-PCIAT_01\", \"PCIAT-PCIAT_02\", \"PCIAT-PCIAT_03\", \"PCIAT-PCIAT_04\", \"PCIAT-PCIAT_05\",\n    \"PCIAT-PCIAT_06\", \"PCIAT-PCIAT_07\", \"PCIAT-PCIAT_08\", \"PCIAT-PCIAT_09\", \"PCIAT-PCIAT_10\", \"PCIAT-PCIAT_11\",\n    \"PCIAT-PCIAT_12\", \"PCIAT-PCIAT_13\", \"PCIAT-PCIAT_14\", \"PCIAT-PCIAT_15\", \"PCIAT-PCIAT_16\", \"PCIAT-PCIAT_17\",\n    \"PCIAT-PCIAT_18\", \"PCIAT-PCIAT_19\", \"PCIAT-PCIAT_20\", \"PCIAT-PCIAT_Total\", \"sii\"]\n\nFEATURE_COLS = [\"Basic_Demos-Enroll_Season\", \"Basic_Demos-Age\", \"Basic_Demos-Sex\", \"CGAS-Season\", \"CGAS-CGAS_Score\",\n    \"Physical-Season\", \"Physical-BMI\", \"Physical-Height\", \"Physical-Weight\", \"Physical-Waist_Circumference\",\n    \"Physical-Diastolic_BP\", \"Physical-HeartRate\", \"Physical-Systolic_BP\", \"Fitness_Endurance-Season\",\n    \"Fitness_Endurance-Max_Stage\", \"Fitness_Endurance-Time_Mins\", \"Fitness_Endurance-Time_Sec\", \"FGC-Season\",\n    \"FGC-FGC_CU\", \"FGC-FGC_CU_Zone\", \"FGC-FGC_GSND\", \"FGC-FGC_GSND_Zone\", \"FGC-FGC_GSD\", \"FGC-FGC_GSD_Zone\",\n    \"FGC-FGC_PU\", \"FGC-FGC_PU_Zone\", \"FGC-FGC_SRL\", \"FGC-FGC_SRL_Zone\", \"FGC-FGC_SRR\", \"FGC-FGC_SRR_Zone\", \"FGC-FGC_TL\",\n    \"FGC-FGC_TL_Zone\", \"BIA-Season\", \"BIA-BIA_Activity_Level_num\", \"BIA-BIA_BMC\", \"BIA-BIA_BMI\", \"BIA-BIA_BMR\",\n    \"BIA-BIA_DEE\", \"BIA-BIA_ECW\", \"BIA-BIA_FFM\", \"BIA-BIA_FFMI\", \"BIA-BIA_FMI\", \"BIA-BIA_Fat\", \"BIA-BIA_Frame_num\",\n    \"BIA-BIA_ICW\", \"BIA-BIA_LDM\", \"BIA-BIA_LST\", \"BIA-BIA_SMM\", \"BIA-BIA_TBW\", \"PAQ_A-Season\", \"PAQ_A-PAQ_A_Total\",\n    \"PAQ_C-Season\", \"PAQ_C-PAQ_C_Total\", \"SDS-Season\", \"SDS-SDS_Total_Raw\", \"SDS-SDS_Total_T\", \"PreInt_EduHx-Season\",\n    \"PreInt_EduHx-computerinternet_hoursday\"]","metadata":{"execution":{"iopub.status.busy":"2024-09-20T19:46:06.955315Z","iopub.execute_input":"2024-09-20T19:46:06.955921Z","iopub.status.idle":"2024-09-20T19:46:06.970014Z","shell.execute_reply.started":"2024-09-20T19:46:06.955860Z","shell.execute_reply":"2024-09-20T19:46:06.968497Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Load Data\n\nWe load the training and test datasets using polars for efficient data manipulation.","metadata":{}},{"cell_type":"code","source":"train = pl.read_csv(DATA_DIR / \"train.csv\")\ntest = pl.read_csv(DATA_DIR / \"test.csv\")","metadata":{"execution":{"iopub.status.busy":"2024-09-20T19:46:06.971582Z","iopub.execute_input":"2024-09-20T19:46:06.972125Z","iopub.status.idle":"2024-09-20T19:46:07.012444Z","shell.execute_reply.started":"2024-09-20T19:46:06.972069Z","shell.execute_reply":"2024-09-20T19:46:07.010219Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Data Preprocessing\n\n* Concatenation: We concatenate the train and test datasets to ensure consistent preprocessing, especially for categorical variables.\n* Categorical Conversion: We convert string columns to categorical types and fill missing values with \"NAN\".\n* Splitting: After preprocessing, we split the data back into train and test sets.\n* Handling Missing Targets: We drop rows in the training set where any of the target columns have missing values.","metadata":{}},{"cell_type":"code","source":"train = pl.read_csv(DATA_DIR / \"train.csv\")\ntest = pl.read_csv(DATA_DIR / \"test.csv\")\n\ntrain_test = pl.concat([train, test], how=\"diagonal\")\n\nIS_TEST = test.height <= 100\n\ntrain_test = train_test.with_columns(cs.string().cast(pl.Categorical).fill_null(\"NAN\"))\n\ntrain = train_test[:train.height]\ntest = train_test[train.height:]","metadata":{"execution":{"iopub.status.busy":"2024-09-20T19:46:07.017654Z","iopub.execute_input":"2024-09-20T19:46:07.018490Z","iopub.status.idle":"2024-09-20T19:46:07.071582Z","shell.execute_reply.started":"2024-09-20T19:46:07.018428Z","shell.execute_reply":"2024-09-20T19:46:07.065221Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Feature and Target Selection\n\nWe select the feature columns for X and the target columns for y. We also select features from the test set. Since some libraries we use later prefer pandas DataFrames, we convert the polars DataFrames to pandas DataFrames.","metadata":{}},{"cell_type":"code","source":"train_without_null = train.drop_nulls(subset=TARGET_COLS)\n\nX = train_without_null.select(FEATURE_COLS)\nX_test = test.select(FEATURE_COLS)\ny = train_without_null.select(TARGET_COLS)\n\nX = X.to_pandas()\nX_test = X_test.to_pandas()\ny = y.to_pandas()","metadata":{"execution":{"iopub.status.busy":"2024-09-20T19:46:07.073346Z","iopub.execute_input":"2024-09-20T19:46:07.073893Z","iopub.status.idle":"2024-09-20T19:46:07.099129Z","shell.execute_reply.started":"2024-09-20T19:46:07.073826Z","shell.execute_reply":"2024-09-20T19:46:07.097950Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Handling Missing Values\n\n* Numerical Features: We fill missing values with the mean of each column.\n* Categorical Features: We ensure that \"NAN\" is a category in all categorical features and fill missing values with \"NAN\".","metadata":{}},{"cell_type":"code","source":"num_features = X.select_dtypes(include=[np.number]).columns.tolist()\nX[num_features] = X[num_features].fillna(X[num_features].mean())\nX_test[num_features] = X_test[num_features].fillna(X[num_features].mean())\n\ncat_features = [col for col in FEATURE_COLS if col not in num_features]\nfor col in cat_features:\n    X[col] = X[col].astype('category')\n    X_test[col] = X_test[col].astype('category')\n\n    if \"NAN\" not in X[col].cat.categories:\n        X[col] = X[col].cat.add_categories(\"NAN\")\n    if \"NAN\" not in X_test[col].cat.categories:\n        X_test[col] = X_test[col].cat.add_categories(\"NAN\")\n\n    X[col] = X[col].fillna(\"NAN\")\n    X_test[col] = X_test[col].fillna(\"NAN\")","metadata":{"execution":{"iopub.status.busy":"2024-09-20T19:46:07.100750Z","iopub.execute_input":"2024-09-20T19:46:07.101254Z","iopub.status.idle":"2024-09-20T19:46:07.182886Z","shell.execute_reply.started":"2024-09-20T19:46:07.101189Z","shell.execute_reply":"2024-09-20T19:46:07.181802Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Encoding Categorical Variables\n\nWe use LabelEncoder to convert categorical variables into numerical values. We fit the encoder on the concatenated data to ensure consistent encoding between train and test sets.","metadata":{}},{"cell_type":"code","source":"label_encoders = {}\nfor col in cat_features:\n    le = LabelEncoder()\n    le.fit(pd.concat([X[col], X_test[col]], axis=0).astype(str))\n    X[col] = le.transform(X[col].astype(str))\n    X_test[col] = le.transform(X_test[col].astype(str))\n    label_encoders[col] = le","metadata":{"execution":{"iopub.status.busy":"2024-09-20T19:46:07.184606Z","iopub.execute_input":"2024-09-20T19:46:07.185087Z","iopub.status.idle":"2024-09-20T19:46:07.233414Z","shell.execute_reply.started":"2024-09-20T19:46:07.185043Z","shell.execute_reply":"2024-09-20T19:46:07.232251Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Feature Scaling and Conversion\n\n* Scaling: We use StandardScaler to standardize numerical features.\n* Conversion to NumPy Arrays: We convert the DataFrames to NumPy arrays for compatibility with PyTorch.\n* Device Configuration: We set the device to GPU if available; otherwise, we use the CPU.","metadata":{}},{"cell_type":"code","source":"scaler = StandardScaler()\nscaler.fit(pd.concat([X[num_features], X_test[num_features]], axis=0))\nX[num_features] = scaler.transform(X[num_features])\nX_test[num_features] = scaler.transform(X_test[num_features])\n\nX_np = X.values\nX_test_np = X_test.values\ny_np = y.values","metadata":{"execution":{"iopub.status.busy":"2024-09-20T19:46:07.234964Z","iopub.execute_input":"2024-09-20T19:46:07.235351Z","iopub.status.idle":"2024-09-20T19:46:07.278989Z","shell.execute_reply.started":"2024-09-20T19:46:07.235309Z","shell.execute_reply":"2024-09-20T19:46:07.277737Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Optimized Rounder Class\n\nWe define the OptimizedRounder class, which uses Optuna to find the optimal thresholds for rounding continuous predictions to discrete class labels, maximizing the Quadratic Weighted Kappa (QWK) score.","metadata":{}},{"cell_type":"code","source":"class OptimizedRounder:\n    \"\"\"\n    Optimizes rounding of continuous predictions into discrete class labels by maximizing\n    the Quadratic Weighted Kappa (QWK) score using Optuna optimization.\n    \"\"\"\n\n    def __init__(self, n_classes: int, n_trials: int = 100):\n        \"\"\"\n        Initializes the OptimizedRounder.\n\n        Args:\n            n_classes (int): The number of discrete class labels to predict.\n            n_trials (int): The number of optimization trials to run. Defaults to 100.\n        \"\"\"\n        self.n_classes = n_classes\n        self.labels = np.arange(n_classes)\n        self.n_trials = n_trials\n        self.metric = partial(cohen_kappa_score, weights=\"quadratic\")\n\n    def fit(self, y_pred: NDArray[np.float_], y_true: NDArray[np.int_]) -> None:\n        \"\"\"\n        Optimizes thresholds that round continuous predictions to the nearest class.\n\n        Args:\n            y_pred (NDArray[np.float_]): Continuous predictions from the model.\n            y_true (NDArray[np.int_]): True target labels for comparison.\n        \"\"\"\n        y_pred = self._normalize(y_pred)\n\n        def objective(trial: optuna.Trial) -> float:\n            thresholds = []\n            for i in range(self.n_classes - 1):\n                low = thresholds[-1] if i > 0 else y_pred.min()\n                high = y_pred.max()\n                th = trial.suggest_float(f\"threshold_{i}\", low, high)\n                thresholds.append(th)\n            try:\n                y_pred_rounded = np.digitize(y_pred, thresholds)\n            except ValueError:\n                return -100\n            return self.metric(y_true, y_pred_rounded)\n\n        optuna.logging.disable_default_handler()\n        study = optuna.create_study(direction=\"maximize\")\n        study.optimize(objective, n_trials=self.n_trials)\n\n        self.thresholds = [study.best_params[f\"threshold_{i}\"] for i in range(self.n_classes - 1)]\n\n    def predict(self, y_pred: NDArray[np.float_]) -> NDArray[np.int_]:\n        \"\"\"\n        Rounds continuous predictions using the optimized thresholds.\n\n        Args:\n            y_pred (NDArray[np.float_]): Continuous predictions.\n\n        Returns:\n            NDArray[np.int_]: Rounded class labels.\n        \"\"\"\n        assert hasattr(self, \"thresholds\"), \"You must call fit() before predict()\"\n        y_pred = self._normalize(y_pred)\n        return np.digitize(y_pred, self.thresholds)\n\n    def _normalize(self, y: NDArray[np.float_]) -> NDArray[np.float_]:\n        \"\"\"\n        Normalizes predictions to a range [0, n_classes - 1].\n\n        Args:\n            y (NDArray[np.float_]): Predictions to normalize.\n\n        Returns:\n            NDArray[np.float_]: Normalized predictions.\n        \"\"\"\n        return (y - y.min()) / (y.max() - y.min()) * (self.n_classes - 1)","metadata":{"execution":{"iopub.status.busy":"2024-09-20T19:46:07.281136Z","iopub.execute_input":"2024-09-20T19:46:07.281555Z","iopub.status.idle":"2024-09-20T19:46:07.298081Z","shell.execute_reply.started":"2024-09-20T19:46:07.281512Z","shell.execute_reply":"2024-09-20T19:46:07.296523Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# TabularDataset Class\n\nWe define the TabularDataset class to handle tabular data in PyTorch. It separates categorical and numerical features and can optionally include targets for supervised learning.","metadata":{}},{"cell_type":"code","source":"class TabularDataset(torch.utils.data.Dataset):\n    \"\"\"\n    Dataset for handling tabular data, separating categorical and numerical features.\n    \"\"\"\n\n    def __init__(self, X_categorical, X_numerical, y=None):\n        \"\"\"\n        Initializes the dataset.\n\n        Args:\n            X_categorical: Categorical input features.\n            X_numerical: Numerical input features.\n            y: Target values. Defaults to None for test data.\n        \"\"\"\n        self.X_categorical = torch.tensor(X_categorical, dtype=torch.long)\n        self.X_numerical = torch.tensor(X_numerical, dtype=torch.float32)\n        self.y = torch.tensor(y, dtype=torch.float32) if y is not None else None\n\n    def __len__(self):\n        \"\"\"Returns the length of the dataset.\"\"\"\n        return len(self.X_categorical)\n\n    def __getitem__(self, idx):\n        \"\"\"\n        Retrieves the feature values and optionally the target for a given index.\n\n        Args:\n            idx: The index of the sample to retrieve.\n\n        Returns:\n            Tuple of categorical and numerical features, and the target if available.\n        \"\"\"\n        if self.y is not None:\n            return self.X_categorical[idx], self.X_numerical[idx], self.y[idx]\n        else:\n            return self.X_categorical[idx], self.X_numerical[idx]","metadata":{"execution":{"iopub.status.busy":"2024-09-20T19:46:07.299886Z","iopub.execute_input":"2024-09-20T19:46:07.300330Z","iopub.status.idle":"2024-09-20T19:46:07.316935Z","shell.execute_reply.started":"2024-09-20T19:46:07.300268Z","shell.execute_reply":"2024-09-20T19:46:07.315521Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Neural Network Model\n\nWe define the MultiTargetRegressionModel class, which is a neural network model for multi-target regression. The model uses embeddings for categorical features and fully connected layers for regression.","metadata":{}},{"cell_type":"code","source":"class MultiTargetRegressionModel(nn.Module):\n    \"\"\"\n    Neural network model for multi-target regression using embeddings for categorical features\n    and fully connected layers for regression.\n    \"\"\"\n\n    def __init__(self, embedding_sizes, num_numerical_features, output_size):\n        \"\"\"\n        Initializes the model architecture.\n\n        Args:\n            embedding_sizes: List of tuples containing the number of categories and embedding sizes for each categorical feature.\n            num_numerical_features: The number of numerical input features.\n            output_size: The number of target variables to predict.\n        \"\"\"\n        super(MultiTargetRegressionModel, self).__init__()\n        self.embeddings = nn.ModuleList(\n            [nn.Embedding(num_embeddings=categories, embedding_dim=size) for categories, size in embedding_sizes])\n        self.embedding_output_size = sum([size for _, size in embedding_sizes])\n        self.fc1 = nn.Linear(self.embedding_output_size + num_numerical_features, 128)\n        self.fc2 = nn.Linear(128, 64)\n        self.output = nn.Linear(64, output_size)\n        self.dropout = nn.Dropout(0.1)\n        self.relu = nn.ReLU()\n\n    def forward(self, x_categorical, x_numerical):\n        \"\"\"\n        Forward pass of the model.\n\n        Args:\n            x_categorical: Categorical input features.\n            x_numerical: Numerical input features.\n\n        Returns:\n            Predicted values for all targets.\n        \"\"\"\n        x = [emb(x_categorical[:, i]) for i, emb in enumerate(self.embeddings)]\n        x = torch.cat(x, 1)\n        x = torch.cat([x, x_numerical], 1)\n        x = self.dropout(self.relu(self.fc1(x)))\n        x = self.dropout(self.relu(self.fc2(x)))\n        x = self.output(x)\n        return x","metadata":{"execution":{"iopub.status.busy":"2024-09-20T19:46:07.318641Z","iopub.execute_input":"2024-09-20T19:46:07.319661Z","iopub.status.idle":"2024-09-20T19:46:07.332363Z","shell.execute_reply.started":"2024-09-20T19:46:07.319603Z","shell.execute_reply":"2024-09-20T19:46:07.330996Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Preparing Embeddings and Model Parameters\n\n* Embedding Sizes: For each categorical feature, we determine the number of categories and define an embedding size.\n* Model Parameters: We set the number of numerical features and the output size (number of target variables).","metadata":{}},{"cell_type":"code","source":"embedding_sizes = []\nfor col in cat_features:\n    num_categories = len(label_encoders[col].classes_)\n    embedding_size = min(50, (num_categories + 1) // 2)\n    embedding_sizes.append((num_categories, embedding_size))\n\nnum_numerical_features = len(num_features)\noutput_size = len(TARGET_COLS)","metadata":{"execution":{"iopub.status.busy":"2024-09-20T19:46:07.333971Z","iopub.execute_input":"2024-09-20T19:46:07.334385Z","iopub.status.idle":"2024-09-20T19:46:07.348287Z","shell.execute_reply.started":"2024-09-20T19:46:07.334327Z","shell.execute_reply":"2024-09-20T19:46:07.346687Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Cross-Validation Setup\n\n* Labels for Stratification: We extract the 'sii' column to use as labels for stratification.\n* Stratified K-Fold: We set up a 5-fold stratified cross-validation to maintain the distribution of classes.\n* Initialization: We initialize lists to store the trained models and predictions.","metadata":{}},{"cell_type":"code","source":"y_sii = y[\"sii\"].to_numpy()\n\nskf = StratifiedKFold(n_splits=5, shuffle=True, random_state=52)\nmodels = []\ny_pred = np.full((X.shape[0], len(TARGET_COLS)), fill_value=np.nan)","metadata":{"execution":{"iopub.status.busy":"2024-09-20T19:46:07.352589Z","iopub.execute_input":"2024-09-20T19:46:07.353295Z","iopub.status.idle":"2024-09-20T19:46:07.366973Z","shell.execute_reply.started":"2024-09-20T19:46:07.353249Z","shell.execute_reply":"2024-09-20T19:46:07.365552Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Model Training with Cross-Validation\n\nWe perform model training using stratified K-Fold cross-validation:\n\n* Data Splitting: For each fold, we split the data into training and validation sets.\n* Dataset and DataLoader: We create PyTorch datasets and loaders for training and validation data.\n* Model Initialization: We initialize the neural network model and move it to the appropriate device.\n* Training Loop: We train the model for a specified number of epochs, using early stopping based on validation loss.\n* Model Saving: We save the best model state for each fold.\n* Prediction Storage: We store the validation predictions for later evaluation.","metadata":{}},{"cell_type":"code","source":"for fold, (train_idx, val_idx) in enumerate(skf.split(X, y_sii)):\n    print(f\"Fold {fold + 1}\")\n    X_train, X_val = X.iloc[train_idx], X.iloc[val_idx]\n    y_train, y_val = y.iloc[train_idx], y.iloc[val_idx]\n\n    X_train_categorical = X_train[cat_features].values.astype(np.int64)\n    X_val_categorical = X_val[cat_features].values.astype(np.int64)\n    X_train_numerical = X_train[num_features].values.astype(np.float32)\n    X_val_numerical = X_val[num_features].values.astype(np.float32)\n\n    y_train_values = y_train.values.astype(np.float32)\n    y_val_values = y_val.values.astype(np.float32)\n\n    train_dataset = TabularDataset(X_train_categorical, X_train_numerical, y_train_values)\n    val_dataset = TabularDataset(X_val_categorical, X_val_numerical, y_val_values)\n\n    train_loader = torch.utils.data.DataLoader(train_dataset, batch_size=256, shuffle=True)\n    val_loader = torch.utils.data.DataLoader(val_dataset, batch_size=256, shuffle=False)\n\n    model = MultiTargetRegressionModel(embedding_sizes, num_numerical_features, output_size)\n    model = model.to(device)\n\n    criterion = nn.MSELoss()\n    optimizer = torch.optim.Adam(model.parameters(), lr=0.001)\n\n    num_epochs = 1 if IS_TEST else 100\n    best_val_loss = np.inf\n    patience = 10\n    counter = 0\n    best_model_state = None\n\n    for epoch in range(num_epochs):\n        model.train()\n        train_losses = []\n        for X_cat_batch, X_num_batch, y_batch in train_loader:\n            X_cat_batch = X_cat_batch.to(device)\n            X_num_batch = X_num_batch.to(device)\n            y_batch = y_batch.to(device)\n\n            optimizer.zero_grad()\n            outputs = model(X_cat_batch, X_num_batch)\n            loss = criterion(outputs, y_batch)\n            loss.backward()\n            optimizer.step()\n            train_losses.append(loss.item())\n\n        model.eval()\n        val_losses = []\n        with torch.no_grad():\n            for X_cat_batch, X_num_batch, y_batch in val_loader:\n                X_cat_batch = X_cat_batch.to(device)\n                X_num_batch = X_num_batch.to(device)\n                y_batch = y_batch.to(device)\n\n                outputs = model(X_cat_batch, X_num_batch)\n                loss = criterion(outputs, y_batch)\n                val_losses.append(loss.item())\n\n        train_loss = np.mean(train_losses)\n        val_loss = np.mean(val_losses)\n\n        print(f\"Epoch {epoch + 1}/{num_epochs}, Train Loss: {train_loss:.4f}, Val Loss: {val_loss:.4f}\")\n\n        if val_loss < best_val_loss:\n            best_val_loss = val_loss\n            best_model_state = model.state_dict()\n            counter = 0\n        else:\n            counter += 1\n            if counter >= patience:\n                print(\"Early stopping\")\n                break\n\n    if best_model_state is None:\n        print(\"No valid model was trained in this fold due to NaNs.\")\n        continue\n\n    model.load_state_dict(best_model_state)\n    models.append(model)\n\n    model.eval()\n    val_preds = []\n    with torch.no_grad():\n        for X_cat_batch, X_num_batch, y_batch in val_loader:\n            X_cat_batch = X_cat_batch.to(device)\n            X_num_batch = X_num_batch.to(device)\n            outputs = model(X_cat_batch, X_num_batch)\n            val_preds.append(outputs.cpu().numpy())\n    val_preds = np.vstack(val_preds)\n\n    y_pred[val_idx] = val_preds","metadata":{"execution":{"iopub.status.busy":"2024-09-20T19:46:07.368577Z","iopub.execute_input":"2024-09-20T19:46:07.369021Z","iopub.status.idle":"2024-09-20T19:46:07.942160Z","shell.execute_reply.started":"2024-09-20T19:46:07.368979Z","shell.execute_reply":"2024-09-20T19:46:07.940924Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Evaluation\n\n* Prediction Validation: We ensure there are no missing values in the predictions.\n* Threshold Optimization: We use the OptimizedRounder to find the best thresholds for rounding.\n* Rounding Predictions: We round the continuous predictions to discrete classes.\n* Calculating QWK Score: We calculate the Quadratic Weighted Kappa score to evaluate model performance.","metadata":{}},{"cell_type":"code","source":"assert np.isnan(y_pred).sum() == 0\n\noptimizer_rounder = OptimizedRounder(n_classes=4, n_trials=300)\ny_pred_total = y_pred[:, TARGET_COLS.index(\"PCIAT-PCIAT_Total\")]\noptimizer_rounder.fit(y_pred_total, y_sii)\n\ny_pred_rounded = optimizer_rounder.predict(y_pred_total)\n\nqwk = cohen_kappa_score(y_sii, y_pred_rounded, weights=\"quadratic\")\nprint(f\"Cross-Validated QWK Score: {qwk}\")","metadata":{"execution":{"iopub.status.busy":"2024-09-20T19:46:07.943581Z","iopub.execute_input":"2024-09-20T19:46:07.944008Z","iopub.status.idle":"2024-09-20T19:46:16.231584Z","shell.execute_reply.started":"2024-09-20T19:46:07.943966Z","shell.execute_reply":"2024-09-20T19:46:16.230477Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Making Predictions on Test Data\n\n* Data Preparation: We prepare the test data in the same way as the training data.\n* Predictions: We make predictions using each of the trained models and average them.\n* Rounding: We round the averaged predictions using the optimized thresholds.\n* Submission File: We create a submission file with the required format.","metadata":{}},{"cell_type":"code","source":"X_test_categorical = X_test[cat_features].values.astype(np.int64)\nX_test_numerical = X_test[num_features].values.astype(np.float32)\n\nassert not np.isnan(X_test_numerical).any(), \"NaNs found in X_test_numerical\"\nassert not np.isnan(X_test_categorical).any(), \"NaNs found in X_test_categorical\"\n\ntest_dataset = TabularDataset(X_test_categorical, X_test_numerical)\ntest_loader = torch.utils.data.DataLoader(test_dataset, batch_size=256, shuffle=False)\n\ntest_preds_list = []\nfor model in models:\n    model.eval()\n    test_preds = []\n    with torch.no_grad():\n        for X_cat_batch, X_num_batch in test_loader:\n            X_cat_batch = X_cat_batch.to(device)\n            X_num_batch = X_num_batch.to(device)\n            outputs = model(X_cat_batch, X_num_batch)\n            test_preds.append(outputs.cpu().numpy())\n    test_preds = np.vstack(test_preds)\n    test_preds_list.append(test_preds)\n\ntest_preds_mean = np.mean(test_preds_list, axis=0)\n\ntest_pred_total = test_preds_mean[:, TARGET_COLS.index(\"PCIAT-PCIAT_Total\")]\n\ntest_pred_rounded = optimizer_rounder.predict(test_pred_total)\n\nsubmission = test.select(\"id\").to_pandas()\nsubmission[\"sii\"] = test_pred_rounded\nsubmission.to_csv(\"submission.csv\", index=False)\n\nprint(\"Submission file successfully generated\")","metadata":{"execution":{"iopub.status.busy":"2024-09-20T19:46:38.885765Z","iopub.execute_input":"2024-09-20T19:46:38.886209Z","iopub.status.idle":"2024-09-20T19:46:38.922322Z","shell.execute_reply.started":"2024-09-20T19:46:38.886167Z","shell.execute_reply":"2024-09-20T19:46:38.920756Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"","metadata":{},"execution_count":null,"outputs":[]}]}