{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"name":"python","version":"3.11.13","mimetype":"text/x-python","codemirror_mode":{"name":"ipython","version":3},"pygments_lexer":"ipython3","nbconvert_exporter":"python","file_extension":".py"},"kaggle":{"accelerator":"none","dataSources":[{"sourceId":96164,"databundleVersionId":11418275,"sourceType":"competition"}],"dockerImageVersionId":31089,"isInternetEnabled":true,"language":"python","sourceType":"notebook","isGpuEnabled":false}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"code","source":"# This Python 3 environment comes with many helpful analytics libraries installed\n# It is defined by the kaggle/python Docker image: https://github.com/kaggle/docker-python\n# For example, here's several helpful packages to load\n\nimport pandas as pd\nimport numpy as np\nfrom sklearn.linear_model import Ridge\nfrom sklearn.preprocessing import StandardScaler\nfrom sklearn.model_selection import TimeSeriesSplit\n\n# Input data files are available in the read-only \"../input/\" directory\n# For example, running this (by clicking run or pressing Shift+Enter) will list all files under the input directory\n\nimport os\nfor dirname, _, filenames in os.walk('/kaggle/input'):\n    for filename in filenames:\n        print(os.path.join(dirname, filename))\n\n# You can write up to 20GB to the current directory (/kaggle/working/) that gets preserved as output when you create a version using \"Save & Run All\" \n# You can also write temporary files to /kaggle/temp/, but they won't be saved outside of the current session","metadata":{"_uuid":"8f2839f25d086af736a60e9eeb907d3b93b6e0e5","_cell_guid":"b1076dfc-b9ad-4769-8c92-a6c4dae69d19","trusted":true,"execution":{"iopub.status.busy":"2025-07-09T12:11:09.054322Z","iopub.execute_input":"2025-07-09T12:11:09.054659Z","iopub.status.idle":"2025-07-09T12:11:09.062850Z","shell.execute_reply.started":"2025-07-09T12:11:09.054633Z","shell.execute_reply":"2025-07-09T12:11:09.061993Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"I define several global configuration parameters to control the behavior of the model and data loading process:\n\n- `Data paths`: Set for compatibility with the Kaggle environment\n- `Selected features`: A list of raw features used in modeling\n- `Lag windows (LAGS)`: Used to create lagged versions of key features (1, 5, and 10 steps)\n- `Model parameters`: The Ridge regression regularization coefficient (`ALPHA = 0.1`) and number of folds (`N_FOLDS = 3`)\n","metadata":{}},{"cell_type":"code","source":"class Config:\n    # Data paths (Kaggle environment)\n    TRAIN_PATH = \"/kaggle/input/drw-crypto-market-prediction/train.parquet\"\n    TEST_PATH = \"/kaggle/input/drw-crypto-market-prediction/test.parquet\"\n    SUBMISSION_PATH = \"/kaggle/input/drw-crypto-market-prediction/sample_submission.csv\"\n\n    # Selected features\n    FEATURES = [\n        \"X863\", \"X345\", \"X612\", \"X855\", \"X174\",\n        \"bid_qty\", \"ask_qty\", \"buy_qty\", \"sell_qty\", \"volume\"\n    ]\n\n    # Lag windows\n    LAGS = [1, 5, 10]\n\n    # Model parameters\n    ALPHA = 0.1\n    N_FOLDS = 3","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-07-09T12:11:09.064308Z","iopub.execute_input":"2025-07-09T12:11:09.064552Z","iopub.status.idle":"2025-07-09T12:11:09.080487Z","shell.execute_reply.started":"2025-07-09T12:11:09.064531Z","shell.execute_reply":"2025-07-09T12:11:09.079534Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"To emphasize that more recent samples are more relevant than older ones, I define a function `create_time_weights()` that generates **linearly increasing weights**. This ensures that the most recent samples have the highest influence during training.\n\nThis approach avoids issues such as numerical underflow that can occur with exponential decay methods.","metadata":{}},{"cell_type":"code","source":"def create_time_weights(n):\n    \"\"\"Linearly increasing weights for recent samples\"\"\"\n    weights = np.linspace(0.1, 1.0, n)  # make sure != 0\n    return weights / weights.sum()","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-07-09T12:11:09.081489Z","iopub.execute_input":"2025-07-09T12:11:09.081795Z","iopub.status.idle":"2025-07-09T12:11:09.097205Z","shell.execute_reply.started":"2025-07-09T12:11:09.081774Z","shell.execute_reply":"2025-07-09T12:11:09.096301Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"I enhance the original dataset with the following derived features:\n\n1. `Liquidity Imbalance`\n\n   \n   Measures the difference between bid and ask quantities relative to total volume.\n\n2. `Order Flow`\n\n   \n   Reflects the net buying or selling pressure over time.\n\n3. `Lagged Features`\n\n   \n   I generate lagged versions of key columns at different intervals (1, 5, and 10 steps) to capture temporal dependencies.\n\nMissing values are filled using forward and backward fill methods to ensure completeness.","metadata":{}},{"cell_type":"code","source":"def add_features(df):\n    \"\"\"\n    Add simple derived features and lagged versions of core columns\n    \n    Args:\n        df: Input DataFrame\n        \n    Returns:\n        DataFrame with engineered features\n    \"\"\"\n    # Feature interactions\n    df['liquidity_imbalance'] = (df['bid_qty'] - df['ask_qty']) / (df['bid_qty'] + df['ask_qty'] + 1e-8)\n    df['order_flow'] = (df['buy_qty'] - df['sell_qty']) / (df['volume'] + 1e-8)\n\n    # Lagged features\n    for col in ['bid_qty', 'ask_qty', 'buy_qty', 'sell_qty', 'volume']:\n        for lag in Config.LAGS:\n            df[f\"{col}_lag_{lag}\"] = df[col].shift(lag)\n\n    # Fill missing values efficiently\n    return df.ffill().bfill()","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-07-09T12:11:09.097870Z","iopub.execute_input":"2025-07-09T12:11:09.098155Z","iopub.status.idle":"2025-07-09T12:11:09.114984Z","shell.execute_reply.started":"2025-07-09T12:11:09.098124Z","shell.execute_reply":"2025-07-09T12:11:09.114121Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"I follow this pipeline for data loading and preprocessing:\n\n1. Load training and test data from parquet files.\n2. Apply feature engineering using the `add_features()` function.\n3. Replace infinite values with NaN and perform forward/backward fill.\n4. Standardize features using `StandardScaler`. The same scaler is applied to both training and test sets to avoid information leakage.\n\nThis results in clean, normalized feature matrices and target arrays ready for model training.","metadata":{}},{"cell_type":"code","source":"def load_data():\n    \"\"\"\n    Load and preprocess training and test data\n    \n    Returns:\n        X_train, y_train, X_test arrays\n    \"\"\"\n    train_df = pd.read_parquet(Config.TRAIN_PATH, columns=Config.FEATURES + [\"label\"])\n    test_df = pd.read_parquet(Config.TEST_PATH, columns=Config.FEATURES)\n\n    print(\"Train shape:\", train_df.shape)\n    print(\"Test shape:\", test_df.shape)\n\n    # Feature engineering\n    train_df = add_features(train_df)\n    test_df = add_features(test_df)\n\n    feature_cols = [col for col in train_df.columns if col != \"label\"]\n\n    # Clean inf values\n    train_df[feature_cols] = train_df[feature_cols].replace([np.inf, -np.inf], np.nan).ffill().bfill()\n    test_df[feature_cols] = test_df[feature_cols].replace([np.inf, -np.inf], np.nan).ffill().bfill()\n\n    # Normalize\n    scaler = StandardScaler()\n    X_train = scaler.fit_transform(train_df[feature_cols])\n    X_test = scaler.transform(test_df[feature_cols])\n\n    y_train = train_df[\"label\"].values\n\n    return X_train, y_train, X_test","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-07-09T12:11:09.116791Z","iopub.execute_input":"2025-07-09T12:11:09.117093Z","iopub.status.idle":"2025-07-09T12:11:09.136412Z","shell.execute_reply.started":"2025-07-09T12:11:09.117071Z","shell.execute_reply":"2025-07-09T12:11:09.135536Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"I use `TimeSeriesSplit` for cross-validation to ensure that the training set always precedes the validation set in time, simulating real-world forecasting conditions.\n\nDuring training:\n- I apply sample weights if their sum is greater than zero.\n- If the weights sum to zero for a fold, I skip weighting and train without it.\n\nFinal predictions are obtained by averaging predictions from all folds, improving robustness and generalization.","metadata":{}},{"cell_type":"code","source":"def train_model(X, y, X_test):\n    tscv = TimeSeriesSplit(n_splits=Config.N_FOLDS)\n    weights = create_time_weights(len(X))\n    preds = np.zeros(len(X_test))\n\n    for fold, (train_idx, val_idx) in enumerate(tscv.split(X)):\n        print(f\"Training Fold {fold+1}/{Config.N_FOLDS}\")\n\n        fold_weights = weights[train_idx]\n\n        model = Ridge(alpha=Config.ALPHA)\n        if fold_weights.sum() > 0:\n            model.fit(X[train_idx], y[train_idx], sample_weight=fold_weights)\n        else:\n            print(\"Skipping weights for this fold.\")\n            model.fit(X[train_idx], y[train_idx])\n\n        preds += model.predict(X_test)\n\n    return preds / Config.N_FOLDS","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-07-09T12:11:09.137637Z","iopub.execute_input":"2025-07-09T12:11:09.137965Z","iopub.status.idle":"2025-07-09T12:11:09.154653Z","shell.execute_reply.started":"2025-07-09T12:11:09.137938Z","shell.execute_reply":"2025-07-09T12:11:09.153627Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"I generate the final submission file using the provided `sample_submission.csv` template. Predicted values are written into the `prediction` column.\n\nThe output file is saved as `submission.csv`, ready for submission to the competition platform.","metadata":{}},{"cell_type":"code","source":"def make_submission(preds):\n    \"\"\"\n    Generate submission file\n    \n    Args:\n        preds: Predicted values\n    \"\"\"\n    submission = pd.read_csv(Config.SUBMISSION_PATH)\n    submission[\"prediction\"] = preds\n    submission.to_csv(\"submission.csv\", index=False)\n    print(\"Submission file created\")","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-07-09T12:11:09.258901Z","iopub.execute_input":"2025-07-09T12:11:09.259189Z","iopub.status.idle":"2025-07-09T12:11:09.263807Z","shell.execute_reply.started":"2025-07-09T12:11:09.259170Z","shell.execute_reply":"2025-07-09T12:11:09.262876Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"Here is the complete execution flow of the script:\n\n1. Load and preprocess the training and test data.\n2. Train the Ridge regression model using time-aware cross-validation and optional sample weighting.\n3. Generate predictions for the test set.\n4. Save the predictions in the required submission format.","metadata":{}},{"cell_type":"code","source":"if __name__ == \"__main__\":\n    X_train, y_train, X_test = load_data()\n    predictions = train_model(X_train, y_train, X_test)\n    make_submission(predictions)","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-07-09T12:11:09.265312Z","iopub.execute_input":"2025-07-09T12:11:09.265618Z","iopub.status.idle":"2025-07-09T12:11:15.025959Z","shell.execute_reply.started":"2025-07-09T12:11:09.265587Z","shell.execute_reply":"2025-07-09T12:11:15.025031Z"}},"outputs":[],"execution_count":null}]}