{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"name":"python","version":"3.10.14","mimetype":"text/x-python","codemirror_mode":{"name":"ipython","version":3},"pygments_lexer":"ipython3","nbconvert_exporter":"python","file_extension":".py"},"kaggle":{"accelerator":"nvidiaTeslaT4","dataSources":[{"sourceId":84493,"databundleVersionId":9871156,"sourceType":"competition"},{"sourceId":10207949,"sourceType":"datasetVersion","datasetId":6308700},{"sourceId":10209302,"sourceType":"datasetVersion","datasetId":6309744}],"dockerImageVersionId":30805,"isInternetEnabled":false,"language":"python","sourceType":"notebook","isGpuEnabled":true}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"# Review\n\nIn the first note: [https://www.kaggle.com/code/nicolesy/data-pre-processing-memory-and-missing-value](http://)\n>We discuss systematic and random (non-systematic) missing values, where non-systematic missing values include complete and partial missing under a certain date.\n\n><font color=blue>Systematic missing values</font>：Data is generated after a certain time of day, and previous times are shown as missing values.\n>They are ***feature_15,17,32,33,39,41,42,44,50,52,53,55,58,73,74***.\n>\n><font color=blue>Random missing values</font>：\n\n>* Completely missing\n>\n>|**Parquet**|**Number of Missing Days**|***feature***|\n>|---|---|---|\n>|partition_id=1|all|***21,26,27,31***|\n>|partition_id=2|all|***21,26,27,31***|\n>|partition_id=3|first_17|***21,26,27,31***|\n>|partition_id=7|random: 2 date|***08***|\n>|partition_id=8|random: 5 date|***08***|\n>|partition_id=9|random: 1 date|***08***|\n>\n>* Partial missing ---- The remaining data...\n\nIn the second note: [https://www.kaggle.com/code/nicolesy/data-pre-processing-cv-performance](http://)\n>In order to reduce the use of memory in the process of cross-validation, we re-split the data into independent parquet files according to 5 folds, then completed the local training and model saving of LGBM model. Here are three experiments:\n>1. compress memory float 16 -> 32\n>1. compress memory + fill systematic missing with \"1\"\n>1. compress memory + fill systematic missing with \"1\" + delete feature_08 in 8 completely missing date + filling feature_21,26,27,31 in completely missing date with \"0\" + ffill & bfill partially missing\n\nIn the third note: [https://www.kaggle.com/code/nicolesy/data-pre-processing-lb-performance](http://)\n>We submitted the model trained in the second note and obtained LB scores to verify the ideas in the first ntoe. The experiment shows the followingthe table.\n>Based on the approximate location of the data we can infer:\n>* **“delete feature_08 in 8 completely missing date”** mainly impact on **fold_3** and **fold_4**\n>* **\"filling feature_21,26,27,31 in completely missing date with \"0\"\"** mainly impact on **fold_0**\n>|Data Pre-Processing| | | |CV| | |LB|\n>|---|---|---|---|---|---|---|---|\n>|***methods***|***fold_0***|***fold_1***|***fold_2***|***fold_3***|***fold_4***|***mean***|***fold_2***|\n>|Experiment 1|0.01478|<font color=blue>0.00801</font>|0.00865|<font color=blue>0.00844</font>|<font color=blue>0.00588</font>|0.00915|0.0043|\n>|Experiment 2|0.01453|0.00725|0.00906|0.00768|0.00562|0.00883|0.0041|\n>|Experiment 3|<font color=blue>0.01503</font>|0.00772|<font color=blue>0.01008</font>|0.00808|0.00504|<font color=blue>0.00919</font>|<font color=blue>0.0045</font>|\n\n\n>The results of Experiment 2 are significantly reduced, so the treatment of systematic missing values may not be good. Compared with Experiment 2, fold_3 and fold_4 increased and decreased respectively in experiment 3, so the effectiveness of feature_08 processing could not be judged. The score of fold_0 is clearly improved, which indicates that the treatment for feature_21,26,27,31 is likely to be efficient. Since the overall score is significantly improved, it shows that filling random partially missing values has great potential.\n\n# Results\nIn this note, our experiments hope to proof the efficient of handling feauture_08 and filling random partially missing values.\n|Methods|Delete feature_08 in 8 completely missing date|Ffill & Bfill partially missing|Groupby(symbol_id) + ffill & bfill partially missing|\n|---|---|---|---|\n|Experiment 1|--------------------------√--------------------------|   |   |\n|Experiment 2|   |--------------√--------------|   |\n|Experiment 3|   |   |----------------------------√----------------------------|\n\nThe results are:\n|Pre-Processing| | |CV| | | |LB|\n|---|---|---|---|---|---|---|---|\n|***methods***|***fold_0***|***fold_1***|***fold_2***|***fold_3***|***fold_4***|***mean***|***fold_2***|\r|compress memory|0.01478|0.00801|0.00865|0.00844|0.00588|0.00915|0.0043|\r|Experiment 1|0.01478|0.00802|0.00865|0.00760|0.00571|0.00895|lb|\r|Experiment 2|0.01450|0.00812|0.00805|0.00770|0.00572|0.00882|lb|\r|Experiment 3|0.01474|0.00805|0.00859|0.00822|0.00563|0.00905|lb|","metadata":{}},{"cell_type":"code","source":"import os\nfrom pathlib import Path\nimport numpy as np\nimport polars as pl\nimport pandas as pd\nimport matplotlib.pyplot as plt\n\nimport lightgbm as lgb\n\nfrom tqdm import tqdm\nimport pyarrow.parquet as pq\nimport shutil\nimport time\n\nimport warnings\nwarnings.filterwarnings(\"ignore\")\nimport kaggle_evaluation.jane_street_inference_server\nimport joblib","metadata":{"_uuid":"8f2839f25d086af736a60e9eeb907d3b93b6e0e5","_cell_guid":"b1076dfc-b9ad-4769-8c92-a6c4dae69d19","trusted":true,"execution":{"iopub.status.busy":"2024-12-15T16:17:58.371418Z","iopub.execute_input":"2024-12-15T16:17:58.371771Z","iopub.status.idle":"2024-12-15T16:18:01.751136Z","shell.execute_reply.started":"2024-12-15T16:17:58.371735Z","shell.execute_reply":"2024-12-15T16:18:01.750094Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"# LGBM baseline","metadata":{}},{"cell_type":"code","source":"TARGET = 'responder_6'\nFEAT_COLS = [f\"feature_{i:02d}\" for i in range(79)]","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-12-15T16:18:01.753153Z","iopub.execute_input":"2024-12-15T16:18:01.753796Z","iopub.status.idle":"2024-12-15T16:18:01.759309Z","shell.execute_reply.started":"2024-12-15T16:18:01.753747Z","shell.execute_reply":"2024-12-15T16:18:01.757967Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"def calculate_r2(y_true, y_pred, weights):\n    \"\"\"Calculate the R2 score, weighted by the provided weights.\"\"\"\n    numerator = np.sum(weights * (y_true - y_pred) ** 2)\n    denominator = np.sum(weights * (y_true ** 2))\n    r2_score = 1 - (numerator / denominator)\n    return r2_score","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-12-15T16:18:01.760657Z","iopub.execute_input":"2024-12-15T16:18:01.760962Z","iopub.status.idle":"2024-12-15T16:18:01.772989Z","shell.execute_reply.started":"2024-12-15T16:18:01.760932Z","shell.execute_reply":"2024-12-15T16:18:01.771993Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"def train_single_fold(fold_idx, data_dir, n_splits, save_path, FEAT_COLS, TARGET, save_model=True):\n    \"\"\"\n    Train a single fold for cross-validation.\n\n    Parameters:\n    - fold_idx: the index of the fold (0 to n_splits-1)\n    - data_dir: the directory containing the .parquet files\n    - n_splits: total number of splits\n    - save_path: path to save the model\n    - FEAT_COLS: list of feature columns\n    - TARGET: target column\n    - save_model: whether to save the model after training\n    \"\"\"\n    # For each fold, load the current fold file\n    valid_file = os.path.join(data_dir, f\"combined_part{fold_idx}.parquet\")\n    print(f\"Fold {fold_idx}: validation file {valid_file}\")\n\n    # Load the current fold's data\n    fold_data = pd.read_parquet(valid_file)\n\n    # Split the data: the `i`-th part is validation, the rest is training\n    valid_data = fold_data.iloc[int(len(fold_data) * (fold_idx / n_splits)): int(len(fold_data) * ((fold_idx + 1) / n_splits))]\n    train_data = pd.concat([fold_data.iloc[:int(len(fold_data) * (fold_idx / n_splits))], \n                            fold_data.iloc[int(len(fold_data) * ((fold_idx + 1) / n_splits)):]],\n                           ignore_index=True)\n    \n\n    # Print date_id range for each of the 5 parts (4 training and 1 validation)\n    print(f\"Fold {fold_idx}:\")\n    for i in range(n_splits):\n        # For each fold, get the date_id range for each part\n        if i == fold_idx:\n            part_data = valid_data\n            part_type = 'Validation'\n        else:\n            part_data = train_data.iloc[int(len(train_data) * (i / n_splits)): int(len(train_data) * ((i + 1) / n_splits))]\n            part_type = f\"Training part {i}\"\n        \n        print(f\"  {part_type} date_id range: ({part_data['date_id'].min()}, {part_data['date_id'].max()})\")\n        \n\n    train_weight = train_data['weight']\n    valid_weight = valid_data['weight']\n\n    # Default LightGBM parameters\n    LGB_PARAMS = {\n        'objective': 'regression_l2',\n        'metric': 'rmse',\n        'learning_rate': 0.05,\n        'num_leaves': 31,\n        'max_depth': -1,\n        'random_state': 42,\n        'device': 'gpu',\n        'verbosity': -1,  # Suppress LightGBM logs\n    }\n\n    # Build the LightGBM dataset\n    train_ds = lgb.Dataset(train_data[FEAT_COLS + ['weight']], label=train_data[TARGET], weight=train_weight)\n    valid_ds = lgb.Dataset(valid_data[FEAT_COLS + ['weight']], label=valid_data[TARGET], weight=valid_weight, reference=train_ds)\n\n    # Callback functions\n    early_stopping_callback = lgb.early_stopping(100)\n    verbose_eval_callback = lgb.log_evaluation(period=50)\n\n    # Train the model\n    model = lgb.train(\n        LGB_PARAMS,\n        train_ds,\n        num_boost_round=1000,\n        valid_sets=[train_ds, valid_ds],  \n        valid_names=['train', 'valid'],\n        callbacks=[early_stopping_callback, verbose_eval_callback],\n    )\n\n    # Save the model for this fold\n    if save_model:\n        model_file = os.path.join(save_path, f\"lgb_model_fold_{fold_idx}.pkl\")\n        joblib.dump(model, model_file)\n        print(f\"Saved model for fold {fold_idx} to {model_file}\")\n    \n    # Evaluate the model\n    y_valid_pred = model.predict(valid_data[FEAT_COLS + ['weight']])\n    r2_score = calculate_r2(valid_data[TARGET], y_valid_pred, valid_weight)\n    print(f\"Fold {fold_idx} validation R2 score: {r2_score}\")\n\n    return model, r2_score","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-12-15T16:23:59.054488Z","iopub.execute_input":"2024-12-15T16:23:59.054854Z","iopub.status.idle":"2024-12-15T16:23:59.069102Z","shell.execute_reply.started":"2024-12-15T16:23:59.054820Z","shell.execute_reply":"2024-12-15T16:23:59.068059Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"base_dir = '/kaggle/working'\nsubdirs = ['del_f8', 'fbfill', 'gp_syb_fbfill']\n\n# Create the directories\nfor subdir in subdirs:\n    dir_path = os.path.join(base_dir, subdir)\n    os.makedirs(dir_path, exist_ok=True)\n\nprint(\"Directories created successfully!\")","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-12-15T16:20:54.169231Z","iopub.execute_input":"2024-12-15T16:20:54.169653Z","iopub.status.idle":"2024-12-15T16:20:54.176202Z","shell.execute_reply.started":"2024-12-15T16:20:54.169620Z","shell.execute_reply":"2024-12-15T16:20:54.175118Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"def load_model(file_path):\n    model = joblib.load(file_path)\n    print(f\"Loaded the model group from {file_path}\")\n    return model","metadata":{"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"# Experiment 1","metadata":{}},{"cell_type":"code","source":"# %%time\n\n# cv_scores = []\n\n# for fold_idx in range(5):\n#     model, r2_score = train_single_fold(fold_idx, data_dir=\"/kaggle/input/js-low-memory-rpmv/dele_f8_8dates\", n_splits=5, \n#                                     save_path=\"/kaggle/working/del_f8\", \n#                                     FEAT_COLS=FEAT_COLS, TARGET=TARGET, save_model=True)\n#     cv_scores.append(r2_score)\n\n# print(f\"Mean R2 score: {np.mean(cv_scores)}, Std: {np.std(cv_scores)}\")","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-12-15T16:24:01.109953Z","iopub.execute_input":"2024-12-15T16:24:01.110700Z","iopub.status.idle":"2024-12-15T16:31:31.453845Z","shell.execute_reply.started":"2024-12-15T16:24:01.110661Z","shell.execute_reply":"2024-12-15T16:31:31.452492Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"# lgb_model = load_model(file_path=\"/kaggle/input/js-rpmv-models/del_f8/lgb_model_fold_2.pkl\")\n# print(f\"Loaded model from the saved file.\")\n# print(f\"Type of lgb_model: {type(lgb_model)}\")","metadata":{"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"# Experiment 2","metadata":{}},{"cell_type":"code","source":"# %%time\n\n# cv_scores = []\n\n# for fold_idx in range(5):\n#     model, r2_score = train_single_fold(fold_idx, data_dir=\"/kaggle/input/js-low-memory-rpmv/fbfill_random_partial\", n_splits=5, \n#                                     save_path=\"/kaggle/working/fbfill\", \n#                                     FEAT_COLS=FEAT_COLS, TARGET=TARGET, save_model=True)\n#     cv_scores.append(r2_score)\n\n# print(f\"Mean R2 score: {np.mean(cv_scores)}, Std: {np.std(cv_scores)}\")","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-12-15T16:31:35.380771Z","iopub.execute_input":"2024-12-15T16:31:35.381988Z","iopub.status.idle":"2024-12-15T16:37:59.236208Z","shell.execute_reply.started":"2024-12-15T16:31:35.381879Z","shell.execute_reply":"2024-12-15T16:37:59.235091Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"# lgb_model = load_model(file_path=\"/kaggle/input/js-rpmv-models/fbfill/lgb_model_fold_2.pkl\")\n# print(f\"Loaded model from the saved file.\")\n# print(f\"Type of lgb_model: {type(lgb_model)}\")","metadata":{"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"# Experiment 3","metadata":{}},{"cell_type":"code","source":"# %%time\n\n# cv_scores = []\n\n# for fold_idx in range(5):\n#     model, r2_score = train_single_fold(fold_idx, data_dir=\"/kaggle/input/js-low-memory-rpmv/fbfill_gp_symbol_random_partial\", n_splits=5, \n#                                     save_path=\"/kaggle/working/gp_syb_fbfill\", \n#                                     FEAT_COLS=FEAT_COLS, TARGET=TARGET, save_model=True)\n#     cv_scores.append(r2_score)\n\n# print(f\"Mean R2 score: {np.mean(cv_scores)}, Std: {np.std(cv_scores)}\")","metadata":{"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"code","source":"lgb_model = load_model(file_path=\"/kaggle/input/js-rpmv-models/gp_syb_fbfill/lgb_model_fold_2.pkl\")\nprint(f\"Loaded model from the saved file.\")\nprint(f\"Type of lgb_model: {type(lgb_model)}\")","metadata":{"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"# Prediction","metadata":{}},{"cell_type":"code","source":"lags_ : pl.DataFrame | None = None\n\ndef predict(test: pl.DataFrame, lags: pl.DataFrame | None) -> pl.DataFrame | pd.DataFrame:\n\n    global lags_\n    if lags is not None:\n        lags_ = lags\n\n    predictions = test.select(\n        'row_id',\n        pl.lit(0.0).alias('responder_6'),\n    )\n    \n    feat = test[FEAT_COLS+['weight']].to_pandas()\n\n    pred = lgb_model.predict(feat)\n\n    \n    predictions = predictions.with_columns(pl.Series('responder_6', pred.ravel()))\n    print(predictions)\n\n    assert isinstance(predictions, pl.DataFrame | pd.DataFrame)\n\n    assert list(predictions.columns) == ['row_id', 'responder_6']\n\n    assert len(predictions) == len(test)\n\n    return predictions","metadata":{"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"code","source":"inference_server = kaggle_evaluation.jane_street_inference_server.JSInferenceServer(predict)\n\nif os.getenv('KAGGLE_IS_COMPETITION_RERUN'):\n    inference_server.serve()\nelse:\n    inference_server.run_local_gateway(\n        (\n            '/kaggle/input/jane-street-real-time-market-data-forecasting/test.parquet',\n            '/kaggle/input/jane-street-real-time-market-data-forecasting/lags.parquet',\n        )\n    )","metadata":{"trusted":true},"outputs":[],"execution_count":null}]}