{
  "id": 603767,
  "title": "[Private 58th] TabM, AutoencoderMLP with online training & GBDT offline models",
  "url": "/competitions/jane-street-real-time-market-data-forecasting/discussion/603767",
  "author_name": "I2nfinit3y",
  "post_date": "2025-09-04T08:47:56.290000",
  "votes": 14,
  "comment_count": 0,
  "views": 0,
  "content": "<p>Thanks to this amazing competition and the efforts of every one of my teammates ! <a href=\"https://www.kaggle.com/lechengyan\" target=\"_blank\">@lechengyan</a> <a href=\"https://www.kaggle.com/chronoscop\" target=\"_blank\">@chronoscop</a><br>\nOur final solution of 0.0092 lb is ensembling NN models with online learning , GBDT offline models and a ridge model. My part is mainly for TabM model and some GBDT models, which is what I am gonna to talk about. AutoencoderMLP and online learning are <a href=\"https://www.kaggle.com/lechengyan\" target=\"_blank\">@lechengyan</a> 's part and <a href=\"https://www.kaggle.com/chronoscop\" target=\"_blank\">@chronoscop</a> is in charge of one of XGB.</p>\n<h1>1. Cross-Validation</h1>\n<p>I simply used the last 120 dates as my validation and it shows good correlations with LB.</p>\n<h1>2. TabM Model</h1>\n<p>Actually, in the early time of this competition, I noticed TabM and found it is a tabular NN model with great potential. I also public a baseline <a href=\"https://www.kaggle.com/code/i2nfinit3y/jane-street-tabm-ft-transformer-inference\" target=\"_blank\">notebook</a>.</p>\n<h3>Model Architecture</h3>\n<p>The model parts I mainly adjust are feature embed layers and backbone.<br>\nFor category features, I use onehot encoding for every category tensor, which contains feature_09, feature_10, feature_11, symbol_id and time_id. I remap these category features with ordinal encode and every new category  in the test data will be mapped to the same category. Apart from one-hot encoding, I also use embedding layers but it didn't help much. It's worth noting that taking time_id as category features improve my model al lot.<br>\nBut for continuous features, I have tried LinearEmbeddings, PeriodicEmbeddings and PiecewiseLinearEmbeddings, which didn't work well. So I don't use any embedding layer for continuous features.<br>\nThe backbone of TabM is just 3 layers MLP with the same 512 dimensions. I haven't experimented with too many dimensional combinations, but 512 seems to be the most stable.</p>\n<h3>Loss function</h3>\n<p>I have used huber loss, logcosh loss, mae loss, zero-mean R2 loss and mse loss, and R2 loss and mse loss work best. Specifically, mse loss can get higher score in the cv but R2 loss is more robust in every validation epoch. So I choose R2 loss as loss function.</p>\n<h3>Hyperparameter</h3>\n<p>Dropout : 0.25. Higher dropout rate (0.5) will result in slower convergence and lower (0.1) will be overfitting, so 0.25 seems to be a good choice for me.<br>\nLearning rate: 1e-3.<br>\nWeight decay: 8e-4.<br>\nBatch size : 8192<br>\nepoch : 5 or 6. Generally, I will submit multiple epoch to confirm the best model checkpoint and most of time 5 or 6 is the best.<br>\noptimizer : AdamW<br>\nk : 16. k is the number of ensemble in the finial output. I tried 8, 16, 24, 32, and I found 16 can takes both scores and training time into account.</p>\n<h3>Datapreprocessing</h3>\n<p>Simple mean-std standardize and fill zero for nan values.</p>\n<h3>Feature Engeneering &amp; Auxiliary targets</h3>\n<p>I use time_id, symbol_id, 78 original features (besides feature_61, because it just change with dates) and 9 responder lags 1 of last record. Additionally, I also use sin and cos of time_id  with 967 periods or 483 periods and feature_61 with 20 period to catch some periodical change. Because though feature_61 only changes with the date, it also changes in about 20 dates. Also, I tried to use more lags features but it work worse in cv and lb.<br>\nBesides responder_6, I use responder_3 as my auxiliary target, because it has high correlations with responder_6. I tried to add more auxiliary targets in my training, but there was no obvious improvement in my model.</p>\n<h3>Data Augmentation</h3>\n<p>I add some gaussian noise in continuous features with 0.02 std.</p>\n<h3>Training Sample</h3>\n<p>I use the data with dates after 252, because there are too many Nan values in the first 252 dates. I have tried use the dates after 750 to train, which improve my cv but decrease my lb.</p>\n<h3>Scores</h3>\n<p>without online learning    cv : 0.0106  lb : 0.0077<br>\nwith online learning    cv : 0.0116  lb : 0.0083</p>\n<h1>3. AutoencoderMLP</h1>\n<h3>Datapreprocessing</h3>\n<p>Fill 3 for Nan values without standardization.</p>\n<h3>Feature Engeneering</h3>\n<p>We use time_id, symbol_id, 79 original features and responder_6 lags 1 of last. We didn't use auxilary target in the AutoencoderMLP</p>\n<h3>Model Architecture</h3>\n<p>Encoder and Decoder : Both are a lieanr layer with 96 hidden dimension.<br>\nMLP : 5 layers with hidden dimension 96, 896, 448, 448, 256</p>\n<h3>Hyperparameter</h3>\n<p>Dropout : [0.035, 0.038, 0.424, 0.104, 0.25, 0.32, 0.271, 0.25]<br>\nLearning rate: 1e-3<br>\nBatch size : 8192<br>\nEpoch : about 15<br>\noptimizer : Adam</p>\n<h3>Loss function</h3>\n<p>We use mse loss as reconstruction loss, and weighted-mse loss as prediction loss</p>\n<h3>Scores</h3>\n<p>without online learning    cv : 0.0103  lb : 0.0072<br>\nwith online learning    cv : 0.0110  lb : 0.0078</p>\n<h1>4. My GBDT and Ridge offline model</h1>\n<h3>Feature Engeneering</h3>\n<p>It is more simple now. For LGB, XGB and Ridge, I just use time_id, symbol_id, 79 features, responder lags 1 of last, mean, std and max. And I did not specify category features in the GBDT model.</p>\n<h3>Training Sample</h3>\n<p>The dates after 750 for XGB and Ridge and dates after 678 for LGB </p>\n<h3>Hyperparameter</h3>\n<p>I didn't optimize my hyperparameter just use fixed one.</p>\n<pre><code>LGB_Params = {\n        ': ,\n        ': ,\n        ': ,\n        ': ,\n        ': ,\n        ': ,\n        ': ,\n        ': ,\n        ': ,\n        ' : '',\n        ': ,\n        ': ',\n    }\n\nXGB_Params = {\n    ': ,\n    ': ,\n    ': ,\n    ': ,\n    ': ,\n    ': ,\n    ': ,\n    ': ,\n    ': ',\n    ' : '',\n    ': ,\n    ' : ':squarederror',\n}\n</code></pre>\n<h3>Scores</h3>\n<p>LGB   cv: 0.0096 lb : 0.0072<br>\nXGB   cv: 0.0102  lb : 0.0073<br>\nRidge cv: 0.0035 lb : 0.0044<br>\nAfter ensembling these 3 models, the lb is 0.0076.</p>\n<h1>5. <a href=\"https://www.kaggle.com/chronoscop\" target=\"_blank\">@chronoscop</a>'s XGB</h1>\n<h3>Datapreprocessing</h3>\n<p>Fill 0 for Nan values without standardization</p>\n<h3>Feature Engeneering</h3>\n<p>use symbol_id, time_id, 79 original feature and responder lags1 of last. If add more lags features, it would be easy to overfit.</p>\n<h3>Training Sample</h3>\n<p>use dates after 917 to train the model.</p>\n<h3>Hyperparameter</h3>\n<pre><code> = {\n: ,\n: ,\n: ,\n : ,\n: ,\n: ,\n: ,\n: ,\n: ,\n: \n}\n</code></pre>\n<h3>Scores</h3>\n<p>cv : 0.0102   lb : 0.0072</p>\n<h1>6. Online Learning</h1>\n<p>Online learning play an important role in this competition.<br>\nWe collect data  for each previous date and then update model in every date. For every date, we will combine the data of the last date with partial data from the previous dates (using random sampling) for online training. In the training part, instead of batch processing, we input all the data into the model and then train 5 epochs.</p>\n<pre><code>if previous_test is not  and lags is not :\n    train = previous_test.(\n        lags_.([, , pl.().()]),\n        on=[, ],\n        how=\n    )\n\n    if pre_train is not None and (pre_train) &gt; :\n        pre_train = pre_train.(n=, seed=)\n    if pre_train is None:\n        pre_train = train\n    else:\n        pre_train = pl.([pre_train, train])\n    X = pre_train.()[cols].().values\n    y = pre_train[].()\n    weights = pre_train[].()\n\n\n    model.()\n    optimizer = torch.optim.(model.(), lr=e-)\n    criterion = nn.()\n\n\n    for epoch in ():\n        optimizer.()\n        _, y_pred = (torch.(X).(device))\n        loss = (y_pred.(), torch.(y.()).(device))\n        loss.()\n        optimizer.()\n        (f)\n\nif previous_test is None:\n    previous_test = test\nelse:\n    previous_test = pl.([previous_test, test])\n</code></pre>\n<p>We tried to add online learning into GBDT, but it didn't work well and cost too many time.</p>\n<h1>Thanks !</h1>\n<p>I think TabM can get higher score if adjust different kind of backbones, but I have no time to experiment. I also tried GRU, LSTM or Transformer in the early time, but they all failed, so then I focus on tabular model. Hoping top teams can share more trick and I indeed learn a lot in the competition. Thanks to all of you! :)<br>\nPS: Our code will be released when it is sorted out.</p>\n<hr>\n<p>UPDATE : <br>\nTabM training code <a href=\"https://www.kaggle.com/code/i2nfinit3y/jane-street-tabm-training?scriptVersionId=217873214\" target=\"_blank\">https://www.kaggle.com/code/i2nfinit3y/jane-street-tabm-training?scriptVersionId=217873214</a><br>\nSingle TabM online learning : <a href=\"https://www.kaggle.com/code/i2nfinit3y/jane-street-tabm-online-learning\" target=\"_blank\">https://www.kaggle.com/code/i2nfinit3y/jane-street-tabm-online-learning</a></p>\n<hr>\n<p>UPDATE :<br>\nwe’ve open-sourced both our inference and training code.</p>",
  "messages": [
    {
      "id": 3281421,
      "postDate": "2025-09-04T08:47:56.290Z",
      "content": "<p>Thanks to this amazing competition and the efforts of every one of my teammates ! <a href=\"https://www.kaggle.com/lechengyan\" target=\"_blank\">@lechengyan</a> <a href=\"https://www.kaggle.com/chronoscop\" target=\"_blank\">@chronoscop</a><br>\nOur final solution of 0.0092 lb is ensembling NN models with online learning , GBDT offline models and a ridge model. My part is mainly for TabM model and some GBDT models, which is what I am gonna to talk about. AutoencoderMLP and online learning are <a href=\"https://www.kaggle.com/lechengyan\" target=\"_blank\">@lechengyan</a> 's part and <a href=\"https://www.kaggle.com/chronoscop\" target=\"_blank\">@chronoscop</a> is in charge of one of XGB.</p>\n<h1>1. Cross-Validation</h1>\n<p>I simply used the last 120 dates as my validation and it shows good correlations with LB.</p>\n<h1>2. TabM Model</h1>\n<p>Actually, in the early time of this competition, I noticed TabM and found it is a tabular NN model with great potential. I also public a baseline <a href=\"https://www.kaggle.com/code/i2nfinit3y/jane-street-tabm-ft-transformer-inference\" target=\"_blank\">notebook</a>.</p>\n<h3>Model Architecture</h3>\n<p>The model parts I mainly adjust are feature embed layers and backbone.<br>\nFor category features, I use onehot encoding for every category tensor, which contains feature_09, feature_10, feature_11, symbol_id and time_id. I remap these category features with ordinal encode and every new category  in the test data will be mapped to the same category. Apart from one-hot encoding, I also use embedding layers but it didn't help much. It's worth noting that taking time_id as category features improve my model al lot.<br>\nBut for continuous features, I have tried LinearEmbeddings, PeriodicEmbeddings and PiecewiseLinearEmbeddings, which didn't work well. So I don't use any embedding layer for continuous features.<br>\nThe backbone of TabM is just 3 layers MLP with the same 512 dimensions. I haven't experimented with too many dimensional combinations, but 512 seems to be the most stable.</p>\n<h3>Loss function</h3>\n<p>I have used huber loss, logcosh loss, mae loss, zero-mean R2 loss and mse loss, and R2 loss and mse loss work best. Specifically, mse loss can get higher score in the cv but R2 loss is more robust in every validation epoch. So I choose R2 loss as loss function.</p>\n<h3>Hyperparameter</h3>\n<p>Dropout : 0.25. Higher dropout rate (0.5) will result in slower convergence and lower (0.1) will be overfitting, so 0.25 seems to be a good choice for me.<br>\nLearning rate: 1e-3.<br>\nWeight decay: 8e-4.<br>\nBatch size : 8192<br>\nepoch : 5 or 6. Generally, I will submit multiple epoch to confirm the best model checkpoint and most of time 5 or 6 is the best.<br>\noptimizer : AdamW<br>\nk : 16. k is the number of ensemble in the finial output. I tried 8, 16, 24, 32, and I found 16 can takes both scores and training time into account.</p>\n<h3>Datapreprocessing</h3>\n<p>Simple mean-std standardize and fill zero for nan values.</p>\n<h3>Feature Engeneering &amp; Auxiliary targets</h3>\n<p>I use time_id, symbol_id, 78 original features (besides feature_61, because it just change with dates) and 9 responder lags 1 of last record. Additionally, I also use sin and cos of time_id  with 967 periods or 483 periods and feature_61 with 20 period to catch some periodical change. Because though feature_61 only changes with the date, it also changes in about 20 dates. Also, I tried to use more lags features but it work worse in cv and lb.<br>\nBesides responder_6, I use responder_3 as my auxiliary target, because it has high correlations with responder_6. I tried to add more auxiliary targets in my training, but there was no obvious improvement in my model.</p>\n<h3>Data Augmentation</h3>\n<p>I add some gaussian noise in continuous features with 0.02 std.</p>\n<h3>Training Sample</h3>\n<p>I use the data with dates after 252, because there are too many Nan values in the first 252 dates. I have tried use the dates after 750 to train, which improve my cv but decrease my lb.</p>\n<h3>Scores</h3>\n<p>without online learning    cv : 0.0106  lb : 0.0077<br>\nwith online learning    cv : 0.0116  lb : 0.0083</p>\n<h1>3. AutoencoderMLP</h1>\n<h3>Datapreprocessing</h3>\n<p>Fill 3 for Nan values without standardization.</p>\n<h3>Feature Engeneering</h3>\n<p>We use time_id, symbol_id, 79 original features and responder_6 lags 1 of last. We didn't use auxilary target in the AutoencoderMLP</p>\n<h3>Model Architecture</h3>\n<p>Encoder and Decoder : Both are a lieanr layer with 96 hidden dimension.<br>\nMLP : 5 layers with hidden dimension 96, 896, 448, 448, 256</p>\n<h3>Hyperparameter</h3>\n<p>Dropout : [0.035, 0.038, 0.424, 0.104, 0.25, 0.32, 0.271, 0.25]<br>\nLearning rate: 1e-3<br>\nBatch size : 8192<br>\nEpoch : about 15<br>\noptimizer : Adam</p>\n<h3>Loss function</h3>\n<p>We use mse loss as reconstruction loss, and weighted-mse loss as prediction loss</p>\n<h3>Scores</h3>\n<p>without online learning    cv : 0.0103  lb : 0.0072<br>\nwith online learning    cv : 0.0110  lb : 0.0078</p>\n<h1>4. My GBDT and Ridge offline model</h1>\n<h3>Feature Engeneering</h3>\n<p>It is more simple now. For LGB, XGB and Ridge, I just use time_id, symbol_id, 79 features, responder lags 1 of last, mean, std and max. And I did not specify category features in the GBDT model.</p>\n<h3>Training Sample</h3>\n<p>The dates after 750 for XGB and Ridge and dates after 678 for LGB </p>\n<h3>Hyperparameter</h3>\n<p>I didn't optimize my hyperparameter just use fixed one.</p>\n<pre><code>LGB_Params = {\n        ': ,\n        ': ,\n        ': ,\n        ': ,\n        ': ,\n        ': ,\n        ': ,\n        ': ,\n        ': ,\n        ' : '',\n        ': ,\n        ': ',\n    }\n\nXGB_Params = {\n    ': ,\n    ': ,\n    ': ,\n    ': ,\n    ': ,\n    ': ,\n    ': ,\n    ': ,\n    ': ',\n    ' : '',\n    ': ,\n    ' : ':squarederror',\n}\n</code></pre>\n<h3>Scores</h3>\n<p>LGB   cv: 0.0096 lb : 0.0072<br>\nXGB   cv: 0.0102  lb : 0.0073<br>\nRidge cv: 0.0035 lb : 0.0044<br>\nAfter ensembling these 3 models, the lb is 0.0076.</p>\n<h1>5. <a href=\"https://www.kaggle.com/chronoscop\" target=\"_blank\">@chronoscop</a>'s XGB</h1>\n<h3>Datapreprocessing</h3>\n<p>Fill 0 for Nan values without standardization</p>\n<h3>Feature Engeneering</h3>\n<p>use symbol_id, time_id, 79 original feature and responder lags1 of last. If add more lags features, it would be easy to overfit.</p>\n<h3>Training Sample</h3>\n<p>use dates after 917 to train the model.</p>\n<h3>Hyperparameter</h3>\n<pre><code> = {\n: ,\n: ,\n: ,\n : ,\n: ,\n: ,\n: ,\n: ,\n: ,\n: \n}\n</code></pre>\n<h3>Scores</h3>\n<p>cv : 0.0102   lb : 0.0072</p>\n<h1>6. Online Learning</h1>\n<p>Online learning play an important role in this competition.<br>\nWe collect data  for each previous date and then update model in every date. For every date, we will combine the data of the last date with partial data from the previous dates (using random sampling) for online training. In the training part, instead of batch processing, we input all the data into the model and then train 5 epochs.</p>\n<pre><code>if previous_test is not  and lags is not :\n    train = previous_test.(\n        lags_.([, , pl.().()]),\n        on=[, ],\n        how=\n    )\n\n    if pre_train is not None and (pre_train) &gt; :\n        pre_train = pre_train.(n=, seed=)\n    if pre_train is None:\n        pre_train = train\n    else:\n        pre_train = pl.([pre_train, train])\n    X = pre_train.()[cols].().values\n    y = pre_train[].()\n    weights = pre_train[].()\n\n\n    model.()\n    optimizer = torch.optim.(model.(), lr=e-)\n    criterion = nn.()\n\n\n    for epoch in ():\n        optimizer.()\n        _, y_pred = (torch.(X).(device))\n        loss = (y_pred.(), torch.(y.()).(device))\n        loss.()\n        optimizer.()\n        (f)\n\nif previous_test is None:\n    previous_test = test\nelse:\n    previous_test = pl.([previous_test, test])\n</code></pre>\n<p>We tried to add online learning into GBDT, but it didn't work well and cost too many time.</p>\n<h1>Thanks !</h1>\n<p>I think TabM can get higher score if adjust different kind of backbones, but I have no time to experiment. I also tried GRU, LSTM or Transformer in the early time, but they all failed, so then I focus on tabular model. Hoping top teams can share more trick and I indeed learn a lot in the competition. Thanks to all of you! :)<br>\nPS: Our code will be released when it is sorted out.</p>\n<hr>\n<p>UPDATE : <br>\nTabM training code <a href=\"https://www.kaggle.com/code/i2nfinit3y/jane-street-tabm-training?scriptVersionId=217873214\" target=\"_blank\">https://www.kaggle.com/code/i2nfinit3y/jane-street-tabm-training?scriptVersionId=217873214</a><br>\nSingle TabM online learning : <a href=\"https://www.kaggle.com/code/i2nfinit3y/jane-street-tabm-online-learning\" target=\"_blank\">https://www.kaggle.com/code/i2nfinit3y/jane-street-tabm-online-learning</a></p>\n<hr>\n<p>UPDATE :<br>\nwe’ve open-sourced both our inference and training code.</p>",
      "rawMarkdown": "Thanks to this amazing competition and the efforts of every one of my teammates ! @lechengyan @chronoscop\nOur final solution of 0.0092 lb is ensembling NN models with online learning , GBDT offline models and a ridge model. My part is mainly for TabM model and some GBDT models, which is what I am gonna to talk about. AutoencoderMLP and online learning are @lechengyan 's part and @chronoscop is in charge of one of XGB.\n\n# 1. Cross-Validation\nI simply used the last 120 dates as my validation and it shows good correlations with LB.\n\n# 2. TabM Model\nActually, in the early time of this competition, I noticed TabM and found it is a tabular NN model with great potential. I also public a baseline [notebook](https://www.kaggle.com/code/i2nfinit3y/jane-street-tabm-ft-transformer-inference).\n### Model Architecture\nThe model parts I mainly adjust are feature embed layers and backbone.\nFor category features, I use onehot encoding for every category tensor, which contains feature_09, feature_10, feature_11, symbol_id and time_id. I remap these category features with ordinal encode and every new category  in the test data will be mapped to the same category. Apart from one-hot encoding, I also use embedding layers but it didn't help much. It's worth noting that taking time_id as category features improve my model al lot.\nBut for continuous features, I have tried LinearEmbeddings, PeriodicEmbeddings and PiecewiseLinearEmbeddings, which didn't work well. So I don't use any embedding layer for continuous features.\nThe backbone of TabM is just 3 layers MLP with the same 512 dimensions. I haven't experimented with too many dimensional combinations, but 512 seems to be the most stable.\n### Loss function\nI have used huber loss, logcosh loss, mae loss, zero-mean R2 loss and mse loss, and R2 loss and mse loss work best. Specifically, mse loss can get higher score in the cv but R2 loss is more robust in every validation epoch. So I choose R2 loss as loss function.\n### Hyperparameter\nDropout : 0.25. Higher dropout rate (0.5) will result in slower convergence and lower (0.1) will be overfitting, so 0.25 seems to be a good choice for me.\nLearning rate: 1e-3.\nWeight decay: 8e-4.\nBatch size : 8192\nepoch : 5 or 6. Generally, I will submit multiple epoch to confirm the best model checkpoint and most of time 5 or 6 is the best.\noptimizer : AdamW\nk : 16. k is the number of ensemble in the finial output. I tried 8, 16, 24, 32, and I found 16 can takes both scores and training time into account.\n### Datapreprocessing\nSimple mean-std standardize and fill zero for nan values.\n### Feature Engeneering & Auxiliary targets\nI use time_id, symbol_id, 78 original features (besides feature_61, because it just change with dates) and 9 responder lags 1 of last record. Additionally, I also use sin and cos of time_id  with 967 periods or 483 periods and feature_61 with 20 period to catch some periodical change. Because though feature_61 only changes with the date, it also changes in about 20 dates. Also, I tried to use more lags features but it work worse in cv and lb.\nBesides responder_6, I use responder_3 as my auxiliary target, because it has high correlations with responder_6. I tried to add more auxiliary targets in my training, but there was no obvious improvement in my model.\n### Data Augmentation\nI add some gaussian noise in continuous features with 0.02 std.\n### Training Sample\nI use the data with dates after 252, because there are too many Nan values in the first 252 dates. I have tried use the dates after 750 to train, which improve my cv but decrease my lb.\n### Scores\nwithout online learning    cv : 0.0106  lb : 0.0077\nwith online learning    cv : 0.0116  lb : 0.0083\n\n# 3. AutoencoderMLP\n### Datapreprocessing\nFill 3 for Nan values without standardization.\n### Feature Engeneering\nWe use time_id, symbol_id, 79 original features and responder_6 lags 1 of last. We didn't use auxilary target in the AutoencoderMLP\n### Model Architecture\nEncoder and Decoder : Both are a lieanr layer with 96 hidden dimension.\nMLP : 5 layers with hidden dimension 96, 896, 448, 448, 256\n### Hyperparameter\nDropout : [0.035, 0.038, 0.424, 0.104, 0.25, 0.32, 0.271, 0.25]\nLearning rate: 1e-3\nBatch size : 8192\nEpoch : about 15\noptimizer : Adam\n### Loss function\nWe use mse loss as reconstruction loss, and weighted-mse loss as prediction loss\n### Scores\nwithout online learning    cv : 0.0103  lb : 0.0072\nwith online learning    cv : 0.0110  lb : 0.0078\n\n# 4. My GBDT and Ridge offline model\n### Feature Engeneering\nIt is more simple now. For LGB, XGB and Ridge, I just use time_id, symbol_id, 79 features, responder lags 1 of last, mean, std and max. And I did not specify category features in the GBDT model.\n### Training Sample\nThe dates after 750 for XGB and Ridge and dates after 678 for LGB \n### Hyperparameter\nI didn't optimize my hyperparameter just use fixed one.\n```\nLGB_Params = {\n        'learning_rate': 0.05,\n        'max_depth': 6,\n        'num_leaves': 62,\n        'n_estimators': 200,\n        'subsample': 0.8,\n        'colsample_bytree': 0.8,\n        'reg_alpha': 1,\n        'reg_lambda': 1,\n        'random_state': 42,\n        'device' : 'gpu',\n        'gpu_use_dp': True,\n        'objective': 'l2',\n    }\n\nXGB_Params = {\n    'learning_rate': 0.05,\n    'max_depth': 6,\n    'n_estimators': 300,\n    'subsample': 0.8,\n    'colsample_bytree': 0.6,\n    'reg_alpha': 1,\n    'reg_lambda': 1,\n    'random_state': 42,\n    'tree_method': 'hist',\n    'device' : 'cuda',\n    'n_gpu': 1,\n    'objective' : 'reg:squarederror',\n}\n```\n### Scores\nLGB   cv: 0.0096 lb : 0.0072\nXGB   cv: 0.0102  lb : 0.0073\nRidge cv: 0.0035 lb : 0.0044\nAfter ensembling these 3 models, the lb is 0.0076.\n\n# 5. @chronoscop's XGB\n### Datapreprocessing\nFill 0 for Nan values without standardization\n### Feature Engeneering\nuse symbol_id, time_id, 79 original feature and responder lags1 of last. If add more lags features, it would be easy to overfit.\n### Training Sample\nuse dates after 917 to train the model.\n### Hyperparameter\n```\nparams = {\n'objective': 'reg:squarederror',\n'random_state': 1212,\n'tree_method': 'hist',\n'device' : 'cuda',\n'learning_rate': 0.02156022412857549,\n'max_depth': 8,\n'subsample': 0.7697954003310141,\n'colsample_bytree': 0.5182134365961873,\n'reg_alpha': 0.0032315937370696354,\n'reg_lambda': 0.002663721647776419\n}\n```\n### Scores\ncv : 0.0102   lb : 0.0072\n\n# 6. Online Learning\nOnline learning play an important role in this competition.\nWe collect data  for each previous date and then update model in every date. For every date, we will combine the data of the last date with partial data from the previous dates (using random sampling) for online training. In the training part, instead of batch processing, we input all the data into the model and then train 5 epochs.\n```\nif previous_test is not None and lags is not None:\n    train = previous_test.join(\n        lags_.select([\"time_id\", \"symbol_id\", pl.col(\"responder_6_lag_1\").alias(\"responder_6\")]),\n        on=[\"time_id\", \"symbol_id\"],\n        how=\"left\"\n    )\n\n    if pre_train is not None and len(pre_train) > 300000:\n        pre_train = pre_train.sample(n=300000, seed=2025)\n    if pre_train is None:\n        pre_train = train\n    else:\n        pre_train = pl.concat([pre_train, train])\n    X = pre_train.to_pandas()[cols].fillna(3).values\n    y = pre_train['responder_6'].to_numpy()\n    weights = pre_train['weight'].to_numpy()\n\n\n    model.train()\n    optimizer = torch.optim.Adam(model.parameters(), lr=1e-4)\n    criterion = nn.MSELoss()\n\n\n    for epoch in range(5):\n        optimizer.zero_grad()\n        _, y_pred = model(torch.FloatTensor(X).to(device))\n        loss = criterion(y_pred.squeeze(), torch.FloatTensor(y.copy()).to(device))\n        loss.backward()\n        optimizer.step()\n        print(f\"Epoch {epoch + 1}/5, Loss: {loss.item()}\")\n\nif previous_test is None:\n    previous_test = test\nelse:\n    previous_test = pl.concat([previous_test, test])\n```\nWe tried to add online learning into GBDT, but it didn't work well and cost too many time.\n\n# Thanks !\nI think TabM can get higher score if adjust different kind of backbones, but I have no time to experiment. I also tried GRU, LSTM or Transformer in the early time, but they all failed, so then I focus on tabular model. Hoping top teams can share more trick and I indeed learn a lot in the competition. Thanks to all of you! :)\nPS: Our code will be released when it is sorted out.\n\n----------------------------------------------------------------------\nUPDATE : \nTabM training code https://www.kaggle.com/code/i2nfinit3y/jane-street-tabm-training?scriptVersionId=217873214\nSingle TabM online learning : https://www.kaggle.com/code/i2nfinit3y/jane-street-tabm-online-learning\n\n----------------------------------------------------------------------\nUPDATE :\nwe’ve open-sourced both our inference and training code.",
      "votes": 14
    }
  ],
  "comments": [],
  "raw_markdown_by_id": {
    "3281421": "Thanks to this amazing competition and the efforts of every one of my teammates ! @lechengyan @chronoscop\nOur final solution of 0.0092 lb is ensembling NN models with online learning , GBDT offline models and a ridge model. My part is mainly for TabM model and some GBDT models, which is what I am gonna to talk about. AutoencoderMLP and online learning are @lechengyan 's part and @chronoscop is in charge of one of XGB.\n\n# 1. Cross-Validation\nI simply used the last 120 dates as my validation and it shows good correlations with LB.\n\n# 2. TabM Model\nActually, in the early time of this competition, I noticed TabM and found it is a tabular NN model with great potential. I also public a baseline [notebook](https://www.kaggle.com/code/i2nfinit3y/jane-street-tabm-ft-transformer-inference).\n### Model Architecture\nThe model parts I mainly adjust are feature embed layers and backbone.\nFor category features, I use onehot encoding for every category tensor, which contains feature_09, feature_10, feature_11, symbol_id and time_id. I remap these category features with ordinal encode and every new category  in the test data will be mapped to the same category. Apart from one-hot encoding, I also use embedding layers but it didn't help much. It's worth noting that taking time_id as category features improve my model al lot.\nBut for continuous features, I have tried LinearEmbeddings, PeriodicEmbeddings and PiecewiseLinearEmbeddings, which didn't work well. So I don't use any embedding layer for continuous features.\nThe backbone of TabM is just 3 layers MLP with the same 512 dimensions. I haven't experimented with too many dimensional combinations, but 512 seems to be the most stable.\n### Loss function\nI have used huber loss, logcosh loss, mae loss, zero-mean R2 loss and mse loss, and R2 loss and mse loss work best. Specifically, mse loss can get higher score in the cv but R2 loss is more robust in every validation epoch. So I choose R2 loss as loss function.\n### Hyperparameter\nDropout : 0.25. Higher dropout rate (0.5) will result in slower convergence and lower (0.1) will be overfitting, so 0.25 seems to be a good choice for me.\nLearning rate: 1e-3.\nWeight decay: 8e-4.\nBatch size : 8192\nepoch : 5 or 6. Generally, I will submit multiple epoch to confirm the best model checkpoint and most of time 5 or 6 is the best.\noptimizer : AdamW\nk : 16. k is the number of ensemble in the finial output. I tried 8, 16, 24, 32, and I found 16 can takes both scores and training time into account.\n### Datapreprocessing\nSimple mean-std standardize and fill zero for nan values.\n### Feature Engeneering & Auxiliary targets\nI use time_id, symbol_id, 78 original features (besides feature_61, because it just change with dates) and 9 responder lags 1 of last record. Additionally, I also use sin and cos of time_id  with 967 periods or 483 periods and feature_61 with 20 period to catch some periodical change. Because though feature_61 only changes with the date, it also changes in about 20 dates. Also, I tried to use more lags features but it work worse in cv and lb.\nBesides responder_6, I use responder_3 as my auxiliary target, because it has high correlations with responder_6. I tried to add more auxiliary targets in my training, but there was no obvious improvement in my model.\n### Data Augmentation\nI add some gaussian noise in continuous features with 0.02 std.\n### Training Sample\nI use the data with dates after 252, because there are too many Nan values in the first 252 dates. I have tried use the dates after 750 to train, which improve my cv but decrease my lb.\n### Scores\nwithout online learning    cv : 0.0106  lb : 0.0077\nwith online learning    cv : 0.0116  lb : 0.0083\n\n# 3. AutoencoderMLP\n### Datapreprocessing\nFill 3 for Nan values without standardization.\n### Feature Engeneering\nWe use time_id, symbol_id, 79 original features and responder_6 lags 1 of last. We didn't use auxilary target in the AutoencoderMLP\n### Model Architecture\nEncoder and Decoder : Both are a lieanr layer with 96 hidden dimension.\nMLP : 5 layers with hidden dimension 96, 896, 448, 448, 256\n### Hyperparameter\nDropout : [0.035, 0.038, 0.424, 0.104, 0.25, 0.32, 0.271, 0.25]\nLearning rate: 1e-3\nBatch size : 8192\nEpoch : about 15\noptimizer : Adam\n### Loss function\nWe use mse loss as reconstruction loss, and weighted-mse loss as prediction loss\n### Scores\nwithout online learning    cv : 0.0103  lb : 0.0072\nwith online learning    cv : 0.0110  lb : 0.0078\n\n# 4. My GBDT and Ridge offline model\n### Feature Engeneering\nIt is more simple now. For LGB, XGB and Ridge, I just use time_id, symbol_id, 79 features, responder lags 1 of last, mean, std and max. And I did not specify category features in the GBDT model.\n### Training Sample\nThe dates after 750 for XGB and Ridge and dates after 678 for LGB \n### Hyperparameter\nI didn't optimize my hyperparameter just use fixed one.\n```\nLGB_Params = {\n        'learning_rate': 0.05,\n        'max_depth': 6,\n        'num_leaves': 62,\n        'n_estimators': 200,\n        'subsample': 0.8,\n        'colsample_bytree': 0.8,\n        'reg_alpha': 1,\n        'reg_lambda': 1,\n        'random_state': 42,\n        'device' : 'gpu',\n        'gpu_use_dp': True,\n        'objective': 'l2',\n    }\n\nXGB_Params = {\n    'learning_rate': 0.05,\n    'max_depth': 6,\n    'n_estimators': 300,\n    'subsample': 0.8,\n    'colsample_bytree': 0.6,\n    'reg_alpha': 1,\n    'reg_lambda': 1,\n    'random_state': 42,\n    'tree_method': 'hist',\n    'device' : 'cuda',\n    'n_gpu': 1,\n    'objective' : 'reg:squarederror',\n}\n```\n### Scores\nLGB   cv: 0.0096 lb : 0.0072\nXGB   cv: 0.0102  lb : 0.0073\nRidge cv: 0.0035 lb : 0.0044\nAfter ensembling these 3 models, the lb is 0.0076.\n\n# 5. @chronoscop's XGB\n### Datapreprocessing\nFill 0 for Nan values without standardization\n### Feature Engeneering\nuse symbol_id, time_id, 79 original feature and responder lags1 of last. If add more lags features, it would be easy to overfit.\n### Training Sample\nuse dates after 917 to train the model.\n### Hyperparameter\n```\nparams = {\n'objective': 'reg:squarederror',\n'random_state': 1212,\n'tree_method': 'hist',\n'device' : 'cuda',\n'learning_rate': 0.02156022412857549,\n'max_depth': 8,\n'subsample': 0.7697954003310141,\n'colsample_bytree': 0.5182134365961873,\n'reg_alpha': 0.0032315937370696354,\n'reg_lambda': 0.002663721647776419\n}\n```\n### Scores\ncv : 0.0102   lb : 0.0072\n\n# 6. Online Learning\nOnline learning play an important role in this competition.\nWe collect data  for each previous date and then update model in every date. For every date, we will combine the data of the last date with partial data from the previous dates (using random sampling) for online training. In the training part, instead of batch processing, we input all the data into the model and then train 5 epochs.\n```\nif previous_test is not None and lags is not None:\n    train = previous_test.join(\n        lags_.select([\"time_id\", \"symbol_id\", pl.col(\"responder_6_lag_1\").alias(\"responder_6\")]),\n        on=[\"time_id\", \"symbol_id\"],\n        how=\"left\"\n    )\n\n    if pre_train is not None and len(pre_train) > 300000:\n        pre_train = pre_train.sample(n=300000, seed=2025)\n    if pre_train is None:\n        pre_train = train\n    else:\n        pre_train = pl.concat([pre_train, train])\n    X = pre_train.to_pandas()[cols].fillna(3).values\n    y = pre_train['responder_6'].to_numpy()\n    weights = pre_train['weight'].to_numpy()\n\n\n    model.train()\n    optimizer = torch.optim.Adam(model.parameters(), lr=1e-4)\n    criterion = nn.MSELoss()\n\n\n    for epoch in range(5):\n        optimizer.zero_grad()\n        _, y_pred = model(torch.FloatTensor(X).to(device))\n        loss = criterion(y_pred.squeeze(), torch.FloatTensor(y.copy()).to(device))\n        loss.backward()\n        optimizer.step()\n        print(f\"Epoch {epoch + 1}/5, Loss: {loss.item()}\")\n\nif previous_test is None:\n    previous_test = test\nelse:\n    previous_test = pl.concat([previous_test, test])\n```\nWe tried to add online learning into GBDT, but it didn't work well and cost too many time.\n\n# Thanks !\nI think TabM can get higher score if adjust different kind of backbones, but I have no time to experiment. I also tried GRU, LSTM or Transformer in the early time, but they all failed, so then I focus on tabular model. Hoping top teams can share more trick and I indeed learn a lot in the competition. Thanks to all of you! :)\nPS: Our code will be released when it is sorted out.\n\n----------------------------------------------------------------------\nUPDATE : \nTabM training code https://www.kaggle.com/code/i2nfinit3y/jane-street-tabm-training?scriptVersionId=217873214\nSingle TabM online learning : https://www.kaggle.com/code/i2nfinit3y/jane-street-tabm-online-learning\n\n----------------------------------------------------------------------\nUPDATE :\nwe’ve open-sourced both our inference and training code."
  }
}