{
  "id": 556610,
  "title": "[Public LB 26th] TabM, AutoencoderMLP with online training & GBDT offline models",
  "url": "/competitions/jane-street-real-time-market-data-forecasting/discussion/556610",
  "author_name": "",
  "post_date": "2025-01-14T07:57:11.928259700Z",
  "votes": 33,
  "comment_count": 21,
  "views": 0,
  "content": "<p>Thanks to this amazing competition and the efforts of every one of my teammates ! <a href=\"https://www.kaggle.com/lechengyan\" target=\"_blank\">@lechengyan</a> <a href=\"https://www.kaggle.com/chronoscop\" target=\"_blank\">@chronoscop</a><br>\nOur final solution of 0.0092 lb is ensembling NN models with online learning , GBDT offline models and a ridge model. My part is mainly for TabM model and some GBDT models, which is what I am gonna to talk about. AutoencoderMLP and online learning are <a href=\"https://www.kaggle.com/lechengyan\" target=\"_blank\">@lechengyan</a> 's part and <a href=\"https://www.kaggle.com/chronoscop\" target=\"_blank\">@chronoscop</a> is in charge of one of XGB.</p>\n<h1>1. Cross-Validation</h1>\n<p>I simply used the last 120 dates as my validation and it shows good correlations with LB.</p>\n<h1>2. TabM Model</h1>\n<p>Actually, in the early time of this competition, I noticed TabM and found it is a tabular NN model with great potential. I also public a baseline <a href=\"https://www.kaggle.com/code/i2nfinit3y/jane-street-tabm-ft-transformer-inference\" target=\"_blank\">notebook</a>.</p>\n<h3>Model Architecture</h3>\n<p>The model parts I mainly adjust are feature embed layers and backbone.<br>\nFor category features, I use onehot encoding for every category tensor, which contains feature_09, feature_10, feature_11, symbol_id and time_id. I remap these category features with ordinal encode and every new category  in the test data will be mapped to the same category. Apart from one-hot encoding, I also use embedding layers but it didn't help much. It's worth noting that taking time_id as category features improve my model al lot.<br>\nBut for continuous features, I have tried LinearEmbeddings, PeriodicEmbeddings and PiecewiseLinearEmbeddings, which didn't work well. So I don't use any embedding layer for continuous features.<br>\nThe backbone of TabM is just 3 layers MLP with the same 512 dimensions. I haven't experimented with too many dimensional combinations, but 512 seems to be the most stable.</p>\n<h3>Loss function</h3>\n<p>I have used huber loss, logcosh loss, mae loss, zero-mean R2 loss and mse loss, and R2 loss and mse loss work best. Specifically, mse loss can get higher score in the cv but R2 loss is more robust in every validation epoch. So I choose R2 loss as loss function.</p>\n<h3>Hyperparameter</h3>\n<p>Dropout : 0.25. Higher dropout rate (0.5) will result in slower convergence and lower (0.1) will be overfitting, so 0.25 seems to be a good choice for me.<br>\nLearning rate: 1e-3.<br>\nWeight decay: 8e-4.<br>\nBatch size : 8192<br>\nepoch : 5 or 6. Generally, I will submit multiple epoch to confirm the best model checkpoint and most of time 5 or 6 is the best.<br>\noptimizer : AdamW<br>\nk : 16. k is the number of ensemble in the finial output. I tried 8, 16, 24, 32, and I found 16 can takes both scores and training time into account.</p>\n<h3>Datapreprocessing</h3>\n<p>Simple mean-std standardize and fill zero for nan values.</p>\n<h3>Feature Engeneering &amp; Auxiliary targets</h3>\n<p>I use time_id, symbol_id, 78 original features (besides feature_61, because it just change with dates) and 9 responder lags 1 of last record. Additionally, I also use sin and cos of time_id  with 967 periods or 483 periods and feature_61 with 20 period to catch some periodical change. Because though feature_61 only changes with the date, it also changes in about 20 dates. Also, I tried to use more lags features but it work worse in cv and lb.<br>\nBesides responder_6, I use responder_3 as my auxiliary target, because it has high correlations with responder_6. I tried to add more auxiliary targets in my training, but there was no obvious improvement in my model.</p>\n<h3>Data Augmentation</h3>\n<p>I add some gaussian noise in continuous features with 0.02 std.</p>\n<h3>Training Sample</h3>\n<p>I use the data with dates after 252, because there are too many Nan values in the first 252 dates. I have tried use the dates after 750 to train, which improve my cv but decrease my lb.</p>\n<h3>Scores</h3>\n<p>without online learning    cv : 0.0106  lb : 0.0077<br>\nwith online learning    cv : 0.0116  lb : 0.0083</p>\n<h1>3. AutoencoderMLP</h1>\n<h3>Datapreprocessing</h3>\n<p>Fill 3 for Nan values without standardization.</p>\n<h3>Feature Engeneering</h3>\n<p>We use time_id, symbol_id, 79 original features and responder_6 lags 1 of last. We didn't use auxilary target in the AutoencoderMLP</p>\n<h3>Model Architecture</h3>\n<p>Encoder and Decoder : Both are a lieanr layer with 96 hidden dimension.<br>\nMLP : 5 layers with hidden dimension 96, 896, 448, 448, 256</p>\n<h3>Hyperparameter</h3>\n<p>Dropout : [0.035, 0.038, 0.424, 0.104, 0.25, 0.32, 0.271, 0.25]<br>\nLearning rate: 1e-3<br>\nBatch size : 8192<br>\nEpoch : about 15<br>\noptimizer : Adam</p>\n<h3>Loss function</h3>\n<p>We use mse loss as reconstruction loss, and weighted-mse loss as prediction loss</p>\n<h3>Scores</h3>\n<p>without online learning    cv : 0.0103  lb : 0.0072<br>\nwith online learning    cv : 0.0110  lb : 0.0078</p>\n<h1>4. My GBDT and Ridge offline model</h1>\n<h3>Feature Engeneering</h3>\n<p>It is more simple now. For LGB, XGB and Ridge, I just use time_id, symbol_id, 79 features, responder lags 1 of last, mean, std and max. And I did not specify category features in the GBDT model.</p>\n<h3>Training Sample</h3>\n<p>The dates after 750 for XGB and Ridge and dates after 678 for LGB </p>\n<h3>Hyperparameter</h3>\n<p>I didn't optimize my hyperparameter just use fixed one.</p>\n<pre><code>LGB_Params = {\n        ': ,\n        ': ,\n        ': ,\n        ': ,\n        ': ,\n        ': ,\n        ': ,\n        ': ,\n        ': ,\n        ' : '',\n        ': ,\n        ': ',\n    }\n\nXGB_Params = {\n    ': ,\n    ': ,\n    ': ,\n    ': ,\n    ': ,\n    ': ,\n    ': ,\n    ': ,\n    ': ',\n    ' : '',\n    ': ,\n    ' : ':squarederror',\n}\n</code></pre>\n<h3>Scores</h3>\n<p>LGB   cv: 0.0096 lb : 0.0072<br>\nXGB   cv: 0.0102  lb : 0.0073<br>\nRidge cv: 0.0035 lb : 0.0044<br>\nAfter ensembling these 3 models, the lb is 0.0076.</p>\n<h1>5. <a href=\"https://www.kaggle.com/chronoscop\" target=\"_blank\">@chronoscop</a>'s XGB</h1>\n<h3>Datapreprocessing</h3>\n<p>Fill 0 for Nan values without standardization</p>\n<h3>Feature Engeneering</h3>\n<p>use symbol_id, time_id, 79 original feature and responder lags1 of last. If add more lags features, it would be easy to overfit.</p>\n<h3>Training Sample</h3>\n<p>use dates after 917 to train the model.</p>\n<h3>Hyperparameter</h3>\n<pre><code> = {\n: ,\n: ,\n: ,\n : ,\n: ,\n: ,\n: ,\n: ,\n: ,\n: \n}\n</code></pre>\n<h3>Scores</h3>\n<p>cv : 0.0102   lb : 0.0072</p>\n<h1>6. Online Learning</h1>\n<p>Online learning play an important role in this competition.<br>\nWe collect data  for each previous date and then update model in every date. For every date, we will combine the data of the last date with partial data from the previous dates (using random sampling) for online training. In the training part, instead of batch processing, we input all the data into the model and then train 5 epochs.</p>\n<pre><code>if previous_test is not  and lags is not :\n    train = previous_test.(\n        lags_.([, , pl.().()]),\n        on=[, ],\n        how=\n    )\n\n    if pre_train is not None and (pre_train) &gt; :\n        pre_train = pre_train.(n=, seed=)\n    if pre_train is None:\n        pre_train = train\n    else:\n        pre_train = pl.([pre_train, train])\n    X = pre_train.()[cols].().values\n    y = pre_train[].()\n    weights = pre_train[].()\n\n\n    model.()\n    optimizer = torch.optim.(model.(), lr=e-)\n    criterion = nn.()\n\n\n    for epoch in ():\n        optimizer.()\n        _, y_pred = (torch.(X).(device))\n        loss = (y_pred.(), torch.(y.()).(device))\n        loss.()\n        optimizer.()\n        (f)\n\nif previous_test is None:\n    previous_test = test\nelse:\n    previous_test = pl.([previous_test, test])\n</code></pre>\n<p>We tried to add online learning into GBDT, but it didn't work well and cost too many time.</p>\n<h1>Thanks !</h1>\n<p>I think TabM can get higher score if adjust different kind of backbones, but I have no time to experiment. I also tried GRU, LSTM or Transformer in the early time, but they all failed, so then I focus on tabular model. Hoping top teams can share more trick and I indeed learn a lot in the competition. Thanks to all of you! :)<br>\nPS: Our code will be released when it is sorted out.</p>\n<hr>\n<p>UPDATE : <br>\nTabM training code <a href=\"https://www.kaggle.com/code/i2nfinit3y/jane-street-tabm-training?scriptVersionId=217873214\" target=\"_blank\">https://www.kaggle.com/code/i2nfinit3y/jane-street-tabm-training?scriptVersionId=217873214</a><br>\nSingle TabM online learning : <a href=\"https://www.kaggle.com/code/i2nfinit3y/jane-street-tabm-online-learning\" target=\"_blank\">https://www.kaggle.com/code/i2nfinit3y/jane-street-tabm-online-learning</a></p>",
  "messages": [
    {
      "id": "3096299",
      "postDate": "01/14/2025 07:57:11",
      "content": "<p>Thanks to this amazing competition and the efforts of every one of my teammates ! <a href=\"https://www.kaggle.com/lechengyan\" target=\"_blank\">@lechengyan</a> <a href=\"https://www.kaggle.com/chronoscop\" target=\"_blank\">@chronoscop</a><br>\nOur final solution of 0.0092 lb is ensembling NN models with online learning , GBDT offline models and a ridge model. My part is mainly for TabM model and some GBDT models, which is what I am gonna to talk about. AutoencoderMLP and online learning are <a href=\"https://www.kaggle.com/lechengyan\" target=\"_blank\">@lechengyan</a> 's part and <a href=\"https://www.kaggle.com/chronoscop\" target=\"_blank\">@chronoscop</a> is in charge of one of XGB.</p>\n<h1>1. Cross-Validation</h1>\n<p>I simply used the last 120 dates as my validation and it shows good correlations with LB.</p>\n<h1>2. TabM Model</h1>\n<p>Actually, in the early time of this competition, I noticed TabM and found it is a tabular NN model with great potential. I also public a baseline <a href=\"https://www.kaggle.com/code/i2nfinit3y/jane-street-tabm-ft-transformer-inference\" target=\"_blank\">notebook</a>.</p>\n<h3>Model Architecture</h3>\n<p>The model parts I mainly adjust are feature embed layers and backbone.<br>\nFor category features, I use onehot encoding for every category tensor, which contains feature_09, feature_10, feature_11, symbol_id and time_id. I remap these category features with ordinal encode and every new category  in the test data will be mapped to the same category. Apart from one-hot encoding, I also use embedding layers but it didn't help much. It's worth noting that taking time_id as category features improve my model al lot.<br>\nBut for continuous features, I have tried LinearEmbeddings, PeriodicEmbeddings and PiecewiseLinearEmbeddings, which didn't work well. So I don't use any embedding layer for continuous features.<br>\nThe backbone of TabM is just 3 layers MLP with the same 512 dimensions. I haven't experimented with too many dimensional combinations, but 512 seems to be the most stable.</p>\n<h3>Loss function</h3>\n<p>I have used huber loss, logcosh loss, mae loss, zero-mean R2 loss and mse loss, and R2 loss and mse loss work best. Specifically, mse loss can get higher score in the cv but R2 loss is more robust in every validation epoch. So I choose R2 loss as loss function.</p>\n<h3>Hyperparameter</h3>\n<p>Dropout : 0.25. Higher dropout rate (0.5) will result in slower convergence and lower (0.1) will be overfitting, so 0.25 seems to be a good choice for me.<br>\nLearning rate: 1e-3.<br>\nWeight decay: 8e-4.<br>\nBatch size : 8192<br>\nepoch : 5 or 6. Generally, I will submit multiple epoch to confirm the best model checkpoint and most of time 5 or 6 is the best.<br>\noptimizer : AdamW<br>\nk : 16. k is the number of ensemble in the finial output. I tried 8, 16, 24, 32, and I found 16 can takes both scores and training time into account.</p>\n<h3>Datapreprocessing</h3>\n<p>Simple mean-std standardize and fill zero for nan values.</p>\n<h3>Feature Engeneering &amp; Auxiliary targets</h3>\n<p>I use time_id, symbol_id, 78 original features (besides feature_61, because it just change with dates) and 9 responder lags 1 of last record. Additionally, I also use sin and cos of time_id  with 967 periods or 483 periods and feature_61 with 20 period to catch some periodical change. Because though feature_61 only changes with the date, it also changes in about 20 dates. Also, I tried to use more lags features but it work worse in cv and lb.<br>\nBesides responder_6, I use responder_3 as my auxiliary target, because it has high correlations with responder_6. I tried to add more auxiliary targets in my training, but there was no obvious improvement in my model.</p>\n<h3>Data Augmentation</h3>\n<p>I add some gaussian noise in continuous features with 0.02 std.</p>\n<h3>Training Sample</h3>\n<p>I use the data with dates after 252, because there are too many Nan values in the first 252 dates. I have tried use the dates after 750 to train, which improve my cv but decrease my lb.</p>\n<h3>Scores</h3>\n<p>without online learning    cv : 0.0106  lb : 0.0077<br>\nwith online learning    cv : 0.0116  lb : 0.0083</p>\n<h1>3. AutoencoderMLP</h1>\n<h3>Datapreprocessing</h3>\n<p>Fill 3 for Nan values without standardization.</p>\n<h3>Feature Engeneering</h3>\n<p>We use time_id, symbol_id, 79 original features and responder_6 lags 1 of last. We didn't use auxilary target in the AutoencoderMLP</p>\n<h3>Model Architecture</h3>\n<p>Encoder and Decoder : Both are a lieanr layer with 96 hidden dimension.<br>\nMLP : 5 layers with hidden dimension 96, 896, 448, 448, 256</p>\n<h3>Hyperparameter</h3>\n<p>Dropout : [0.035, 0.038, 0.424, 0.104, 0.25, 0.32, 0.271, 0.25]<br>\nLearning rate: 1e-3<br>\nBatch size : 8192<br>\nEpoch : about 15<br>\noptimizer : Adam</p>\n<h3>Loss function</h3>\n<p>We use mse loss as reconstruction loss, and weighted-mse loss as prediction loss</p>\n<h3>Scores</h3>\n<p>without online learning    cv : 0.0103  lb : 0.0072<br>\nwith online learning    cv : 0.0110  lb : 0.0078</p>\n<h1>4. My GBDT and Ridge offline model</h1>\n<h3>Feature Engeneering</h3>\n<p>It is more simple now. For LGB, XGB and Ridge, I just use time_id, symbol_id, 79 features, responder lags 1 of last, mean, std and max. And I did not specify category features in the GBDT model.</p>\n<h3>Training Sample</h3>\n<p>The dates after 750 for XGB and Ridge and dates after 678 for LGB </p>\n<h3>Hyperparameter</h3>\n<p>I didn't optimize my hyperparameter just use fixed one.</p>\n<pre><code>LGB_Params = {\n        ': ,\n        ': ,\n        ': ,\n        ': ,\n        ': ,\n        ': ,\n        ': ,\n        ': ,\n        ': ,\n        ' : '',\n        ': ,\n        ': ',\n    }\n\nXGB_Params = {\n    ': ,\n    ': ,\n    ': ,\n    ': ,\n    ': ,\n    ': ,\n    ': ,\n    ': ,\n    ': ',\n    ' : '',\n    ': ,\n    ' : ':squarederror',\n}\n</code></pre>\n<h3>Scores</h3>\n<p>LGB   cv: 0.0096 lb : 0.0072<br>\nXGB   cv: 0.0102  lb : 0.0073<br>\nRidge cv: 0.0035 lb : 0.0044<br>\nAfter ensembling these 3 models, the lb is 0.0076.</p>\n<h1>5. <a href=\"https://www.kaggle.com/chronoscop\" target=\"_blank\">@chronoscop</a>'s XGB</h1>\n<h3>Datapreprocessing</h3>\n<p>Fill 0 for Nan values without standardization</p>\n<h3>Feature Engeneering</h3>\n<p>use symbol_id, time_id, 79 original feature and responder lags1 of last. If add more lags features, it would be easy to overfit.</p>\n<h3>Training Sample</h3>\n<p>use dates after 917 to train the model.</p>\n<h3>Hyperparameter</h3>\n<pre><code> = {\n: ,\n: ,\n: ,\n : ,\n: ,\n: ,\n: ,\n: ,\n: ,\n: \n}\n</code></pre>\n<h3>Scores</h3>\n<p>cv : 0.0102   lb : 0.0072</p>\n<h1>6. Online Learning</h1>\n<p>Online learning play an important role in this competition.<br>\nWe collect data  for each previous date and then update model in every date. For every date, we will combine the data of the last date with partial data from the previous dates (using random sampling) for online training. In the training part, instead of batch processing, we input all the data into the model and then train 5 epochs.</p>\n<pre><code>if previous_test is not  and lags is not :\n    train = previous_test.(\n        lags_.([, , pl.().()]),\n        on=[, ],\n        how=\n    )\n\n    if pre_train is not None and (pre_train) &gt; :\n        pre_train = pre_train.(n=, seed=)\n    if pre_train is None:\n        pre_train = train\n    else:\n        pre_train = pl.([pre_train, train])\n    X = pre_train.()[cols].().values\n    y = pre_train[].()\n    weights = pre_train[].()\n\n\n    model.()\n    optimizer = torch.optim.(model.(), lr=e-)\n    criterion = nn.()\n\n\n    for epoch in ():\n        optimizer.()\n        _, y_pred = (torch.(X).(device))\n        loss = (y_pred.(), torch.(y.()).(device))\n        loss.()\n        optimizer.()\n        (f)\n\nif previous_test is None:\n    previous_test = test\nelse:\n    previous_test = pl.([previous_test, test])\n</code></pre>\n<p>We tried to add online learning into GBDT, but it didn't work well and cost too many time.</p>\n<h1>Thanks !</h1>\n<p>I think TabM can get higher score if adjust different kind of backbones, but I have no time to experiment. I also tried GRU, LSTM or Transformer in the early time, but they all failed, so then I focus on tabular model. Hoping top teams can share more trick and I indeed learn a lot in the competition. Thanks to all of you! :)<br>\nPS: Our code will be released when it is sorted out.</p>\n<hr>\n<p>UPDATE : <br>\nTabM training code <a href=\"https://www.kaggle.com/code/i2nfinit3y/jane-street-tabm-training?scriptVersionId=217873214\" target=\"_blank\">https://www.kaggle.com/code/i2nfinit3y/jane-street-tabm-training?scriptVersionId=217873214</a><br>\nSingle TabM online learning : <a href=\"https://www.kaggle.com/code/i2nfinit3y/jane-street-tabm-online-learning\" target=\"_blank\">https://www.kaggle.com/code/i2nfinit3y/jane-street-tabm-online-learning</a></p>",
      "rawMarkdown": "Thanks to this amazing competition and the efforts of every one of my teammates ! @lechengyan @chronoscop\nOur final solution of 0.0092 lb is ensembling NN models with online learning , GBDT offline models and a ridge model. My part is mainly for TabM model and some GBDT models, which is what I am gonna to talk about. AutoencoderMLP and online learning are @lechengyan 's part and @chronoscop is in charge of one of XGB.\n\n# 1. Cross-Validation\nI simply used the last 120 dates as my validation and it shows good correlations with LB.\n\n# 2. TabM Model\nActually, in the early time of this competition, I noticed TabM and found it is a tabular NN model with great potential. I also public a baseline [notebook](https://www.kaggle.com/code/i2nfinit3y/jane-street-tabm-ft-transformer-inference).\n### Model Architecture\nThe model parts I mainly adjust are feature embed layers and backbone.\nFor category features, I use onehot encoding for every category tensor, which contains feature_09, feature_10, feature_11, symbol_id and time_id. I remap these category features with ordinal encode and every new category  in the test data will be mapped to the same category. Apart from one-hot encoding, I also use embedding layers but it didn't help much. It's worth noting that taking time_id as category features improve my model al lot.\nBut for continuous features, I have tried LinearEmbeddings, PeriodicEmbeddings and PiecewiseLinearEmbeddings, which didn't work well. So I don't use any embedding layer for continuous features.\nThe backbone of TabM is just 3 layers MLP with the same 512 dimensions. I haven't experimented with too many dimensional combinations, but 512 seems to be the most stable.\n### Loss function\nI have used huber loss, logcosh loss, mae loss, zero-mean R2 loss and mse loss, and R2 loss and mse loss work best. Specifically, mse loss can get higher score in the cv but R2 loss is more robust in every validation epoch. So I choose R2 loss as loss function.\n### Hyperparameter\nDropout : 0.25. Higher dropout rate (0.5) will result in slower convergence and lower (0.1) will be overfitting, so 0.25 seems to be a good choice for me.\nLearning rate: 1e-3.\nWeight decay: 8e-4.\nBatch size : 8192\nepoch : 5 or 6. Generally, I will submit multiple epoch to confirm the best model checkpoint and most of time 5 or 6 is the best.\noptimizer : AdamW\nk : 16. k is the number of ensemble in the finial output. I tried 8, 16, 24, 32, and I found 16 can takes both scores and training time into account.\n### Datapreprocessing\nSimple mean-std standardize and fill zero for nan values.\n### Feature Engeneering & Auxiliary targets\nI use time_id, symbol_id, 78 original features (besides feature_61, because it just change with dates) and 9 responder lags 1 of last record. Additionally, I also use sin and cos of time_id  with 967 periods or 483 periods and feature_61 with 20 period to catch some periodical change. Because though feature_61 only changes with the date, it also changes in about 20 dates. Also, I tried to use more lags features but it work worse in cv and lb.\nBesides responder_6, I use responder_3 as my auxiliary target, because it has high correlations with responder_6. I tried to add more auxiliary targets in my training, but there was no obvious improvement in my model.\n### Data Augmentation\nI add some gaussian noise in continuous features with 0.02 std.\n### Training Sample\nI use the data with dates after 252, because there are too many Nan values in the first 252 dates. I have tried use the dates after 750 to train, which improve my cv but decrease my lb.\n### Scores\nwithout online learning    cv : 0.0106  lb : 0.0077\nwith online learning    cv : 0.0116  lb : 0.0083\n\n# 3. AutoencoderMLP\n### Datapreprocessing\nFill 3 for Nan values without standardization.\n### Feature Engeneering\nWe use time_id, symbol_id, 79 original features and responder_6 lags 1 of last. We didn't use auxilary target in the AutoencoderMLP\n### Model Architecture\nEncoder and Decoder : Both are a lieanr layer with 96 hidden dimension.\nMLP : 5 layers with hidden dimension 96, 896, 448, 448, 256\n### Hyperparameter\nDropout : [0.035, 0.038, 0.424, 0.104, 0.25, 0.32, 0.271, 0.25]\nLearning rate: 1e-3\nBatch size : 8192\nEpoch : about 15\noptimizer : Adam\n### Loss function\nWe use mse loss as reconstruction loss, and weighted-mse loss as prediction loss\n### Scores\nwithout online learning    cv : 0.0103  lb : 0.0072\nwith online learning    cv : 0.0110  lb : 0.0078\n\n# 4. My GBDT and Ridge offline model\n### Feature Engeneering\nIt is more simple now. For LGB, XGB and Ridge, I just use time_id, symbol_id, 79 features, responder lags 1 of last, mean, std and max. And I did not specify category features in the GBDT model.\n### Training Sample\nThe dates after 750 for XGB and Ridge and dates after 678 for LGB \n### Hyperparameter\nI didn't optimize my hyperparameter just use fixed one.\n```\nLGB_Params = {\n        'learning_rate': 0.05,\n        'max_depth': 6,\n        'num_leaves': 62,\n        'n_estimators': 200,\n        'subsample': 0.8,\n        'colsample_bytree': 0.8,\n        'reg_alpha': 1,\n        'reg_lambda': 1,\n        'random_state': 42,\n        'device' : 'gpu',\n        'gpu_use_dp': True,\n        'objective': 'l2',\n    }\n\nXGB_Params = {\n    'learning_rate': 0.05,\n    'max_depth': 6,\n    'n_estimators': 300,\n    'subsample': 0.8,\n    'colsample_bytree': 0.6,\n    'reg_alpha': 1,\n    'reg_lambda': 1,\n    'random_state': 42,\n    'tree_method': 'hist',\n    'device' : 'cuda',\n    'n_gpu': 1,\n    'objective' : 'reg:squarederror',\n}\n```\n### Scores\nLGB   cv: 0.0096 lb : 0.0072\nXGB   cv: 0.0102  lb : 0.0073\nRidge cv: 0.0035 lb : 0.0044\nAfter ensembling these 3 models, the lb is 0.0076.\n\n# 5. @chronoscop's XGB\n### Datapreprocessing\nFill 0 for Nan values without standardization\n### Feature Engeneering\nuse symbol_id, time_id, 79 original feature and responder lags1 of last. If add more lags features, it would be easy to overfit.\n### Training Sample\nuse dates after 917 to train the model.\n### Hyperparameter\n```\nparams = {\n'objective': 'reg:squarederror',\n'random_state': 1212,\n'tree_method': 'hist',\n'device' : 'cuda',\n'learning_rate': 0.02156022412857549,\n'max_depth': 8,\n'subsample': 0.7697954003310141,\n'colsample_bytree': 0.5182134365961873,\n'reg_alpha': 0.0032315937370696354,\n'reg_lambda': 0.002663721647776419\n}\n```\n### Scores\ncv : 0.0102   lb : 0.0072\n\n# 6. Online Learning\nOnline learning play an important role in this competition.\nWe collect data  for each previous date and then update model in every date. For every date, we will combine the data of the last date with partial data from the previous dates (using random sampling) for online training. In the training part, instead of batch processing, we input all the data into the model and then train 5 epochs.\n```\nif previous_test is not None and lags is not None:\n    train = previous_test.join(\n        lags_.select([\"time_id\", \"symbol_id\", pl.col(\"responder_6_lag_1\").alias(\"responder_6\")]),\n        on=[\"time_id\", \"symbol_id\"],\n        how=\"left\"\n    )\n\n    if pre_train is not None and len(pre_train) > 300000:\n        pre_train = pre_train.sample(n=300000, seed=2025)\n    if pre_train is None:\n        pre_train = train\n    else:\n        pre_train = pl.concat([pre_train, train])\n    X = pre_train.to_pandas()[cols].fillna(3).values\n    y = pre_train['responder_6'].to_numpy()\n    weights = pre_train['weight'].to_numpy()\n\n\n    model.train()\n    optimizer = torch.optim.Adam(model.parameters(), lr=1e-4)\n    criterion = nn.MSELoss()\n\n\n    for epoch in range(5):\n        optimizer.zero_grad()\n        _, y_pred = model(torch.FloatTensor(X).to(device))\n        loss = criterion(y_pred.squeeze(), torch.FloatTensor(y.copy()).to(device))\n        loss.backward()\n        optimizer.step()\n        print(f\"Epoch {epoch + 1}/5, Loss: {loss.item()}\")\n\nif previous_test is None:\n    previous_test = test\nelse:\n    previous_test = pl.concat([previous_test, test])\n```\nWe tried to add online learning into GBDT, but it didn't work well and cost too many time.\n\n# Thanks !\nI think TabM can get higher score if adjust different kind of backbones, but I have no time to experiment. I also tried GRU, LSTM or Transformer in the early time, but they all failed, so then I focus on tabular model. Hoping top teams can share more trick and I indeed learn a lot in the competition. Thanks to all of you! :)\nPS: Our code will be released when it is sorted out.\n\n----------------------------------------------------------------------\nUPDATE : \nTabM training code https://www.kaggle.com/code/i2nfinit3y/jane-street-tabm-training?scriptVersionId=217873214\nSingle TabM online learning : https://www.kaggle.com/code/i2nfinit3y/jane-street-tabm-online-learning",
      "votes": null
    },
    {
      "id": "3096310",
      "postDate": "01/14/2025 08:29:54",
      "content": "<p>Great work! Also thank you for the great public Tabm notebook! That helps a lot.</p>",
      "rawMarkdown": "Great work! Also thank you for the great public Tabm notebook! That helps a lot.",
      "votes": null
    },
    {
      "id": "3096386",
      "postDate": "01/14/2025 10:21:41",
      "content": "<p>Thanks for adding the mlp part now! May i know how you select layers as [96, 896, 448, 448, 256], and related dropout? My hyper takes forever to run.</p>\n<p>I straggled with optuna, and grid search also take too much time for getting optimal dropout rate for my mlp with layer [2048, 1024, 512, 256, 128, 64], especially for this large dataset. Each of my mlp search with 30 epoch and full dataset took me about 8 hours to run. </p>",
      "rawMarkdown": "Thanks for adding the mlp part now! May i know how you select layers as [96, 896, 448, 448, 256], and related dropout? My hyper takes forever to run.\n\nI straggled with optuna, and grid search also take too much time for getting optimal dropout rate for my mlp with layer [2048, 1024, 512, 256, 128, 64], especially for this large dataset. Each of my mlp search with 30 epoch and full dataset took me about 8 hours to run.",
      "votes": null
    },
    {
      "id": "3096405",
      "postDate": "01/14/2025 10:55:01",
      "content": "<p>This we do not have a specific selection rule, start from the number of layers, combined with the convergence speed and cv we determined that there are five hidden layers, as for the specific parameters is experience, there hasn't been much adjustment, so I think the mlp part is definitely more than that.🥲</p>",
      "rawMarkdown": "This we do not have a specific selection rule, start from the number of layers, combined with the convergence speed and cv we determined that there are five hidden layers, as for the specific parameters is experience, there hasn't been much adjustment, so I think the mlp part is definitely more than that.🥲",
      "votes": null
    },
    {
      "id": "3096458",
      "postDate": "01/14/2025 12:20:40",
      "content": "<p>got it, seems like i was focusing on the wrong direction the whole time by selecting different layers and batch size. Thanks for the reply.</p>",
      "rawMarkdown": "got it, seems like i was focusing on the wrong direction the whole time by selecting different layers and batch size. Thanks for the reply.",
      "votes": null
    },
    {
      "id": "3096579",
      "postDate": "01/14/2025 14:22:13",
      "content": "<p>In the Autoencoder MLP model, was the autoencoder applied to the 79 original features, or was it specifically used to encode and decode the auxiliary target feature (responder_6 lag 1)?</p>",
      "rawMarkdown": "In the Autoencoder MLP model, was the autoencoder applied to the 79 original features, or was it specifically used to encode and decode the auxiliary target feature (responder_6 lag 1)?",
      "votes": null
    },
    {
      "id": "3096588",
      "postDate": "01/14/2025 14:26:25",
      "content": "<p>We didn't use auxiliary target in the AutoencoderMLP and responder_6_lag_1 was just taken as input feature. Autoencoder was applied to all the 82 features (symbol_id, time_id, 79 original features and responder_6_lag_1)</p>",
      "rawMarkdown": "We didn't use auxiliary target in the AutoencoderMLP and responder_6_lag_1 was just taken as input feature. Autoencoder was applied to all the 82 features (symbol_id, time_id, 79 original features and responder_6_lag_1)",
      "votes": null
    },
    {
      "id": "3096657",
      "postDate": "01/14/2025 15:08:03",
      "content": "<p>I actually have one more question <a href=\"https://www.kaggle.com/lechengyan\" target=\"_blank\">@lechengyan</a> . For the xgb, did you normalize each column as a whole or by symbol id? Was xgb score 0.0073 (and 0.0072) from esembles of kfold groupby date_id or just one single xgb model?</p>",
      "rawMarkdown": "I actually have one more question @lechengyan . For the xgb, did you normalize each column as a whole or by symbol id? Was xgb score 0.0073 (and 0.0072) from esembles of kfold groupby date_id or just one single xgb model?",
      "votes": null
    },
    {
      "id": "3096663",
      "postDate": "01/14/2025 15:12:53",
      "content": "<p>We train 2 xgb model. And the xgb with score 0.0073 used normalization and the other one didn't. We just normalize each column as a whole and it's one single xgb model without kfold split.</p>",
      "rawMarkdown": "We train 2 xgb model. And the xgb with score 0.0073 used normalization and the other one didn't. We just normalize each column as a whole and it's one single xgb model without kfold split.",
      "votes": null
    },
    {
      "id": "3096678",
      "postDate": "01/14/2025 15:20:02",
      "content": "<p>Got it, thanks for the reply!</p>",
      "rawMarkdown": "Got it, thanks for the reply!",
      "votes": null
    },
    {
      "id": "3098864",
      "postDate": "01/17/2025 03:07:15",
      "content": "<blockquote>\n  <p>params = {<br>\n  'objective': 'reg:squarederror',<br>\n  'random_state': 1212,<br>\n  'tree_method': 'hist',<br>\n  'device' : 'cuda',<br>\n  'learning_rate': 0.02156022412857549,<br>\n  'max_depth': 8,<br>\n  'subsample': 0.7697954003310141,<br>\n  'colsample_bytree': 0.5182134365961873,<br>\n  'reg_alpha': 0.0032315937370696354,<br>\n  'reg_lambda': 0.002663721647776419<br>\n  }</p>\n</blockquote>\n<p>It came to my mind when i use optuna to train best xgb model, my submission score is always negative, which is likely due to overfitting. Have you encounter a similar case when tuning your model and how you avoided overfitting?</p>",
      "rawMarkdown": ">params = {\n'objective': 'reg:squarederror',\n'random_state': 1212,\n'tree_method': 'hist',\n'device' : 'cuda',\n'learning_rate': 0.02156022412857549,\n'max_depth': 8,\n'subsample': 0.7697954003310141,\n'colsample_bytree': 0.5182134365961873,\n'reg_alpha': 0.0032315937370696354,\n'reg_lambda': 0.002663721647776419\n}\n\nIt came to my mind when i use optuna to train best xgb model, my submission score is always negative, which is likely due to overfitting. Have you encounter a similar case when tuning your model and how you avoided overfitting?",
      "votes": null
    },
    {
      "id": "3099016",
      "postDate": "01/17/2025 07:24:13",
      "content": "<p>Amazing work! How did you deal with the issue of having different symbols at each time (since MLP requires a fix size), did you expand the database to fill in every (symbol_id, time_id) pair?</p>",
      "rawMarkdown": "Amazing work! How did you deal with the issue of having different symbols at each time (since MLP requires a fix size), did you expand the database to fill in every (symbol_id, time_id) pair?",
      "votes": null
    },
    {
      "id": "3099143",
      "postDate": "01/17/2025 11:11:59",
      "content": "<p>Most of time, if my cv increase a lot (0.001+) when tuning the model,  my lb can also increase. But if cv just increase a little bit (0.0003), lb maybe dont't change too much, so I just judge if my idea work by this. Considering the data is a time-series, you can check if you use the data including the last dates ( after 1000 date ) to train your model.</p>",
      "rawMarkdown": "Most of time, if my cv increase a lot (0.001+) when tuning the model,  my lb can also increase. But if cv just increase a little bit (0.0003), lb maybe dont't change too much, so I just judge if my idea work by this. Considering the data is a time-series, you can check if you use the data including the last dates ( after 1000 date ) to train your model.",
      "votes": null
    },
    {
      "id": "3099148",
      "postDate": "01/17/2025 11:12:59",
      "content": "<p>We just take symbol_id as a feature and the input dimension of model is (batch_size, feature_dim)</p>",
      "rawMarkdown": "We just take symbol_id as a feature and the input dimension of model is (batch_size, feature_dim)",
      "votes": null
    },
    {
      "id": "3099621",
      "postDate": "01/18/2025 02:57:11",
      "content": "<p>Hi everyone,we’ve open-sourced both our inference and training code.</p>\n<p><strong>Inference Code</strong>: You can access our inference code <a href=\"https://www.kaggle.com/code/chronoscop/fork-of-jane-street-rmf-final-submission?scriptVersionId=218125086\" target=\"_blank\">here</a>.<br>\n<strong>Training Code</strong>: Our training code is available on GitHub <a href=\"https://github.com/chronoscop/JS-Public-LB-26th-training-code\" target=\"_blank\">here</a>.</p>",
      "rawMarkdown": "Hi everyone,we’ve open-sourced both our inference and training code.\n\n**Inference Code**: You can access our inference code [here](https://www.kaggle.com/code/chronoscop/fork-of-jane-street-rmf-final-submission?scriptVersionId=218125086).\n**Training Code**: Our training code is available on GitHub [here](https://github.com/chronoscop/JS-Public-LB-26th-training-code).",
      "votes": null
    },
    {
      "id": "3099785",
      "postDate": "01/18/2025 09:12:02",
      "content": "<p>Hey, thanks for sharing! Have you had a chance to try using the date_id batch for NN models? </p>",
      "rawMarkdown": "Hey, thanks for sharing! Have you had a chance to try using the date_id batch for NN models?",
      "votes": null
    },
    {
      "id": "3099876",
      "postDate": "01/18/2025 12:25:12",
      "content": "<p>I tuned xgb again with the latest 1000 days before cv. It took some time to run but still got negative R2 cv score. I will play around with it but your idea to use more recent data seems to be the right direction for me. Thanks for the reply and posing your training code (not many people did that)!</p>",
      "rawMarkdown": "I tuned xgb again with the latest 1000 days before cv. It took some time to run but still got negative R2 cv score. I will play around with it but your idea to use more recent data seems to be the right direction for me. Thanks for the reply and posing your training code (not many people did that)!",
      "votes": null
    },
    {
      "id": "3099881",
      "postDate": "01/18/2025 12:40:58",
      "content": "<p>Not yet. But I wil try. It seems to be a good idea.</p>",
      "rawMarkdown": "Not yet. But I wil try. It seems to be a good idea.",
      "votes": null
    },
    {
      "id": "3100532",
      "postDate": "01/19/2025 13:56:48",
      "content": "<p>Hi, thanks a lot for sharing your training script along with the inference one for xgb. Please could you also share your optuna tuning script. I couldn't get rid of my negative R2 in cv with my hyper tuned xgb param after several tries after the deadline. Thanks.</p>",
      "rawMarkdown": "Hi, thanks a lot for sharing your training script along with the inference one for xgb. Please could you also share your optuna tuning script. I couldn't get rid of my negative R2 in cv with my hyper tuned xgb param after several tries after the deadline. Thanks.",
      "votes": null
    },
    {
      "id": "3100846",
      "postDate": "01/20/2025 02:07:01",
      "content": "<p>Hi <a href=\"https://www.kaggle.com/yimin218\" target=\"_blank\">@yimin218</a>, I’m just using RMSE as the metric when optimizing with Optuna.</p>\n<p>I open scource the training srcipt you can see at <a href=\"https://www.kaggle.com/code/chronoscop/optuna-find-params/notebook\" target=\"_blank\">here</a></p>",
      "rawMarkdown": "Hi @yimin218, I’m just using RMSE as the metric when optimizing with Optuna.\n\nI open scource the training srcipt you can see at [here](https://www.kaggle.com/code/chronoscop/optuna-find-params/notebook)",
      "votes": null
    },
    {
      "id": "3100856",
      "postDate": "01/20/2025 02:31:58",
      "content": "<p>Figured out. It was due to my improper scaler in prediction function. Thank you very much!</p>",
      "rawMarkdown": "Figured out. It was due to my improper scaler in prediction function. Thank you very much!",
      "votes": null
    },
    {
      "id": "3102424",
      "postDate": "01/22/2025 05:53:02",
      "content": "<p>update : I have tried used a date_id batch for my tabm model, but there is no boost. But when I tried another sequence model, it improve the model a lot. I guess this method is better for sequence model</p>",
      "rawMarkdown": "update : I have tried used a date_id batch for my tabm model, but there is no boost. But when I tried another sequence model, it improve the model a lot. I guess this method is better for sequence model",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 3096310,
      "author_name": "yich723",
      "author_url": "",
      "post_date": "01/14/2025 08:29:54",
      "content": "<p>Great work! Also thank you for the great public Tabm notebook! That helps a lot.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 3096386,
      "author_name": "yimin218",
      "author_url": "",
      "post_date": "01/14/2025 10:21:41",
      "content": "<p>Thanks for adding the mlp part now! May i know how you select layers as [96, 896, 448, 448, 256], and related dropout? My hyper takes forever to run.</p>\n<p>I straggled with optuna, and grid search also take too much time for getting optimal dropout rate for my mlp with layer [2048, 1024, 512, 256, 128, 64], especially for this large dataset. Each of my mlp search with 30 epoch and full dataset took me about 8 hours to run. </p>",
      "votes": null,
      "replies": [
        {
          "id": 3096405,
          "author_name": "lechengyan",
          "author_url": "",
          "post_date": "01/14/2025 10:55:01",
          "content": "<p>This we do not have a specific selection rule, start from the number of layers, combined with the convergence speed and cv we determined that there are five hidden layers, as for the specific parameters is experience, there hasn't been much adjustment, so I think the mlp part is definitely more than that.🥲</p>",
          "votes": null,
          "replies": [
            {
              "id": 3096458,
              "author_name": "yimin218",
              "author_url": "",
              "post_date": "01/14/2025 12:20:40",
              "content": "<p>got it, seems like i was focusing on the wrong direction the whole time by selecting different layers and batch size. Thanks for the reply.</p>",
              "votes": null,
              "replies": []
            },
            {
              "id": 3096657,
              "author_name": "yimin218",
              "author_url": "",
              "post_date": "01/14/2025 15:08:03",
              "content": "<p>I actually have one more question <a href=\"https://www.kaggle.com/lechengyan\" target=\"_blank\">@lechengyan</a> . For the xgb, did you normalize each column as a whole or by symbol id? Was xgb score 0.0073 (and 0.0072) from esembles of kfold groupby date_id or just one single xgb model?</p>",
              "votes": null,
              "replies": [
                {
                  "id": 3096663,
                  "author_name": "i2nfinit3y",
                  "author_url": "",
                  "post_date": "01/14/2025 15:12:53",
                  "content": "<p>We train 2 xgb model. And the xgb with score 0.0073 used normalization and the other one didn't. We just normalize each column as a whole and it's one single xgb model without kfold split.</p>",
                  "votes": null,
                  "replies": [
                    {
                      "id": 3096678,
                      "author_name": "yimin218",
                      "author_url": "",
                      "post_date": "01/14/2025 15:20:02",
                      "content": "<p>Got it, thanks for the reply!</p>",
                      "votes": null,
                      "replies": []
                    }
                  ]
                }
              ]
            }
          ]
        }
      ]
    },
    {
      "id": 3096579,
      "author_name": "byunjins",
      "author_url": "",
      "post_date": "01/14/2025 14:22:13",
      "content": "<p>In the Autoencoder MLP model, was the autoencoder applied to the 79 original features, or was it specifically used to encode and decode the auxiliary target feature (responder_6 lag 1)?</p>",
      "votes": null,
      "replies": [
        {
          "id": 3096588,
          "author_name": "i2nfinit3y",
          "author_url": "",
          "post_date": "01/14/2025 14:26:25",
          "content": "<p>We didn't use auxiliary target in the AutoencoderMLP and responder_6_lag_1 was just taken as input feature. Autoencoder was applied to all the 82 features (symbol_id, time_id, 79 original features and responder_6_lag_1)</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 3098864,
      "author_name": "yimin218",
      "author_url": "",
      "post_date": "01/17/2025 03:07:15",
      "content": "<blockquote>\n  <p>params = {<br>\n  'objective': 'reg:squarederror',<br>\n  'random_state': 1212,<br>\n  'tree_method': 'hist',<br>\n  'device' : 'cuda',<br>\n  'learning_rate': 0.02156022412857549,<br>\n  'max_depth': 8,<br>\n  'subsample': 0.7697954003310141,<br>\n  'colsample_bytree': 0.5182134365961873,<br>\n  'reg_alpha': 0.0032315937370696354,<br>\n  'reg_lambda': 0.002663721647776419<br>\n  }</p>\n</blockquote>\n<p>It came to my mind when i use optuna to train best xgb model, my submission score is always negative, which is likely due to overfitting. Have you encounter a similar case when tuning your model and how you avoided overfitting?</p>",
      "votes": null,
      "replies": [
        {
          "id": 3099143,
          "author_name": "i2nfinit3y",
          "author_url": "",
          "post_date": "01/17/2025 11:11:59",
          "content": "<p>Most of time, if my cv increase a lot (0.001+) when tuning the model,  my lb can also increase. But if cv just increase a little bit (0.0003), lb maybe dont't change too much, so I just judge if my idea work by this. Considering the data is a time-series, you can check if you use the data including the last dates ( after 1000 date ) to train your model.</p>",
          "votes": null,
          "replies": [
            {
              "id": 3099876,
              "author_name": "yimin218",
              "author_url": "",
              "post_date": "01/18/2025 12:25:12",
              "content": "<p>I tuned xgb again with the latest 1000 days before cv. It took some time to run but still got negative R2 cv score. I will play around with it but your idea to use more recent data seems to be the right direction for me. Thanks for the reply and posing your training code (not many people did that)!</p>",
              "votes": null,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 3099016,
      "author_name": "leomok",
      "author_url": "",
      "post_date": "01/17/2025 07:24:13",
      "content": "<p>Amazing work! How did you deal with the issue of having different symbols at each time (since MLP requires a fix size), did you expand the database to fill in every (symbol_id, time_id) pair?</p>",
      "votes": null,
      "replies": [
        {
          "id": 3099148,
          "author_name": "i2nfinit3y",
          "author_url": "",
          "post_date": "01/17/2025 11:12:59",
          "content": "<p>We just take symbol_id as a feature and the input dimension of model is (batch_size, feature_dim)</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 3099621,
      "author_name": "chronoscop",
      "author_url": "",
      "post_date": "01/18/2025 02:57:11",
      "content": "<p>Hi everyone,we’ve open-sourced both our inference and training code.</p>\n<p><strong>Inference Code</strong>: You can access our inference code <a href=\"https://www.kaggle.com/code/chronoscop/fork-of-jane-street-rmf-final-submission?scriptVersionId=218125086\" target=\"_blank\">here</a>.<br>\n<strong>Training Code</strong>: Our training code is available on GitHub <a href=\"https://github.com/chronoscop/JS-Public-LB-26th-training-code\" target=\"_blank\">here</a>.</p>",
      "votes": null,
      "replies": [
        {
          "id": 3100532,
          "author_name": "yimin218",
          "author_url": "",
          "post_date": "01/19/2025 13:56:48",
          "content": "<p>Hi, thanks a lot for sharing your training script along with the inference one for xgb. Please could you also share your optuna tuning script. I couldn't get rid of my negative R2 in cv with my hyper tuned xgb param after several tries after the deadline. Thanks.</p>",
          "votes": null,
          "replies": [
            {
              "id": 3100846,
              "author_name": "chronoscop",
              "author_url": "",
              "post_date": "01/20/2025 02:07:01",
              "content": "<p>Hi <a href=\"https://www.kaggle.com/yimin218\" target=\"_blank\">@yimin218</a>, I’m just using RMSE as the metric when optimizing with Optuna.</p>\n<p>I open scource the training srcipt you can see at <a href=\"https://www.kaggle.com/code/chronoscop/optuna-find-params/notebook\" target=\"_blank\">here</a></p>",
              "votes": null,
              "replies": [
                {
                  "id": 3100856,
                  "author_name": "yimin218",
                  "author_url": "",
                  "post_date": "01/20/2025 02:31:58",
                  "content": "<p>Figured out. It was due to my improper scaler in prediction function. Thank you very much!</p>",
                  "votes": null,
                  "replies": []
                }
              ]
            }
          ]
        }
      ]
    },
    {
      "id": 3099785,
      "author_name": "alexeigor",
      "author_url": "",
      "post_date": "01/18/2025 09:12:02",
      "content": "<p>Hey, thanks for sharing! Have you had a chance to try using the date_id batch for NN models? </p>",
      "votes": null,
      "replies": [
        {
          "id": 3099881,
          "author_name": "i2nfinit3y",
          "author_url": "",
          "post_date": "01/18/2025 12:40:58",
          "content": "<p>Not yet. But I wil try. It seems to be a good idea.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 3102424,
          "author_name": "i2nfinit3y",
          "author_url": "",
          "post_date": "01/22/2025 05:53:02",
          "content": "<p>update : I have tried used a date_id batch for my tabm model, but there is no boost. But when I tried another sequence model, it improve the model a lot. I guess this method is better for sequence model</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "3096299": "Thanks to this amazing competition and the efforts of every one of my teammates ! @lechengyan @chronoscop\nOur final solution of 0.0092 lb is ensembling NN models with online learning , GBDT offline models and a ridge model. My part is mainly for TabM model and some GBDT models, which is what I am gonna to talk about. AutoencoderMLP and online learning are @lechengyan 's part and @chronoscop is in charge of one of XGB.\n\n# 1. Cross-Validation\nI simply used the last 120 dates as my validation and it shows good correlations with LB.\n\n# 2. TabM Model\nActually, in the early time of this competition, I noticed TabM and found it is a tabular NN model with great potential. I also public a baseline [notebook](https://www.kaggle.com/code/i2nfinit3y/jane-street-tabm-ft-transformer-inference).\n### Model Architecture\nThe model parts I mainly adjust are feature embed layers and backbone.\nFor category features, I use onehot encoding for every category tensor, which contains feature_09, feature_10, feature_11, symbol_id and time_id. I remap these category features with ordinal encode and every new category  in the test data will be mapped to the same category. Apart from one-hot encoding, I also use embedding layers but it didn't help much. It's worth noting that taking time_id as category features improve my model al lot.\nBut for continuous features, I have tried LinearEmbeddings, PeriodicEmbeddings and PiecewiseLinearEmbeddings, which didn't work well. So I don't use any embedding layer for continuous features.\nThe backbone of TabM is just 3 layers MLP with the same 512 dimensions. I haven't experimented with too many dimensional combinations, but 512 seems to be the most stable.\n### Loss function\nI have used huber loss, logcosh loss, mae loss, zero-mean R2 loss and mse loss, and R2 loss and mse loss work best. Specifically, mse loss can get higher score in the cv but R2 loss is more robust in every validation epoch. So I choose R2 loss as loss function.\n### Hyperparameter\nDropout : 0.25. Higher dropout rate (0.5) will result in slower convergence and lower (0.1) will be overfitting, so 0.25 seems to be a good choice for me.\nLearning rate: 1e-3.\nWeight decay: 8e-4.\nBatch size : 8192\nepoch : 5 or 6. Generally, I will submit multiple epoch to confirm the best model checkpoint and most of time 5 or 6 is the best.\noptimizer : AdamW\nk : 16. k is the number of ensemble in the finial output. I tried 8, 16, 24, 32, and I found 16 can takes both scores and training time into account.\n### Datapreprocessing\nSimple mean-std standardize and fill zero for nan values.\n### Feature Engeneering & Auxiliary targets\nI use time_id, symbol_id, 78 original features (besides feature_61, because it just change with dates) and 9 responder lags 1 of last record. Additionally, I also use sin and cos of time_id  with 967 periods or 483 periods and feature_61 with 20 period to catch some periodical change. Because though feature_61 only changes with the date, it also changes in about 20 dates. Also, I tried to use more lags features but it work worse in cv and lb.\nBesides responder_6, I use responder_3 as my auxiliary target, because it has high correlations with responder_6. I tried to add more auxiliary targets in my training, but there was no obvious improvement in my model.\n### Data Augmentation\nI add some gaussian noise in continuous features with 0.02 std.\n### Training Sample\nI use the data with dates after 252, because there are too many Nan values in the first 252 dates. I have tried use the dates after 750 to train, which improve my cv but decrease my lb.\n### Scores\nwithout online learning    cv : 0.0106  lb : 0.0077\nwith online learning    cv : 0.0116  lb : 0.0083\n\n# 3. AutoencoderMLP\n### Datapreprocessing\nFill 3 for Nan values without standardization.\n### Feature Engeneering\nWe use time_id, symbol_id, 79 original features and responder_6 lags 1 of last. We didn't use auxilary target in the AutoencoderMLP\n### Model Architecture\nEncoder and Decoder : Both are a lieanr layer with 96 hidden dimension.\nMLP : 5 layers with hidden dimension 96, 896, 448, 448, 256\n### Hyperparameter\nDropout : [0.035, 0.038, 0.424, 0.104, 0.25, 0.32, 0.271, 0.25]\nLearning rate: 1e-3\nBatch size : 8192\nEpoch : about 15\noptimizer : Adam\n### Loss function\nWe use mse loss as reconstruction loss, and weighted-mse loss as prediction loss\n### Scores\nwithout online learning    cv : 0.0103  lb : 0.0072\nwith online learning    cv : 0.0110  lb : 0.0078\n\n# 4. My GBDT and Ridge offline model\n### Feature Engeneering\nIt is more simple now. For LGB, XGB and Ridge, I just use time_id, symbol_id, 79 features, responder lags 1 of last, mean, std and max. And I did not specify category features in the GBDT model.\n### Training Sample\nThe dates after 750 for XGB and Ridge and dates after 678 for LGB \n### Hyperparameter\nI didn't optimize my hyperparameter just use fixed one.\n```\nLGB_Params = {\n        'learning_rate': 0.05,\n        'max_depth': 6,\n        'num_leaves': 62,\n        'n_estimators': 200,\n        'subsample': 0.8,\n        'colsample_bytree': 0.8,\n        'reg_alpha': 1,\n        'reg_lambda': 1,\n        'random_state': 42,\n        'device' : 'gpu',\n        'gpu_use_dp': True,\n        'objective': 'l2',\n    }\n\nXGB_Params = {\n    'learning_rate': 0.05,\n    'max_depth': 6,\n    'n_estimators': 300,\n    'subsample': 0.8,\n    'colsample_bytree': 0.6,\n    'reg_alpha': 1,\n    'reg_lambda': 1,\n    'random_state': 42,\n    'tree_method': 'hist',\n    'device' : 'cuda',\n    'n_gpu': 1,\n    'objective' : 'reg:squarederror',\n}\n```\n### Scores\nLGB   cv: 0.0096 lb : 0.0072\nXGB   cv: 0.0102  lb : 0.0073\nRidge cv: 0.0035 lb : 0.0044\nAfter ensembling these 3 models, the lb is 0.0076.\n\n# 5. @chronoscop's XGB\n### Datapreprocessing\nFill 0 for Nan values without standardization\n### Feature Engeneering\nuse symbol_id, time_id, 79 original feature and responder lags1 of last. If add more lags features, it would be easy to overfit.\n### Training Sample\nuse dates after 917 to train the model.\n### Hyperparameter\n```\nparams = {\n'objective': 'reg:squarederror',\n'random_state': 1212,\n'tree_method': 'hist',\n'device' : 'cuda',\n'learning_rate': 0.02156022412857549,\n'max_depth': 8,\n'subsample': 0.7697954003310141,\n'colsample_bytree': 0.5182134365961873,\n'reg_alpha': 0.0032315937370696354,\n'reg_lambda': 0.002663721647776419\n}\n```\n### Scores\ncv : 0.0102   lb : 0.0072\n\n# 6. Online Learning\nOnline learning play an important role in this competition.\nWe collect data  for each previous date and then update model in every date. For every date, we will combine the data of the last date with partial data from the previous dates (using random sampling) for online training. In the training part, instead of batch processing, we input all the data into the model and then train 5 epochs.\n```\nif previous_test is not None and lags is not None:\n    train = previous_test.join(\n        lags_.select([\"time_id\", \"symbol_id\", pl.col(\"responder_6_lag_1\").alias(\"responder_6\")]),\n        on=[\"time_id\", \"symbol_id\"],\n        how=\"left\"\n    )\n\n    if pre_train is not None and len(pre_train) > 300000:\n        pre_train = pre_train.sample(n=300000, seed=2025)\n    if pre_train is None:\n        pre_train = train\n    else:\n        pre_train = pl.concat([pre_train, train])\n    X = pre_train.to_pandas()[cols].fillna(3).values\n    y = pre_train['responder_6'].to_numpy()\n    weights = pre_train['weight'].to_numpy()\n\n\n    model.train()\n    optimizer = torch.optim.Adam(model.parameters(), lr=1e-4)\n    criterion = nn.MSELoss()\n\n\n    for epoch in range(5):\n        optimizer.zero_grad()\n        _, y_pred = model(torch.FloatTensor(X).to(device))\n        loss = criterion(y_pred.squeeze(), torch.FloatTensor(y.copy()).to(device))\n        loss.backward()\n        optimizer.step()\n        print(f\"Epoch {epoch + 1}/5, Loss: {loss.item()}\")\n\nif previous_test is None:\n    previous_test = test\nelse:\n    previous_test = pl.concat([previous_test, test])\n```\nWe tried to add online learning into GBDT, but it didn't work well and cost too many time.\n\n# Thanks !\nI think TabM can get higher score if adjust different kind of backbones, but I have no time to experiment. I also tried GRU, LSTM or Transformer in the early time, but they all failed, so then I focus on tabular model. Hoping top teams can share more trick and I indeed learn a lot in the competition. Thanks to all of you! :)\nPS: Our code will be released when it is sorted out.\n\n----------------------------------------------------------------------\nUPDATE : \nTabM training code https://www.kaggle.com/code/i2nfinit3y/jane-street-tabm-training?scriptVersionId=217873214\nSingle TabM online learning : https://www.kaggle.com/code/i2nfinit3y/jane-street-tabm-online-learning",
    "3096310": "Great work! Also thank you for the great public Tabm notebook! That helps a lot.",
    "3096386": "Thanks for adding the mlp part now! May i know how you select layers as [96, 896, 448, 448, 256], and related dropout? My hyper takes forever to run.\n\nI straggled with optuna, and grid search also take too much time for getting optimal dropout rate for my mlp with layer [2048, 1024, 512, 256, 128, 64], especially for this large dataset. Each of my mlp search with 30 epoch and full dataset took me about 8 hours to run.",
    "3096405": "This we do not have a specific selection rule, start from the number of layers, combined with the convergence speed and cv we determined that there are five hidden layers, as for the specific parameters is experience, there hasn't been much adjustment, so I think the mlp part is definitely more than that.🥲",
    "3096458": "got it, seems like i was focusing on the wrong direction the whole time by selecting different layers and batch size. Thanks for the reply.",
    "3096579": "In the Autoencoder MLP model, was the autoencoder applied to the 79 original features, or was it specifically used to encode and decode the auxiliary target feature (responder_6 lag 1)?",
    "3096588": "We didn't use auxiliary target in the AutoencoderMLP and responder_6_lag_1 was just taken as input feature. Autoencoder was applied to all the 82 features (symbol_id, time_id, 79 original features and responder_6_lag_1)",
    "3096657": "I actually have one more question @lechengyan . For the xgb, did you normalize each column as a whole or by symbol id? Was xgb score 0.0073 (and 0.0072) from esembles of kfold groupby date_id or just one single xgb model?",
    "3096663": "We train 2 xgb model. And the xgb with score 0.0073 used normalization and the other one didn't. We just normalize each column as a whole and it's one single xgb model without kfold split.",
    "3096678": "Got it, thanks for the reply!",
    "3098864": ">params = {\n'objective': 'reg:squarederror',\n'random_state': 1212,\n'tree_method': 'hist',\n'device' : 'cuda',\n'learning_rate': 0.02156022412857549,\n'max_depth': 8,\n'subsample': 0.7697954003310141,\n'colsample_bytree': 0.5182134365961873,\n'reg_alpha': 0.0032315937370696354,\n'reg_lambda': 0.002663721647776419\n}\n\nIt came to my mind when i use optuna to train best xgb model, my submission score is always negative, which is likely due to overfitting. Have you encounter a similar case when tuning your model and how you avoided overfitting?",
    "3099016": "Amazing work! How did you deal with the issue of having different symbols at each time (since MLP requires a fix size), did you expand the database to fill in every (symbol_id, time_id) pair?",
    "3099143": "Most of time, if my cv increase a lot (0.001+) when tuning the model,  my lb can also increase. But if cv just increase a little bit (0.0003), lb maybe dont't change too much, so I just judge if my idea work by this. Considering the data is a time-series, you can check if you use the data including the last dates ( after 1000 date ) to train your model.",
    "3099148": "We just take symbol_id as a feature and the input dimension of model is (batch_size, feature_dim)",
    "3099621": "Hi everyone,we’ve open-sourced both our inference and training code.\n\n**Inference Code**: You can access our inference code [here](https://www.kaggle.com/code/chronoscop/fork-of-jane-street-rmf-final-submission?scriptVersionId=218125086).\n**Training Code**: Our training code is available on GitHub [here](https://github.com/chronoscop/JS-Public-LB-26th-training-code).",
    "3099785": "Hey, thanks for sharing! Have you had a chance to try using the date_id batch for NN models?",
    "3099876": "I tuned xgb again with the latest 1000 days before cv. It took some time to run but still got negative R2 cv score. I will play around with it but your idea to use more recent data seems to be the right direction for me. Thanks for the reply and posing your training code (not many people did that)!",
    "3099881": "Not yet. But I wil try. It seems to be a good idea.",
    "3100532": "Hi, thanks a lot for sharing your training script along with the inference one for xgb. Please could you also share your optuna tuning script. I couldn't get rid of my negative R2 in cv with my hyper tuned xgb param after several tries after the deadline. Thanks.",
    "3100846": "Hi @yimin218, I’m just using RMSE as the metric when optimizing with Optuna.\n\nI open scource the training srcipt you can see at [here](https://www.kaggle.com/code/chronoscop/optuna-find-params/notebook)",
    "3100856": "Figured out. It was due to my improper scaler in prediction function. Thank you very much!",
    "3102424": "update : I have tried used a date_id batch for my tabm model, but there is no boost. But when I tried another sequence model, it improve the model a lot. I guess this method is better for sequence model"
  },
  "source": "meta"
}