{
  "id": 556558,
  "title": "Public 195th: Our Efforts Behind a Modest Score",
  "url": "/competitions/jane-street-real-time-market-data-forecasting/discussion/556558",
  "author_name": "ganchan",
  "post_date": "2025-01-14T03:03:36.210000",
  "votes": 14,
  "comment_count": 18,
  "views": 0,
  "content": "<p>I found this competition a bit challenging due to the Time Series API.<br>\nHowever, I learned a lot through this experience, and I’d like to express my gratitude to the competition hosts and all the participants!<br>\nOur score isn’t very high (upper bronze), but we’re sharing our solution because comparing different approaches can be incredibly valuable.</p>\n<p>Our final submission is an ensemble of:</p>\n<p>LGBM (online training) LB 0.0070<br>\nMLP (online training) LB 0.0070<br>\nTabM (online training) LB 0.0075<br>\n⇒ Ensemble LB 0.0083<br>\nOnly TabM uses lag features and symbol_id, while the other models rely solely on time_id and feature_00 to feature_79.<br>\nOnline training was performed every 0.4M rows.<br>\nDuring online training, we randomly sampled and added the same amount of data from train.parquet. This proved effective in improving our score.<br>\nMLP and TabM were retrained for only 1 epoch. Increasing the number of epochs didn’t work well for us.<br>\nWithout online training, our score was only around 0.0065.<br>\nWe believe this is due to either insufficient features or suboptimal model architecture, so we look forward to learning from the solutions shared by others.</p>\n<p>Looking back, I think there’s still plenty of room for improvement, and I’m excited to read the top solutions!</p>\n<p>Thank you.</p>",
  "messages": [
    {
      "id": 3096038,
      "postDate": "2025-01-14T03:03:36.210Z",
      "content": "<p>I found this competition a bit challenging due to the Time Series API.<br>\nHowever, I learned a lot through this experience, and I’d like to express my gratitude to the competition hosts and all the participants!<br>\nOur score isn’t very high (upper bronze), but we’re sharing our solution because comparing different approaches can be incredibly valuable.</p>\n<p>Our final submission is an ensemble of:</p>\n<p>LGBM (online training) LB 0.0070<br>\nMLP (online training) LB 0.0070<br>\nTabM (online training) LB 0.0075<br>\n⇒ Ensemble LB 0.0083<br>\nOnly TabM uses lag features and symbol_id, while the other models rely solely on time_id and feature_00 to feature_79.<br>\nOnline training was performed every 0.4M rows.<br>\nDuring online training, we randomly sampled and added the same amount of data from train.parquet. This proved effective in improving our score.<br>\nMLP and TabM were retrained for only 1 epoch. Increasing the number of epochs didn’t work well for us.<br>\nWithout online training, our score was only around 0.0065.<br>\nWe believe this is due to either insufficient features or suboptimal model architecture, so we look forward to learning from the solutions shared by others.</p>\n<p>Looking back, I think there’s still plenty of room for improvement, and I’m excited to read the top solutions!</p>\n<p>Thank you.</p>",
      "rawMarkdown": "I found this competition a bit challenging due to the Time Series API.\nHowever, I learned a lot through this experience, and I’d like to express my gratitude to the competition hosts and all the participants!\nOur score isn’t very high (upper bronze), but we’re sharing our solution because comparing different approaches can be incredibly valuable.\n\nOur final submission is an ensemble of:\n\nLGBM (online training) LB 0.0070\nMLP (online training) LB 0.0070\nTabM (online training) LB 0.0075\n⇒ Ensemble LB 0.0083\nOnly TabM uses lag features and symbol_id, while the other models rely solely on time_id and feature_00 to feature_79.\nOnline training was performed every 0.4M rows.\nDuring online training, we randomly sampled and added the same amount of data from train.parquet. This proved effective in improving our score.\nMLP and TabM were retrained for only 1 epoch. Increasing the number of epochs didn’t work well for us.\nWithout online training, our score was only around 0.0065.\nWe believe this is due to either insufficient features or suboptimal model architecture, so we look forward to learning from the solutions shared by others.\n\nLooking back, I think there’s still plenty of room for improvement, and I’m excited to read the top solutions!\n\nThank you.",
      "votes": 14
    },
    {
      "id": 3096332,
      "postDate": "2025-01-14T09:12:57.890Z",
      "content": "<p>PS.<br>\nIn our case, online training with new data and original train data (sampling) improved score.<br>\nOnline training with only new data did not work well.</p>",
      "rawMarkdown": "PS.\nIn our case, online training with new data and original train data (sampling) improved score.\nOnline training with only new data did not work well.",
      "replies": [
        {
          "id": 3097494,
          "postDate": "2025-01-15T12:37:55.090Z",
          "content": "<p>I also sampled the last 480 days of data and trained the online MLP. My single MLP online score is 0.0077, with hidden layer [512,512,256,256,128,128,64,32], dropout[0.4,0.4,0.3,0.3,0.2,0.2,0.1,0.0], lr 5e-5, and epoch 30, batchsize 8192. However, it is a pity that I submitted this result for the last time and failed to integrate with my offline gbdt. The single score of my previous one integrated with offline gbdt in online MLP was 0.0071 and the fusion score was 0.0083</p>",
          "rawMarkdown": "I also sampled the last 480 days of data and trained the online MLP. My single MLP online score is 0.0077, with hidden layer [512,512,256,256,128,128,64,32], dropout[0.4,0.4,0.3,0.3,0.2,0.2,0.1,0.0], lr 5e-5, and epoch 30, batchsize 8192. However, it is a pity that I submitted this result for the last time and failed to integrate with my offline gbdt. The single score of my previous one integrated with offline gbdt in online MLP was 0.0071 and the fusion score was 0.0083"
        },
        {
          "id": 3097606,
          "postDate": "2025-01-15T14:30:44.740Z",
          "content": "<p>Very interesting, I tried lgb only with new data with many different larameters and indeed did not see any improvements. Did you use init_model? And if I may ask how did you tweak the learning rate and number of boosting rounds?</p>",
          "rawMarkdown": "Very interesting, I tried lgb only with new data with many different larameters and indeed did not see any improvements. Did you use init_model? And if I may ask how did you tweak the learning rate and number of boosting rounds?"
        }
      ]
    },
    {
      "id": 3096210,
      "postDate": "2025-01-14T06:22:16.253Z",
      "content": "<p>how did you train LGBM to 0.007? I can only get up to 0.005. Any insight you could share?</p>",
      "rawMarkdown": "how did you train LGBM to 0.007? I can only get up to 0.005. Any insight you could share?",
      "replies": [
        {
          "id": 3096223,
          "postDate": "2025-01-14T06:28:47.183Z",
          "content": "<p>I used all data for xgb, with large n_estimator and small lr. i used ['time_id', 'sybmol_id'] + 79 features + 8 lags, and normalize all the 89 features. Without online learning, my offline xgb gets 0.0070. </p>\n<p>Hopefully others can bring their findings to improve the gbt offline score here.</p>",
          "rawMarkdown": "I used all data for xgb, with large n_estimator and small lr. i used ['time_id', 'sybmol_id'] + 79 features + 8 lags, and normalize all the 89 features. Without online learning, my offline xgb gets 0.0070. \n\nHopefully others can bring their findings to improve the gbt offline score here.",
          "replies": [
            {
              "id": 3096257,
              "postDate": "2025-01-14T07:11:48.073Z",
              "content": "<p>Appriciate your reply. <br>\nWas your normalization per symbol id or was it just whole norm for the entire column?</p>",
              "rawMarkdown": "Appriciate your reply. \nWas your normalization per symbol id or was it just whole norm for the entire column?\n"
            },
            {
              "id": 3096260,
              "postDate": "2025-01-14T07:18:54.383Z",
              "content": "<p>just the whole norm for the entire column. I used 5 groupk fold  by date_id, otherwise a single model won't get above 0.0068</p>",
              "rawMarkdown": "just the whole norm for the entire column. I used 5 groupk fold  by date_id, otherwise a single model won't get above 0.0068"
            }
          ]
        },
        {
          "id": 3096258,
          "postDate": "2025-01-14T07:13:11.567Z",
          "content": "<p>Actually, there’s no magic here.<br>\nI simply trained the model using all features and time_id, achieving an offline score of 0.0063. (online: 0.0070)</p>\n<p>lgb_params = {<br>\n    \"objective\": \"regression\",<br>\n    \"boosting\": \"gbdt\",<br>\n    \"max_depth\": -1,<br>\n    \"num_leaves\": 40,<br>\n    \"subsample\": 0.8,<br>\n    \"subsample_freq\": 1,<br>\n    \"bagging_seed\": rand,<br>\n    \"learning_rate\": 0.05,<br>\n    \"feature_fraction\": 0.6,<br>\n    \"min_data_in_leaf\": 100,<br>\n    \"lambda_l1\": 0,<br>\n    \"lambda_l2\": 0,<br>\n    \"random_state\": rand,<br>\n    \"metric\": \"rmse\",<br>\n}</p>",
          "rawMarkdown": "Actually, there’s no magic here.\nI simply trained the model using all features and time_id, achieving an offline score of 0.0063. (online: 0.0070)\n\nlgb_params = {\n    \"objective\": \"regression\",\n    \"boosting\": \"gbdt\",\n    \"max_depth\": -1,\n    \"num_leaves\": 40,\n    \"subsample\": 0.8,\n    \"subsample_freq\": 1,\n    \"bagging_seed\": rand,\n    \"learning_rate\": 0.05,\n    \"feature_fraction\": 0.6,\n    \"min_data_in_leaf\": 100,\n    \"lambda_l1\": 0,\n    \"lambda_l2\": 0,\n    \"random_state\": rand,\n    \"metric\": \"rmse\",\n}",
          "votes": 2,
          "replies": [
            {
              "id": 3096277,
              "postDate": "2025-01-14T07:31:28.370Z",
              "content": "<p>thanks for sharing</p>",
              "rawMarkdown": "thanks for sharing"
            }
          ]
        },
        {
          "id": 3098151,
          "postDate": "2025-01-16T06:41:14.090Z",
          "content": "<p>add time_id in features, only offline lgbm  can 0.007 +</p>",
          "rawMarkdown": "add time_id in features, only offline lgbm  can 0.007 +"
        }
      ]
    },
    {
      "id": 3096112,
      "postDate": "2025-01-14T05:05:06.367Z",
      "content": "<p>Nice post. What learning rate, epoch, and batch size for mlp and tabm did you guys used for online learning? Mine was 1e- 3 with 819600 as batch size. Surprisingly a small lr as fine tune didn't work for me</p>",
      "rawMarkdown": "Nice post. What learning rate, epoch, and batch size for mlp and tabm did you guys used for online learning? Mine was 1e- 3 with 819600 as batch size. Surprisingly a small lr as fine tune didn't work for me",
      "replies": [
        {
          "id": 3096126,
          "postDate": "2025-01-14T05:16:58.480Z",
          "content": "<p>Our parameters are as follows:<br>\nThe learning rates were determined through individual experiments but were set to the same value (1e-4).<br>\nIncreasing the learning rate resulted in worse scores.</p>\n<p><strong>NN(MLP)</strong><br>\nlearning rate: 1e-4<br>\nepoch: 1 (retrain every 0.4M new data)<br>\nbatch size: 8192<br>\noptimizer: Adam</p>\n<p><strong>TabM</strong><br>\nlearning rate: 1e-4<br>\nepoch: 1 (retrain every 0.4M new data)<br>\nbatch size: 8192<br>\noptimizer: AdamW</p>",
          "rawMarkdown": "Our parameters are as follows:\nThe learning rates were determined through individual experiments but were set to the same value (1e-4).\nIncreasing the learning rate resulted in worse scores.\n\n**NN(MLP)**\nlearning rate: 1e-4\nepoch: 1 (retrain every 0.4M new data)\nbatch size: 8192\noptimizer: Adam\n\n**TabM**\nlearning rate: 1e-4\nepoch: 1 (retrain every 0.4M new data)\nbatch size: 8192\noptimizer: AdamW",
          "votes": 1,
          "replies": [
            {
              "id": 3096165,
              "postDate": "2025-01-14T06:10:26.207Z",
              "content": "<p>my learning rate and batch size were just way too low it seems which explains my issues with online learning. I was using 5e-7 to 5e-6 as my test ranges and batch sizes of 64 or 128. I was experimenting with larger batch sizes early on but once I added Gru I found that performance just tanked with larger batches because of the batch size mismatch between train and test</p>",
              "rawMarkdown": "my learning rate and batch size were just way too low it seems which explains my issues with online learning. I was using 5e-7 to 5e-6 as my test ranges and batch sizes of 64 or 128. I was experimenting with larger batch sizes early on but once I added Gru I found that performance just tanked with larger batches because of the batch size mismatch between train and test"
            },
            {
              "id": 3096204,
              "postDate": "2025-01-14T06:21:17.503Z",
              "content": "<p>Is your batch sizes number of rows or something else?</p>",
              "rawMarkdown": "Is your batch sizes number of rows or something else?"
            },
            {
              "id": 3096218,
              "postDate": "2025-01-14T06:25:31.420Z",
              "content": "<p>Thanks for the detail! Mine MLP online score is 0.0069, with hidden layer [2048, 1024, 512, 256, 128, 64], lr 1e-3, and epoch 5. Different  MLP structure have different online learning behavior i guess. May i know how much epoch did you train the offline MLP? I have to train 30 epoch to see R2 stop improving, but i stoped training at 20 epoch. Overfitting might be the bottleneck of my lb score. Did you stop training your offline mlp after certain epoch?</p>",
              "rawMarkdown": "Thanks for the detail! Mine MLP online score is 0.0069, with hidden layer [2048, 1024, 512, 256, 128, 64], lr 1e-3, and epoch 5. Different  MLP structure have different online learning behavior i guess. May i know how much epoch did you train the offline MLP? I have to train 30 epoch to see R2 stop improving, but i stoped training at 20 epoch. Overfitting might be the bottleneck of my lb score. Did you stop training your offline mlp after certain epoch?"
            },
            {
              "id": 3096248,
              "postDate": "2025-01-14T07:02:56.250Z",
              "rawMarkdown": "",
              "isDeleted": true
            },
            {
              "id": 3096249,
              "postDate": "2025-01-14T07:03:36.760Z",
              "content": "<blockquote>\n  <p>Is your batch sizes number of rows or something else?</p>\n</blockquote>\n<p>yes, number of rows</p>",
              "rawMarkdown": ">Is your batch sizes number of rows or something else?\n\nyes, number of rows"
            },
            {
              "id": 3096251,
              "postDate": "2025-01-14T07:05:31.810Z",
              "content": "<blockquote>\n  <p>Different MLP structure have different online learning behavior i guess</p>\n</blockquote>\n<p>I agree.</p>\n<blockquote>\n  <p>May i know how much epoch did you train the offline MLP?</p>\n</blockquote>\n<p>I used early stopping with patience 15 epochs. (not certain epoch)<br>\nIt stopped around 50 epochs.</p>",
              "rawMarkdown": ">Different MLP structure have different online learning behavior i guess\n\nI agree.\n\n\n>May i know how much epoch did you train the offline MLP?\n\nI used early stopping with patience 15 epochs. (not certain epoch)\nIt stopped around 50 epochs.",
              "votes": 1
            }
          ]
        }
      ]
    }
  ],
  "comments": [
    {
      "id": 3096332,
      "author_name": "ganchan",
      "author_url": "",
      "post_date": "2025-01-14T09:12:57.890000",
      "content": "<p>PS.<br>\nIn our case, online training with new data and original train data (sampling) improved score.<br>\nOnline training with only new data did not work well.</p>",
      "votes": 0,
      "replies": [
        {
          "id": 3097494,
          "author_name": "Shiqiang Lee",
          "author_url": "",
          "post_date": "2025-01-15T12:37:55.090000",
          "content": "<p>I also sampled the last 480 days of data and trained the online MLP. My single MLP online score is 0.0077, with hidden layer [512,512,256,256,128,128,64,32], dropout[0.4,0.4,0.3,0.3,0.2,0.2,0.1,0.0], lr 5e-5, and epoch 30, batchsize 8192. However, it is a pity that I submitted this result for the last time and failed to integrate with my offline gbdt. The single score of my previous one integrated with offline gbdt in online MLP was 0.0071 and the fusion score was 0.0083</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 3097606,
          "author_name": "pl4$m4",
          "author_url": "",
          "post_date": "2025-01-15T14:30:44.740000",
          "content": "<p>Very interesting, I tried lgb only with new data with many different larameters and indeed did not see any improvements. Did you use init_model? And if I may ask how did you tweak the learning rate and number of boosting rounds?</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 3096210,
      "author_name": "ZT",
      "author_url": "",
      "post_date": "2025-01-14T06:22:16.253000",
      "content": "<p>how did you train LGBM to 0.007? I can only get up to 0.005. Any insight you could share?</p>",
      "votes": 0,
      "replies": [
        {
          "id": 3096223,
          "author_name": "Yimin Hu",
          "author_url": "",
          "post_date": "2025-01-14T06:28:47.183000",
          "content": "<p>I used all data for xgb, with large n_estimator and small lr. i used ['time_id', 'sybmol_id'] + 79 features + 8 lags, and normalize all the 89 features. Without online learning, my offline xgb gets 0.0070. </p>\n<p>Hopefully others can bring their findings to improve the gbt offline score here.</p>",
          "votes": 0,
          "replies": [
            {
              "id": 3096257,
              "author_name": "ZT",
              "author_url": "",
              "post_date": "2025-01-14T07:11:48.073000",
              "content": "<p>Appriciate your reply. <br>\nWas your normalization per symbol id or was it just whole norm for the entire column?</p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 3096260,
              "author_name": "Yimin Hu",
              "author_url": "",
              "post_date": "2025-01-14T07:18:54.383000",
              "content": "<p>just the whole norm for the entire column. I used 5 groupk fold  by date_id, otherwise a single model won't get above 0.0068</p>",
              "votes": 0,
              "replies": []
            }
          ]
        },
        {
          "id": 3096258,
          "author_name": "ganchan",
          "author_url": "",
          "post_date": "2025-01-14T07:13:11.567000",
          "content": "<p>Actually, there’s no magic here.<br>\nI simply trained the model using all features and time_id, achieving an offline score of 0.0063. (online: 0.0070)</p>\n<p>lgb_params = {<br>\n    \"objective\": \"regression\",<br>\n    \"boosting\": \"gbdt\",<br>\n    \"max_depth\": -1,<br>\n    \"num_leaves\": 40,<br>\n    \"subsample\": 0.8,<br>\n    \"subsample_freq\": 1,<br>\n    \"bagging_seed\": rand,<br>\n    \"learning_rate\": 0.05,<br>\n    \"feature_fraction\": 0.6,<br>\n    \"min_data_in_leaf\": 100,<br>\n    \"lambda_l1\": 0,<br>\n    \"lambda_l2\": 0,<br>\n    \"random_state\": rand,<br>\n    \"metric\": \"rmse\",<br>\n}</p>",
          "votes": 2,
          "replies": [
            {
              "id": 3096277,
              "author_name": "ZT",
              "author_url": "",
              "post_date": "2025-01-14T07:31:28.370000",
              "content": "<p>thanks for sharing</p>",
              "votes": 0,
              "replies": []
            }
          ]
        },
        {
          "id": 3098151,
          "author_name": "Shiqiang Lee",
          "author_url": "",
          "post_date": "2025-01-16T06:41:14.090000",
          "content": "<p>add time_id in features, only offline lgbm  can 0.007 +</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 3096112,
      "author_name": "Yimin Hu",
      "author_url": "",
      "post_date": "2025-01-14T05:05:06.367000",
      "content": "<p>Nice post. What learning rate, epoch, and batch size for mlp and tabm did you guys used for online learning? Mine was 1e- 3 with 819600 as batch size. Surprisingly a small lr as fine tune didn't work for me</p>",
      "votes": 0,
      "replies": [
        {
          "id": 3096126,
          "author_name": "ganchan",
          "author_url": "",
          "post_date": "2025-01-14T05:16:58.480000",
          "content": "<p>Our parameters are as follows:<br>\nThe learning rates were determined through individual experiments but were set to the same value (1e-4).<br>\nIncreasing the learning rate resulted in worse scores.</p>\n<p><strong>NN(MLP)</strong><br>\nlearning rate: 1e-4<br>\nepoch: 1 (retrain every 0.4M new data)<br>\nbatch size: 8192<br>\noptimizer: Adam</p>\n<p><strong>TabM</strong><br>\nlearning rate: 1e-4<br>\nepoch: 1 (retrain every 0.4M new data)<br>\nbatch size: 8192<br>\noptimizer: AdamW</p>",
          "votes": 1,
          "replies": [
            {
              "id": 3096165,
              "author_name": "Michael Timbs",
              "author_url": "",
              "post_date": "2025-01-14T06:10:26.207000",
              "content": "<p>my learning rate and batch size were just way too low it seems which explains my issues with online learning. I was using 5e-7 to 5e-6 as my test ranges and batch sizes of 64 or 128. I was experimenting with larger batch sizes early on but once I added Gru I found that performance just tanked with larger batches because of the batch size mismatch between train and test</p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 3096204,
              "author_name": "Lu Bin Liu",
              "author_url": "",
              "post_date": "2025-01-14T06:21:17.503000",
              "content": "<p>Is your batch sizes number of rows or something else?</p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 3096218,
              "author_name": "Yimin Hu",
              "author_url": "",
              "post_date": "2025-01-14T06:25:31.420000",
              "content": "<p>Thanks for the detail! Mine MLP online score is 0.0069, with hidden layer [2048, 1024, 512, 256, 128, 64], lr 1e-3, and epoch 5. Different  MLP structure have different online learning behavior i guess. May i know how much epoch did you train the offline MLP? I have to train 30 epoch to see R2 stop improving, but i stoped training at 20 epoch. Overfitting might be the bottleneck of my lb score. Did you stop training your offline mlp after certain epoch?</p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 3096248,
              "author_name": "",
              "author_url": "",
              "post_date": "2025-01-14T07:02:56.250000",
              "content": "",
              "votes": 0,
              "replies": []
            },
            {
              "id": 3096249,
              "author_name": "ganchan",
              "author_url": "",
              "post_date": "2025-01-14T07:03:36.760000",
              "content": "<blockquote>\n  <p>Is your batch sizes number of rows or something else?</p>\n</blockquote>\n<p>yes, number of rows</p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 3096251,
              "author_name": "ganchan",
              "author_url": "",
              "post_date": "2025-01-14T07:05:31.810000",
              "content": "<blockquote>\n  <p>Different MLP structure have different online learning behavior i guess</p>\n</blockquote>\n<p>I agree.</p>\n<blockquote>\n  <p>May i know how much epoch did you train the offline MLP?</p>\n</blockquote>\n<p>I used early stopping with patience 15 epochs. (not certain epoch)<br>\nIt stopped around 50 epochs.</p>",
              "votes": 1,
              "replies": []
            }
          ]
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "3096038": "I found this competition a bit challenging due to the Time Series API.\nHowever, I learned a lot through this experience, and I’d like to express my gratitude to the competition hosts and all the participants!\nOur score isn’t very high (upper bronze), but we’re sharing our solution because comparing different approaches can be incredibly valuable.\n\nOur final submission is an ensemble of:\n\nLGBM (online training) LB 0.0070\nMLP (online training) LB 0.0070\nTabM (online training) LB 0.0075\n⇒ Ensemble LB 0.0083\nOnly TabM uses lag features and symbol_id, while the other models rely solely on time_id and feature_00 to feature_79.\nOnline training was performed every 0.4M rows.\nDuring online training, we randomly sampled and added the same amount of data from train.parquet. This proved effective in improving our score.\nMLP and TabM were retrained for only 1 epoch. Increasing the number of epochs didn’t work well for us.\nWithout online training, our score was only around 0.0065.\nWe believe this is due to either insufficient features or suboptimal model architecture, so we look forward to learning from the solutions shared by others.\n\nLooking back, I think there’s still plenty of room for improvement, and I’m excited to read the top solutions!\n\nThank you.",
    "3096332": "PS.\nIn our case, online training with new data and original train data (sampling) improved score.\nOnline training with only new data did not work well.",
    "3096210": "how did you train LGBM to 0.007? I can only get up to 0.005. Any insight you could share?",
    "3096112": "Nice post. What learning rate, epoch, and batch size for mlp and tabm did you guys used for online learning? Mine was 1e- 3 with 819600 as batch size. Surprisingly a small lr as fine tune didn't work for me"
  }
}