{
  "id": 553561,
  "title": "Is Catboost much worse than LightGBM and XGBoost?",
  "url": "/competitions/jane-street-real-time-market-data-forecasting/discussion/553561",
  "author_name": "xsong2020",
  "post_date": "2024-12-26T21:11:51.693000",
  "votes": 3,
  "comment_count": 35,
  "views": 0,
  "content": "<p>I find the Catboost is much worse than LightGBM and XGBoost in both CV and LB. I'm wondering if anyone else also finds a similar problem?</p>",
  "messages": [
    {
      "id": 3081541,
      "postDate": "2024-12-26T21:21:51.697Z",
      "content": "<p>Interesting… CatBoost is worse than LGB in local CV? I got a completely different conclusion.</p>",
      "rawMarkdown": "Interesting... CatBoost is worse than LGB in local CV? I got a completely different conclusion.",
      "votes": 3,
      "replies": [
        {
          "id": 3081559,
          "postDate": "2024-12-26T22:18:31.263Z",
          "content": "<p>Then your catboost must be trained very well. I tried to tune hyperparameyers for both catboost and lightgbm, but turns out catboost performs much worse than LGB in local CV and I have no idea what went wrong.</p>",
          "rawMarkdown": "Then your catboost must be trained very well. I tried to tune hyperparameyers for both catboost and lightgbm, but turns out catboost performs much worse than LGB in local CV and I have no idea what went wrong.",
          "votes": 1,
          "replies": [
            {
              "id": 3081560,
              "postDate": "2024-12-26T22:23:19.400Z",
              "content": "<p>Funny thing is for CatBoost model, I used default HPs only tuning lr (0.05), while for lgb, I did HPT, but cbt is way better than LGB models…</p>",
              "rawMarkdown": "Funny thing is for CatBoost model, I used default HPs only tuning lr (0.05), while for lgb, I did HPT, but cbt is way better than LGB models...",
              "votes": 2
            },
            {
              "id": 3081570,
              "postDate": "2024-12-26T23:07:50.167Z",
              "content": "<p>Wow, i can't believe you just used default HPs in catboost and achieve way better than LGB models. May I ask how many iterations did you train your catboost model?</p>",
              "rawMarkdown": "Wow, i can't believe you just used default HPs in catboost and achieve way better than LGB models. May I ask how many iterations did you train your catboost model?"
            },
            {
              "id": 3081572,
              "postDate": "2024-12-26T23:12:14.583Z",
              "content": "<p>600 iters with 0.05 lr, while 1000 iters with 0.03 lr gives a bit of better result. </p>",
              "rawMarkdown": "600 iters with 0.05 lr, while 1000 iters with 0.03 lr gives a bit of better result. ",
              "votes": 2
            },
            {
              "id": 3082080,
              "postDate": "2024-12-27T16:33:30.623Z",
              "content": "<p>Thank you and I'll probably have another try</p>",
              "rawMarkdown": "Thank you and I'll probably have another try"
            }
          ]
        },
        {
          "id": 3082700,
          "postDate": "2024-12-28T13:49:24.550Z",
          "content": "<p>can tree model compete with nn?</p>",
          "rawMarkdown": "can tree model compete with nn?",
          "replies": [
            {
              "id": 3082710,
              "postDate": "2024-12-28T14:09:08.287Z",
              "content": "<p>Not really. I spent a lot of time on tree models in early stage and they did play a big role for ensembling when my nn models were relatively bad. But after my NN models break 0.01, ensembling with tree models (CatBoost models) didn't bring much benefit any more, which is not surprising. </p>",
              "rawMarkdown": "Not really. I spent a lot of time on tree models in early stage and they did play a big role for ensembling when my nn models were relatively bad. But after my NN models break 0.01, ensembling with tree models (CatBoost models) didn't bring much benefit any more, which is not surprising. ",
              "votes": 4
            },
            {
              "id": 3082930,
              "postDate": "2024-12-28T19:22:58.040Z",
              "content": "<p>seems like gbdt is a dead end</p>",
              "rawMarkdown": "seems like gbdt is a dead end",
              "votes": 1
            },
            {
              "id": 3083291,
              "postDate": "2024-12-29T08:57:40.937Z",
              "rawMarkdown": "",
              "isDeleted": true
            },
            {
              "id": 3083293,
              "postDate": "2024-12-29T08:58:55.233Z",
              "content": "<p>maybe try transformer? I remember a competition where transformer did it very well</p>",
              "rawMarkdown": "maybe try transformer? I remember a competition where transformer did it very well"
            },
            {
              "id": 3083490,
              "postDate": "2024-12-29T14:40:59.110Z",
              "content": "<p>i could get over 0.006 with xgbs w/o much tuning, but it seems to get higher scores, nn is necessary</p>",
              "rawMarkdown": "i could get over 0.006 with xgbs w/o much tuning, but it seems to get higher scores, nn is necessary"
            },
            {
              "id": 3083494,
              "postDate": "2024-12-29T14:49:42.147Z",
              "content": "<p>From what I see, the limit of GBDT is 0.008 in current PB. I would be very surprised if anyone can get a higher score using pure GBDT solution. </p>",
              "rawMarkdown": "From what I see, the limit of GBDT is 0.008 in current PB. I would be very surprised if anyone can get a higher score using pure GBDT solution. ",
              "votes": 2
            },
            {
              "id": 3083501,
              "postDate": "2024-12-29T15:02:11.710Z",
              "content": "<p>I think that's probably accurate. I did xgb but didn't fine tune and only used small portion of the data and got to 0.0061 using 3 xgb ensemble without tuning. But top scores will definitely use nn.</p>",
              "rawMarkdown": "I think that's probably accurate. I did xgb but didn't fine tune and only used small portion of the data and got to 0.0061 using 3 xgb ensemble without tuning. But top scores will definitely use nn."
            },
            {
              "id": 3083506,
              "postDate": "2024-12-29T15:09:58.827Z",
              "content": "<p>To reach 0.008 using gbdt, fine tuning HPs is not enough at all. And if you did more experiemnts by adding more data, soon you will find most times it didn't help either. You have to manage to create a robust gbdt pipeline to do online learning, just like NN methods, which is not an easy task especially for GBDT models in this competition setup, but feasible. </p>",
              "rawMarkdown": "To reach 0.008 using gbdt, fine tuning HPs is not enough at all. And if you did more experiemnts by adding more data, soon you will find most times it didn't help either. You have to manage to create a robust gbdt pipeline to do online learning, just like NN methods, which is not an easy task especially for GBDT models in this competition setup, but feasible. ",
              "votes": 5
            },
            {
              "id": 3083508,
              "postDate": "2024-12-29T15:12:56.513Z",
              "content": "<p>I see. how much do you think NN can get to without online learning?</p>",
              "rawMarkdown": "I see. how much do you think NN can get to without online learning?"
            },
            {
              "id": 3083509,
              "postDate": "2024-12-29T15:15:06.797Z",
              "content": "<p>For me, it's around 0.009, but I'm pretty sure for current top 1/2, the number could be 0.01+. But as you may see, it's not easy to get there. </p>",
              "rawMarkdown": "For me, it's around 0.009, but I'm pretty sure for current top 1/2, the number could be 0.01+. But as you may see, it's not easy to get there. ",
              "votes": 1
            },
            {
              "id": 3083511,
              "postDate": "2024-12-29T15:21:33.527Z",
              "content": "<p>can i ask if that's with a simple mlp or a more complicated structure (transformer, lstm, etc)?</p>",
              "rawMarkdown": "can i ask if that's with a simple mlp or a more complicated structure (transformer, lstm, etc)?"
            },
            {
              "id": 3083528,
              "postDate": "2024-12-29T15:54:20.233Z",
              "content": "<p>It will be more fun finding it by yourself, right?😀</p>",
              "rawMarkdown": "It will be more fun finding it by yourself, right?😀",
              "votes": 11
            },
            {
              "id": 3091870,
              "postDate": "2025-01-08T21:02:38.613Z",
              "content": "<p>May I know whether the score of 0.009 was obtained from the public leaderboard or from your local CV?</p>",
              "rawMarkdown": "May I know whether the score of 0.009 was obtained from the public leaderboard or from your local CV?"
            },
            {
              "id": 3091880,
              "postDate": "2025-01-08T21:19:10.420Z",
              "content": "<p>Public leaderboard</p>",
              "rawMarkdown": "Public leaderboard",
              "votes": 1
            },
            {
              "id": 3091967,
              "postDate": "2025-01-09T02:50:16.320Z",
              "content": "<p>Thanks for the response! The reason I am asking is because my NN only scored 0.0047, seems like I still got a lot of space to improve.</p>",
              "rawMarkdown": "Thanks for the response! The reason I am asking is because my NN only scored 0.0047, seems like I still got a lot of space to improve."
            },
            {
              "id": 3092905,
              "postDate": "2025-01-10T08:01:54.673Z",
              "content": "<p>but how do we keep the online learning &lt;60s? seems like an impossible task, but you guys certainly did it looks like</p>",
              "rawMarkdown": "but how do we keep the online learning <60s? seems like an impossible task, but you guys certainly did it looks like",
              "votes": 1
            }
          ]
        }
      ]
    },
    {
      "id": 3081540,
      "postDate": "2024-12-26T21:11:51.693Z",
      "content": "<p>I find the Catboost is much worse than LightGBM and XGBoost in both CV and LB. I'm wondering if anyone else also finds a similar problem?</p>",
      "rawMarkdown": "I find the Catboost is much worse than LightGBM and XGBoost in both CV and LB. I'm wondering if anyone else also finds a similar problem?",
      "votes": 3
    },
    {
      "id": 3095219,
      "postDate": "2025-01-13T04:34:19.030Z",
      "content": "<p>All the gradient boosting models seems overfitting and kinds of linear regression is not utterly fitted. Therefore NN remains. I suppose that. What do you think about my supposition?</p>",
      "rawMarkdown": "All the gradient boosting models seems overfitting and kinds of linear regression is not utterly fitted. Therefore NN remains. I suppose that. What do you think about my supposition?"
    },
    {
      "id": 3087331,
      "postDate": "2025-01-03T11:34:47.467Z",
      "content": "<p>I spent a couple weeks on the Catboost path, even hiring some H100s and could not get better than 0.0032 after all that experimentation. An out the box LightGBM with default params was 0.005</p>",
      "rawMarkdown": "I spent a couple weeks on the Catboost path, even hiring some H100s and could not get better than 0.0032 after all that experimentation. An out the box LightGBM with default params was 0.005",
      "replies": [
        {
          "id": 3087357,
          "postDate": "2025-01-03T12:00:11.777Z",
          "content": "<p>Have you tried a pure online way? E.g using last N days to train a small model and predict the next day. N is small enough to fit in the 1min time constraints  and big enough to provide enough data for training. In this way, there is no “ base” model trained using super old data, and the model is always updated to accommodate the non-stationary nature of the data. For validation, every predicted date is your oof valid set. </p>\n<p>I don’t know if it’s guaranteed to work but maybe worth a try…</p>",
          "rawMarkdown": "Have you tried a pure online way? E.g using last N days to train a small model and predict the next day. N is small enough to fit in the 1min time constraints  and big enough to provide enough data for training. In this way, there is no “ base” model trained using super old data, and the model is always updated to accommodate the non-stationary nature of the data. For validation, every predicted date is your oof valid set. \n\nI don’t know if it’s guaranteed to work but maybe worth a try…",
          "votes": 2,
          "replies": [
            {
              "id": 3087483,
              "postDate": "2025-01-03T14:23:09.303Z",
              "content": "<p>I didn't only because lags seemed somewhat anti-predictive. In that looking at previous days data seemed to make predictions worse rather than better so I was worried that there would be no signal there.</p>",
              "rawMarkdown": "I didn't only because lags seemed somewhat anti-predictive. In that looking at previous days data seemed to make predictions worse rather than better so I was worried that there would be no signal there."
            },
            {
              "id": 3092906,
              "postDate": "2025-01-10T08:03:38.277Z",
              "content": "<p>what would be a good N in our case? I tried to keep 100_000, 10_000 in the window and they all timed out :(</p>",
              "rawMarkdown": "what would be a good N in our case? I tried to keep 100_000, 10_000 in the window and they all timed out :("
            }
          ]
        },
        {
          "id": 3087464,
          "postDate": "2025-01-03T14:12:48.830Z",
          "content": "<p>69bps (0.0069) LB for 6-fold static CB blend is doable. </p>",
          "rawMarkdown": "69bps (0.0069) LB for 6-fold static CB blend is doable. ",
          "replies": [
            {
              "id": 3087476,
              "postDate": "2025-01-03T14:19:27.443Z",
              "content": "<p>I didn't try ensembling*. I tried a single model trained on the entire dataset (and entire minus first 500 days).</p>\n<p>Maybe would have had more luck training on subsets of the data and then combining predictions.</p>\n<p>Edit: I did actually try an \"ensemble\" but it was auxiliary models. I had 5 models predict some other responders and then fed those predictions, along with the base features + feature engineering in to predict responder 6 but this ended up worse because of error multiplication. The auxiliary predictions tanked in LB data which made the auxiliary ensemble worse than having no auxiliary models</p>",
              "rawMarkdown": "I didn't try ensembling*. I tried a single model trained on the entire dataset (and entire minus first 500 days).\n\nMaybe would have had more luck training on subsets of the data and then combining predictions.\n\nEdit: I did actually try an \"ensemble\" but it was auxiliary models. I had 5 models predict some other responders and then fed those predictions, along with the base features + feature engineering in to predict responder 6 but this ended up worse because of error multiplication. The auxiliary predictions tanked in LB data which made the auxiliary ensemble worse than having no auxiliary models"
            },
            {
              "id": 3090188,
              "postDate": "2025-01-07T02:09:03.570Z",
              "content": "<p>I actually went back and tried LightGBM again and managed to get a 0.007 LB score from a simple and shallow model</p>",
              "rawMarkdown": "I actually went back and tried LightGBM again and managed to get a 0.007 LB score from a simple and shallow model"
            }
          ]
        }
      ]
    },
    {
      "id": 3083780,
      "postDate": "2024-12-30T02:53:25.100Z",
      "content": "<p>May I ask what your CV strategy is (if you’re able to share)? From my personal experience, CatBoost tends to outperform LGBM on most categorical datasets. However, when it comes to time-series data, the results can vary. I think the performance largely depends on feature engineering and the CV strategy.</p>",
      "rawMarkdown": "May I ask what your CV strategy is (if you’re able to share)? From my personal experience, CatBoost tends to outperform LGBM on most categorical datasets. However, when it comes to time-series data, the results can vary. I think the performance largely depends on feature engineering and the CV strategy.",
      "replies": [
        {
          "id": 3084561,
          "postDate": "2024-12-31T02:51:35.310Z",
          "content": "<p>Nothing special but just use the last 100 days as validation</p>",
          "rawMarkdown": "Nothing special but just use the last 100 days as validation"
        }
      ]
    },
    {
      "id": 3081689,
      "postDate": "2024-12-27T04:48:24.917Z",
      "content": "<p>May I ask how long do your catboost take to train?I just use some simple hyperparameyers but it takes at least a few hours to train, which is much slower than my lgb and xgb. (I use about 1500 days to train)</p>",
      "rawMarkdown": "May I ask how long do your catboost take to train?I just use some simple hyperparameyers but it takes at least a few hours to train, which is much slower than my lgb and xgb. (I use about 1500 days to train)",
      "replies": [
        {
          "id": 3082081,
          "postDate": "2024-12-27T16:33:54.763Z",
          "content": "<p>If you enable the GPU acceleration, it should be fast.</p>",
          "rawMarkdown": "If you enable the GPU acceleration, it should be fast."
        }
      ]
    }
  ],
  "comments": [
    {
      "id": 3081541,
      "author_name": "HAO",
      "author_url": "",
      "post_date": "2024-12-26T21:21:51.697000",
      "content": "<p>Interesting… CatBoost is worse than LGB in local CV? I got a completely different conclusion.</p>",
      "votes": 3,
      "replies": [
        {
          "id": 3081559,
          "author_name": "xsong2020",
          "author_url": "",
          "post_date": "2024-12-26T22:18:31.263000",
          "content": "<p>Then your catboost must be trained very well. I tried to tune hyperparameyers for both catboost and lightgbm, but turns out catboost performs much worse than LGB in local CV and I have no idea what went wrong.</p>",
          "votes": 1,
          "replies": [
            {
              "id": 3081560,
              "author_name": "HAO",
              "author_url": "",
              "post_date": "2024-12-26T22:23:19.400000",
              "content": "<p>Funny thing is for CatBoost model, I used default HPs only tuning lr (0.05), while for lgb, I did HPT, but cbt is way better than LGB models…</p>",
              "votes": 2,
              "replies": []
            },
            {
              "id": 3081570,
              "author_name": "xsong2020",
              "author_url": "",
              "post_date": "2024-12-26T23:07:50.167000",
              "content": "<p>Wow, i can't believe you just used default HPs in catboost and achieve way better than LGB models. May I ask how many iterations did you train your catboost model?</p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 3081572,
              "author_name": "HAO",
              "author_url": "",
              "post_date": "2024-12-26T23:12:14.583000",
              "content": "<p>600 iters with 0.05 lr, while 1000 iters with 0.03 lr gives a bit of better result. </p>",
              "votes": 2,
              "replies": []
            },
            {
              "id": 3082080,
              "author_name": "xsong2020",
              "author_url": "",
              "post_date": "2024-12-27T16:33:30.623000",
              "content": "<p>Thank you and I'll probably have another try</p>",
              "votes": 0,
              "replies": []
            }
          ]
        },
        {
          "id": 3082700,
          "author_name": "yuanzhe zhou",
          "author_url": "",
          "post_date": "2024-12-28T13:49:24.550000",
          "content": "<p>can tree model compete with nn?</p>",
          "votes": 0,
          "replies": [
            {
              "id": 3082710,
              "author_name": "HAO",
              "author_url": "",
              "post_date": "2024-12-28T14:09:08.287000",
              "content": "<p>Not really. I spent a lot of time on tree models in early stage and they did play a big role for ensembling when my nn models were relatively bad. But after my NN models break 0.01, ensembling with tree models (CatBoost models) didn't bring much benefit any more, which is not surprising. </p>",
              "votes": 4,
              "replies": []
            },
            {
              "id": 3082930,
              "author_name": "yuanzhe zhou",
              "author_url": "",
              "post_date": "2024-12-28T19:22:58.040000",
              "content": "<p>seems like gbdt is a dead end</p>",
              "votes": 1,
              "replies": []
            },
            {
              "id": 3083291,
              "author_name": "",
              "author_url": "",
              "post_date": "2024-12-29T08:57:40.937000",
              "content": "",
              "votes": 0,
              "replies": []
            },
            {
              "id": 3083293,
              "author_name": "M4XD",
              "author_url": "",
              "post_date": "2024-12-29T08:58:55.233000",
              "content": "<p>maybe try transformer? I remember a competition where transformer did it very well</p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 3083490,
              "author_name": "rulan",
              "author_url": "",
              "post_date": "2024-12-29T14:40:59.110000",
              "content": "<p>i could get over 0.006 with xgbs w/o much tuning, but it seems to get higher scores, nn is necessary</p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 3083494,
              "author_name": "HAO",
              "author_url": "",
              "post_date": "2024-12-29T14:49:42.147000",
              "content": "<p>From what I see, the limit of GBDT is 0.008 in current PB. I would be very surprised if anyone can get a higher score using pure GBDT solution. </p>",
              "votes": 2,
              "replies": []
            },
            {
              "id": 3083501,
              "author_name": "rulan",
              "author_url": "",
              "post_date": "2024-12-29T15:02:11.710000",
              "content": "<p>I think that's probably accurate. I did xgb but didn't fine tune and only used small portion of the data and got to 0.0061 using 3 xgb ensemble without tuning. But top scores will definitely use nn.</p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 3083506,
              "author_name": "HAO",
              "author_url": "",
              "post_date": "2024-12-29T15:09:58.827000",
              "content": "<p>To reach 0.008 using gbdt, fine tuning HPs is not enough at all. And if you did more experiemnts by adding more data, soon you will find most times it didn't help either. You have to manage to create a robust gbdt pipeline to do online learning, just like NN methods, which is not an easy task especially for GBDT models in this competition setup, but feasible. </p>",
              "votes": 5,
              "replies": []
            },
            {
              "id": 3083508,
              "author_name": "rulan",
              "author_url": "",
              "post_date": "2024-12-29T15:12:56.513000",
              "content": "<p>I see. how much do you think NN can get to without online learning?</p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 3083509,
              "author_name": "HAO",
              "author_url": "",
              "post_date": "2024-12-29T15:15:06.797000",
              "content": "<p>For me, it's around 0.009, but I'm pretty sure for current top 1/2, the number could be 0.01+. But as you may see, it's not easy to get there. </p>",
              "votes": 1,
              "replies": []
            },
            {
              "id": 3083511,
              "author_name": "rulan",
              "author_url": "",
              "post_date": "2024-12-29T15:21:33.527000",
              "content": "<p>can i ask if that's with a simple mlp or a more complicated structure (transformer, lstm, etc)?</p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 3083528,
              "author_name": "HAO",
              "author_url": "",
              "post_date": "2024-12-29T15:54:20.233000",
              "content": "<p>It will be more fun finding it by yourself, right?😀</p>",
              "votes": 11,
              "replies": []
            },
            {
              "id": 3091870,
              "author_name": "Maaax",
              "author_url": "",
              "post_date": "2025-01-08T21:02:38.613000",
              "content": "<p>May I know whether the score of 0.009 was obtained from the public leaderboard or from your local CV?</p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 3091880,
              "author_name": "HAO",
              "author_url": "",
              "post_date": "2025-01-08T21:19:10.420000",
              "content": "<p>Public leaderboard</p>",
              "votes": 1,
              "replies": []
            },
            {
              "id": 3091967,
              "author_name": "Maaax",
              "author_url": "",
              "post_date": "2025-01-09T02:50:16.320000",
              "content": "<p>Thanks for the response! The reason I am asking is because my NN only scored 0.0047, seems like I still got a lot of space to improve.</p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 3092905,
              "author_name": "Creative-Ataraxia",
              "author_url": "",
              "post_date": "2025-01-10T08:01:54.673000",
              "content": "<p>but how do we keep the online learning &lt;60s? seems like an impossible task, but you guys certainly did it looks like</p>",
              "votes": 1,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 3095219,
      "author_name": "Daisy",
      "author_url": "",
      "post_date": "2025-01-13T04:34:19.030000",
      "content": "<p>All the gradient boosting models seems overfitting and kinds of linear regression is not utterly fitted. Therefore NN remains. I suppose that. What do you think about my supposition?</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 3087331,
      "author_name": "Michael Timbs",
      "author_url": "",
      "post_date": "2025-01-03T11:34:47.467000",
      "content": "<p>I spent a couple weeks on the Catboost path, even hiring some H100s and could not get better than 0.0032 after all that experimentation. An out the box LightGBM with default params was 0.005</p>",
      "votes": 0,
      "replies": [
        {
          "id": 3087357,
          "author_name": "SLi",
          "author_url": "",
          "post_date": "2025-01-03T12:00:11.777000",
          "content": "<p>Have you tried a pure online way? E.g using last N days to train a small model and predict the next day. N is small enough to fit in the 1min time constraints  and big enough to provide enough data for training. In this way, there is no “ base” model trained using super old data, and the model is always updated to accommodate the non-stationary nature of the data. For validation, every predicted date is your oof valid set. </p>\n<p>I don’t know if it’s guaranteed to work but maybe worth a try…</p>",
          "votes": 2,
          "replies": [
            {
              "id": 3087483,
              "author_name": "Michael Timbs",
              "author_url": "",
              "post_date": "2025-01-03T14:23:09.303000",
              "content": "<p>I didn't only because lags seemed somewhat anti-predictive. In that looking at previous days data seemed to make predictions worse rather than better so I was worried that there would be no signal there.</p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 3092906,
              "author_name": "Creative-Ataraxia",
              "author_url": "",
              "post_date": "2025-01-10T08:03:38.277000",
              "content": "<p>what would be a good N in our case? I tried to keep 100_000, 10_000 in the window and they all timed out :(</p>",
              "votes": 0,
              "replies": []
            }
          ]
        },
        {
          "id": 3087464,
          "author_name": "Nicholas Lyu",
          "author_url": "",
          "post_date": "2025-01-03T14:12:48.830000",
          "content": "<p>69bps (0.0069) LB for 6-fold static CB blend is doable. </p>",
          "votes": 0,
          "replies": [
            {
              "id": 3087476,
              "author_name": "Michael Timbs",
              "author_url": "",
              "post_date": "2025-01-03T14:19:27.443000",
              "content": "<p>I didn't try ensembling*. I tried a single model trained on the entire dataset (and entire minus first 500 days).</p>\n<p>Maybe would have had more luck training on subsets of the data and then combining predictions.</p>\n<p>Edit: I did actually try an \"ensemble\" but it was auxiliary models. I had 5 models predict some other responders and then fed those predictions, along with the base features + feature engineering in to predict responder 6 but this ended up worse because of error multiplication. The auxiliary predictions tanked in LB data which made the auxiliary ensemble worse than having no auxiliary models</p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 3090188,
              "author_name": "Michael Timbs",
              "author_url": "",
              "post_date": "2025-01-07T02:09:03.570000",
              "content": "<p>I actually went back and tried LightGBM again and managed to get a 0.007 LB score from a simple and shallow model</p>",
              "votes": 0,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 3083780,
      "author_name": "W.Z. Kendrick",
      "author_url": "",
      "post_date": "2024-12-30T02:53:25.100000",
      "content": "<p>May I ask what your CV strategy is (if you’re able to share)? From my personal experience, CatBoost tends to outperform LGBM on most categorical datasets. However, when it comes to time-series data, the results can vary. I think the performance largely depends on feature engineering and the CV strategy.</p>",
      "votes": 0,
      "replies": [
        {
          "id": 3084561,
          "author_name": "xsong2020",
          "author_url": "",
          "post_date": "2024-12-31T02:51:35.310000",
          "content": "<p>Nothing special but just use the last 100 days as validation</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 3081689,
      "author_name": "I2nfinit3y",
      "author_url": "",
      "post_date": "2024-12-27T04:48:24.917000",
      "content": "<p>May I ask how long do your catboost take to train?I just use some simple hyperparameyers but it takes at least a few hours to train, which is much slower than my lgb and xgb. (I use about 1500 days to train)</p>",
      "votes": 0,
      "replies": [
        {
          "id": 3082081,
          "author_name": "xsong2020",
          "author_url": "",
          "post_date": "2024-12-27T16:33:54.763000",
          "content": "<p>If you enable the GPU acceleration, it should be fast.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "3081541": "Interesting... CatBoost is worse than LGB in local CV? I got a completely different conclusion.",
    "3081540": "I find the Catboost is much worse than LightGBM and XGBoost in both CV and LB. I'm wondering if anyone else also finds a similar problem?",
    "3095219": "All the gradient boosting models seems overfitting and kinds of linear regression is not utterly fitted. Therefore NN remains. I suppose that. What do you think about my supposition?",
    "3087331": "I spent a couple weeks on the Catboost path, even hiring some H100s and could not get better than 0.0032 after all that experimentation. An out the box LightGBM with default params was 0.005",
    "3083780": "May I ask what your CV strategy is (if you’re able to share)? From my personal experience, CatBoost tends to outperform LGBM on most categorical datasets. However, when it comes to time-series data, the results can vary. I think the performance largely depends on feature engineering and the CV strategy.",
    "3081689": "May I ask how long do your catboost take to train?I just use some simple hyperparameyers but it takes at least a few hours to train, which is much slower than my lgb and xgb. (I use about 1500 days to train)"
  }
}