{
  "id": 548405,
  "title": "Context Importance",
  "url": "/competitions/jane-street-real-time-market-data-forecasting/discussion/548405",
  "author_name": "",
  "post_date": "2024-11-26T15:50:37.109567600Z",
  "votes": 2,
  "comment_count": 13,
  "views": 0,
  "content": "<p>Hey everyone,</p>\n<p>I'm curious about the role of context in this type of problem. I've tried extensive feature engineering and employed some of the most robust and computationally expensive methods you can imagine. However, adding features that capture the context of previous instances consistently resulted in performance similar to using the original features, with only minor variations.</p>\n<p>This outcome was unexpected, especially since I’m using models like standard MLPs and GBMs, which are typically blind to context. I had assumed that incorporating features capturing the context would significantly enhance performance, but that hasn’t been the case.</p>\n<p>Has anyone succeeded in boosting performance through feature engineering in this competition? I would greatly appreciate it if someone could explain how context might be irrelevant or underutilized in this time series problem. Any insights or suggestions are welcome!</p>",
  "messages": [
    {
      "id": "3056147",
      "postDate": "11/26/2024 15:50:37",
      "content": "<p>Hey everyone,</p>\n<p>I'm curious about the role of context in this type of problem. I've tried extensive feature engineering and employed some of the most robust and computationally expensive methods you can imagine. However, adding features that capture the context of previous instances consistently resulted in performance similar to using the original features, with only minor variations.</p>\n<p>This outcome was unexpected, especially since I’m using models like standard MLPs and GBMs, which are typically blind to context. I had assumed that incorporating features capturing the context would significantly enhance performance, but that hasn’t been the case.</p>\n<p>Has anyone succeeded in boosting performance through feature engineering in this competition? I would greatly appreciate it if someone could explain how context might be irrelevant or underutilized in this time series problem. Any insights or suggestions are welcome!</p>",
      "rawMarkdown": "Hey everyone,\n\nI'm curious about the role of context in this type of problem. I've tried extensive feature engineering and employed some of the most robust and computationally expensive methods you can imagine. However, adding features that capture the context of previous instances consistently resulted in performance similar to using the original features, with only minor variations.\n\nThis outcome was unexpected, especially since I’m using models like standard MLPs and GBMs, which are typically blind to context. I had assumed that incorporating features capturing the context would significantly enhance performance, but that hasn’t been the case.\n\nHas anyone succeeded in boosting performance through feature engineering in this competition? I would greatly appreciate it if someone could explain how context might be irrelevant or underutilized in this time series problem. Any insights or suggestions are welcome!",
      "votes": null
    },
    {
      "id": "3056213",
      "postDate": "11/26/2024 17:12:15",
      "content": "<p>In my case, feature engineering (additional ~30 features) could boost my lgb model from 0.0048 to 0.0059 (LB), while CatBoost model from 0.0043 to 0.0052 (without online learning) and 0.0060 (with online learning). But as I said in another thread, I think in this competition, because of the anonymous features and 1 minute time limitation, GBDT is kind of a dead end. You can't get a gold medal using pure GBDT methods (at least from my knowledge), which is not the case in previous Optiver competition (I won 10th using a single CatBoost model). </p>",
      "rawMarkdown": "In my case, feature engineering (additional ~30 features) could boost my lgb model from 0.0048 to 0.0059 (LB), while CatBoost model from 0.0043 to 0.0052 (without online learning) and 0.0060 (with online learning). But as I said in another thread, I think in this competition, because of the anonymous features and 1 minute time limitation, GBDT is kind of a dead end. You can't get a gold medal using pure GBDT methods (at least from my knowledge), which is not the case in previous Optiver competition (I won 10th using a single CatBoost model).",
      "votes": null
    },
    {
      "id": "3056216",
      "postDate": "11/26/2024 17:17:49",
      "content": "<p>are the features you added like lagged and moving window features or did you did a more complex feature engineering. I am only curious if those features that enhance your performance are contextual features or something else </p>",
      "rawMarkdown": "are the features you added like lagged and moving window features or did you did a more complex feature engineering. I am only curious if those features that enhance your performance are contextual features or something else",
      "votes": null
    },
    {
      "id": "3056222",
      "postDate": "11/26/2024 17:30:03",
      "content": "<p>Yes, they are simple moving window features. I mean those provided features are abonymous, what else feature engineering we could do about them?😃</p>",
      "rawMarkdown": "Yes, they are simple moving window features. I mean those provided features are abonymous, what else feature engineering we could do about them?😃",
      "votes": null
    },
    {
      "id": "3056224",
      "postDate": "11/26/2024 17:32:13",
      "content": "<p>man that make me want to cry you wouldn't imagine how much time and effort i put in feature engineering with no result so far 😢<br>\nhowever, good job and thanks for answering </p>",
      "rawMarkdown": "man that make me want to cry you wouldn't imagine how much time and effort i put in feature engineering with no result so far 😢\nhowever, good job and thanks for answering",
      "votes": null
    },
    {
      "id": "3056231",
      "postDate": "11/26/2024 17:49:30",
      "content": "<p>Thanks for providing that insight, one thing I'm curious about is your choice of CatBoost. I reckon that it is specialized in categorical features, which we do not have (Assume you don't have a lot of engineered categorical features), and since it performs worse than LightGBM on LB, what's the motive of using that for online training instead of lightGBM? </p>",
      "rawMarkdown": "Thanks for providing that insight, one thing I'm curious about is your choice of CatBoost. I reckon that it is specialized in categorical features, which we do not have (Assume you don't have a lot of engineered categorical features), and since it performs worse than LightGBM on LB, what's the motive of using that for online training instead of lightGBM?",
      "votes": null
    },
    {
      "id": "3056242",
      "postDate": "11/26/2024 18:02:54",
      "content": "<p><code>I reckon that it is specialized in categorical features</code><br>\nI'm not sure about this actually. In previous optiver competition, CatBoost is much better compared to LGB in CV and LB, even though all my features are numerical (same in this comp). In this competition, actually from local CV, the CatBoost is always better than LGB with a clear margin (tested with multiple validation vintages), but worse than LGB in LB (couldn't figure out why yet, maybe it's just coincidence). <br>\n<code>what's the motive of using that for online training instead of lightGBM</code><br>\nUsing CatBoost for online learning has two reasons, one is it's better and more stable from local CV, another reason is it's simply faster and consuming lower ram compared to LGB (I still couldn't figure out a stable way to do online training with 110+ features using comparable data amount which I could easily make it for catboost online learning). </p>",
      "rawMarkdown": "`I reckon that it is specialized in categorical features`\nI'm not sure about this actually. In previous optiver competition, CatBoost is much better compared to LGB in CV and LB, even though all my features are numerical (same in this comp). In this competition, actually from local CV, the CatBoost is always better than LGB with a clear margin (tested with multiple validation vintages), but worse than LGB in LB (couldn't figure out why yet, maybe it's just coincidence). \n`what's the motive of using that for online training instead of lightGBM`\nUsing CatBoost for online learning has two reasons, one is it's better and more stable from local CV, another reason is it's simply faster and consuming lower ram compared to LGB (I still couldn't figure out a stable way to do online training with 110+ features using comparable data amount which I could easily make it for catboost online learning).",
      "votes": null
    },
    {
      "id": "3056248",
      "postDate": "11/26/2024 18:13:49",
      "content": "<p>Thanks for the reply! That is true, doing online training with LGBM seems to consume a lot of RAM, and a lot of the time leads to memory leakage, I haven't tested with Catboost yet, but this gives me a lot to think about. I still think it is possible to squeeze a few drops out of GBM, which could also be my delusion😆.</p>",
      "rawMarkdown": "Thanks for the reply! That is true, doing online training with LGBM seems to consume a lot of RAM, and a lot of the time leads to memory leakage, I haven't tested with Catboost yet, but this gives me a lot to think about. I still think it is possible to squeeze a few drops out of GBM, which could also be my delusion😆.",
      "votes": null
    },
    {
      "id": "3056411",
      "postDate": "11/27/2024 00:26:10",
      "content": "<p>What I gather is in this context whatever pattern you find is generally very time-specific. Autoencoders for example combine the features in unique ways that may be very useful for some time intervals but cause whatever model you use them in to overfit. Take one of the pilars of this competition, for example - online learning. If you're retraining your model but not your autoencoder, the features it creates may very well be useless after some point. Now I'm not saying feature engineering is not applicable here, but I think the gain is on the margins, and I would test performance for each new created feature.</p>",
      "rawMarkdown": "What I gather is in this context whatever pattern you find is generally very time-specific. Autoencoders for example combine the features in unique ways that may be very useful for some time intervals but cause whatever model you use them in to overfit. Take one of the pilars of this competition, for example - online learning. If you're retraining your model but not your autoencoder, the features it creates may very well be useless after some point. Now I'm not saying feature engineering is not applicable here, but I think the gain is on the margins, and I would test performance for each new created feature.",
      "votes": null
    },
    {
      "id": "3056500",
      "postDate": "11/27/2024 03:21:16",
      "content": "<p>based on your insight, I suspect these features are already constructed by feature engineering and may have involved previous information.</p>",
      "rawMarkdown": "based on your insight, I suspect these features are already constructed by feature engineering and may have involved previous information.",
      "votes": null
    },
    {
      "id": "3056618",
      "postDate": "11/27/2024 07:24:56",
      "content": "<p>yeah exactly!<br>\nsome features would enhance performance for some folds(intervals) while for other intervals it might make performance even worse. </p>",
      "rawMarkdown": "yeah exactly!\nsome features would enhance performance for some folds(intervals) while for other intervals it might make performance even worse.",
      "votes": null
    },
    {
      "id": "3057149",
      "postDate": "11/27/2024 18:57:46",
      "content": "<p>please Can you share your pipeline for online learning?</p>",
      "rawMarkdown": "please Can you share your pipeline for online learning?",
      "votes": null
    },
    {
      "id": "3066207",
      "postDate": "12/07/2024 20:17:44",
      "content": "<p>How much time does catboost model only cost with online learning? 4 hours?</p>",
      "rawMarkdown": "How much time does catboost model only cost with online learning? 4 hours?",
      "votes": null
    },
    {
      "id": "3066232",
      "postDate": "12/07/2024 21:13:07",
      "content": "<p>Depending on the model update frequency and online training strategy. Yes, 4 hours, sometimes. </p>",
      "rawMarkdown": "Depending on the model update frequency and online training strategy. Yes, 4 hours, sometimes.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 3056213,
      "author_name": "lihaorocky",
      "author_url": "",
      "post_date": "11/26/2024 17:12:15",
      "content": "<p>In my case, feature engineering (additional ~30 features) could boost my lgb model from 0.0048 to 0.0059 (LB), while CatBoost model from 0.0043 to 0.0052 (without online learning) and 0.0060 (with online learning). But as I said in another thread, I think in this competition, because of the anonymous features and 1 minute time limitation, GBDT is kind of a dead end. You can't get a gold medal using pure GBDT methods (at least from my knowledge), which is not the case in previous Optiver competition (I won 10th using a single CatBoost model). </p>",
      "votes": null,
      "replies": [
        {
          "id": 3056216,
          "author_name": "aymanallawi",
          "author_url": "",
          "post_date": "11/26/2024 17:17:49",
          "content": "<p>are the features you added like lagged and moving window features or did you did a more complex feature engineering. I am only curious if those features that enhance your performance are contextual features or something else </p>",
          "votes": null,
          "replies": [
            {
              "id": 3056222,
              "author_name": "lihaorocky",
              "author_url": "",
              "post_date": "11/26/2024 17:30:03",
              "content": "<p>Yes, they are simple moving window features. I mean those provided features are abonymous, what else feature engineering we could do about them?😃</p>",
              "votes": null,
              "replies": [
                {
                  "id": 3056224,
                  "author_name": "aymanallawi",
                  "author_url": "",
                  "post_date": "11/26/2024 17:32:13",
                  "content": "<p>man that make me want to cry you wouldn't imagine how much time and effort i put in feature engineering with no result so far 😢<br>\nhowever, good job and thanks for answering </p>",
                  "votes": null,
                  "replies": []
                },
                {
                  "id": 3056231,
                  "author_name": "woprime",
                  "author_url": "",
                  "post_date": "11/26/2024 17:49:30",
                  "content": "<p>Thanks for providing that insight, one thing I'm curious about is your choice of CatBoost. I reckon that it is specialized in categorical features, which we do not have (Assume you don't have a lot of engineered categorical features), and since it performs worse than LightGBM on LB, what's the motive of using that for online training instead of lightGBM? </p>",
                  "votes": null,
                  "replies": [
                    {
                      "id": 3056242,
                      "author_name": "lihaorocky",
                      "author_url": "",
                      "post_date": "11/26/2024 18:02:54",
                      "content": "<p><code>I reckon that it is specialized in categorical features</code><br>\nI'm not sure about this actually. In previous optiver competition, CatBoost is much better compared to LGB in CV and LB, even though all my features are numerical (same in this comp). In this competition, actually from local CV, the CatBoost is always better than LGB with a clear margin (tested with multiple validation vintages), but worse than LGB in LB (couldn't figure out why yet, maybe it's just coincidence). <br>\n<code>what's the motive of using that for online training instead of lightGBM</code><br>\nUsing CatBoost for online learning has two reasons, one is it's better and more stable from local CV, another reason is it's simply faster and consuming lower ram compared to LGB (I still couldn't figure out a stable way to do online training with 110+ features using comparable data amount which I could easily make it for catboost online learning). </p>",
                      "votes": null,
                      "replies": [
                        {
                          "id": 3056248,
                          "author_name": "woprime",
                          "author_url": "",
                          "post_date": "11/26/2024 18:13:49",
                          "content": "<p>Thanks for the reply! That is true, doing online training with LGBM seems to consume a lot of RAM, and a lot of the time leads to memory leakage, I haven't tested with Catboost yet, but this gives me a lot to think about. I still think it is possible to squeeze a few drops out of GBM, which could also be my delusion😆.</p>",
                          "votes": null,
                          "replies": []
                        }
                      ]
                    }
                  ]
                }
              ]
            }
          ]
        },
        {
          "id": 3057149,
          "author_name": "pearsejim01",
          "author_url": "",
          "post_date": "11/27/2024 18:57:46",
          "content": "<p>please Can you share your pipeline for online learning?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 3066207,
          "author_name": "yuanzhezhou",
          "author_url": "",
          "post_date": "12/07/2024 20:17:44",
          "content": "<p>How much time does catboost model only cost with online learning? 4 hours?</p>",
          "votes": null,
          "replies": [
            {
              "id": 3066232,
              "author_name": "lihaorocky",
              "author_url": "",
              "post_date": "12/07/2024 21:13:07",
              "content": "<p>Depending on the model update frequency and online training strategy. Yes, 4 hours, sometimes. </p>",
              "votes": null,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 3056411,
      "author_name": "natanlabarrere",
      "author_url": "",
      "post_date": "11/27/2024 00:26:10",
      "content": "<p>What I gather is in this context whatever pattern you find is generally very time-specific. Autoencoders for example combine the features in unique ways that may be very useful for some time intervals but cause whatever model you use them in to overfit. Take one of the pilars of this competition, for example - online learning. If you're retraining your model but not your autoencoder, the features it creates may very well be useless after some point. Now I'm not saying feature engineering is not applicable here, but I think the gain is on the margins, and I would test performance for each new created feature.</p>",
      "votes": null,
      "replies": [
        {
          "id": 3056618,
          "author_name": "aymanallawi",
          "author_url": "",
          "post_date": "11/27/2024 07:24:56",
          "content": "<p>yeah exactly!<br>\nsome features would enhance performance for some folds(intervals) while for other intervals it might make performance even worse. </p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 3056500,
      "author_name": "metropolis40",
      "author_url": "",
      "post_date": "11/27/2024 03:21:16",
      "content": "<p>based on your insight, I suspect these features are already constructed by feature engineering and may have involved previous information.</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "3056147": "Hey everyone,\n\nI'm curious about the role of context in this type of problem. I've tried extensive feature engineering and employed some of the most robust and computationally expensive methods you can imagine. However, adding features that capture the context of previous instances consistently resulted in performance similar to using the original features, with only minor variations.\n\nThis outcome was unexpected, especially since I’m using models like standard MLPs and GBMs, which are typically blind to context. I had assumed that incorporating features capturing the context would significantly enhance performance, but that hasn’t been the case.\n\nHas anyone succeeded in boosting performance through feature engineering in this competition? I would greatly appreciate it if someone could explain how context might be irrelevant or underutilized in this time series problem. Any insights or suggestions are welcome!",
    "3056213": "In my case, feature engineering (additional ~30 features) could boost my lgb model from 0.0048 to 0.0059 (LB), while CatBoost model from 0.0043 to 0.0052 (without online learning) and 0.0060 (with online learning). But as I said in another thread, I think in this competition, because of the anonymous features and 1 minute time limitation, GBDT is kind of a dead end. You can't get a gold medal using pure GBDT methods (at least from my knowledge), which is not the case in previous Optiver competition (I won 10th using a single CatBoost model).",
    "3056216": "are the features you added like lagged and moving window features or did you did a more complex feature engineering. I am only curious if those features that enhance your performance are contextual features or something else",
    "3056222": "Yes, they are simple moving window features. I mean those provided features are abonymous, what else feature engineering we could do about them?😃",
    "3056224": "man that make me want to cry you wouldn't imagine how much time and effort i put in feature engineering with no result so far 😢\nhowever, good job and thanks for answering",
    "3056231": "Thanks for providing that insight, one thing I'm curious about is your choice of CatBoost. I reckon that it is specialized in categorical features, which we do not have (Assume you don't have a lot of engineered categorical features), and since it performs worse than LightGBM on LB, what's the motive of using that for online training instead of lightGBM?",
    "3056242": "`I reckon that it is specialized in categorical features`\nI'm not sure about this actually. In previous optiver competition, CatBoost is much better compared to LGB in CV and LB, even though all my features are numerical (same in this comp). In this competition, actually from local CV, the CatBoost is always better than LGB with a clear margin (tested with multiple validation vintages), but worse than LGB in LB (couldn't figure out why yet, maybe it's just coincidence). \n`what's the motive of using that for online training instead of lightGBM`\nUsing CatBoost for online learning has two reasons, one is it's better and more stable from local CV, another reason is it's simply faster and consuming lower ram compared to LGB (I still couldn't figure out a stable way to do online training with 110+ features using comparable data amount which I could easily make it for catboost online learning).",
    "3056248": "Thanks for the reply! That is true, doing online training with LGBM seems to consume a lot of RAM, and a lot of the time leads to memory leakage, I haven't tested with Catboost yet, but this gives me a lot to think about. I still think it is possible to squeeze a few drops out of GBM, which could also be my delusion😆.",
    "3056411": "What I gather is in this context whatever pattern you find is generally very time-specific. Autoencoders for example combine the features in unique ways that may be very useful for some time intervals but cause whatever model you use them in to overfit. Take one of the pilars of this competition, for example - online learning. If you're retraining your model but not your autoencoder, the features it creates may very well be useless after some point. Now I'm not saying feature engineering is not applicable here, but I think the gain is on the margins, and I would test performance for each new created feature.",
    "3056500": "based on your insight, I suspect these features are already constructed by feature engineering and may have involved previous information.",
    "3056618": "yeah exactly!\nsome features would enhance performance for some folds(intervals) while for other intervals it might make performance even worse.",
    "3057149": "please Can you share your pipeline for online learning?",
    "3066207": "How much time does catboost model only cost with online learning? 4 hours?",
    "3066232": "Depending on the model update frequency and online training strategy. Yes, 4 hours, sometimes."
  },
  "source": "meta"
}