{
  "id": 543081,
  "title": "How to keep historical cache in the submission for time-series modelling",
  "url": "/competitions/jane-street-real-time-market-data-forecasting/discussion/543081",
  "author_name": "SLi",
  "post_date": "2024-10-28T14:07:20.201000",
  "votes": 30,
  "comment_count": 9,
  "views": 0,
  "content": "<p>For this competition, while not many existing works leverage the time-series aspect of features, I believe incorporating temporal data will become increasingly valuable over time. One key technical challenge with time-series input is maintaining and managing the sequential data within the submission API—a task I found neither straightforward nor trivial. So, I’d like to share my approach to implementing a historical cache, and I’d appreciate any advice from the community to refine it further!</p>\n<p>To start, I’ve outlined my setup in this notebook: <a href=\"https://www.kaggle.com/code/shiyili/js2024-rmf-gru-inference\" target=\"_blank\"><strong>Link</strong></a>.</p>\n<p>In the notebook, I explored using a GRU model on the time-series of features with a sequence length of 100. (The model didn’t perform well so don’t mind the private dataset.) </p>\n<p>It was suggested using global vars to store the historical data (e.g. <code>lags_</code>) in the pinned demo. However, I found object-oriented design (OOD) to be a more structured approach for managing variables, so I created the <code>JaneStreetPredictor</code> class, which organizes the historical cache, previous day lags, and progress bars for easier control.</p>\n<p>For caching, I used a dictionary where each key is a symbol ID, and each value holds a time-series of features in an <code>np.array</code>. The dictionary updates dynamically in the predict method, trimming the cache to the required sequence length to prevent memory overflow.</p>\n<p>If you’re curious about runtime, you can copy the code and try it with any model of your choice. Setting <code>USE_SYNTHETIC = True</code> to run inference on synthetic test data would allow for more accessible debugging. In general, scoring on real submission takes ~30x the synthetic runtime (e.g., if synthetic data takes 2 minutes, expect around 1 hour for actual scoring).</p>\n<p>One thing worth mentioning: the dates in the train and test sets are likely continuous (though the test set starts with <code>date_id=0</code>). If that’s the case, we can initialize the cache with data from the last date in the train set, which could provide a smoother transition into the test phase.</p>\n<p>I’d love to hear your thoughts on making this approach more robust and efficient. Happy Kaggling!</p>",
  "messages": [
    {
      "id": 3030416,
      "postDate": "2024-10-28T14:07:20.203Z",
      "content": "<p>For this competition, while not many existing works leverage the time-series aspect of features, I believe incorporating temporal data will become increasingly valuable over time. One key technical challenge with time-series input is maintaining and managing the sequential data within the submission API—a task I found neither straightforward nor trivial. So, I’d like to share my approach to implementing a historical cache, and I’d appreciate any advice from the community to refine it further!</p>\n<p>To start, I’ve outlined my setup in this notebook: <a href=\"https://www.kaggle.com/code/shiyili/js2024-rmf-gru-inference\" target=\"_blank\"><strong>Link</strong></a>.</p>\n<p>In the notebook, I explored using a GRU model on the time-series of features with a sequence length of 100. (The model didn’t perform well so don’t mind the private dataset.) </p>\n<p>It was suggested using global vars to store the historical data (e.g. <code>lags_</code>) in the pinned demo. However, I found object-oriented design (OOD) to be a more structured approach for managing variables, so I created the <code>JaneStreetPredictor</code> class, which organizes the historical cache, previous day lags, and progress bars for easier control.</p>\n<p>For caching, I used a dictionary where each key is a symbol ID, and each value holds a time-series of features in an <code>np.array</code>. The dictionary updates dynamically in the predict method, trimming the cache to the required sequence length to prevent memory overflow.</p>\n<p>If you’re curious about runtime, you can copy the code and try it with any model of your choice. Setting <code>USE_SYNTHETIC = True</code> to run inference on synthetic test data would allow for more accessible debugging. In general, scoring on real submission takes ~30x the synthetic runtime (e.g., if synthetic data takes 2 minutes, expect around 1 hour for actual scoring).</p>\n<p>One thing worth mentioning: the dates in the train and test sets are likely continuous (though the test set starts with <code>date_id=0</code>). If that’s the case, we can initialize the cache with data from the last date in the train set, which could provide a smoother transition into the test phase.</p>\n<p>I’d love to hear your thoughts on making this approach more robust and efficient. Happy Kaggling!</p>",
      "rawMarkdown": "For this competition, while not many existing works leverage the time-series aspect of features, I believe incorporating temporal data will become increasingly valuable over time. One key technical challenge with time-series input is maintaining and managing the sequential data within the submission API—a task I found neither straightforward nor trivial. So, I’d like to share my approach to implementing a historical cache, and I’d appreciate any advice from the community to refine it further!\n\nTo start, I’ve outlined my setup in this notebook: [**Link**](https://www.kaggle.com/code/shiyili/js2024-rmf-gru-inference).\n\nIn the notebook, I explored using a GRU model on the time-series of features with a sequence length of 100. (The model didn’t perform well so don’t mind the private dataset.) \n\nIt was suggested using global vars to store the historical data (e.g. `lags_`) in the pinned demo. However, I found object-oriented design (OOD) to be a more structured approach for managing variables, so I created the `JaneStreetPredictor` class, which organizes the historical cache, previous day lags, and progress bars for easier control.\n\nFor caching, I used a dictionary where each key is a symbol ID, and each value holds a time-series of features in an `np.array`. The dictionary updates dynamically in the predict method, trimming the cache to the required sequence length to prevent memory overflow.\n\nIf you’re curious about runtime, you can copy the code and try it with any model of your choice. Setting `USE_SYNTHETIC = True` to run inference on synthetic test data would allow for more accessible debugging. In general, scoring on real submission takes ~30x the synthetic runtime (e.g., if synthetic data takes 2 minutes, expect around 1 hour for actual scoring).\n\nOne thing worth mentioning: the dates in the train and test sets are likely continuous (though the test set starts with `date_id=0`). If that’s the case, we can initialize the cache with data from the last date in the train set, which could provide a smoother transition into the test phase.\n\nI’d love to hear your thoughts on making this approach more robust and efficient. Happy Kaggling!",
      "votes": 30
    },
    {
      "id": 3034825,
      "postDate": "2024-11-02T16:55:08.510Z",
      "content": "<p>Thank you very much for this insight, you mentioned continuity in the date IDs from the training set to the test set while forecasting(although the sample given only contains  date_id=0) and I believe this should be the only logical way to ensure continuity in the data flow, does the host confirm this that the date_id in the test continues from where the train_data stops or otherwise, Would appreciate your input. Thanks once again for sharing this information</p>",
      "rawMarkdown": "Thank you very much for this insight, you mentioned continuity in the date IDs from the training set to the test set while forecasting(although the sample given only contains  date_id=0) and I believe this should be the only logical way to ensure continuity in the data flow, does the host confirm this that the date_id in the test continues from where the train_data stops or otherwise, Would appreciate your input. Thanks once again for sharing this information",
      "votes": 1,
      "replies": [
        {
          "id": 3034890,
          "postDate": "2024-11-02T18:20:12.103Z",
          "content": "<p>Yes. This was described in the data card.</p>",
          "rawMarkdown": "Yes. This was described in the data card.",
          "replies": [
            {
              "id": 3035715,
              "postDate": "2024-11-03T19:42:12.650Z",
              "content": "<p><a href=\"https://www.kaggle.com/shiyili\" target=\"_blank\">@shiyili</a> I don't see it from data card. Could you quote the description here? thanks!</p>",
              "rawMarkdown": "@shiyili I don't see it from data card. Could you quote the description here? thanks!"
            },
            {
              "id": 3035804,
              "postDate": "2024-11-03T21:22:33.200Z",
              "content": "<p>Thank you very much for this,I guess I will work with this idea then</p>",
              "rawMarkdown": "Thank you very much for this,I guess I will work with this idea then"
            }
          ]
        }
      ]
    },
    {
      "id": 3032144,
      "postDate": "2024-10-30T15:34:28.413Z",
      "content": "<p>Is this allowed during the forecasting phase? I wasn't sure if this was against the rules, since I haven't seen an approach like this yet (not even in last years challenge)</p>",
      "rawMarkdown": "Is this allowed during the forecasting phase? I wasn't sure if this was against the rules, since I haven't seen an approach like this yet (not even in last years challenge)",
      "replies": [
        {
          "id": 3032154,
          "postDate": "2024-10-30T15:54:37.630Z",
          "content": "<p>It is allowed. You can check the Optiver trading at the close competition. </p>",
          "rawMarkdown": "It is allowed. You can check the Optiver trading at the close competition. ",
          "votes": 1
        }
      ]
    },
    {
      "id": 3032006,
      "postDate": "2024-10-30T12:51:21.870Z",
      "content": "<p>Thank you for sharing this example, awesome! I was just looking for a way to do historical caching. </p>\n<p>Have you been able to leverage this setup with GBMs? For example <code>lag_2</code>, <code>lag_3</code> features on the responders?</p>",
      "rawMarkdown": "Thank you for sharing this example, awesome! I was just looking for a way to do historical caching. \n\nHave you been able to leverage this setup with GBMs? For example `lag_2`, `lag_3` features on the responders?",
      "replies": [
        {
          "id": 3032037,
          "postDate": "2024-10-30T13:32:10.157Z",
          "content": "<p>Yes. For example here: <a href=\"https://www.kaggle.com/code/shiyili/janestreet-2024-gbdt-inference|\" target=\"_blank\">https://www.kaggle.com/code/shiyili/janestreet-2024-gbdt-inference|</a></p>\n<p>I generated a responder_6_lag feature using the last record from the previous day. You can easily apply a similar approach to any other type of lagged features as you like. If you look at the commented code in the predict function, you will noticed that I also tried using the statistics of the lagged responders as features (but didn't work well).</p>",
          "rawMarkdown": "Yes. For example here: https://www.kaggle.com/code/shiyili/janestreet-2024-gbdt-inference|\n\nI generated a responder_6_lag feature using the last record from the previous day. You can easily apply a similar approach to any other type of lagged features as you like. If you look at the commented code in the predict function, you will noticed that I also tried using the statistics of the lagged responders as features (but didn't work well)."
        }
      ]
    },
    {
      "id": 3035925,
      "postDate": "2024-11-04T02:43:15.763Z",
      "rawMarkdown": "",
      "votes": -1,
      "isDeleted": true
    }
  ],
  "comments": [
    {
      "id": 3034825,
      "author_name": "Oluwatobi Betiku",
      "author_url": "",
      "post_date": "2024-11-02T16:55:08.510000",
      "content": "<p>Thank you very much for this insight, you mentioned continuity in the date IDs from the training set to the test set while forecasting(although the sample given only contains  date_id=0) and I believe this should be the only logical way to ensure continuity in the data flow, does the host confirm this that the date_id in the test continues from where the train_data stops or otherwise, Would appreciate your input. Thanks once again for sharing this information</p>",
      "votes": 1,
      "replies": [
        {
          "id": 3034890,
          "author_name": "SLi",
          "author_url": "",
          "post_date": "2024-11-02T18:20:12.103000",
          "content": "<p>Yes. This was described in the data card.</p>",
          "votes": 0,
          "replies": [
            {
              "id": 3035715,
              "author_name": "san tan",
              "author_url": "",
              "post_date": "2024-11-03T19:42:12.650000",
              "content": "<p><a href=\"https://www.kaggle.com/shiyili\" target=\"_blank\">@shiyili</a> I don't see it from data card. Could you quote the description here? thanks!</p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 3035804,
              "author_name": "Oluwatobi Betiku",
              "author_url": "",
              "post_date": "2024-11-03T21:22:33.200000",
              "content": "<p>Thank you very much for this,I guess I will work with this idea then</p>",
              "votes": 0,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 3032144,
      "author_name": "chr",
      "author_url": "",
      "post_date": "2024-10-30T15:34:28.413000",
      "content": "<p>Is this allowed during the forecasting phase? I wasn't sure if this was against the rules, since I haven't seen an approach like this yet (not even in last years challenge)</p>",
      "votes": 0,
      "replies": [
        {
          "id": 3032154,
          "author_name": "SLi",
          "author_url": "",
          "post_date": "2024-10-30T15:54:37.630000",
          "content": "<p>It is allowed. You can check the Optiver trading at the close competition. </p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 3032006,
      "author_name": "Carlo",
      "author_url": "",
      "post_date": "2024-10-30T12:51:21.870000",
      "content": "<p>Thank you for sharing this example, awesome! I was just looking for a way to do historical caching. </p>\n<p>Have you been able to leverage this setup with GBMs? For example <code>lag_2</code>, <code>lag_3</code> features on the responders?</p>",
      "votes": 0,
      "replies": [
        {
          "id": 3032037,
          "author_name": "SLi",
          "author_url": "",
          "post_date": "2024-10-30T13:32:10.157000",
          "content": "<p>Yes. For example here: <a href=\"https://www.kaggle.com/code/shiyili/janestreet-2024-gbdt-inference|\" target=\"_blank\">https://www.kaggle.com/code/shiyili/janestreet-2024-gbdt-inference|</a></p>\n<p>I generated a responder_6_lag feature using the last record from the previous day. You can easily apply a similar approach to any other type of lagged features as you like. If you look at the commented code in the predict function, you will noticed that I also tried using the statistics of the lagged responders as features (but didn't work well).</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 3035925,
      "author_name": "",
      "author_url": "",
      "post_date": "2024-11-04T02:43:15.763000",
      "content": "",
      "votes": -1,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "3030416": "For this competition, while not many existing works leverage the time-series aspect of features, I believe incorporating temporal data will become increasingly valuable over time. One key technical challenge with time-series input is maintaining and managing the sequential data within the submission API—a task I found neither straightforward nor trivial. So, I’d like to share my approach to implementing a historical cache, and I’d appreciate any advice from the community to refine it further!\n\nTo start, I’ve outlined my setup in this notebook: [**Link**](https://www.kaggle.com/code/shiyili/js2024-rmf-gru-inference).\n\nIn the notebook, I explored using a GRU model on the time-series of features with a sequence length of 100. (The model didn’t perform well so don’t mind the private dataset.) \n\nIt was suggested using global vars to store the historical data (e.g. `lags_`) in the pinned demo. However, I found object-oriented design (OOD) to be a more structured approach for managing variables, so I created the `JaneStreetPredictor` class, which organizes the historical cache, previous day lags, and progress bars for easier control.\n\nFor caching, I used a dictionary where each key is a symbol ID, and each value holds a time-series of features in an `np.array`. The dictionary updates dynamically in the predict method, trimming the cache to the required sequence length to prevent memory overflow.\n\nIf you’re curious about runtime, you can copy the code and try it with any model of your choice. Setting `USE_SYNTHETIC = True` to run inference on synthetic test data would allow for more accessible debugging. In general, scoring on real submission takes ~30x the synthetic runtime (e.g., if synthetic data takes 2 minutes, expect around 1 hour for actual scoring).\n\nOne thing worth mentioning: the dates in the train and test sets are likely continuous (though the test set starts with `date_id=0`). If that’s the case, we can initialize the cache with data from the last date in the train set, which could provide a smoother transition into the test phase.\n\nI’d love to hear your thoughts on making this approach more robust and efficient. Happy Kaggling!",
    "3034825": "Thank you very much for this insight, you mentioned continuity in the date IDs from the training set to the test set while forecasting(although the sample given only contains  date_id=0) and I believe this should be the only logical way to ensure continuity in the data flow, does the host confirm this that the date_id in the test continues from where the train_data stops or otherwise, Would appreciate your input. Thanks once again for sharing this information",
    "3032144": "Is this allowed during the forecasting phase? I wasn't sure if this was against the rules, since I haven't seen an approach like this yet (not even in last years challenge)",
    "3032006": "Thank you for sharing this example, awesome! I was just looking for a way to do historical caching. \n\nHave you been able to leverage this setup with GBMs? For example `lag_2`, `lag_3` features on the responders?",
    "3035925": ""
  }
}