{
  "id": 549550,
  "title": "ARIMA Online Learning",
  "url": "/competitions/jane-street-real-time-market-data-forecasting/discussion/549550",
  "author_name": "Sergio Henrique",
  "post_date": "2024-12-02T20:36:57.625000",
  "votes": 14,
  "comment_count": 7,
  "views": 0,
  "content": "<p>Hello</p>\n<p>I just made a proof of concept of how to address this challenge using online learning. <br>\n<strong>Notebook</strong>: <a href=\"https://www.kaggle.com/code/serjhenrique/jrtsmdf-arima-online-learning\" target=\"_blank\">https://www.kaggle.com/code/serjhenrique/jrtsmdf-arima-online-learning</a></p>\n<p>I wrote a blog post about it <a href=\"https://serjhenrique.com/arima-and-online-learning-in-financial-forecasting/\" target=\"_blank\">https://serjhenrique.com/arima-and-online-learning-in-financial-forecasting/</a>.</p>\n<p>For those who don't want to read the full post the following extraction can give a good idea of the work done:</p>\n<p>From this <a href=\"https://www.kaggle.com/code/chumajin/janestreet-easy-to-understand-new-time-series-api/notebook\" target=\"_blank\">notebook</a> we understood that new data is served by time_id. In other words, we receive a batch related from time_id = 0 and must return the predictions for this batch. According to the competition host, the lags parameter will be provided every time time_id = 0. Therefore, whenever the lags are not null, we can assume it is a new&nbsp;date_id. Based on this information, I chose to train the ARIMA model each time we encounter a new&nbsp;date_id.</p>\n<p>The main idea is to train an ARIMA model for each&nbsp;symbol_id&nbsp;whenever we have a new&nbsp;date_id. We predict an entire day (968 time units) for every symbol and store these predictions in the Jane Predictor object. Then, for every prediction call that doesn't have lags (meaning we are still within the same already predicted&nbsp;date_id), we retrieve the first prediction for each symbol from our Jane Predictor's&nbsp;symbol_arr&nbsp;and&nbsp;pred_arr, remove it from the Jane Predictor's attributes, join it with the test DataFrame, and return it as the prediction. In other words,&nbsp;symbol_arr&nbsp;and&nbsp;pred_arr&nbsp;function like a stack that we consume with each new&nbsp;time_id&nbsp;until we reach a new&nbsp;date_id.</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1342094%2Fe4e8e747646c636f9c1e23eaaf2de07c%2FARIMA-and-Online-Learning-in-Financial-Forecasting-img-001.png?generation=1733171761409554&amp;alt=media\" alt=\"\"></p>\n<p>I hope you enjoy!!</p>",
  "messages": [
    {
      "id": 3061588,
      "postDate": "2024-12-02T20:36:57.627Z",
      "content": "<p>Hello</p>\n<p>I just made a proof of concept of how to address this challenge using online learning. <br>\n<strong>Notebook</strong>: <a href=\"https://www.kaggle.com/code/serjhenrique/jrtsmdf-arima-online-learning\" target=\"_blank\">https://www.kaggle.com/code/serjhenrique/jrtsmdf-arima-online-learning</a></p>\n<p>I wrote a blog post about it <a href=\"https://serjhenrique.com/arima-and-online-learning-in-financial-forecasting/\" target=\"_blank\">https://serjhenrique.com/arima-and-online-learning-in-financial-forecasting/</a>.</p>\n<p>For those who don't want to read the full post the following extraction can give a good idea of the work done:</p>\n<p>From this <a href=\"https://www.kaggle.com/code/chumajin/janestreet-easy-to-understand-new-time-series-api/notebook\" target=\"_blank\">notebook</a> we understood that new data is served by time_id. In other words, we receive a batch related from time_id = 0 and must return the predictions for this batch. According to the competition host, the lags parameter will be provided every time time_id = 0. Therefore, whenever the lags are not null, we can assume it is a new&nbsp;date_id. Based on this information, I chose to train the ARIMA model each time we encounter a new&nbsp;date_id.</p>\n<p>The main idea is to train an ARIMA model for each&nbsp;symbol_id&nbsp;whenever we have a new&nbsp;date_id. We predict an entire day (968 time units) for every symbol and store these predictions in the Jane Predictor object. Then, for every prediction call that doesn't have lags (meaning we are still within the same already predicted&nbsp;date_id), we retrieve the first prediction for each symbol from our Jane Predictor's&nbsp;symbol_arr&nbsp;and&nbsp;pred_arr, remove it from the Jane Predictor's attributes, join it with the test DataFrame, and return it as the prediction. In other words,&nbsp;symbol_arr&nbsp;and&nbsp;pred_arr&nbsp;function like a stack that we consume with each new&nbsp;time_id&nbsp;until we reach a new&nbsp;date_id.</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1342094%2Fe4e8e747646c636f9c1e23eaaf2de07c%2FARIMA-and-Online-Learning-in-Financial-Forecasting-img-001.png?generation=1733171761409554&amp;alt=media\" alt=\"\"></p>\n<p>I hope you enjoy!!</p>",
      "rawMarkdown": "Hello\n\nI just made a proof of concept of how to address this challenge using online learning. \n**Notebook**: [https://www.kaggle.com/code/serjhenrique/jrtsmdf-arima-online-learning](https://www.kaggle.com/code/serjhenrique/jrtsmdf-arima-online-learning)\n\nI wrote a blog post about it [https://serjhenrique.com/arima-and-online-learning-in-financial-forecasting/](https://serjhenrique.com/arima-and-online-learning-in-financial-forecasting/).\n\nFor those who don't want to read the full post the following extraction can give a good idea of the work done:\n\nFrom this [notebook](https://www.kaggle.com/code/chumajin/janestreet-easy-to-understand-new-time-series-api/notebook) we understood that new data is served by time_id. In other words, we receive a batch related from time_id = 0 and must return the predictions for this batch. According to the competition host, the lags parameter will be provided every time time_id = 0. Therefore, whenever the lags are not null, we can assume it is a new date_id. Based on this information, I chose to train the ARIMA model each time we encounter a new date_id.\n\nThe main idea is to train an ARIMA model for each symbol_id whenever we have a new date_id. We predict an entire day (968 time units) for every symbol and store these predictions in the Jane Predictor object. Then, for every prediction call that doesn't have lags (meaning we are still within the same already predicted date_id), we retrieve the first prediction for each symbol from our Jane Predictor's symbol_arr and pred_arr, remove it from the Jane Predictor's attributes, join it with the test DataFrame, and return it as the prediction. In other words, symbol_arr and pred_arr function like a stack that we consume with each new time_id until we reach a new date_id.\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1342094%2Fe4e8e747646c636f9c1e23eaaf2de07c%2FARIMA-and-Online-Learning-in-Financial-Forecasting-img-001.png?generation=1733171761409554&alt=media)\n\nI hope you enjoy!!",
      "votes": 14
    },
    {
      "id": 3062212,
      "postDate": "2024-12-03T11:39:08.677Z",
      "content": "<p>It seems a graceful work! I wonder whether you try to submit. In my experiment, spliting by <code>symbol_id</code> and concating them may consume a lot of time leading to timeout.</p>",
      "rawMarkdown": "It seems a graceful work! I wonder whether you try to submit. In my experiment, spliting by `symbol_id` and concating them may consume a lot of time leading to timeout.",
      "replies": [
        {
          "id": 3062250,
          "postDate": "2024-12-03T12:19:32.870Z",
          "content": "<p>I made successful submissions with this code.</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1342094%2Fbcf36677623453ffc433af92caed099e%2Fsubmissions.png?generation=1733227824369502&amp;alt=media\" alt=\"\"></p>\n<p>I created a copy of my private notebook, cleaned up the code, and added comments before sharing it. In this approach, I am discarding features and only using responder_6 to predict itself. The next step is to change ARIMA to a more powerful model and use a small subset of features to improve the forecast. I think it serves as a baseline for building an online learning pipeline. I also tried to create classes to facilitate modifications and experimentation.</p>",
          "rawMarkdown": "I made successful submissions with this code.\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1342094%2Fbcf36677623453ffc433af92caed099e%2Fsubmissions.png?generation=1733227824369502&alt=media)\n\nI created a copy of my private notebook, cleaned up the code, and added comments before sharing it. In this approach, I am discarding features and only using responder_6 to predict itself. The next step is to change ARIMA to a more powerful model and use a small subset of features to improve the forecast. I think it serves as a baseline for building an online learning pipeline. I also tried to create classes to facilitate modifications and experimentation.",
          "votes": 1,
          "replies": [
            {
              "id": 3062274,
              "postDate": "2024-12-03T12:53:42.360Z",
              "content": "<p>Great work and thank you for your sharing! </p>",
              "rawMarkdown": "Great work and thank you for your sharing! "
            }
          ]
        }
      ]
    },
    {
      "id": 3064892,
      "postDate": "2024-12-06T05:10:10.857Z",
      "content": "<p>Thanks for sharing. I’m curious about the performance of the whole-day prediction. I considered using time-series models initially but was concerned that making hundreds of sequential predictions without being able to take into account intraday information would lead to significant error compounding.</p>",
      "rawMarkdown": "Thanks for sharing. I’m curious about the performance of the whole-day prediction. I considered using time-series models initially but was concerned that making hundreds of sequential predictions without being able to take into account intraday information would lead to significant error compounding.",
      "isDeleted": true,
      "replies": [
        {
          "id": 3065144,
          "postDate": "2024-12-06T12:16:50.167Z",
          "content": "<p>You are right. Discard intraday features hurts the forecast power of the model. I am refining this approach and currently changed the ARIMA model for a LGBM model:</p>\n<p><a href=\"https://www.kaggle.com/code/serjhenrique/jsrtmdf-lgbm-online-learning-evaluation-pb\" target=\"_blank\">https://www.kaggle.com/code/serjhenrique/jsrtmdf-lgbm-online-learning-evaluation-pb</a></p>",
          "rawMarkdown": "You are right. Discard intraday features hurts the forecast power of the model. I am refining this approach and currently changed the ARIMA model for a LGBM model:\n\n[https://www.kaggle.com/code/serjhenrique/jsrtmdf-lgbm-online-learning-evaluation-pb](https://www.kaggle.com/code/serjhenrique/jsrtmdf-lgbm-online-learning-evaluation-pb)"
        }
      ]
    },
    {
      "id": 3063559,
      "postDate": "2024-12-04T15:57:43.217Z",
      "content": "<p>thank you!</p>",
      "rawMarkdown": "thank you!"
    },
    {
      "id": 3062073,
      "postDate": "2024-12-03T09:03:05.233Z",
      "content": "<p>thanks for sharing!</p>",
      "rawMarkdown": "thanks for sharing!"
    }
  ],
  "comments": [
    {
      "id": 3062212,
      "author_name": "#1BuBu",
      "author_url": "",
      "post_date": "2024-12-03T11:39:08.677000",
      "content": "<p>It seems a graceful work! I wonder whether you try to submit. In my experiment, spliting by <code>symbol_id</code> and concating them may consume a lot of time leading to timeout.</p>",
      "votes": 0,
      "replies": [
        {
          "id": 3062250,
          "author_name": "Sergio Henrique",
          "author_url": "",
          "post_date": "2024-12-03T12:19:32.870000",
          "content": "<p>I made successful submissions with this code.</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1342094%2Fbcf36677623453ffc433af92caed099e%2Fsubmissions.png?generation=1733227824369502&amp;alt=media\" alt=\"\"></p>\n<p>I created a copy of my private notebook, cleaned up the code, and added comments before sharing it. In this approach, I am discarding features and only using responder_6 to predict itself. The next step is to change ARIMA to a more powerful model and use a small subset of features to improve the forecast. I think it serves as a baseline for building an online learning pipeline. I also tried to create classes to facilitate modifications and experimentation.</p>",
          "votes": 1,
          "replies": [
            {
              "id": 3062274,
              "author_name": "#1BuBu",
              "author_url": "",
              "post_date": "2024-12-03T12:53:42.360000",
              "content": "<p>Great work and thank you for your sharing! </p>",
              "votes": 0,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 3064892,
      "author_name": "",
      "author_url": "",
      "post_date": "2024-12-06T05:10:10.857000",
      "content": "<p>Thanks for sharing. I’m curious about the performance of the whole-day prediction. I considered using time-series models initially but was concerned that making hundreds of sequential predictions without being able to take into account intraday information would lead to significant error compounding.</p>",
      "votes": 0,
      "replies": [
        {
          "id": 3065144,
          "author_name": "Sergio Henrique",
          "author_url": "",
          "post_date": "2024-12-06T12:16:50.167000",
          "content": "<p>You are right. Discard intraday features hurts the forecast power of the model. I am refining this approach and currently changed the ARIMA model for a LGBM model:</p>\n<p><a href=\"https://www.kaggle.com/code/serjhenrique/jsrtmdf-lgbm-online-learning-evaluation-pb\" target=\"_blank\">https://www.kaggle.com/code/serjhenrique/jsrtmdf-lgbm-online-learning-evaluation-pb</a></p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 3063559,
      "author_name": "ironrro",
      "author_url": "",
      "post_date": "2024-12-04T15:57:43.217000",
      "content": "<p>thank you!</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 3062073,
      "author_name": "ZT",
      "author_url": "",
      "post_date": "2024-12-03T09:03:05.233000",
      "content": "<p>thanks for sharing!</p>",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "3061588": "Hello\n\nI just made a proof of concept of how to address this challenge using online learning. \n**Notebook**: [https://www.kaggle.com/code/serjhenrique/jrtsmdf-arima-online-learning](https://www.kaggle.com/code/serjhenrique/jrtsmdf-arima-online-learning)\n\nI wrote a blog post about it [https://serjhenrique.com/arima-and-online-learning-in-financial-forecasting/](https://serjhenrique.com/arima-and-online-learning-in-financial-forecasting/).\n\nFor those who don't want to read the full post the following extraction can give a good idea of the work done:\n\nFrom this [notebook](https://www.kaggle.com/code/chumajin/janestreet-easy-to-understand-new-time-series-api/notebook) we understood that new data is served by time_id. In other words, we receive a batch related from time_id = 0 and must return the predictions for this batch. According to the competition host, the lags parameter will be provided every time time_id = 0. Therefore, whenever the lags are not null, we can assume it is a new date_id. Based on this information, I chose to train the ARIMA model each time we encounter a new date_id.\n\nThe main idea is to train an ARIMA model for each symbol_id whenever we have a new date_id. We predict an entire day (968 time units) for every symbol and store these predictions in the Jane Predictor object. Then, for every prediction call that doesn't have lags (meaning we are still within the same already predicted date_id), we retrieve the first prediction for each symbol from our Jane Predictor's symbol_arr and pred_arr, remove it from the Jane Predictor's attributes, join it with the test DataFrame, and return it as the prediction. In other words, symbol_arr and pred_arr function like a stack that we consume with each new time_id until we reach a new date_id.\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1342094%2Fe4e8e747646c636f9c1e23eaaf2de07c%2FARIMA-and-Online-Learning-in-Financial-Forecasting-img-001.png?generation=1733171761409554&alt=media)\n\nI hope you enjoy!!",
    "3062212": "It seems a graceful work! I wonder whether you try to submit. In my experiment, spliting by `symbol_id` and concating them may consume a lot of time leading to timeout.",
    "3064892": "Thanks for sharing. I’m curious about the performance of the whole-day prediction. I considered using time-series models initially but was concerned that making hundreds of sequential predictions without being able to take into account intraday information would lead to significant error compounding.",
    "3063559": "thank you!",
    "3062073": "thanks for sharing!"
  }
}