{
  "id": 548334,
  "title": "What is the forecast horizon?",
  "url": "/competitions/jane-street-real-time-market-data-forecasting/discussion/548334",
  "author_name": "",
  "post_date": "2024-11-26T07:40:06.437258300Z",
  "votes": 2,
  "comment_count": 2,
  "views": 0,
  "content": "<p>I am confused by the forecast horizon and I am struggling with the API.</p>\n<p>In my understanding, we have:</p>\n<p>[~-----  public set --------], [-----test set------] ,[ ---------private set-----------]</p>\n<p>During the training phase, we train our model (e.g., say, model A) based on the public set, and we get scores based on the unseen test set.</p>\n<p>During the forecasting phase, we use model A to predict based on the data from the private set (i.e., the undisclosed test set plus future data)</p>\n<p>so here are the questions:</p>\n<p>Q1: the test set will never be seen, however, when the forecasting phase starts, do we get the chance (via the code) to re-train model A (e.g., update the parameters) using the data in the test set?</p>\n<p>Q2: I am a bit confused by 'lags. parquet', does it mean when we conduct the forecast, we DO get access to the data for the same symbol_id for the first row of the time_id in the previous date? so are we forecasting one day ahead?</p>\n<p>If there are gaps between each set, and if we have a very long forecasting horizon - then I think it looks more like a cross-section problem rather than a time series problem.</p>\n<p>Thanks for your comments in advance.</p>",
  "messages": [
    {
      "id": "3055855",
      "postDate": "11/26/2024 07:40:06",
      "content": "<p>I am confused by the forecast horizon and I am struggling with the API.</p>\n<p>In my understanding, we have:</p>\n<p>[~-----  public set --------], [-----test set------] ,[ ---------private set-----------]</p>\n<p>During the training phase, we train our model (e.g., say, model A) based on the public set, and we get scores based on the unseen test set.</p>\n<p>During the forecasting phase, we use model A to predict based on the data from the private set (i.e., the undisclosed test set plus future data)</p>\n<p>so here are the questions:</p>\n<p>Q1: the test set will never be seen, however, when the forecasting phase starts, do we get the chance (via the code) to re-train model A (e.g., update the parameters) using the data in the test set?</p>\n<p>Q2: I am a bit confused by 'lags. parquet', does it mean when we conduct the forecast, we DO get access to the data for the same symbol_id for the first row of the time_id in the previous date? so are we forecasting one day ahead?</p>\n<p>If there are gaps between each set, and if we have a very long forecasting horizon - then I think it looks more like a cross-section problem rather than a time series problem.</p>\n<p>Thanks for your comments in advance.</p>",
      "rawMarkdown": "I am confused by the forecast horizon and I am struggling with the API.\n\nIn my understanding, we have:\n\n[~-----  public set --------], [-----test set------] ,[ ---------private set-----------]\n          \n\nDuring the training phase, we train our model (e.g., say, model A) based on the public set, and we get scores based on the unseen test set.\n\nDuring the forecasting phase, we use model A to predict based on the data from the private set (i.e., the undisclosed test set plus future data)\n\nso here are the questions:\n\nQ1: the test set will never be seen, however, when the forecasting phase starts, do we get the chance (via the code) to re-train model A (e.g., update the parameters) using the data in the test set?\n\nQ2: I am a bit confused by 'lags. parquet', does it mean when we conduct the forecast, we DO get access to the data for the same symbol_id for the first row of the time_id in the previous date? so are we forecasting one day ahead?\n\n\n\nIf there are gaps between each set, and if we have a very long forecasting horizon - then I think it looks more like a cross-section problem rather than a time series problem.\n\n\nThanks for your comments in advance.",
      "votes": null
    },
    {
      "id": "3056701",
      "postDate": "11/27/2024 09:49:47",
      "content": "<p>I think in the forecasting phase, a chunk of what you're calling \"private set\" will become \"test set\", that is, we won't have access to it but it will count to our public scores, so that we can validate if our model is still working on a longer time interval, closer to the actual private set, that will be used for the final scores. So you get a chance not to retrain your model with new data, but to validate it, and update your premises.<br>\nIn the lags you get the true responders for the previous day, that's all it is. You can use it as features (though most have been unsuccessful), or you can use them to calibrate your model with new data i.e. online learning.</p>",
      "rawMarkdown": "I think in the forecasting phase, a chunk of what you're calling \"private set\" will become \"test set\", that is, we won't have access to it but it will count to our public scores, so that we can validate if our model is still working on a longer time interval, closer to the actual private set, that will be used for the final scores. So you get a chance not to retrain your model with new data, but to validate it, and update your premises.\nIn the lags you get the true responders for the previous day, that's all it is. You can use it as features (though most have been unsuccessful), or you can use them to calibrate your model with new data i.e. online learning.",
      "votes": null
    },
    {
      "id": "3057336",
      "postDate": "11/28/2024 03:28:33",
      "content": "<p>Thanks for the clarification.</p>\n<p>Following your comments, suppose we have submitted our models and the time has come to 4 May 2025- Can we use any 'true value' of responder_6 until 3 May 2025 to predict the responder_6 on 4 May 2025?</p>\n<p>and then, suppose one day after,  on 5 May 2025, Can we use any 'true value' of responder_6 on 4 May 2025 to predict the responder_6 on 5 May 2025?</p>",
      "rawMarkdown": "Thanks for the clarification.\n\nFollowing your comments, suppose we have submitted our models and the time has come to 4 May 2025- Can we use any 'true value' of responder_6 until 3 May 2025 to predict the responder_6 on 4 May 2025?\n\nand then, suppose one day after,  on 5 May 2025, Can we use any 'true value' of responder_6 on 4 May 2025 to predict the responder_6 on 5 May 2025?",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 3056701,
      "author_name": "natanlabarrere",
      "author_url": "",
      "post_date": "11/27/2024 09:49:47",
      "content": "<p>I think in the forecasting phase, a chunk of what you're calling \"private set\" will become \"test set\", that is, we won't have access to it but it will count to our public scores, so that we can validate if our model is still working on a longer time interval, closer to the actual private set, that will be used for the final scores. So you get a chance not to retrain your model with new data, but to validate it, and update your premises.<br>\nIn the lags you get the true responders for the previous day, that's all it is. You can use it as features (though most have been unsuccessful), or you can use them to calibrate your model with new data i.e. online learning.</p>",
      "votes": null,
      "replies": [
        {
          "id": 3057336,
          "author_name": "metropolis40",
          "author_url": "",
          "post_date": "11/28/2024 03:28:33",
          "content": "<p>Thanks for the clarification.</p>\n<p>Following your comments, suppose we have submitted our models and the time has come to 4 May 2025- Can we use any 'true value' of responder_6 until 3 May 2025 to predict the responder_6 on 4 May 2025?</p>\n<p>and then, suppose one day after,  on 5 May 2025, Can we use any 'true value' of responder_6 on 4 May 2025 to predict the responder_6 on 5 May 2025?</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "3055855": "I am confused by the forecast horizon and I am struggling with the API.\n\nIn my understanding, we have:\n\n[~-----  public set --------], [-----test set------] ,[ ---------private set-----------]\n          \n\nDuring the training phase, we train our model (e.g., say, model A) based on the public set, and we get scores based on the unseen test set.\n\nDuring the forecasting phase, we use model A to predict based on the data from the private set (i.e., the undisclosed test set plus future data)\n\nso here are the questions:\n\nQ1: the test set will never be seen, however, when the forecasting phase starts, do we get the chance (via the code) to re-train model A (e.g., update the parameters) using the data in the test set?\n\nQ2: I am a bit confused by 'lags. parquet', does it mean when we conduct the forecast, we DO get access to the data for the same symbol_id for the first row of the time_id in the previous date? so are we forecasting one day ahead?\n\n\n\nIf there are gaps between each set, and if we have a very long forecasting horizon - then I think it looks more like a cross-section problem rather than a time series problem.\n\n\nThanks for your comments in advance.",
    "3056701": "I think in the forecasting phase, a chunk of what you're calling \"private set\" will become \"test set\", that is, we won't have access to it but it will count to our public scores, so that we can validate if our model is still working on a longer time interval, closer to the actual private set, that will be used for the final scores. So you get a chance not to retrain your model with new data, but to validate it, and update your premises.\nIn the lags you get the true responders for the previous day, that's all it is. You can use it as features (though most have been unsuccessful), or you can use them to calibrate your model with new data i.e. online learning.",
    "3057336": "Thanks for the clarification.\n\nFollowing your comments, suppose we have submitted our models and the time has come to 4 May 2025- Can we use any 'true value' of responder_6 until 3 May 2025 to predict the responder_6 on 4 May 2025?\n\nand then, suppose one day after,  on 5 May 2025, Can we use any 'true value' of responder_6 on 4 May 2025 to predict the responder_6 on 5 May 2025?"
  },
  "source": "meta"
}