{
  "id": 543327,
  "title": "Initial condition issues (day 0) due to no historical data",
  "url": "/competitions/jane-street-real-time-market-data-forecasting/discussion/543327",
  "author_name": "",
  "post_date": "2024-10-30T00:12:47.693186300Z",
  "votes": 2,
  "comment_count": 6,
  "views": 0,
  "content": "<p>on date_id = 0 for the test/evaluation period we are going to get a lags dataframe at time_id = 0 with the responder values for the entire previous day, however we will not have any of the feature values at this point so we have nothing to match them up against.</p>\n<p>That means on the very first test/eval day we won't have any data at all for models that look back at previous values. Assuming your model looks back at the features as well as responder values.</p>\n<p>I thought maybe the <code>is_scored</code> value might be set to false for the first day to allow us to cache a history of data but in the example sets it is set to true right from the start.</p>\n<p>Am i right in thinking that this means we cannot use models with this architecture for at least the first day? From day 1 onwards we will have roughly 600-1000 previous known values to feed into our model but day 0 we might have to run a simpler model (maybe even just guess 0 for all values). Anyone have any other suggestions?</p>",
  "messages": [
    {
      "id": "3031633",
      "postDate": "10/30/2024 00:12:47",
      "content": "<p>on date_id = 0 for the test/evaluation period we are going to get a lags dataframe at time_id = 0 with the responder values for the entire previous day, however we will not have any of the feature values at this point so we have nothing to match them up against.</p>\n<p>That means on the very first test/eval day we won't have any data at all for models that look back at previous values. Assuming your model looks back at the features as well as responder values.</p>\n<p>I thought maybe the <code>is_scored</code> value might be set to false for the first day to allow us to cache a history of data but in the example sets it is set to true right from the start.</p>\n<p>Am i right in thinking that this means we cannot use models with this architecture for at least the first day? From day 1 onwards we will have roughly 600-1000 previous known values to feed into our model but day 0 we might have to run a simpler model (maybe even just guess 0 for all values). Anyone have any other suggestions?</p>",
      "rawMarkdown": "on date_id = 0 for the test/evaluation period we are going to get a lags dataframe at time_id = 0 with the responder values for the entire previous day, however we will not have any of the feature values at this point so we have nothing to match them up against.\n\nThat means on the very first test/eval day we won't have any data at all for models that look back at previous values. Assuming your model looks back at the features as well as responder values.\n\nI thought maybe the `is_scored` value might be set to false for the first day to allow us to cache a history of data but in the example sets it is set to true right from the start.\n\nAm i right in thinking that this means we cannot use models with this architecture for at least the first day? From day 1 onwards we will have roughly 600-1000 previous known values to feed into our model but day 0 we might have to run a simpler model (maybe even just guess 0 for all values). Anyone have any other suggestions?",
      "votes": null
    },
    {
      "id": "3031635",
      "postDate": "10/30/2024 00:28:14",
      "content": "<p>you can use the last day in the training set to initialize the history cache.</p>",
      "rawMarkdown": "you can use the last day in the training set to initialize the history cache.",
      "votes": null
    },
    {
      "id": "3031640",
      "postDate": "10/30/2024 00:42:42",
      "content": "<p>That may work for the test if we are guaranteed to know it starts from the next day but it wouldn't be valid for the eval period as far as I can tell. Will we have an opportunity to refresh the cache and update the model? Is the eval guaranteed to immediately pick up the day after test period?</p>",
      "rawMarkdown": "That may work for the test if we are guaranteed to know it starts from the next day but it wouldn't be valid for the eval period as far as I can tell. Will we have an opportunity to refresh the cache and update the model? Is the eval guaranteed to immediately pick up the day after test period?",
      "votes": null
    },
    {
      "id": "3031651",
      "postDate": "10/30/2024 01:09:18",
      "content": "<p>it is guaranteed as stated in the data card. If you still have doubts, it is better to directly ask <a href=\"https://www.kaggle.com/sohier\" target=\"_blank\">@sohier</a> </p>",
      "rawMarkdown": "it is guaranteed as stated in the data card. If you still have doubts, it is better to directly ask @sohier",
      "votes": null
    },
    {
      "id": "3031675",
      "postDate": "10/30/2024 01:46:21",
      "content": "<p>You are right. I missed this part</p>\n<blockquote>\n  <p>During the forecasting phase, the evaluation API will serve test data from the beginning of the public set to the end of the private set. You must make predictions at every timestep, but, in this phase, only predictions on the private set are scored. (You may predict 0.0 on the unscored segments, if you like.)</p>\n</blockquote>\n<p>So we can essentially build up a historical cache during the public set and predict 0 in this period</p>",
      "rawMarkdown": "You are right. I missed this part\n\n> During the forecasting phase, the evaluation API will serve test data from the beginning of the public set to the end of the private set. You must make predictions at every timestep, but, in this phase, only predictions on the private set are scored. (You may predict 0.0 on the unscored segments, if you like.)\n\nSo we can essentially build up a historical cache during the public set and predict 0 in this period",
      "votes": null
    },
    {
      "id": "3031949",
      "postDate": "10/30/2024 10:43:59",
      "content": "<p>The only issue with this is it says that the test phase will start from date_id 0 which means you'll need to either shift everything arbitrarily or use negative date ranges for the cache which im not sure will work great </p>",
      "rawMarkdown": "The only issue with this is it says that the test phase will start from date_id 0 which means you'll need to either shift everything arbitrarily or use negative date ranges for the cache which im not sure will work great",
      "votes": null
    },
    {
      "id": "3032040",
      "postDate": "10/30/2024 13:35:07",
      "content": "<p>Yes, date_id in test starts from 0, so some negative date_ids in the initialized cache would be expected.</p>",
      "rawMarkdown": "Yes, date_id in test starts from 0, so some negative date_ids in the initialized cache would be expected.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 3031635,
      "author_name": "shiyili",
      "author_url": "",
      "post_date": "10/30/2024 00:28:14",
      "content": "<p>you can use the last day in the training set to initialize the history cache.</p>",
      "votes": null,
      "replies": [
        {
          "id": 3031640,
          "author_name": "michaeltimbs",
          "author_url": "",
          "post_date": "10/30/2024 00:42:42",
          "content": "<p>That may work for the test if we are guaranteed to know it starts from the next day but it wouldn't be valid for the eval period as far as I can tell. Will we have an opportunity to refresh the cache and update the model? Is the eval guaranteed to immediately pick up the day after test period?</p>",
          "votes": null,
          "replies": [
            {
              "id": 3031651,
              "author_name": "shiyili",
              "author_url": "",
              "post_date": "10/30/2024 01:09:18",
              "content": "<p>it is guaranteed as stated in the data card. If you still have doubts, it is better to directly ask <a href=\"https://www.kaggle.com/sohier\" target=\"_blank\">@sohier</a> </p>",
              "votes": null,
              "replies": [
                {
                  "id": 3031675,
                  "author_name": "michaeltimbs",
                  "author_url": "",
                  "post_date": "10/30/2024 01:46:21",
                  "content": "<p>You are right. I missed this part</p>\n<blockquote>\n  <p>During the forecasting phase, the evaluation API will serve test data from the beginning of the public set to the end of the private set. You must make predictions at every timestep, but, in this phase, only predictions on the private set are scored. (You may predict 0.0 on the unscored segments, if you like.)</p>\n</blockquote>\n<p>So we can essentially build up a historical cache during the public set and predict 0 in this period</p>",
                  "votes": null,
                  "replies": [
                    {
                      "id": 3031949,
                      "author_name": "michaeltimbs",
                      "author_url": "",
                      "post_date": "10/30/2024 10:43:59",
                      "content": "<p>The only issue with this is it says that the test phase will start from date_id 0 which means you'll need to either shift everything arbitrarily or use negative date ranges for the cache which im not sure will work great </p>",
                      "votes": null,
                      "replies": [
                        {
                          "id": 3032040,
                          "author_name": "shiyili",
                          "author_url": "",
                          "post_date": "10/30/2024 13:35:07",
                          "content": "<p>Yes, date_id in test starts from 0, so some negative date_ids in the initialized cache would be expected.</p>",
                          "votes": null,
                          "replies": []
                        }
                      ]
                    }
                  ]
                }
              ]
            }
          ]
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "3031633": "on date_id = 0 for the test/evaluation period we are going to get a lags dataframe at time_id = 0 with the responder values for the entire previous day, however we will not have any of the feature values at this point so we have nothing to match them up against.\n\nThat means on the very first test/eval day we won't have any data at all for models that look back at previous values. Assuming your model looks back at the features as well as responder values.\n\nI thought maybe the `is_scored` value might be set to false for the first day to allow us to cache a history of data but in the example sets it is set to true right from the start.\n\nAm i right in thinking that this means we cannot use models with this architecture for at least the first day? From day 1 onwards we will have roughly 600-1000 previous known values to feed into our model but day 0 we might have to run a simpler model (maybe even just guess 0 for all values). Anyone have any other suggestions?",
    "3031635": "you can use the last day in the training set to initialize the history cache.",
    "3031640": "That may work for the test if we are guaranteed to know it starts from the next day but it wouldn't be valid for the eval period as far as I can tell. Will we have an opportunity to refresh the cache and update the model? Is the eval guaranteed to immediately pick up the day after test period?",
    "3031651": "it is guaranteed as stated in the data card. If you still have doubts, it is better to directly ask @sohier",
    "3031675": "You are right. I missed this part\n\n> During the forecasting phase, the evaluation API will serve test data from the beginning of the public set to the end of the private set. You must make predictions at every timestep, but, in this phase, only predictions on the private set are scored. (You may predict 0.0 on the unscored segments, if you like.)\n\nSo we can essentially build up a historical cache during the public set and predict 0 in this period",
    "3031949": "The only issue with this is it says that the test phase will start from date_id 0 which means you'll need to either shift everything arbitrarily or use negative date ranges for the cache which im not sure will work great",
    "3032040": "Yes, date_id in test starts from 0, so some negative date_ids in the initialized cache would be expected."
  },
  "source": "meta"
}