{
  "id": 540510,
  "title": "Questions Regarding the Test data (both public and private(Forecasting Phase) ).",
  "url": "/competitions/jane-street-real-time-market-data-forecasting/discussion/540510",
  "author_name": "",
  "post_date": "2024-10-15T01:29:19.461975900Z",
  "votes": 17,
  "comment_count": 14,
  "views": 0,
  "content": "<p>It's been a while since I've participated in a financial competition, and it seems really fun. Thank you for hosting it!</p>\n<p>I have a question regarding the test data(both public and private(Forecasting Phase) ).</p>\n<ol>\n<li><p>Is my understanding correct that for the test data, we are expected to predict all symbol IDs from 0 to 38 in each batch? Specifically, I would like to know if there will be any missing symbol IDs or if new symbol IDs will appear in each batch.</p></li>\n<li><p>Does the number of time IDs change across batches?</p></li>\n<li><p>Is there a possibility that the data types of each column might change mid-way? In past competitions, I believe submission errors were caused by changes in data types due to NaN values being present.</p></li>\n</ol>\n<p>I would really appreciate it if the host could answer this! Thank you in advance.</p>",
  "messages": [
    {
      "id": "3017489",
      "postDate": "10/15/2024 01:29:19",
      "content": "<p>It's been a while since I've participated in a financial competition, and it seems really fun. Thank you for hosting it!</p>\n<p>I have a question regarding the test data(both public and private(Forecasting Phase) ).</p>\n<ol>\n<li><p>Is my understanding correct that for the test data, we are expected to predict all symbol IDs from 0 to 38 in each batch? Specifically, I would like to know if there will be any missing symbol IDs or if new symbol IDs will appear in each batch.</p></li>\n<li><p>Does the number of time IDs change across batches?</p></li>\n<li><p>Is there a possibility that the data types of each column might change mid-way? In past competitions, I believe submission errors were caused by changes in data types due to NaN values being present.</p></li>\n</ol>\n<p>I would really appreciate it if the host could answer this! Thank you in advance.</p>",
      "rawMarkdown": "It's been a while since I've participated in a financial competition, and it seems really fun. Thank you for hosting it!\n\nI have a question regarding the test data(both public and private(Forecasting Phase) ).\n\n1.  Is my understanding correct that for the test data, we are expected to predict all symbol IDs from 0 to 38 in each batch? Specifically, I would like to know if there will be any missing symbol IDs or if new symbol IDs will appear in each batch.\n\n\n2. Does the number of time IDs change across batches?\n\n\n3. Is there a possibility that the data types of each column might change mid-way? In past competitions, I believe submission errors were caused by changes in data types due to NaN values being present.\n\nI would really appreciate it if the host could answer this! Thank you in advance.",
      "votes": null
    },
    {
      "id": "3017616",
      "postDate": "10/15/2024 03:53:21",
      "content": "<p>impossible is Debug 😴</p>",
      "rawMarkdown": "impossible is Debug 😴",
      "votes": null
    },
    {
      "id": "3017644",
      "postDate": "10/15/2024 05:00:36",
      "content": "<p>Thank you for comments. I hope the host or kaggle reply here !</p>",
      "rawMarkdown": "Thank you for comments. I hope the host or kaggle reply here !",
      "votes": null
    },
    {
      "id": "3017776",
      "postDate": "10/15/2024 07:35:21",
      "content": "<p>For question 1, check the last paragraph in the Data card: </p>\n<blockquote>\n  <p>Each symbol_id is not guaranteed to appear in all time_id and date_id combinations. Additionally, new symbol_id values may appear in future test sets.</p>\n</blockquote>",
      "rawMarkdown": "For question 1, check the last paragraph in the Data card: \n\n>Each symbol_id is not guaranteed to appear in all time_id and date_id combinations. Additionally, new symbol_id values may appear in future test sets.",
      "votes": null
    },
    {
      "id": "3017784",
      "postDate": "10/15/2024 07:41:47",
      "content": "<p><a href=\"https://www.kaggle.com/shiyili\" target=\"_blank\">@shiyili</a> I missed it! Thank you very much!</p>",
      "rawMarkdown": "shiyili I missed it! Thank you very much!",
      "votes": null
    },
    {
      "id": "3017950",
      "postDate": "10/15/2024 11:55:56",
      "content": "<p>Hi <a href=\"https://www.kaggle.com/chumajin\" target=\"_blank\">@chumajin</a>,</p>\n<ol>\n<li>You should predict one response value for each row in the test set. The symbol_ids can vary from batch to batch.</li>\n<li>The time_ids generally don't change from batch to batch, but we aren't making an explicit guarantee here. You should make your submission robust to a varying number of ids.</li>\n<li>The data types will generally not change from batch to batch, but again we aren't making an explicit guarantee.</li>\n</ol>\n<p>As a rule, you should try to make your submissions robust to failure. Market data is messy and uncertain. We will make a good-faith effort to make the data reasonably well-behaved, but the nature of a forecasting competition means that we can't be certain all of our assumptions will hold into the final rerun.</p>",
      "rawMarkdown": "Hi @chumajin,\n\n1. You should predict one response value for each row in the test set. The symbol_ids can vary from batch to batch.\n2. The time_ids generally don't change from batch to batch, but we aren't making an explicit guarantee here. You should make your submission robust to a varying number of ids.\n3. The data types will generally not change from batch to batch, but again we aren't making an explicit guarantee.\n\nAs a rule, you should try to make your submissions robust to failure. Market data is messy and uncertain. We will make a good-faith effort to make the data reasonably well-behaved, but the nature of a forecasting competition means that we can't be certain all of our assumptions will hold into the final rerun.",
      "votes": null
    },
    {
      "id": "3018043",
      "postDate": "10/15/2024 13:31:52",
      "content": "<p><a href=\"https://www.kaggle.com/ryanholbrook\" target=\"_blank\">@ryanholbrook</a> Thank you for your response! I'm glad to have a conversation with you again after a while. Understood! </p>",
      "rawMarkdown": "ryanholbrook Thank you for your response! I'm glad to have a conversation with you again after a while. Understood!",
      "votes": null
    },
    {
      "id": "3018047",
      "postDate": "10/15/2024 13:32:57",
      "content": "<p>Additionally for the second question, the forecasting stage is very likely to maintain a consistent number of unique <code>time_id</code> entries per <code>date_id</code>, mirroring the pattern in the public test set. This consistency is expected unless significant real-world market changes occur. As Ryan said, it's advisable to factor this assumption into your analysis and modelling.</p>",
      "rawMarkdown": "Additionally for the second question, the forecasting stage is very likely to maintain a consistent number of unique `time_id` entries per `date_id`, mirroring the pattern in the public test set. This consistency is expected unless significant real-world market changes occur. As Ryan said, it's advisable to factor this assumption into your analysis and modelling.",
      "votes": null
    },
    {
      "id": "3018073",
      "postDate": "10/15/2024 13:44:40",
      "content": "<p><a href=\"https://www.kaggle.com/gogo827jz\" target=\"_blank\">@gogo827jz</a> Thank you! I see. That really deepened my understanding!</p>",
      "rawMarkdown": "gogo827jz Thank you! I see. That really deepened my understanding!",
      "votes": null
    },
    {
      "id": "3018617",
      "postDate": "10/15/2024 23:01:14",
      "content": "<p>Similar to question #2, but I don't see an answer for it: I'll assume from other answers there's no explicit guarantee, but can we typically expect to get all rows with a given date and time_id in a single batch?</p>",
      "rawMarkdown": "Similar to question #2, but I don't see an answer for it: I'll assume from other answers there's no explicit guarantee, but can we typically expect to get all rows with a given date and time_id in a single batch?",
      "votes": null
    },
    {
      "id": "3018629",
      "postDate": "10/15/2024 23:35:45",
      "content": "<p>Never mind, found the answer: \"which serves test set data one timestep by timestep\"</p>",
      "rawMarkdown": "Never mind, found the answer: \"which serves test set data one timestep by timestep\"",
      "votes": null
    },
    {
      "id": "3018645",
      "postDate": "10/16/2024 00:31:43",
      "content": "<p><a href=\"https://www.kaggle.com/roberthatch\" target=\"_blank\">@roberthatch</a> OK. I'm glad you found the answer.</p>",
      "rawMarkdown": "roberthatch OK. I'm glad you found the answer.",
      "votes": null
    },
    {
      "id": "3021470",
      "postDate": "10/18/2024 14:32:37",
      "content": "<p>is the date_id range consistant too ?</p>",
      "rawMarkdown": "is the date_id range consistant too ?",
      "votes": null
    },
    {
      "id": "3021970",
      "postDate": "10/19/2024 05:43:21",
      "content": "<p><a href=\"https://www.kaggle.com/srilakshmivelagapudi\" target=\"_blank\">@srilakshmivelagapudi</a> Thank you for comment. I think it's probably about the same by understanding from the data explanation page.</p>\n<pre><code>Competition Phases  Data Updates\nIn    forecasting task,  competition will proceed   phases:\n\n A model training phase   test   historical data. This test  has about  million rows.\n A forecasting phase   test   be collected  submissions . You should expect this test   be about  same size   test     phase.\n</code></pre>",
      "rawMarkdown": "srilakshmivelagapudi Thank you for comment. I think it's probably about the same by understanding from the data explanation page.\n\n~~~\nCompetition Phases and Data Updates\nIn line with the forecasting task, the competition will proceed in two phases:\n\n1. A model training phase with a test set of historical data. This test set has about 4.5 million rows.\n2. A forecasting phase with a test set to be collected after submissions close. You should expect this test set to be about the same size as the test set in the first phase.\n~~~",
      "votes": null
    },
    {
      "id": "3021984",
      "postDate": "10/19/2024 06:13:10",
      "content": "<p><a href=\"https://www.kaggle.com/srilakshmivelagapudi\" target=\"_blank\">@srilakshmivelagapudi</a> </p>\n<p>I estimated the number of date_id. Please take a look here in detail.</p>\n<p><a href=\"https://www.kaggle.com/competitions/jane-street-real-time-market-data-forecasting/discussion/541314#3021981\" target=\"_blank\">https://www.kaggle.com/competitions/jane-street-real-time-market-data-forecasting/discussion/541314#3021981</a></p>\n<p>Maybe 119 day ~ 120 day for the public data, 119 day = 120 day for the private data. </p>\n<p>In the total test data, it’s likely that there are 6 months of public data and 6 months of private data, making the test data span exactly one year. </p>\n<p>However, this is just my speculation!</p>",
      "rawMarkdown": "srilakshmivelagapudi \n\nI estimated the number of date_id. Please take a look here in detail.\n\nhttps://www.kaggle.com/competitions/jane-street-real-time-market-data-forecasting/discussion/541314#3021981\n\nMaybe 119 day ~ 120 day for the public data, 119 day = 120 day for the private data. \n\nIn the total test data, it’s likely that there are 6 months of public data and 6 months of private data, making the test data span exactly one year. \n\nHowever, this is just my speculation!",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 3017616,
      "author_name": "yuanzhezhou",
      "author_url": "",
      "post_date": "10/15/2024 03:53:21",
      "content": "<p>impossible is Debug 😴</p>",
      "votes": null,
      "replies": [
        {
          "id": 3017644,
          "author_name": "chumajin",
          "author_url": "",
          "post_date": "10/15/2024 05:00:36",
          "content": "<p>Thank you for comments. I hope the host or kaggle reply here !</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 3017776,
      "author_name": "shiyili",
      "author_url": "",
      "post_date": "10/15/2024 07:35:21",
      "content": "<p>For question 1, check the last paragraph in the Data card: </p>\n<blockquote>\n  <p>Each symbol_id is not guaranteed to appear in all time_id and date_id combinations. Additionally, new symbol_id values may appear in future test sets.</p>\n</blockquote>",
      "votes": null,
      "replies": [
        {
          "id": 3017784,
          "author_name": "chumajin",
          "author_url": "",
          "post_date": "10/15/2024 07:41:47",
          "content": "<p><a href=\"https://www.kaggle.com/shiyili\" target=\"_blank\">@shiyili</a> I missed it! Thank you very much!</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 3017950,
      "author_name": "ryanholbrook",
      "author_url": "",
      "post_date": "10/15/2024 11:55:56",
      "content": "<p>Hi <a href=\"https://www.kaggle.com/chumajin\" target=\"_blank\">@chumajin</a>,</p>\n<ol>\n<li>You should predict one response value for each row in the test set. The symbol_ids can vary from batch to batch.</li>\n<li>The time_ids generally don't change from batch to batch, but we aren't making an explicit guarantee here. You should make your submission robust to a varying number of ids.</li>\n<li>The data types will generally not change from batch to batch, but again we aren't making an explicit guarantee.</li>\n</ol>\n<p>As a rule, you should try to make your submissions robust to failure. Market data is messy and uncertain. We will make a good-faith effort to make the data reasonably well-behaved, but the nature of a forecasting competition means that we can't be certain all of our assumptions will hold into the final rerun.</p>",
      "votes": null,
      "replies": [
        {
          "id": 3018043,
          "author_name": "chumajin",
          "author_url": "",
          "post_date": "10/15/2024 13:31:52",
          "content": "<p><a href=\"https://www.kaggle.com/ryanholbrook\" target=\"_blank\">@ryanholbrook</a> Thank you for your response! I'm glad to have a conversation with you again after a while. Understood! </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 3018047,
          "author_name": "gogo827jz",
          "author_url": "",
          "post_date": "10/15/2024 13:32:57",
          "content": "<p>Additionally for the second question, the forecasting stage is very likely to maintain a consistent number of unique <code>time_id</code> entries per <code>date_id</code>, mirroring the pattern in the public test set. This consistency is expected unless significant real-world market changes occur. As Ryan said, it's advisable to factor this assumption into your analysis and modelling.</p>",
          "votes": null,
          "replies": [
            {
              "id": 3018073,
              "author_name": "chumajin",
              "author_url": "",
              "post_date": "10/15/2024 13:44:40",
              "content": "<p><a href=\"https://www.kaggle.com/gogo827jz\" target=\"_blank\">@gogo827jz</a> Thank you! I see. That really deepened my understanding!</p>",
              "votes": null,
              "replies": []
            },
            {
              "id": 3021470,
              "author_name": "srilakshmivelagapudi",
              "author_url": "",
              "post_date": "10/18/2024 14:32:37",
              "content": "<p>is the date_id range consistant too ?</p>",
              "votes": null,
              "replies": [
                {
                  "id": 3021970,
                  "author_name": "chumajin",
                  "author_url": "",
                  "post_date": "10/19/2024 05:43:21",
                  "content": "<p><a href=\"https://www.kaggle.com/srilakshmivelagapudi\" target=\"_blank\">@srilakshmivelagapudi</a> Thank you for comment. I think it's probably about the same by understanding from the data explanation page.</p>\n<pre><code>Competition Phases  Data Updates\nIn    forecasting task,  competition will proceed   phases:\n\n A model training phase   test   historical data. This test  has about  million rows.\n A forecasting phase   test   be collected  submissions . You should expect this test   be about  same size   test     phase.\n</code></pre>",
                  "votes": null,
                  "replies": [
                    {
                      "id": 3021984,
                      "author_name": "chumajin",
                      "author_url": "",
                      "post_date": "10/19/2024 06:13:10",
                      "content": "<p><a href=\"https://www.kaggle.com/srilakshmivelagapudi\" target=\"_blank\">@srilakshmivelagapudi</a> </p>\n<p>I estimated the number of date_id. Please take a look here in detail.</p>\n<p><a href=\"https://www.kaggle.com/competitions/jane-street-real-time-market-data-forecasting/discussion/541314#3021981\" target=\"_blank\">https://www.kaggle.com/competitions/jane-street-real-time-market-data-forecasting/discussion/541314#3021981</a></p>\n<p>Maybe 119 day ~ 120 day for the public data, 119 day = 120 day for the private data. </p>\n<p>In the total test data, it’s likely that there are 6 months of public data and 6 months of private data, making the test data span exactly one year. </p>\n<p>However, this is just my speculation!</p>",
                      "votes": null,
                      "replies": []
                    }
                  ]
                }
              ]
            }
          ]
        }
      ]
    },
    {
      "id": 3018617,
      "author_name": "roberthatch",
      "author_url": "",
      "post_date": "10/15/2024 23:01:14",
      "content": "<p>Similar to question #2, but I don't see an answer for it: I'll assume from other answers there's no explicit guarantee, but can we typically expect to get all rows with a given date and time_id in a single batch?</p>",
      "votes": null,
      "replies": [
        {
          "id": 3018629,
          "author_name": "roberthatch",
          "author_url": "",
          "post_date": "10/15/2024 23:35:45",
          "content": "<p>Never mind, found the answer: \"which serves test set data one timestep by timestep\"</p>",
          "votes": null,
          "replies": [
            {
              "id": 3018645,
              "author_name": "chumajin",
              "author_url": "",
              "post_date": "10/16/2024 00:31:43",
              "content": "<p><a href=\"https://www.kaggle.com/roberthatch\" target=\"_blank\">@roberthatch</a> OK. I'm glad you found the answer.</p>",
              "votes": null,
              "replies": []
            }
          ]
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "3017489": "It's been a while since I've participated in a financial competition, and it seems really fun. Thank you for hosting it!\n\nI have a question regarding the test data(both public and private(Forecasting Phase) ).\n\n1.  Is my understanding correct that for the test data, we are expected to predict all symbol IDs from 0 to 38 in each batch? Specifically, I would like to know if there will be any missing symbol IDs or if new symbol IDs will appear in each batch.\n\n\n2. Does the number of time IDs change across batches?\n\n\n3. Is there a possibility that the data types of each column might change mid-way? In past competitions, I believe submission errors were caused by changes in data types due to NaN values being present.\n\nI would really appreciate it if the host could answer this! Thank you in advance.",
    "3017616": "impossible is Debug 😴",
    "3017644": "Thank you for comments. I hope the host or kaggle reply here !",
    "3017776": "For question 1, check the last paragraph in the Data card: \n\n>Each symbol_id is not guaranteed to appear in all time_id and date_id combinations. Additionally, new symbol_id values may appear in future test sets.",
    "3017784": "shiyili I missed it! Thank you very much!",
    "3017950": "Hi @chumajin,\n\n1. You should predict one response value for each row in the test set. The symbol_ids can vary from batch to batch.\n2. The time_ids generally don't change from batch to batch, but we aren't making an explicit guarantee here. You should make your submission robust to a varying number of ids.\n3. The data types will generally not change from batch to batch, but again we aren't making an explicit guarantee.\n\nAs a rule, you should try to make your submissions robust to failure. Market data is messy and uncertain. We will make a good-faith effort to make the data reasonably well-behaved, but the nature of a forecasting competition means that we can't be certain all of our assumptions will hold into the final rerun.",
    "3018043": "ryanholbrook Thank you for your response! I'm glad to have a conversation with you again after a while. Understood!",
    "3018047": "Additionally for the second question, the forecasting stage is very likely to maintain a consistent number of unique `time_id` entries per `date_id`, mirroring the pattern in the public test set. This consistency is expected unless significant real-world market changes occur. As Ryan said, it's advisable to factor this assumption into your analysis and modelling.",
    "3018073": "gogo827jz Thank you! I see. That really deepened my understanding!",
    "3018617": "Similar to question #2, but I don't see an answer for it: I'll assume from other answers there's no explicit guarantee, but can we typically expect to get all rows with a given date and time_id in a single batch?",
    "3018629": "Never mind, found the answer: \"which serves test set data one timestep by timestep\"",
    "3018645": "roberthatch OK. I'm glad you found the answer.",
    "3021470": "is the date_id range consistant too ?",
    "3021970": "srilakshmivelagapudi Thank you for comment. I think it's probably about the same by understanding from the data explanation page.\n\n~~~\nCompetition Phases and Data Updates\nIn line with the forecasting task, the competition will proceed in two phases:\n\n1. A model training phase with a test set of historical data. This test set has about 4.5 million rows.\n2. A forecasting phase with a test set to be collected after submissions close. You should expect this test set to be about the same size as the test set in the first phase.\n~~~",
    "3021984": "srilakshmivelagapudi \n\nI estimated the number of date_id. Please take a look here in detail.\n\nhttps://www.kaggle.com/competitions/jane-street-real-time-market-data-forecasting/discussion/541314#3021981\n\nMaybe 119 day ~ 120 day for the public data, 119 day = 120 day for the private data. \n\nIn the total test data, it’s likely that there are 6 months of public data and 6 months of private data, making the test data span exactly one year. \n\nHowever, this is just my speculation!"
  },
  "source": "meta"
}