{
  "id": 542022,
  "title": "Big difference between local submission time estimate and actual API submission",
  "url": "/competitions/jane-street-real-time-market-data-forecasting/discussion/542022",
  "author_name": "",
  "post_date": "2024-10-22T15:45:42.079995400Z",
  "votes": 21,
  "comment_count": 23,
  "views": 0,
  "content": "<p>Hello all,</p>\n<p>I usually simulate the submission using a local test set that resembles the API structure to test my submission script execution time in such forecasting competitions. This is quite necessary for such competitions to prevent time-outs and OOM issues in the forecasting phase. </p>\n<p>I usually do this on Kaggle using the same GPU and RAM that I use for the LB submission to simulate the execution exactly as the public/ private leaderboard as well. I find that in this competition, my local submission estimated time is much much shorter than the actual API based submission to the leaderboard. I can confirm that my code does not have a bug and that my local simulation data also has 4.5 million rows in chunks as per the API structure. I am using a private version of the <a href=\"https://www.kaggle.com/competitions/jane-street-real-time-market-data-forecasting/discussion/541619\" target=\"_blank\">API simulator</a> discussed here. </p>\n<p>As an example, my baseline model took less than 10 minutes to submit locally over the simulated public leaderboard (locally) while it took more than 1.5 hours on the leaderboard. Another model (private) took about 20 minutes locally and about 2.5 hours on the actual submission. </p>\n<p>I wish to know if you too are facing this issue as this could be fatal for the actual forecast and can result in OOM/ timeout issues later on.</p>",
  "messages": [
    {
      "id": "3025288",
      "postDate": "10/22/2024 15:45:42",
      "content": "<p>Hello all,</p>\n<p>I usually simulate the submission using a local test set that resembles the API structure to test my submission script execution time in such forecasting competitions. This is quite necessary for such competitions to prevent time-outs and OOM issues in the forecasting phase. </p>\n<p>I usually do this on Kaggle using the same GPU and RAM that I use for the LB submission to simulate the execution exactly as the public/ private leaderboard as well. I find that in this competition, my local submission estimated time is much much shorter than the actual API based submission to the leaderboard. I can confirm that my code does not have a bug and that my local simulation data also has 4.5 million rows in chunks as per the API structure. I am using a private version of the <a href=\"https://www.kaggle.com/competitions/jane-street-real-time-market-data-forecasting/discussion/541619\" target=\"_blank\">API simulator</a> discussed here. </p>\n<p>As an example, my baseline model took less than 10 minutes to submit locally over the simulated public leaderboard (locally) while it took more than 1.5 hours on the leaderboard. Another model (private) took about 20 minutes locally and about 2.5 hours on the actual submission. </p>\n<p>I wish to know if you too are facing this issue as this could be fatal for the actual forecast and can result in OOM/ timeout issues later on.</p>",
      "rawMarkdown": "Hello all,\n\nI usually simulate the submission using a local test set that resembles the API structure to test my submission script execution time in such forecasting competitions. This is quite necessary for such competitions to prevent time-outs and OOM issues in the forecasting phase. \n\nI usually do this on Kaggle using the same GPU and RAM that I use for the LB submission to simulate the execution exactly as the public/ private leaderboard as well. I find that in this competition, my local submission estimated time is much much shorter than the actual API based submission to the leaderboard. I can confirm that my code does not have a bug and that my local simulation data also has 4.5 million rows in chunks as per the API structure. I am using a private version of the [API simulator](https://www.kaggle.com/competitions/jane-street-real-time-market-data-forecasting/discussion/541619) discussed here. \n\nAs an example, my baseline model took less than 10 minutes to submit locally over the simulated public leaderboard (locally) while it took more than 1.5 hours on the leaderboard. Another model (private) took about 20 minutes locally and about 2.5 hours on the actual submission. \n\nI wish to know if you too are facing this issue as this could be fatal for the actual forecast and can result in OOM/ timeout issues later on.",
      "votes": null
    },
    {
      "id": "3025316",
      "postDate": "10/22/2024 16:26:54",
      "content": "<p>Hi <a href=\"https://www.kaggle.com/ravi20076\" target=\"_blank\">@ravi20076</a> </p>\n<p>In the API, both test and lags for each date_id are loaded from files for each date. Assume 120 days in the LB, that would be 240 times file IO. If you already load data in the memory and used group_by() to get the partition for each date_id, it should be faster than loading from disk. Maybe you can quickly run a test with your local AIP simulator, comparing group_by('date_id') with read_parquet(). </p>\n<p>Another reason for the performance difference might be the <code>write_submission()</code> method. During the API run, the prediction for each batch is aggregated in a list. All predictions gets concatenated in the <code>write_submission()</code> method in the gateway. The concat dataframe is then wrote into the <code>submission.parquet</code>. I can imagine it is also time-consuming to concatenate 4.5million rows. </p>\n<p>I am copying the <code>generate_data_batches</code> method in the API script for a reference. </p>\n<pre><code>def (self):\n        date_ids = (\n            pl.(self.test_path)\n            .(pl.().())\n            .()\n            .()\n        )\n        assert date_ids[] == \n\n        for date_id in date_ids:\n            test_batches = pl.(\n                os.path.(self.test_path, f),\n            ).(, maintain_order=True)\n\n            lags = pl.(\n                os.path.(self.lags_path, f),\n            )\n\n            for (time_id,), test in test_batches:\n                test_data = (test, lags if time_id ==  else None)\n                validation_data = test.()\n                yield test_data, validation_data\n</code></pre>",
      "rawMarkdown": "Hi @ravi20076 \n\nIn the API, both test and lags for each date_id are loaded from files for each date. Assume 120 days in the LB, that would be 240 times file IO. If you already load data in the memory and used group_by() to get the partition for each date_id, it should be faster than loading from disk. Maybe you can quickly run a test with your local AIP simulator, comparing group_by('date_id') with read_parquet(). \n\nAnother reason for the performance difference might be the `write_submission()` method. During the API run, the prediction for each batch is aggregated in a list. All predictions gets concatenated in the `write_submission()` method in the gateway. The concat dataframe is then wrote into the `submission.parquet`. I can imagine it is also time-consuming to concatenate 4.5million rows. \n\nI am copying the `generate_data_batches` method in the API script for a reference. \n\n```\ndef generate_data_batches(self):\n        date_ids = sorted(\n            pl.scan_parquet(self.test_path)\n            .select(pl.col(\"date_id\").unique())\n            .collect()\n            .get_column(\"date_id\")\n        )\n        assert date_ids[0] == 0\n\n        for date_id in date_ids:\n            test_batches = pl.read_parquet(\n                os.path.join(self.test_path, f\"date_id={date_id}\"),\n            ).group_by(\"time_id\", maintain_order=True)\n\n            lags = pl.read_parquet(\n                os.path.join(self.lags_path, f\"date_id={date_id}\"),\n            )\n\n            for (time_id,), test in test_batches:\n                test_data = (test, lags if time_id == 0 else None)\n                validation_data = test.select('row_id')\n                yield test_data, validation_data\n```",
      "votes": null
    },
    {
      "id": "3025753",
      "postDate": "10/23/2024 06:03:35",
      "content": "<p><a href=\"https://www.kaggle.com/shiyili\" target=\"_blank\">@shiyili</a> thanks for the update and revert!<br>\nI am also loading the data from disk and am not using RAM. I am still unsure where the bottleneck is located, but will find out for sure.<br>\nThanks for the revert!</p>",
      "rawMarkdown": "shiyili thanks for the update and revert!\nI am also loading the data from disk and am not using RAM. I am still unsure where the bottleneck is located, but will find out for sure.\nThanks for the revert!",
      "votes": null
    },
    {
      "id": "3026139",
      "postDate": "10/23/2024 13:31:52",
      "content": "<p><a href=\"https://www.kaggle.com/ravi20076\" target=\"_blank\">@ravi20076</a> OK. I will check in my case. Just a moment.</p>",
      "rawMarkdown": "ravi20076 OK. I will check in my case. Just a moment.",
      "votes": null
    },
    {
      "id": "3026182",
      "postDate": "10/23/2024 14:11:13",
      "content": "<p><a href=\"https://www.kaggle.com/ravi20076\" target=\"_blank\">@ravi20076</a> Yes, you are right. In my case, the result is as follows:</p>\n<p>edit mode in my simulator : 9:35<br>\ncommit in my simulator : 8:44<br>\nsubmission : 32 min</p>\n<p>My / your code might have been too fast, or there could be a possibility that a dummy file for the private part is included during submission (but 2 times more…) .</p>",
      "rawMarkdown": "ravi20076 Yes, you are right. In my case, the result is as follows:\n\nedit mode in my simulator : 9:35\ncommit in my simulator : 8:44\nsubmission : 32 min\n\nMy / your code might have been too fast, or there could be a possibility that a dummy file for the private part is included during submission (but 2 times more...) .",
      "votes": null
    },
    {
      "id": "3026185",
      "postDate": "10/23/2024 14:16:01",
      "content": "<p>I hope the API does not have a bug. Else we are collectively doomed in the forecasting period <a href=\"https://www.kaggle.com/chumajin\" target=\"_blank\">@chumajin</a> </p>",
      "rawMarkdown": "I hope the API does not have a bug. Else we are collectively doomed in the forecasting period @chumajin",
      "votes": null
    },
    {
      "id": "3026194",
      "postDate": "10/23/2024 14:25:21",
      "content": "<p><a href=\"https://www.kaggle.com/ravi20076\" target=\"_blank\">@ravi20076</a>  Yes, we hope that. </p>\n<p>Now that you mention it, when I was debugging his code as stated in this comment, </p>\n<p><a href=\"https://www.kaggle.com/competitions/jane-street-real-time-market-data-forecasting/discussion/541619#3023920\" target=\"_blank\">https://www.kaggle.com/competitions/jane-street-real-time-market-data-forecasting/discussion/541619#3023920</a></p>\n<p>there didn’t seem to be a significant difference between the Simulator and the Submission. I'll re-submit his code to measure the time and compare it with the Simulation.</p>\n<p>I'll report it tomorrow. (because it takes over 2.5 hours)</p>\n<p>※ His code uses only a neural network (NN).</p>",
      "rawMarkdown": "ravi20076  Yes, we hope that. \n\nNow that you mention it, when I was debugging his code as stated in this comment, \n\nhttps://www.kaggle.com/competitions/jane-street-real-time-market-data-forecasting/discussion/541619#3023920\n\nthere didn’t seem to be a significant difference between the Simulator and the Submission. I'll re-submit his code to measure the time and compare it with the Simulation.\n\nI'll report it tomorrow. (because it takes over 2.5 hours)\n\n※ His code uses only a neural network (NN).",
      "votes": null
    },
    {
      "id": "3026228",
      "postDate": "10/23/2024 14:51:07",
      "content": "<p>Maybe we can directly use 'kaggle_evaluation' provided in the competition folder:</p>\n<pre><code> kaggle_evaluation.jane_street_inference_server\n\ninference_server = kaggle_evaluation.jane_street_inference_server.JSInferenceServer(predict)\n\ninference_server.run_local_gateway(\n    (\n        , \n        ,\n    )\n)\n</code></pre>",
      "rawMarkdown": "Maybe we can directly use 'kaggle_evaluation' provided in the competition folder:\n\n```python\nimport kaggle_evaluation.jane_street_inference_server\n\ninference_server = kaggle_evaluation.jane_street_inference_server.JSInferenceServer(predict)\n\ninference_server.run_local_gateway(\n    (\n        'test.parquet', # replaced by our custom test-set\n        'lags.parquet',\n    )\n)\n```",
      "votes": null
    },
    {
      "id": "3026248",
      "postDate": "10/23/2024 15:01:40",
      "content": "<p><a href=\"https://www.kaggle.com/zzy990106\" target=\"_blank\">@zzy990106</a> I see. After converting the test data and lag data into the desired format, I'll give it a try!</p>",
      "rawMarkdown": "zzy990106 I see. After converting the test data and lag data into the desired format, I'll give it a try!",
      "votes": null
    },
    {
      "id": "3026354",
      "postDate": "10/23/2024 17:36:27",
      "content": "<p>For more accurate timings I would recommend using the evaluation API directly rather than a simulator. The most basic setup will still run marginally faster than the real thing due to the lack of network calls, but it will be much closer. You'll need to:</p>\n<ul>\n<li>Reformat a subset of the train data into local copies of <code>test.parquet</code> and <code>lags.parquet</code>, partitioned by <code>date_id</code>.</li>\n<li>Update the data paths passed to <code>inference_server.run_local_gateway</code>.</li>\n</ul>",
      "rawMarkdown": "For more accurate timings I would recommend using the evaluation API directly rather than a simulator. The most basic setup will still run marginally faster than the real thing due to the lack of network calls, but it will be much closer. You'll need to:\n- Reformat a subset of the train data into local copies of `test.parquet` and `lags.parquet`, partitioned by `date_id`.\n- Update the data paths passed to `inference_server.run_local_gateway`.",
      "votes": null
    },
    {
      "id": "3026394",
      "postDate": "10/23/2024 18:38:05",
      "content": "<p>Hi <a href=\"https://www.kaggle.com/sohier\" target=\"_blank\">@sohier</a> , I would like to confirm something about the response time between two continuous predict call. </p>\n<p>In this official demo notebook \"<a href=\"https://www.kaggle.com/code/ryanholbrook/jane-street-rmf-demo-submission?scriptVersionId=201350220&amp;cellId=4\" target=\"_blank\">Jane Street RMF Demo Submission</a>\", it's mentioned \"<strong><em>If you need more than 15 minutes to load your model you can do so during the very first predict call, which does not have the usual 10 minute response deadline</em></strong>.\", does it mean between two continous predict call, we have 10 minutes for the batch inference? </p>\n<p>But in the kaggle evaluation file \"<strong><em>jane_street_gateway.py</em></strong>\" line 14 \"<strong><em>self.set_response_timeout_seconds(60)</em></strong>\", it's clearly showing the time limitation between each two continuous call is 1 minute. Is this a typo in this kaggle evaluation code? How long maximumly are we allowed to spend between two continous predict call?</p>\n<p><strong><em></em></strong> I have been suffering from this time limitation for my latest ~25 submissions. Hope get confirmation about this. Thanks! (BTW, I have been using exact setup you suggested. From what I see, it's exactly 60 seconds. Once it exceeds this number, I will get \"Deadline Exceeded\" error\")</p>",
      "rawMarkdown": "Hi @sohier , I would like to confirm something about the response time between two continuous predict call. \n\nIn this official demo notebook \"[Jane Street RMF Demo Submission](https://www.kaggle.com/code/ryanholbrook/jane-street-rmf-demo-submission?scriptVersionId=201350220&cellId=4)\", it's mentioned \"***If you need more than 15 minutes to load your model you can do so during the very first predict call, which does not have the usual 10 minute response deadline***.\", does it mean between two continous predict call, we have 10 minutes for the batch inference? \n\nBut in the kaggle evaluation file \"***jane_street_gateway.py***\" line 14 \"***self.set_response_timeout_seconds(60)***\", it's clearly showing the time limitation between each two continuous call is 1 minute. Is this a typo in this kaggle evaluation code? How long maximumly are we allowed to spend between two continous predict call?\n\n****** I have been suffering from this time limitation for my latest ~25 submissions. Hope get confirmation about this. Thanks! (BTW, I have been using exact setup you suggested. From what I see, it's exactly 60 seconds. Once it exceeds this number, I will get \"Deadline Exceeded\" error\")",
      "votes": null
    },
    {
      "id": "3026540",
      "postDate": "10/24/2024 00:20:38",
      "content": "<p><a href=\"https://www.kaggle.com/ravi20076\" target=\"_blank\">@ravi20076</a> Surprisingly, in this case, it was almost correct as follows. Is it slow with decision trees? That's strange. </p>\n<p>model : <a href=\"https://www.kaggle.com/code/ulrich07/js-tensorflow-debug\" target=\"_blank\">https://www.kaggle.com/code/ulrich07/js-tensorflow-debug</a></p>\n<p>I add and fix like this.</p>\n<pre><code>  test.select(FE).to_pandas().fillna().values\n・・・\n  predictions.fill_nan()\n</code></pre>\n<p>commit time of simulation : 2 hours 38 min<br>\nsubmission time : 2 hours 47 min<br>\nenvironment : cpu only</p>\n<p>almost same…</p>\n<p>I'll try using the emulator.</p>",
      "rawMarkdown": "ravi20076 Surprisingly, in this case, it was almost correct as follows. Is it slow with decision trees? That's strange. \n\nmodel : https://www.kaggle.com/code/ulrich07/js-tensorflow-debug\n\nI add and fix like this.\n~~~\nx = test.select(FE).to_pandas().fillna(3).values\n・・・\npredictions = predictions.fill_nan(0)\n~~~\n\ncommit time of simulation : 2 hours 38 min\nsubmission time : 2 hours 47 min\nenvironment : cpu only\n\nalmost same...\n\nI'll try using the emulator.",
      "votes": null
    },
    {
      "id": "3026543",
      "postDate": "10/24/2024 00:25:15",
      "content": "<p><a href=\"https://www.kaggle.com/sohier\" target=\"_blank\">@sohier</a> In this case, how should the date_id be set for test.parquet and lag.parquet, which are reformatted from the train data?<br>\nIn the sample, everything starts from 0, but should it also start from 0?</p>\n<p>Additionally, it would be helpful if you could explain how the date_id is handled in the actual submission.</p>",
      "rawMarkdown": "sohier In this case, how should the date_id be set for test.parquet and lag.parquet, which are reformatted from the train data?\nIn the sample, everything starts from 0, but should it also start from 0?\n\nAdditionally, it would be helpful if you could explain how the date_id is handled in the actual submission.",
      "votes": null
    },
    {
      "id": "3027063",
      "postDate": "10/24/2024 13:48:36",
      "content": "<p>Attaching to the point mentioned above. The time limitation in the evaluation API code is 60 instead of 600. It seems to be a bug or typo. <code>self.set_response_timeout_seconds(60)</code></p>\n<p><a href=\"https://www.kaggle.com/ryanholbrook\" target=\"_blank\">@ryanholbrook</a> </p>",
      "rawMarkdown": "Attaching to the point mentioned above. The time limitation in the evaluation API code is 60 instead of 600. It seems to be a bug or typo. `self.set_response_timeout_seconds(60)`\n\n@ryanholbrook",
      "votes": null
    },
    {
      "id": "3027250",
      "postDate": "10/24/2024 15:57:46",
      "content": "<p><a href=\"https://www.kaggle.com/sohier\" target=\"_blank\">@sohier</a> I understood. Thank you for every time. I will update my simulation notebook.</p>",
      "rawMarkdown": "sohier I understood. Thank you for every time. I will update my simulation notebook.",
      "votes": null
    },
    {
      "id": "3027258",
      "postDate": "10/24/2024 16:03:58",
      "content": "<p>I tried it.</p>\n<p>real submission : 32 min</p>\n<p>[old] my simulator result (date_id &gt;= 1577)</p>\n<p>edit mode  : 9:35<br>\ncommit  : 8:44</p>\n<p>[new] using kaggle_evaluation.jane_street_inference_server (date_id &gt;= 1577)</p>\n<p>edit mode  : 24:05<br>\ncommit  : 24:42</p>\n<p>It's getting a bit closer to 32 min.<br>\nI will update my notebook.</p>",
      "rawMarkdown": "I tried it.\n\nreal submission : 32 min\n\n[old] my simulator result (date_id >= 1577)\n\nedit mode  : 9:35\ncommit  : 8:44\n\n[new] using kaggle_evaluation.jane_street_inference_server (date_id >= 1577)\n\nedit mode  : 24:05\ncommit  : 24:42\n\nIt's getting a bit closer to 32 min.\nI will update my notebook.",
      "votes": null
    },
    {
      "id": "3027301",
      "postDate": "10/24/2024 16:44:12",
      "content": "<ul>\n<li>The <code>date_id</code> must start from zero in both train and test. Similarly, the synthetic local copies of <code>test.parquet</code> and <code>lags.parquet</code> must be partitioned by <code>date_id</code>.</li>\n<li>What's your question about how the <code>date_id</code> is handled? The new evaluation API code is provided in plain text, unlike the old time series API, so you can review the code directly if that helps.</li>\n</ul>",
      "rawMarkdown": "The `date_id` must start from zero in both train and test. Similarly, the synthetic local copies of `test.parquet` and `lags.parquet` must be partitioned by `date_id`.\n- What's your question about how the `date_id` is handled? The new evaluation API code is provided in plain text, unlike the old time series API, so you can review the code directly if that helps.",
      "votes": null
    },
    {
      "id": "3027305",
      "postDate": "10/24/2024 16:52:00",
      "content": "<p><a href=\"https://www.kaggle.com/lihaorocky\" target=\"_blank\">@lihaorocky</a> thanks for flagging this. I'll need to confer with Ryan to verify which was correct, but that may take some time due to travel. For the time being I've updated the demo submission notebook comments to reflect the 60 second limit.</p>",
      "rawMarkdown": "lihaorocky thanks for flagging this. I'll need to confer with Ryan to verify which was correct, but that may take some time due to travel. For the time being I've updated the demo submission notebook comments to reflect the 60 second limit.",
      "votes": null
    },
    {
      "id": "3027333",
      "postDate": "10/24/2024 17:21:51",
      "content": "<p><a href=\"https://www.kaggle.com/sohier\" target=\"_blank\">@sohier</a> OK. I understood by seeing that as below in plain text.  Thank you very much. </p>\n<pre><code>def (self):\n        date_ids = (\n            pl.(self.test_path)\n            .(pl.().())\n            .()\n            .()\n        )\n        assert date_ids[] == \n\n        for date_id in date_ids:\n            test_batches = pl.(\n                os.path.(self.test_path, f),\n            ).(, maintain_order=True)\n\n            lags = pl.(\n                os.path.(self.lags_path, f),\n            )\n\n            for (time_id,), test in test_batches:\n                test_data = (test, lags if time_id ==  else None)\n                validation_data = test.()\n                yield test_data, validation_data\n</code></pre>\n<p>I'm a bit curious about how the date_id is handled within the lag file, but I think it can be processed without it.</p>",
      "rawMarkdown": "sohier OK. I understood by seeing that as below in plain text.  Thank you very much. \n\n~~~\ndef generate_data_batches(self):\n        date_ids = sorted(\n            pl.scan_parquet(self.test_path)\n            .select(pl.col(\"date_id\").unique())\n            .collect()\n            .get_column(\"date_id\")\n        )\n        assert date_ids[0] == 0\n\n        for date_id in date_ids:\n            test_batches = pl.read_parquet(\n                os.path.join(self.test_path, f\"date_id={date_id}\"),\n            ).group_by(\"time_id\", maintain_order=True)\n\n            lags = pl.read_parquet(\n                os.path.join(self.lags_path, f\"date_id={date_id}\"),\n            )\n\n            for (time_id,), test in test_batches:\n                test_data = (test, lags if time_id == 0 else None)\n                validation_data = test.select('row_id')\n                yield test_data, validation_data\n~~~\n\nI'm a bit curious about how the date_id is handled within the lag file, but I think it can be processed without it.",
      "votes": null
    },
    {
      "id": "3033802",
      "postDate": "11/01/2024 14:24:14",
      "content": "<p>Hi <a href=\"https://www.kaggle.com/lihaorocky\" target=\"_blank\">@lihaorocky</a>, Thanks for the heads up about the discrepancy, and apologies for the delayed response. The 60 second timeout is the correct limit.</p>\n<p>There's a reasonable argument that this should be extended to allow time for processing the lag responders, but after review, we've decided to keep the 60 second limit. During the initial phase of the competition, this might indeed seem overly restrictive. Keep in mind however that the test data is going to be extended going into the final forecasting phase and there will be many more lag batches to process during that time.</p>",
      "rawMarkdown": "Hi @lihaorocky, Thanks for the heads up about the discrepancy, and apologies for the delayed response. The 60 second timeout is the correct limit.\n\nThere's a reasonable argument that this should be extended to allow time for processing the lag responders, but after review, we've decided to keep the 60 second limit. During the initial phase of the competition, this might indeed seem overly restrictive. Keep in mind however that the test data is going to be extended going into the final forecasting phase and there will be many more lag batches to process during that time.",
      "votes": null
    },
    {
      "id": "3033863",
      "postDate": "11/01/2024 15:22:40",
      "content": "<p><a href=\"https://www.kaggle.com/ryanholbrook\" target=\"_blank\">@ryanholbrook</a> Thank you for discussing and clarifying this. I believe this is important information, so could we make it more noticeable? For example, could you create a new topic and pin it? I think some people who have been here from the beginning might mistakenly believe it’s only 10 minutes.</p>",
      "rawMarkdown": "ryanholbrook Thank you for discussing and clarifying this. I believe this is important information, so could we make it more noticeable? For example, could you create a new topic and pin it? I think some people who have been here from the beginning might mistakenly believe it’s only 10 minutes.",
      "votes": null
    },
    {
      "id": "3038879",
      "postDate": "11/07/2024 13:38:39",
      "content": "<p>So the current API is not serving equal amounts of test data that the forecasting phase will be? We need to know the differences to avoid timeouts…?</p>",
      "rawMarkdown": "So the current API is not serving equal amounts of test data that the forecasting phase will be? We need to know the differences to avoid timeouts...?",
      "votes": null
    },
    {
      "id": "3039075",
      "postDate": "11/07/2024 16:13:07",
      "content": "<p><a href=\"https://www.kaggle.com/julianmukaj\" target=\"_blank\">@julianmukaj</a> This is documented on the <em>Competition Phases and Data Updates</em> section on the Data page.</p>",
      "rawMarkdown": "julianmukaj This is documented on the *Competition Phases and Data Updates* section on the Data page.",
      "votes": null
    },
    {
      "id": "3040068",
      "postDate": "11/08/2024 16:43:07",
      "content": "<p>Is it possible for online training within 60s?</p>",
      "rawMarkdown": "Is it possible for online training within 60s?",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 3025316,
      "author_name": "shiyili",
      "author_url": "",
      "post_date": "10/22/2024 16:26:54",
      "content": "<p>Hi <a href=\"https://www.kaggle.com/ravi20076\" target=\"_blank\">@ravi20076</a> </p>\n<p>In the API, both test and lags for each date_id are loaded from files for each date. Assume 120 days in the LB, that would be 240 times file IO. If you already load data in the memory and used group_by() to get the partition for each date_id, it should be faster than loading from disk. Maybe you can quickly run a test with your local AIP simulator, comparing group_by('date_id') with read_parquet(). </p>\n<p>Another reason for the performance difference might be the <code>write_submission()</code> method. During the API run, the prediction for each batch is aggregated in a list. All predictions gets concatenated in the <code>write_submission()</code> method in the gateway. The concat dataframe is then wrote into the <code>submission.parquet</code>. I can imagine it is also time-consuming to concatenate 4.5million rows. </p>\n<p>I am copying the <code>generate_data_batches</code> method in the API script for a reference. </p>\n<pre><code>def (self):\n        date_ids = (\n            pl.(self.test_path)\n            .(pl.().())\n            .()\n            .()\n        )\n        assert date_ids[] == \n\n        for date_id in date_ids:\n            test_batches = pl.(\n                os.path.(self.test_path, f),\n            ).(, maintain_order=True)\n\n            lags = pl.(\n                os.path.(self.lags_path, f),\n            )\n\n            for (time_id,), test in test_batches:\n                test_data = (test, lags if time_id ==  else None)\n                validation_data = test.()\n                yield test_data, validation_data\n</code></pre>",
      "votes": null,
      "replies": [
        {
          "id": 3025753,
          "author_name": "ravi20076",
          "author_url": "",
          "post_date": "10/23/2024 06:03:35",
          "content": "<p><a href=\"https://www.kaggle.com/shiyili\" target=\"_blank\">@shiyili</a> thanks for the update and revert!<br>\nI am also loading the data from disk and am not using RAM. I am still unsure where the bottleneck is located, but will find out for sure.<br>\nThanks for the revert!</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 3026139,
      "author_name": "chumajin",
      "author_url": "",
      "post_date": "10/23/2024 13:31:52",
      "content": "<p><a href=\"https://www.kaggle.com/ravi20076\" target=\"_blank\">@ravi20076</a> OK. I will check in my case. Just a moment.</p>",
      "votes": null,
      "replies": [
        {
          "id": 3026182,
          "author_name": "chumajin",
          "author_url": "",
          "post_date": "10/23/2024 14:11:13",
          "content": "<p><a href=\"https://www.kaggle.com/ravi20076\" target=\"_blank\">@ravi20076</a> Yes, you are right. In my case, the result is as follows:</p>\n<p>edit mode in my simulator : 9:35<br>\ncommit in my simulator : 8:44<br>\nsubmission : 32 min</p>\n<p>My / your code might have been too fast, or there could be a possibility that a dummy file for the private part is included during submission (but 2 times more…) .</p>",
          "votes": null,
          "replies": [
            {
              "id": 3026185,
              "author_name": "ravi20076",
              "author_url": "",
              "post_date": "10/23/2024 14:16:01",
              "content": "<p>I hope the API does not have a bug. Else we are collectively doomed in the forecasting period <a href=\"https://www.kaggle.com/chumajin\" target=\"_blank\">@chumajin</a> </p>",
              "votes": null,
              "replies": [
                {
                  "id": 3026194,
                  "author_name": "chumajin",
                  "author_url": "",
                  "post_date": "10/23/2024 14:25:21",
                  "content": "<p><a href=\"https://www.kaggle.com/ravi20076\" target=\"_blank\">@ravi20076</a>  Yes, we hope that. </p>\n<p>Now that you mention it, when I was debugging his code as stated in this comment, </p>\n<p><a href=\"https://www.kaggle.com/competitions/jane-street-real-time-market-data-forecasting/discussion/541619#3023920\" target=\"_blank\">https://www.kaggle.com/competitions/jane-street-real-time-market-data-forecasting/discussion/541619#3023920</a></p>\n<p>there didn’t seem to be a significant difference between the Simulator and the Submission. I'll re-submit his code to measure the time and compare it with the Simulation.</p>\n<p>I'll report it tomorrow. (because it takes over 2.5 hours)</p>\n<p>※ His code uses only a neural network (NN).</p>",
                  "votes": null,
                  "replies": [
                    {
                      "id": 3026540,
                      "author_name": "chumajin",
                      "author_url": "",
                      "post_date": "10/24/2024 00:20:38",
                      "content": "<p><a href=\"https://www.kaggle.com/ravi20076\" target=\"_blank\">@ravi20076</a> Surprisingly, in this case, it was almost correct as follows. Is it slow with decision trees? That's strange. </p>\n<p>model : <a href=\"https://www.kaggle.com/code/ulrich07/js-tensorflow-debug\" target=\"_blank\">https://www.kaggle.com/code/ulrich07/js-tensorflow-debug</a></p>\n<p>I add and fix like this.</p>\n<pre><code>  test.select(FE).to_pandas().fillna().values\n・・・\n  predictions.fill_nan()\n</code></pre>\n<p>commit time of simulation : 2 hours 38 min<br>\nsubmission time : 2 hours 47 min<br>\nenvironment : cpu only</p>\n<p>almost same…</p>\n<p>I'll try using the emulator.</p>",
                      "votes": null,
                      "replies": []
                    }
                  ]
                }
              ]
            }
          ]
        },
        {
          "id": 3026228,
          "author_name": "zzy990106",
          "author_url": "",
          "post_date": "10/23/2024 14:51:07",
          "content": "<p>Maybe we can directly use 'kaggle_evaluation' provided in the competition folder:</p>\n<pre><code> kaggle_evaluation.jane_street_inference_server\n\ninference_server = kaggle_evaluation.jane_street_inference_server.JSInferenceServer(predict)\n\ninference_server.run_local_gateway(\n    (\n        , \n        ,\n    )\n)\n</code></pre>",
          "votes": null,
          "replies": [
            {
              "id": 3026248,
              "author_name": "chumajin",
              "author_url": "",
              "post_date": "10/23/2024 15:01:40",
              "content": "<p><a href=\"https://www.kaggle.com/zzy990106\" target=\"_blank\">@zzy990106</a> I see. After converting the test data and lag data into the desired format, I'll give it a try!</p>",
              "votes": null,
              "replies": [
                {
                  "id": 3027258,
                  "author_name": "chumajin",
                  "author_url": "",
                  "post_date": "10/24/2024 16:03:58",
                  "content": "<p>I tried it.</p>\n<p>real submission : 32 min</p>\n<p>[old] my simulator result (date_id &gt;= 1577)</p>\n<p>edit mode  : 9:35<br>\ncommit  : 8:44</p>\n<p>[new] using kaggle_evaluation.jane_street_inference_server (date_id &gt;= 1577)</p>\n<p>edit mode  : 24:05<br>\ncommit  : 24:42</p>\n<p>It's getting a bit closer to 32 min.<br>\nI will update my notebook.</p>",
                  "votes": null,
                  "replies": []
                }
              ]
            }
          ]
        }
      ]
    },
    {
      "id": 3026354,
      "author_name": "sohier",
      "author_url": "",
      "post_date": "10/23/2024 17:36:27",
      "content": "<p>For more accurate timings I would recommend using the evaluation API directly rather than a simulator. The most basic setup will still run marginally faster than the real thing due to the lack of network calls, but it will be much closer. You'll need to:</p>\n<ul>\n<li>Reformat a subset of the train data into local copies of <code>test.parquet</code> and <code>lags.parquet</code>, partitioned by <code>date_id</code>.</li>\n<li>Update the data paths passed to <code>inference_server.run_local_gateway</code>.</li>\n</ul>",
      "votes": null,
      "replies": [
        {
          "id": 3026394,
          "author_name": "lihaorocky",
          "author_url": "",
          "post_date": "10/23/2024 18:38:05",
          "content": "<p>Hi <a href=\"https://www.kaggle.com/sohier\" target=\"_blank\">@sohier</a> , I would like to confirm something about the response time between two continuous predict call. </p>\n<p>In this official demo notebook \"<a href=\"https://www.kaggle.com/code/ryanholbrook/jane-street-rmf-demo-submission?scriptVersionId=201350220&amp;cellId=4\" target=\"_blank\">Jane Street RMF Demo Submission</a>\", it's mentioned \"<strong><em>If you need more than 15 minutes to load your model you can do so during the very first predict call, which does not have the usual 10 minute response deadline</em></strong>.\", does it mean between two continous predict call, we have 10 minutes for the batch inference? </p>\n<p>But in the kaggle evaluation file \"<strong><em>jane_street_gateway.py</em></strong>\" line 14 \"<strong><em>self.set_response_timeout_seconds(60)</em></strong>\", it's clearly showing the time limitation between each two continuous call is 1 minute. Is this a typo in this kaggle evaluation code? How long maximumly are we allowed to spend between two continous predict call?</p>\n<p><strong><em></em></strong> I have been suffering from this time limitation for my latest ~25 submissions. Hope get confirmation about this. Thanks! (BTW, I have been using exact setup you suggested. From what I see, it's exactly 60 seconds. Once it exceeds this number, I will get \"Deadline Exceeded\" error\")</p>",
          "votes": null,
          "replies": [
            {
              "id": 3027305,
              "author_name": "sohier",
              "author_url": "",
              "post_date": "10/24/2024 16:52:00",
              "content": "<p><a href=\"https://www.kaggle.com/lihaorocky\" target=\"_blank\">@lihaorocky</a> thanks for flagging this. I'll need to confer with Ryan to verify which was correct, but that may take some time due to travel. For the time being I've updated the demo submission notebook comments to reflect the 60 second limit.</p>",
              "votes": null,
              "replies": []
            },
            {
              "id": 3033802,
              "author_name": "ryanholbrook",
              "author_url": "",
              "post_date": "11/01/2024 14:24:14",
              "content": "<p>Hi <a href=\"https://www.kaggle.com/lihaorocky\" target=\"_blank\">@lihaorocky</a>, Thanks for the heads up about the discrepancy, and apologies for the delayed response. The 60 second timeout is the correct limit.</p>\n<p>There's a reasonable argument that this should be extended to allow time for processing the lag responders, but after review, we've decided to keep the 60 second limit. During the initial phase of the competition, this might indeed seem overly restrictive. Keep in mind however that the test data is going to be extended going into the final forecasting phase and there will be many more lag batches to process during that time.</p>",
              "votes": null,
              "replies": [
                {
                  "id": 3033863,
                  "author_name": "chumajin",
                  "author_url": "",
                  "post_date": "11/01/2024 15:22:40",
                  "content": "<p><a href=\"https://www.kaggle.com/ryanholbrook\" target=\"_blank\">@ryanholbrook</a> Thank you for discussing and clarifying this. I believe this is important information, so could we make it more noticeable? For example, could you create a new topic and pin it? I think some people who have been here from the beginning might mistakenly believe it’s only 10 minutes.</p>",
                  "votes": null,
                  "replies": []
                },
                {
                  "id": 3038879,
                  "author_name": "julianmukaj",
                  "author_url": "",
                  "post_date": "11/07/2024 13:38:39",
                  "content": "<p>So the current API is not serving equal amounts of test data that the forecasting phase will be? We need to know the differences to avoid timeouts…?</p>",
                  "votes": null,
                  "replies": [
                    {
                      "id": 3039075,
                      "author_name": "ryanholbrook",
                      "author_url": "",
                      "post_date": "11/07/2024 16:13:07",
                      "content": "<p><a href=\"https://www.kaggle.com/julianmukaj\" target=\"_blank\">@julianmukaj</a> This is documented on the <em>Competition Phases and Data Updates</em> section on the Data page.</p>",
                      "votes": null,
                      "replies": []
                    }
                  ]
                }
              ]
            }
          ]
        },
        {
          "id": 3026543,
          "author_name": "chumajin",
          "author_url": "",
          "post_date": "10/24/2024 00:25:15",
          "content": "<p><a href=\"https://www.kaggle.com/sohier\" target=\"_blank\">@sohier</a> In this case, how should the date_id be set for test.parquet and lag.parquet, which are reformatted from the train data?<br>\nIn the sample, everything starts from 0, but should it also start from 0?</p>\n<p>Additionally, it would be helpful if you could explain how the date_id is handled in the actual submission.</p>",
          "votes": null,
          "replies": [
            {
              "id": 3027250,
              "author_name": "chumajin",
              "author_url": "",
              "post_date": "10/24/2024 15:57:46",
              "content": "<p><a href=\"https://www.kaggle.com/sohier\" target=\"_blank\">@sohier</a> I understood. Thank you for every time. I will update my simulation notebook.</p>",
              "votes": null,
              "replies": [
                {
                  "id": 3027301,
                  "author_name": "sohier",
                  "author_url": "",
                  "post_date": "10/24/2024 16:44:12",
                  "content": "<ul>\n<li>The <code>date_id</code> must start from zero in both train and test. Similarly, the synthetic local copies of <code>test.parquet</code> and <code>lags.parquet</code> must be partitioned by <code>date_id</code>.</li>\n<li>What's your question about how the <code>date_id</code> is handled? The new evaluation API code is provided in plain text, unlike the old time series API, so you can review the code directly if that helps.</li>\n</ul>",
                  "votes": null,
                  "replies": [
                    {
                      "id": 3027333,
                      "author_name": "chumajin",
                      "author_url": "",
                      "post_date": "10/24/2024 17:21:51",
                      "content": "<p><a href=\"https://www.kaggle.com/sohier\" target=\"_blank\">@sohier</a> OK. I understood by seeing that as below in plain text.  Thank you very much. </p>\n<pre><code>def (self):\n        date_ids = (\n            pl.(self.test_path)\n            .(pl.().())\n            .()\n            .()\n        )\n        assert date_ids[] == \n\n        for date_id in date_ids:\n            test_batches = pl.(\n                os.path.(self.test_path, f),\n            ).(, maintain_order=True)\n\n            lags = pl.(\n                os.path.(self.lags_path, f),\n            )\n\n            for (time_id,), test in test_batches:\n                test_data = (test, lags if time_id ==  else None)\n                validation_data = test.()\n                yield test_data, validation_data\n</code></pre>\n<p>I'm a bit curious about how the date_id is handled within the lag file, but I think it can be processed without it.</p>",
                      "votes": null,
                      "replies": []
                    }
                  ]
                }
              ]
            }
          ]
        },
        {
          "id": 3027063,
          "author_name": "shiyili",
          "author_url": "",
          "post_date": "10/24/2024 13:48:36",
          "content": "<p>Attaching to the point mentioned above. The time limitation in the evaluation API code is 60 instead of 600. It seems to be a bug or typo. <code>self.set_response_timeout_seconds(60)</code></p>\n<p><a href=\"https://www.kaggle.com/ryanholbrook\" target=\"_blank\">@ryanholbrook</a> </p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 3040068,
      "author_name": "voix97",
      "author_url": "",
      "post_date": "11/08/2024 16:43:07",
      "content": "<p>Is it possible for online training within 60s?</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "3025288": "Hello all,\n\nI usually simulate the submission using a local test set that resembles the API structure to test my submission script execution time in such forecasting competitions. This is quite necessary for such competitions to prevent time-outs and OOM issues in the forecasting phase. \n\nI usually do this on Kaggle using the same GPU and RAM that I use for the LB submission to simulate the execution exactly as the public/ private leaderboard as well. I find that in this competition, my local submission estimated time is much much shorter than the actual API based submission to the leaderboard. I can confirm that my code does not have a bug and that my local simulation data also has 4.5 million rows in chunks as per the API structure. I am using a private version of the [API simulator](https://www.kaggle.com/competitions/jane-street-real-time-market-data-forecasting/discussion/541619) discussed here. \n\nAs an example, my baseline model took less than 10 minutes to submit locally over the simulated public leaderboard (locally) while it took more than 1.5 hours on the leaderboard. Another model (private) took about 20 minutes locally and about 2.5 hours on the actual submission. \n\nI wish to know if you too are facing this issue as this could be fatal for the actual forecast and can result in OOM/ timeout issues later on.",
    "3025316": "Hi @ravi20076 \n\nIn the API, both test and lags for each date_id are loaded from files for each date. Assume 120 days in the LB, that would be 240 times file IO. If you already load data in the memory and used group_by() to get the partition for each date_id, it should be faster than loading from disk. Maybe you can quickly run a test with your local AIP simulator, comparing group_by('date_id') with read_parquet(). \n\nAnother reason for the performance difference might be the `write_submission()` method. During the API run, the prediction for each batch is aggregated in a list. All predictions gets concatenated in the `write_submission()` method in the gateway. The concat dataframe is then wrote into the `submission.parquet`. I can imagine it is also time-consuming to concatenate 4.5million rows. \n\nI am copying the `generate_data_batches` method in the API script for a reference. \n\n```\ndef generate_data_batches(self):\n        date_ids = sorted(\n            pl.scan_parquet(self.test_path)\n            .select(pl.col(\"date_id\").unique())\n            .collect()\n            .get_column(\"date_id\")\n        )\n        assert date_ids[0] == 0\n\n        for date_id in date_ids:\n            test_batches = pl.read_parquet(\n                os.path.join(self.test_path, f\"date_id={date_id}\"),\n            ).group_by(\"time_id\", maintain_order=True)\n\n            lags = pl.read_parquet(\n                os.path.join(self.lags_path, f\"date_id={date_id}\"),\n            )\n\n            for (time_id,), test in test_batches:\n                test_data = (test, lags if time_id == 0 else None)\n                validation_data = test.select('row_id')\n                yield test_data, validation_data\n```",
    "3025753": "shiyili thanks for the update and revert!\nI am also loading the data from disk and am not using RAM. I am still unsure where the bottleneck is located, but will find out for sure.\nThanks for the revert!",
    "3026139": "ravi20076 OK. I will check in my case. Just a moment.",
    "3026182": "ravi20076 Yes, you are right. In my case, the result is as follows:\n\nedit mode in my simulator : 9:35\ncommit in my simulator : 8:44\nsubmission : 32 min\n\nMy / your code might have been too fast, or there could be a possibility that a dummy file for the private part is included during submission (but 2 times more...) .",
    "3026185": "I hope the API does not have a bug. Else we are collectively doomed in the forecasting period @chumajin",
    "3026194": "ravi20076  Yes, we hope that. \n\nNow that you mention it, when I was debugging his code as stated in this comment, \n\nhttps://www.kaggle.com/competitions/jane-street-real-time-market-data-forecasting/discussion/541619#3023920\n\nthere didn’t seem to be a significant difference between the Simulator and the Submission. I'll re-submit his code to measure the time and compare it with the Simulation.\n\nI'll report it tomorrow. (because it takes over 2.5 hours)\n\n※ His code uses only a neural network (NN).",
    "3026228": "Maybe we can directly use 'kaggle_evaluation' provided in the competition folder:\n\n```python\nimport kaggle_evaluation.jane_street_inference_server\n\ninference_server = kaggle_evaluation.jane_street_inference_server.JSInferenceServer(predict)\n\ninference_server.run_local_gateway(\n    (\n        'test.parquet', # replaced by our custom test-set\n        'lags.parquet',\n    )\n)\n```",
    "3026248": "zzy990106 I see. After converting the test data and lag data into the desired format, I'll give it a try!",
    "3026354": "For more accurate timings I would recommend using the evaluation API directly rather than a simulator. The most basic setup will still run marginally faster than the real thing due to the lack of network calls, but it will be much closer. You'll need to:\n- Reformat a subset of the train data into local copies of `test.parquet` and `lags.parquet`, partitioned by `date_id`.\n- Update the data paths passed to `inference_server.run_local_gateway`.",
    "3026394": "Hi @sohier , I would like to confirm something about the response time between two continuous predict call. \n\nIn this official demo notebook \"[Jane Street RMF Demo Submission](https://www.kaggle.com/code/ryanholbrook/jane-street-rmf-demo-submission?scriptVersionId=201350220&cellId=4)\", it's mentioned \"***If you need more than 15 minutes to load your model you can do so during the very first predict call, which does not have the usual 10 minute response deadline***.\", does it mean between two continous predict call, we have 10 minutes for the batch inference? \n\nBut in the kaggle evaluation file \"***jane_street_gateway.py***\" line 14 \"***self.set_response_timeout_seconds(60)***\", it's clearly showing the time limitation between each two continuous call is 1 minute. Is this a typo in this kaggle evaluation code? How long maximumly are we allowed to spend between two continous predict call?\n\n****** I have been suffering from this time limitation for my latest ~25 submissions. Hope get confirmation about this. Thanks! (BTW, I have been using exact setup you suggested. From what I see, it's exactly 60 seconds. Once it exceeds this number, I will get \"Deadline Exceeded\" error\")",
    "3026540": "ravi20076 Surprisingly, in this case, it was almost correct as follows. Is it slow with decision trees? That's strange. \n\nmodel : https://www.kaggle.com/code/ulrich07/js-tensorflow-debug\n\nI add and fix like this.\n~~~\nx = test.select(FE).to_pandas().fillna(3).values\n・・・\npredictions = predictions.fill_nan(0)\n~~~\n\ncommit time of simulation : 2 hours 38 min\nsubmission time : 2 hours 47 min\nenvironment : cpu only\n\nalmost same...\n\nI'll try using the emulator.",
    "3026543": "sohier In this case, how should the date_id be set for test.parquet and lag.parquet, which are reformatted from the train data?\nIn the sample, everything starts from 0, but should it also start from 0?\n\nAdditionally, it would be helpful if you could explain how the date_id is handled in the actual submission.",
    "3027063": "Attaching to the point mentioned above. The time limitation in the evaluation API code is 60 instead of 600. It seems to be a bug or typo. `self.set_response_timeout_seconds(60)`\n\n@ryanholbrook",
    "3027250": "sohier I understood. Thank you for every time. I will update my simulation notebook.",
    "3027258": "I tried it.\n\nreal submission : 32 min\n\n[old] my simulator result (date_id >= 1577)\n\nedit mode  : 9:35\ncommit  : 8:44\n\n[new] using kaggle_evaluation.jane_street_inference_server (date_id >= 1577)\n\nedit mode  : 24:05\ncommit  : 24:42\n\nIt's getting a bit closer to 32 min.\nI will update my notebook.",
    "3027301": "The `date_id` must start from zero in both train and test. Similarly, the synthetic local copies of `test.parquet` and `lags.parquet` must be partitioned by `date_id`.\n- What's your question about how the `date_id` is handled? The new evaluation API code is provided in plain text, unlike the old time series API, so you can review the code directly if that helps.",
    "3027305": "lihaorocky thanks for flagging this. I'll need to confer with Ryan to verify which was correct, but that may take some time due to travel. For the time being I've updated the demo submission notebook comments to reflect the 60 second limit.",
    "3027333": "sohier OK. I understood by seeing that as below in plain text.  Thank you very much. \n\n~~~\ndef generate_data_batches(self):\n        date_ids = sorted(\n            pl.scan_parquet(self.test_path)\n            .select(pl.col(\"date_id\").unique())\n            .collect()\n            .get_column(\"date_id\")\n        )\n        assert date_ids[0] == 0\n\n        for date_id in date_ids:\n            test_batches = pl.read_parquet(\n                os.path.join(self.test_path, f\"date_id={date_id}\"),\n            ).group_by(\"time_id\", maintain_order=True)\n\n            lags = pl.read_parquet(\n                os.path.join(self.lags_path, f\"date_id={date_id}\"),\n            )\n\n            for (time_id,), test in test_batches:\n                test_data = (test, lags if time_id == 0 else None)\n                validation_data = test.select('row_id')\n                yield test_data, validation_data\n~~~\n\nI'm a bit curious about how the date_id is handled within the lag file, but I think it can be processed without it.",
    "3033802": "Hi @lihaorocky, Thanks for the heads up about the discrepancy, and apologies for the delayed response. The 60 second timeout is the correct limit.\n\nThere's a reasonable argument that this should be extended to allow time for processing the lag responders, but after review, we've decided to keep the 60 second limit. During the initial phase of the competition, this might indeed seem overly restrictive. Keep in mind however that the test data is going to be extended going into the final forecasting phase and there will be many more lag batches to process during that time.",
    "3033863": "ryanholbrook Thank you for discussing and clarifying this. I believe this is important information, so could we make it more noticeable? For example, could you create a new topic and pin it? I think some people who have been here from the beginning might mistakenly believe it’s only 10 minutes.",
    "3038879": "So the current API is not serving equal amounts of test data that the forecasting phase will be? We need to know the differences to avoid timeouts...?",
    "3039075": "julianmukaj This is documented on the *Competition Phases and Data Updates* section on the Data page.",
    "3040068": "Is it possible for online training within 60s?"
  },
  "source": "meta"
}