{
  "id": 554604,
  "title": "Update to the Synthetic Test Data",
  "url": "/competitions/jane-street-real-time-market-data-forecasting/discussion/554604",
  "author_name": "",
  "post_date": "2025-01-02T10:42:38.483665300Z",
  "votes": 14,
  "comment_count": 19,
  "views": 0,
  "content": "<p>In the beginning of this competition, I made a public kernel [<a href=\"https://www.kaggle.com/code/shiyili/js24-rmf-submission-api-debug-with-synthetic-test?scriptVersionId=215778373\" target=\"_blank\">link</a>] that generates synthetic test data to help us better debugging our submission pipeline. I am very happy to see that many people have upvoted and copied this work, which means it was found useful by many of you. </p>\n<p>As the competition approaching to the end, I finally have time to make an update to this synthetic test data. </p>\n<p>The major update is to add the <code>is_scored</code> flag in the dataset. In this updated version (version 3), half of the dates are marked with <code>is_scored=False</code>, allowing us to debugging our pipeline considering if the data is scored or not. This is ultimately an important part to save time for inference. </p>\n<p><strong>Note: Extended Sample size</strong></p>\n<p><strong>In order to accommodate the non-scored dates, dates included in the dataset have been extended. The start date_id of the new version is from 1690, and only the last 5 days are marked with</strong> <code>is_scored=True</code>.</p>\n<p>If you are using this dataset in your pipeline, please consider use this updated version. If there is any question, suggestion or bugs, please let me know :)</p>\n<p>Happy kaggling and wish everyone a big success! </p>",
  "messages": [
    {
      "id": "3086447",
      "postDate": "01/02/2025 10:42:38",
      "content": "<p>In the beginning of this competition, I made a public kernel [<a href=\"https://www.kaggle.com/code/shiyili/js24-rmf-submission-api-debug-with-synthetic-test?scriptVersionId=215778373\" target=\"_blank\">link</a>] that generates synthetic test data to help us better debugging our submission pipeline. I am very happy to see that many people have upvoted and copied this work, which means it was found useful by many of you. </p>\n<p>As the competition approaching to the end, I finally have time to make an update to this synthetic test data. </p>\n<p>The major update is to add the <code>is_scored</code> flag in the dataset. In this updated version (version 3), half of the dates are marked with <code>is_scored=False</code>, allowing us to debugging our pipeline considering if the data is scored or not. This is ultimately an important part to save time for inference. </p>\n<p><strong>Note: Extended Sample size</strong></p>\n<p><strong>In order to accommodate the non-scored dates, dates included in the dataset have been extended. The start date_id of the new version is from 1690, and only the last 5 days are marked with</strong> <code>is_scored=True</code>.</p>\n<p>If you are using this dataset in your pipeline, please consider use this updated version. If there is any question, suggestion or bugs, please let me know :)</p>\n<p>Happy kaggling and wish everyone a big success! </p>",
      "rawMarkdown": "In the beginning of this competition, I made a public kernel [[link](https://www.kaggle.com/code/shiyili/js24-rmf-submission-api-debug-with-synthetic-test?scriptVersionId=215778373)] that generates synthetic test data to help us better debugging our submission pipeline. I am very happy to see that many people have upvoted and copied this work, which means it was found useful by many of you. \n\nAs the competition approaching to the end, I finally have time to make an update to this synthetic test data. \n\nThe major update is to add the `is_scored` flag in the dataset. In this updated version (version 3), half of the dates are marked with `is_scored=False`, allowing us to debugging our pipeline considering if the data is scored or not. This is ultimately an important part to save time for inference. \n\n**Note: Extended Sample size**\n\n**In order to accommodate the non-scored dates, dates included in the dataset have been extended. The start date_id of the new version is from 1690, and only the last 5 days are marked with** `is_scored=True`.\n\nIf you are using this dataset in your pipeline, please consider use this updated version. If there is any question, suggestion or bugs, please let me know :)\n\nHappy kaggling and wish everyone a big success!",
      "votes": null
    },
    {
      "id": "3086871",
      "postDate": "01/02/2025 19:20:16",
      "content": "<p>Thanks for your work, the dataset is extremely helpful! But please mention that you've also changed the sample size. There's an important test case where the number of symbols is 38 instead of 39, but its date_id has now been altered.</p>",
      "rawMarkdown": "Thanks for your work, the dataset is extremely helpful! But please mention that you've also changed the sample size. There's an important test case where the number of symbols is 38 instead of 39, but its date_id has now been altered.",
      "votes": null
    },
    {
      "id": "3086886",
      "postDate": "01/02/2025 19:43:12",
      "content": "<p>Thanks for letting me know :) I have added notes in my post and also in the notebook about the extension of the sample size. In the notebook, I also added a cell displaying the number of unique symbols per date. I hope this would make it clear how the data looks like. </p>\n<p>The date with 38 symbols is now associated with <code>date_id=6</code></p>",
      "rawMarkdown": "Thanks for letting me know :) I have added notes in my post and also in the notebook about the extension of the sample size. In the notebook, I also added a cell displaying the number of unique symbols per date. I hope this would make it clear how the data looks like. \n\nThe date with 38 symbols is now associated with `date_id=6`",
      "votes": null
    },
    {
      "id": "3087240",
      "postDate": "01/03/2025 09:09:02",
      "content": "<p>this is quite helpful, I have been using your set for inference all the time, thanks!</p>",
      "rawMarkdown": "this is quite helpful, I have been using your set for inference all the time, thanks!",
      "votes": null
    },
    {
      "id": "3087261",
      "postDate": "01/03/2025 09:37:53",
      "content": "<p>This is super helpful, I think I've figured out why my online learning has negatively impacted my submission API scores for over a month.</p>\n<p>On a given date (date_id = x) with time_id = 0, the lags are provided with date_id = x, not x-1.(The lagged responders are responders for date_id x-1). However, I have been directly merging these lags into my cache, which seems to have caused misalignments.</p>\n<p>Please let me know if my understanding is incorrect.</p>",
      "rawMarkdown": "This is super helpful, I think I've figured out why my online learning has negatively impacted my submission API scores for over a month.\n\nOn a given date (date_id = x) with time_id = 0, the lags are provided with date_id = x, not x-1.(The lagged responders are responders for date_id x-1). However, I have been directly merging these lags into my cache, which seems to have caused misalignments.\n\nPlease let me know if my understanding is incorrect.",
      "votes": null
    },
    {
      "id": "3087314",
      "postDate": "01/03/2025 11:12:58",
      "content": "<p>Yes you are right. The date_id in lags are the same as the current day <code>x</code>. If you want to merge/join them with the history cache (in which the date_id are <code>x-1</code>), you need to shit them by <code>-1</code>. </p>",
      "rawMarkdown": "Yes you are right. The date_id in lags are the same as the current day `x`. If you want to merge/join them with the history cache (in which the date_id are `x-1`), you need to shit them by `-1`.",
      "votes": null
    },
    {
      "id": "3087320",
      "postDate": "01/03/2025 11:18:17",
      "content": "<p>Thanks for helping me confirm, this isn't an easy bug to fix as the problem lies only with numerical value. I tried plenty of ways/validations and almost gave up…</p>",
      "rawMarkdown": "Thanks for helping me confirm, this isn't an easy bug to fix as the problem lies only with numerical value. I tried plenty of ways/validations and almost gave up...",
      "votes": null
    },
    {
      "id": "3092673",
      "postDate": "01/09/2025 20:34:33",
      "content": "<p>Hi, your notebook has been a lifesaver for me :) Thank you so much for this resource!</p>\n<p>I was wondering if there's a way to see how long each predict call takes? I want to make sure I'm not cutting it too close when doing my OL.</p>",
      "rawMarkdown": "Hi, your notebook has been a lifesaver for me :) Thank you so much for this resource!\n\nI was wondering if there's a way to see how long each predict call takes? I want to make sure I'm not cutting it too close when doing my OL.",
      "votes": null
    },
    {
      "id": "3092688",
      "postDate": "01/09/2025 21:04:48",
      "content": "<p>You can add time.time() points in your predict function to measure. </p>\n<pre><code>data_prep_time = \nprediction_time = \nfunction_time = \n\n ():\n     data_prep_time, prediction_time, function_time\n\n    t1 = time.time()\n    test = preprocess(test)\n    t2 = time.time()\n    predictions = model.predict(test)\n    t3 = time.time()\n\n    data_prep_time += t2 - t1\n    prediction_time += t3 - t2\n    function_time += t3 - t1\n\n     predictions\n</code></pre>",
      "rawMarkdown": "You can add time.time() points in your predict function to measure. \n\n```python\ndata_prep_time = 0.0\nprediction_time = 0.0\nfunction_time = 0.0\n\ndef predict():\n    global data_prep_time, prediction_time, function_time\n    \n    t1 = time.time()\n    test = preprocess(test)\n    t2 = time.time()\n    predictions = model.predict(test)\n    t3 = time.time()\n\n    data_prep_time += t2 - t1\n    prediction_time += t3 - t2\n    function_time += t3 - t1\n\n    return predictions\n```",
      "votes": null
    },
    {
      "id": "3092694",
      "postDate": "01/09/2025 21:21:55",
      "content": "<p>Thanks! I tried that out. What do you think is a \"safe\" amount of time to use? I'm doing online learning but I'm worried that if we get new symbols, I might be overtime.</p>",
      "rawMarkdown": "Thanks! I tried that out. What do you think is a \"safe\" amount of time to use? I'm doing online learning but I'm worried that if we get new symbols, I might be overtime.",
      "votes": null
    },
    {
      "id": "3092717",
      "postDate": "01/09/2025 22:34:42",
      "content": "<p>Hi glad to know it helped you! </p>\n<p>Add <code>.time()</code> to profile the code is definitely useful. An alternative way is to use a progress bar. You can check this notebook [<a href=\"https://www.kaggle.com/code/shiyili/js2024-rmf-gru-inference\" target=\"_blank\">Link</a>] to see how to setup a <code>tqdm</code> progress bar and integrate the synthetic data into your pipeline. </p>\n<p>If you have no intention to create an extra class (i.e. the <code>JaneStreetPredictor</code> in the above notebook), you can simply make a global variable for the progress bar and call the <code>.update(1)</code> method in your predict function.</p>",
      "rawMarkdown": "Hi glad to know it helped you! \n\nAdd `.time()` to profile the code is definitely useful. An alternative way is to use a progress bar. You can check this notebook [[Link](https://www.kaggle.com/code/shiyili/js2024-rmf-gru-inference)] to see how to setup a `tqdm` progress bar and integrate the synthetic data into your pipeline. \n\nIf you have no intention to create an extra class (i.e. the `JaneStreetPredictor` in the above notebook), you can simply make a global variable for the progress bar and call the `.update(1)` method in your predict function.",
      "votes": null
    },
    {
      "id": "3093306",
      "postDate": "01/10/2025 18:05:59",
      "content": "<p>I'm wondering what the maximum per predict time is as well. I've only hit timeouts so far so it'd be nice to know what the threshold is.</p>",
      "rawMarkdown": "I'm wondering what the maximum per predict time is as well. I've only hit timeouts so far so it'd be nice to know what the threshold is.",
      "votes": null
    },
    {
      "id": "3093339",
      "postDate": "01/10/2025 18:46:48",
      "content": "<p>In a winning strategy of Optiver competition last year, the online learning was no longer used once the running time reached a preset limit.</p>",
      "rawMarkdown": "In a winning strategy of Optiver competition last year, the online learning was no longer used once the running time reached a preset limit.",
      "votes": null
    },
    {
      "id": "3093349",
      "postDate": "01/10/2025 19:10:03",
      "content": "<p>From what I have read from other Kagglers who have done experimentation, 60 seconds/1 minute is the limit.  </p>",
      "rawMarkdown": "From what I have read from other Kagglers who have done experimentation, 60 seconds/1 minute is the limit.",
      "votes": null
    },
    {
      "id": "3093360",
      "postDate": "01/10/2025 19:33:15",
      "content": "<p>The max predict time is one minute. My code is running well under a minute in the synthetic test data but going overtime when I submit it</p>",
      "rawMarkdown": "The max predict time is one minute. My code is running well under a minute in the synthetic test data but going overtime when I submit it",
      "votes": null
    },
    {
      "id": "3093362",
      "postDate": "01/10/2025 19:36:48",
      "content": "<p>Same. When I test locally, I get something like 15-50 ms per time_id prediction which must be way too slow.</p>",
      "rawMarkdown": "Same. When I test locally, I get something like 15-50 ms per time_id prediction which must be way too slow.",
      "votes": null
    },
    {
      "id": "3093366",
      "postDate": "01/10/2025 19:41:59",
      "content": "<p>i only time out when using my online learning, but my online learning takes like 20seconds in the synthetic test but still times out in submission. my other predictions are like 80ms but that works fine when i submit it without the OL.</p>",
      "rawMarkdown": "i only time out when using my online learning, but my online learning takes like 20seconds in the synthetic test but still times out in submission. my other predictions are like 80ms but that works fine when i submit it without the OL.",
      "votes": null
    },
    {
      "id": "3093427",
      "postDate": "01/10/2025 22:20:40",
      "content": "<p>There is an additional 8 or 9 hour total runtime limit for every notebook that is submitted on top of the 1 minute limit between predictions. </p>",
      "rawMarkdown": "There is an additional 8 or 9 hour total runtime limit for every notebook that is submitted on top of the 1 minute limit between predictions.",
      "votes": null
    },
    {
      "id": "3094342",
      "postDate": "01/12/2025 00:58:03",
      "content": "<p>What would be the best way to use this to simulate an entire submission for testing run time and memory? Is doing a 300? iteration for loop over the following an accurate way of measuring this:</p>\n<pre><code>inference_server = kaggle_evaluation.jane_street_inference_server.JSInferenceServer(predict)\n\n os.getenv():\n    inference_server.serve()\n:\n    inference_server.run_local_gateway(\n        (\n            ,\n            ,\n        )\n    )\n</code></pre>\n<p>I'm not sure if you've observed some multiplier between the local runtime and server runtime, for either CPU / GPU instances.</p>",
      "rawMarkdown": "What would be the best way to use this to simulate an entire submission for testing run time and memory? Is doing a 300? iteration for loop over the following an accurate way of measuring this:\n\n```python\n\ninference_server = kaggle_evaluation.jane_street_inference_server.JSInferenceServer(predict)\n\nif os.getenv('KAGGLE_IS_COMPETITION_RERUN'):\n    inference_server.serve()\nelse:\n    inference_server.run_local_gateway(\n        (\n            '/kaggle/working/synthetic_test.parquet',\n            '/kaggle/working/synthetic_lag.parquet',\n        )\n    )\n\n```\nI'm not sure if you've observed some multiplier between the local runtime and server runtime, for either CPU / GPU instances.",
      "votes": null
    },
    {
      "id": "3094345",
      "postDate": "01/12/2025 01:10:32",
      "content": "<p>In the current version of the synthetic data, there are 4 days with <code>is_score==False</code> and 5 days with <code>is_score==True</code>. The real test dataset has around 200 days for the public LB and another 120 days for the private LB. You can do a simple math based on your own configuration. </p>\n<p>For instance, if you don't skip any dates and run inference using all 9 days  with the synthetic test, you can expect the real test takes about 20x longer to run. If you skip the 4 days using the <code>is_score==False</code> flag, you can expect the real test takes about 40x longer to finish.</p>\n<p>For memory test, this synthetic dataset is not ideal. You can use the train.parquet to test it, by loading the expected amount of data and check the memory usage.</p>",
      "rawMarkdown": "In the current version of the synthetic data, there are 4 days with `is_score==False` and 5 days with `is_score==True`. The real test dataset has around 200 days for the public LB and another 120 days for the private LB. You can do a simple math based on your own configuration. \n\nFor instance, if you don't skip any dates and run inference using all 9 days  with the synthetic test, you can expect the real test takes about 20x longer to run. If you skip the 4 days using the `is_score==False` flag, you can expect the real test takes about 40x longer to finish.\n\nFor memory test, this synthetic dataset is not ideal. You can use the train.parquet to test it, by loading the expected amount of data and check the memory usage.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 3086871,
      "author_name": "eivolkova",
      "author_url": "",
      "post_date": "01/02/2025 19:20:16",
      "content": "<p>Thanks for your work, the dataset is extremely helpful! But please mention that you've also changed the sample size. There's an important test case where the number of symbols is 38 instead of 39, but its date_id has now been altered.</p>",
      "votes": null,
      "replies": [
        {
          "id": 3086886,
          "author_name": "shiyili",
          "author_url": "",
          "post_date": "01/02/2025 19:43:12",
          "content": "<p>Thanks for letting me know :) I have added notes in my post and also in the notebook about the extension of the sample size. In the notebook, I also added a cell displaying the number of unique symbols per date. I hope this would make it clear how the data looks like. </p>\n<p>The date with 38 symbols is now associated with <code>date_id=6</code></p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 3087240,
      "author_name": "zhangyue199",
      "author_url": "",
      "post_date": "01/03/2025 09:09:02",
      "content": "<p>this is quite helpful, I have been using your set for inference all the time, thanks!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 3087261,
      "author_name": "herryxie",
      "author_url": "",
      "post_date": "01/03/2025 09:37:53",
      "content": "<p>This is super helpful, I think I've figured out why my online learning has negatively impacted my submission API scores for over a month.</p>\n<p>On a given date (date_id = x) with time_id = 0, the lags are provided with date_id = x, not x-1.(The lagged responders are responders for date_id x-1). However, I have been directly merging these lags into my cache, which seems to have caused misalignments.</p>\n<p>Please let me know if my understanding is incorrect.</p>",
      "votes": null,
      "replies": [
        {
          "id": 3087314,
          "author_name": "shiyili",
          "author_url": "",
          "post_date": "01/03/2025 11:12:58",
          "content": "<p>Yes you are right. The date_id in lags are the same as the current day <code>x</code>. If you want to merge/join them with the history cache (in which the date_id are <code>x-1</code>), you need to shit them by <code>-1</code>. </p>",
          "votes": null,
          "replies": [
            {
              "id": 3087320,
              "author_name": "herryxie",
              "author_url": "",
              "post_date": "01/03/2025 11:18:17",
              "content": "<p>Thanks for helping me confirm, this isn't an easy bug to fix as the problem lies only with numerical value. I tried plenty of ways/validations and almost gave up…</p>",
              "votes": null,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 3092673,
      "author_name": "rulangu",
      "author_url": "",
      "post_date": "01/09/2025 20:34:33",
      "content": "<p>Hi, your notebook has been a lifesaver for me :) Thank you so much for this resource!</p>\n<p>I was wondering if there's a way to see how long each predict call takes? I want to make sure I'm not cutting it too close when doing my OL.</p>",
      "votes": null,
      "replies": [
        {
          "id": 3092688,
          "author_name": "turkenm",
          "author_url": "",
          "post_date": "01/09/2025 21:04:48",
          "content": "<p>You can add time.time() points in your predict function to measure. </p>\n<pre><code>data_prep_time = \nprediction_time = \nfunction_time = \n\n ():\n     data_prep_time, prediction_time, function_time\n\n    t1 = time.time()\n    test = preprocess(test)\n    t2 = time.time()\n    predictions = model.predict(test)\n    t3 = time.time()\n\n    data_prep_time += t2 - t1\n    prediction_time += t3 - t2\n    function_time += t3 - t1\n\n     predictions\n</code></pre>",
          "votes": null,
          "replies": [
            {
              "id": 3092694,
              "author_name": "rulangu",
              "author_url": "",
              "post_date": "01/09/2025 21:21:55",
              "content": "<p>Thanks! I tried that out. What do you think is a \"safe\" amount of time to use? I'm doing online learning but I'm worried that if we get new symbols, I might be overtime.</p>",
              "votes": null,
              "replies": [
                {
                  "id": 3093306,
                  "author_name": "xraygoth",
                  "author_url": "",
                  "post_date": "01/10/2025 18:05:59",
                  "content": "<p>I'm wondering what the maximum per predict time is as well. I've only hit timeouts so far so it'd be nice to know what the threshold is.</p>",
                  "votes": null,
                  "replies": [
                    {
                      "id": 3093349,
                      "author_name": "johncaresio",
                      "author_url": "",
                      "post_date": "01/10/2025 19:10:03",
                      "content": "<p>From what I have read from other Kagglers who have done experimentation, 60 seconds/1 minute is the limit.  </p>",
                      "votes": null,
                      "replies": []
                    },
                    {
                      "id": 3093360,
                      "author_name": "rulangu",
                      "author_url": "",
                      "post_date": "01/10/2025 19:33:15",
                      "content": "<p>The max predict time is one minute. My code is running well under a minute in the synthetic test data but going overtime when I submit it</p>",
                      "votes": null,
                      "replies": [
                        {
                          "id": 3093362,
                          "author_name": "xraygoth",
                          "author_url": "",
                          "post_date": "01/10/2025 19:36:48",
                          "content": "<p>Same. When I test locally, I get something like 15-50 ms per time_id prediction which must be way too slow.</p>",
                          "votes": null,
                          "replies": [
                            {
                              "id": 3093366,
                              "author_name": "rulangu",
                              "author_url": "",
                              "post_date": "01/10/2025 19:41:59",
                              "content": "<p>i only time out when using my online learning, but my online learning takes like 20seconds in the synthetic test but still times out in submission. my other predictions are like 80ms but that works fine when i submit it without the OL.</p>",
                              "votes": null,
                              "replies": []
                            }
                          ]
                        },
                        {
                          "id": 3093427,
                          "author_name": "johncaresio",
                          "author_url": "",
                          "post_date": "01/10/2025 22:20:40",
                          "content": "<p>There is an additional 8 or 9 hour total runtime limit for every notebook that is submitted on top of the 1 minute limit between predictions. </p>",
                          "votes": null,
                          "replies": []
                        }
                      ]
                    }
                  ]
                },
                {
                  "id": 3093339,
                  "author_name": "tapioca",
                  "author_url": "",
                  "post_date": "01/10/2025 18:46:48",
                  "content": "<p>In a winning strategy of Optiver competition last year, the online learning was no longer used once the running time reached a preset limit.</p>",
                  "votes": null,
                  "replies": []
                }
              ]
            }
          ]
        },
        {
          "id": 3092717,
          "author_name": "shiyili",
          "author_url": "",
          "post_date": "01/09/2025 22:34:42",
          "content": "<p>Hi glad to know it helped you! </p>\n<p>Add <code>.time()</code> to profile the code is definitely useful. An alternative way is to use a progress bar. You can check this notebook [<a href=\"https://www.kaggle.com/code/shiyili/js2024-rmf-gru-inference\" target=\"_blank\">Link</a>] to see how to setup a <code>tqdm</code> progress bar and integrate the synthetic data into your pipeline. </p>\n<p>If you have no intention to create an extra class (i.e. the <code>JaneStreetPredictor</code> in the above notebook), you can simply make a global variable for the progress bar and call the <code>.update(1)</code> method in your predict function.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 3094342,
      "author_name": "redfoongus",
      "author_url": "",
      "post_date": "01/12/2025 00:58:03",
      "content": "<p>What would be the best way to use this to simulate an entire submission for testing run time and memory? Is doing a 300? iteration for loop over the following an accurate way of measuring this:</p>\n<pre><code>inference_server = kaggle_evaluation.jane_street_inference_server.JSInferenceServer(predict)\n\n os.getenv():\n    inference_server.serve()\n:\n    inference_server.run_local_gateway(\n        (\n            ,\n            ,\n        )\n    )\n</code></pre>\n<p>I'm not sure if you've observed some multiplier between the local runtime and server runtime, for either CPU / GPU instances.</p>",
      "votes": null,
      "replies": [
        {
          "id": 3094345,
          "author_name": "shiyili",
          "author_url": "",
          "post_date": "01/12/2025 01:10:32",
          "content": "<p>In the current version of the synthetic data, there are 4 days with <code>is_score==False</code> and 5 days with <code>is_score==True</code>. The real test dataset has around 200 days for the public LB and another 120 days for the private LB. You can do a simple math based on your own configuration. </p>\n<p>For instance, if you don't skip any dates and run inference using all 9 days  with the synthetic test, you can expect the real test takes about 20x longer to run. If you skip the 4 days using the <code>is_score==False</code> flag, you can expect the real test takes about 40x longer to finish.</p>\n<p>For memory test, this synthetic dataset is not ideal. You can use the train.parquet to test it, by loading the expected amount of data and check the memory usage.</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "3086447": "In the beginning of this competition, I made a public kernel [[link](https://www.kaggle.com/code/shiyili/js24-rmf-submission-api-debug-with-synthetic-test?scriptVersionId=215778373)] that generates synthetic test data to help us better debugging our submission pipeline. I am very happy to see that many people have upvoted and copied this work, which means it was found useful by many of you. \n\nAs the competition approaching to the end, I finally have time to make an update to this synthetic test data. \n\nThe major update is to add the `is_scored` flag in the dataset. In this updated version (version 3), half of the dates are marked with `is_scored=False`, allowing us to debugging our pipeline considering if the data is scored or not. This is ultimately an important part to save time for inference. \n\n**Note: Extended Sample size**\n\n**In order to accommodate the non-scored dates, dates included in the dataset have been extended. The start date_id of the new version is from 1690, and only the last 5 days are marked with** `is_scored=True`.\n\nIf you are using this dataset in your pipeline, please consider use this updated version. If there is any question, suggestion or bugs, please let me know :)\n\nHappy kaggling and wish everyone a big success!",
    "3086871": "Thanks for your work, the dataset is extremely helpful! But please mention that you've also changed the sample size. There's an important test case where the number of symbols is 38 instead of 39, but its date_id has now been altered.",
    "3086886": "Thanks for letting me know :) I have added notes in my post and also in the notebook about the extension of the sample size. In the notebook, I also added a cell displaying the number of unique symbols per date. I hope this would make it clear how the data looks like. \n\nThe date with 38 symbols is now associated with `date_id=6`",
    "3087240": "this is quite helpful, I have been using your set for inference all the time, thanks!",
    "3087261": "This is super helpful, I think I've figured out why my online learning has negatively impacted my submission API scores for over a month.\n\nOn a given date (date_id = x) with time_id = 0, the lags are provided with date_id = x, not x-1.(The lagged responders are responders for date_id x-1). However, I have been directly merging these lags into my cache, which seems to have caused misalignments.\n\nPlease let me know if my understanding is incorrect.",
    "3087314": "Yes you are right. The date_id in lags are the same as the current day `x`. If you want to merge/join them with the history cache (in which the date_id are `x-1`), you need to shit them by `-1`.",
    "3087320": "Thanks for helping me confirm, this isn't an easy bug to fix as the problem lies only with numerical value. I tried plenty of ways/validations and almost gave up...",
    "3092673": "Hi, your notebook has been a lifesaver for me :) Thank you so much for this resource!\n\nI was wondering if there's a way to see how long each predict call takes? I want to make sure I'm not cutting it too close when doing my OL.",
    "3092688": "You can add time.time() points in your predict function to measure. \n\n```python\ndata_prep_time = 0.0\nprediction_time = 0.0\nfunction_time = 0.0\n\ndef predict():\n    global data_prep_time, prediction_time, function_time\n    \n    t1 = time.time()\n    test = preprocess(test)\n    t2 = time.time()\n    predictions = model.predict(test)\n    t3 = time.time()\n\n    data_prep_time += t2 - t1\n    prediction_time += t3 - t2\n    function_time += t3 - t1\n\n    return predictions\n```",
    "3092694": "Thanks! I tried that out. What do you think is a \"safe\" amount of time to use? I'm doing online learning but I'm worried that if we get new symbols, I might be overtime.",
    "3092717": "Hi glad to know it helped you! \n\nAdd `.time()` to profile the code is definitely useful. An alternative way is to use a progress bar. You can check this notebook [[Link](https://www.kaggle.com/code/shiyili/js2024-rmf-gru-inference)] to see how to setup a `tqdm` progress bar and integrate the synthetic data into your pipeline. \n\nIf you have no intention to create an extra class (i.e. the `JaneStreetPredictor` in the above notebook), you can simply make a global variable for the progress bar and call the `.update(1)` method in your predict function.",
    "3093306": "I'm wondering what the maximum per predict time is as well. I've only hit timeouts so far so it'd be nice to know what the threshold is.",
    "3093339": "In a winning strategy of Optiver competition last year, the online learning was no longer used once the running time reached a preset limit.",
    "3093349": "From what I have read from other Kagglers who have done experimentation, 60 seconds/1 minute is the limit.",
    "3093360": "The max predict time is one minute. My code is running well under a minute in the synthetic test data but going overtime when I submit it",
    "3093362": "Same. When I test locally, I get something like 15-50 ms per time_id prediction which must be way too slow.",
    "3093366": "i only time out when using my online learning, but my online learning takes like 20seconds in the synthetic test but still times out in submission. my other predictions are like 80ms but that works fine when i submit it without the OL.",
    "3093427": "There is an additional 8 or 9 hour total runtime limit for every notebook that is submitted on top of the 1 minute limit between predictions.",
    "3094342": "What would be the best way to use this to simulate an entire submission for testing run time and memory? Is doing a 300? iteration for loop over the following an accurate way of measuring this:\n\n```python\n\ninference_server = kaggle_evaluation.jane_street_inference_server.JSInferenceServer(predict)\n\nif os.getenv('KAGGLE_IS_COMPETITION_RERUN'):\n    inference_server.serve()\nelse:\n    inference_server.run_local_gateway(\n        (\n            '/kaggle/working/synthetic_test.parquet',\n            '/kaggle/working/synthetic_lag.parquet',\n        )\n    )\n\n```\nI'm not sure if you've observed some multiplier between the local runtime and server runtime, for either CPU / GPU instances.",
    "3094345": "In the current version of the synthetic data, there are 4 days with `is_score==False` and 5 days with `is_score==True`. The real test dataset has around 200 days for the public LB and another 120 days for the private LB. You can do a simple math based on your own configuration. \n\nFor instance, if you don't skip any dates and run inference using all 9 days  with the synthetic test, you can expect the real test takes about 20x longer to run. If you skip the 4 days using the `is_score==False` flag, you can expect the real test takes about 40x longer to finish.\n\nFor memory test, this synthetic dataset is not ideal. You can use the train.parquet to test it, by loading the expected amount of data and check the memory usage."
  },
  "source": "meta"
}