{
  "id": 554902,
  "title": "prediction time for each time id is only 0.16s ?",
  "url": "/competitions/jane-street-real-time-market-data-forecasting/discussion/554902",
  "author_name": "",
  "post_date": "2025-01-04T03:09:08.190778900Z",
  "votes": 1,
  "comment_count": 7,
  "views": 0,
  "content": "<p>I got a question here. And much thanks to anyone could answer my question:</p>\n<p>Since for a successful public leaderboard submission, we got a limit of 8-hours running time. And the validation set in leaderboard has nearly 200 days. So the actual inference time should be completed in 8<em>60</em>60 = 28800s / （900 time id for each date *200） = 0.16s. Am I right?</p>",
  "messages": [
    {
      "id": "3087886",
      "postDate": "01/04/2025 03:09:08",
      "content": "<p>I got a question here. And much thanks to anyone could answer my question:</p>\n<p>Since for a successful public leaderboard submission, we got a limit of 8-hours running time. And the validation set in leaderboard has nearly 200 days. So the actual inference time should be completed in 8<em>60</em>60 = 28800s / （900 time id for each date *200） = 0.16s. Am I right?</p>",
      "rawMarkdown": "I got a question here. And much thanks to anyone could answer my question:\n\nSince for a successful public leaderboard submission, we got a limit of 8-hours running time. And the validation set in leaderboard has nearly 200 days. So the actual inference time should be completed in 8*60*60 = 28800s / （900 time id for each date *200） = 0.16s. Am I right?",
      "votes": null
    },
    {
      "id": "3090165",
      "postDate": "01/07/2025 01:00:33",
      "content": "<p>Actually, it‘s 9 hours. Therefore, on a daily basis, the maximum would be 9 * 60 / 200 = 2.7 minutes. Hence, if you don't engage in online learning, 2 minutes each day should be a safe bet. This implies that (120 seconds / 968 time_ids per day = 0.12 seconds per time_id).</p>",
      "rawMarkdown": "Actually, it‘s 9 hours. Therefore, on a daily basis, the maximum would be 9 * 60 / 200 = 2.7 minutes. Hence, if you don't engage in online learning, 2 minutes each day should be a safe bet. This implies that (120 seconds / 968 time_ids per day = 0.12 seconds per time_id).",
      "votes": null
    },
    {
      "id": "3090443",
      "postDate": "01/07/2025 09:54:21",
      "content": "<p>It's a very tight limit. 😂 Thx for your reply!</p>",
      "rawMarkdown": "It's a very tight limit. 😂 Thx for your reply!",
      "votes": null
    },
    {
      "id": "3090445",
      "postDate": "01/07/2025 09:58:13",
      "content": "<p>Btw, I encountered another problem when using polars to do data processing: The speed is continuously dropping for evey new time id. But the local simulation does not have the problem. The memory is quiet sufficient. Does anyone know what the wrong part could be?</p>",
      "rawMarkdown": "Btw, I encountered another problem when using polars to do data processing: The speed is continuously dropping for evey new time id. But the local simulation does not have the problem. The memory is quiet sufficient. Does anyone know what the wrong part could be?",
      "votes": null
    },
    {
      "id": "3091197",
      "postDate": "01/08/2025 07:04:02",
      "content": "<p>Sounds like you're accumulating data throughout the day -- concatenating test DataFrames perhaps -- and then doing some work on the combined DF on every time_id.  Since the DF size grows with every time_id, performance will suffer.  Make sure you have a local test environment that can simulate calling predict() with a single time_id.  I'm sure you'll spot the problem once you get a test harness going.  Good luck.</p>",
      "rawMarkdown": "Sounds like you're accumulating data throughout the day -- concatenating test DataFrames perhaps -- and then doing some work on the combined DF on every time_id.  Since the DF size grows with every time_id, performance will suffer.  Make sure you have a local test environment that can simulate calling predict() with a single time_id.  I'm sure you'll spot the problem once you get a test harness going.  Good luck.",
      "votes": null
    },
    {
      "id": "3091647",
      "postDate": "01/08/2025 15:58:16",
      "content": "<p>Thanks for your reply😀 But I guess it's not the case. I only maintain a minimum of timeid cache for my feature engineering. Once I got a new one timeid, the oldest one will be dropped. Also, I have simulated the local environment that could call predict() and found no problem. But in kaggle, the submission speed will drop continuously.</p>",
      "rawMarkdown": "Thanks for your reply😀 But I guess it's not the case. I only maintain a minimum of timeid cache for my feature engineering. Once I got a new one timeid, the oldest one will be dropped. Also, I have simulated the local environment that could call predict() and found no problem. But in kaggle, the submission speed will drop continuously.",
      "votes": null
    },
    {
      "id": "3091670",
      "postDate": "01/08/2025 16:33:01",
      "content": "<p>How do you know that the speed drops on submissions?  There is pretty much no ability to get information out of a submission run, so what are you going on? Submission run times are not going to match run times on your local environment — submissions get queued and that wait time is part of the total run-time; they tun on hardware that is probably different from your local hardware, etc. </p>",
      "rawMarkdown": "How do you know that the speed drops on submissions?  There is pretty much no ability to get information out of a submission run, so what are you going on? Submission run times are not going to match run times on your local environment — submissions get queued and that wait time is part of the total run-time; they tun on hardware that is probably different from your local hardware, etc.",
      "votes": null
    },
    {
      "id": "3091750",
      "postDate": "01/08/2025 18:05:01",
      "content": "<p>I see! Sorry for a misleading mistake here. I want to say \"Kaggle notebook server simulation speed \" rather than \"the submission speed\". But your information about submission speed calculation is very insightful to me. 🥰</p>",
      "rawMarkdown": "I see! Sorry for a misleading mistake here. I want to say \"Kaggle notebook server simulation speed \" rather than \"the submission speed\". But your information about submission speed calculation is very insightful to me. 🥰",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 3090165,
      "author_name": "carrotwait",
      "author_url": "",
      "post_date": "01/07/2025 01:00:33",
      "content": "<p>Actually, it‘s 9 hours. Therefore, on a daily basis, the maximum would be 9 * 60 / 200 = 2.7 minutes. Hence, if you don't engage in online learning, 2 minutes each day should be a safe bet. This implies that (120 seconds / 968 time_ids per day = 0.12 seconds per time_id).</p>",
      "votes": null,
      "replies": [
        {
          "id": 3090443,
          "author_name": "melodyd",
          "author_url": "",
          "post_date": "01/07/2025 09:54:21",
          "content": "<p>It's a very tight limit. 😂 Thx for your reply!</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 3090445,
      "author_name": "melodyd",
      "author_url": "",
      "post_date": "01/07/2025 09:58:13",
      "content": "<p>Btw, I encountered another problem when using polars to do data processing: The speed is continuously dropping for evey new time id. But the local simulation does not have the problem. The memory is quiet sufficient. Does anyone know what the wrong part could be?</p>",
      "votes": null,
      "replies": [
        {
          "id": 3091197,
          "author_name": "maciejzawadzki",
          "author_url": "",
          "post_date": "01/08/2025 07:04:02",
          "content": "<p>Sounds like you're accumulating data throughout the day -- concatenating test DataFrames perhaps -- and then doing some work on the combined DF on every time_id.  Since the DF size grows with every time_id, performance will suffer.  Make sure you have a local test environment that can simulate calling predict() with a single time_id.  I'm sure you'll spot the problem once you get a test harness going.  Good luck.</p>",
          "votes": null,
          "replies": [
            {
              "id": 3091647,
              "author_name": "melodyd",
              "author_url": "",
              "post_date": "01/08/2025 15:58:16",
              "content": "<p>Thanks for your reply😀 But I guess it's not the case. I only maintain a minimum of timeid cache for my feature engineering. Once I got a new one timeid, the oldest one will be dropped. Also, I have simulated the local environment that could call predict() and found no problem. But in kaggle, the submission speed will drop continuously.</p>",
              "votes": null,
              "replies": [
                {
                  "id": 3091670,
                  "author_name": "maciejzawadzki",
                  "author_url": "",
                  "post_date": "01/08/2025 16:33:01",
                  "content": "<p>How do you know that the speed drops on submissions?  There is pretty much no ability to get information out of a submission run, so what are you going on? Submission run times are not going to match run times on your local environment — submissions get queued and that wait time is part of the total run-time; they tun on hardware that is probably different from your local hardware, etc. </p>",
                  "votes": null,
                  "replies": [
                    {
                      "id": 3091750,
                      "author_name": "melodyd",
                      "author_url": "",
                      "post_date": "01/08/2025 18:05:01",
                      "content": "<p>I see! Sorry for a misleading mistake here. I want to say \"Kaggle notebook server simulation speed \" rather than \"the submission speed\". But your information about submission speed calculation is very insightful to me. 🥰</p>",
                      "votes": null,
                      "replies": []
                    }
                  ]
                }
              ]
            }
          ]
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "3087886": "I got a question here. And much thanks to anyone could answer my question:\n\nSince for a successful public leaderboard submission, we got a limit of 8-hours running time. And the validation set in leaderboard has nearly 200 days. So the actual inference time should be completed in 8*60*60 = 28800s / （900 time id for each date *200） = 0.16s. Am I right?",
    "3090165": "Actually, it‘s 9 hours. Therefore, on a daily basis, the maximum would be 9 * 60 / 200 = 2.7 minutes. Hence, if you don't engage in online learning, 2 minutes each day should be a safe bet. This implies that (120 seconds / 968 time_ids per day = 0.12 seconds per time_id).",
    "3090443": "It's a very tight limit. 😂 Thx for your reply!",
    "3090445": "Btw, I encountered another problem when using polars to do data processing: The speed is continuously dropping for evey new time id. But the local simulation does not have the problem. The memory is quiet sufficient. Does anyone know what the wrong part could be?",
    "3091197": "Sounds like you're accumulating data throughout the day -- concatenating test DataFrames perhaps -- and then doing some work on the combined DF on every time_id.  Since the DF size grows with every time_id, performance will suffer.  Make sure you have a local test environment that can simulate calling predict() with a single time_id.  I'm sure you'll spot the problem once you get a test harness going.  Good luck.",
    "3091647": "Thanks for your reply😀 But I guess it's not the case. I only maintain a minimum of timeid cache for my feature engineering. Once I got a new one timeid, the oldest one will be dropped. Also, I have simulated the local environment that could call predict() and found no problem. But in kaggle, the submission speed will drop continuously.",
    "3091670": "How do you know that the speed drops on submissions?  There is pretty much no ability to get information out of a submission run, so what are you going on? Submission run times are not going to match run times on your local environment — submissions get queued and that wait time is part of the total run-time; they tun on hardware that is probably different from your local hardware, etc.",
    "3091750": "I see! Sorry for a misleading mistake here. I want to say \"Kaggle notebook server simulation speed \" rather than \"the submission speed\". But your information about submission speed calculation is very insightful to me. 🥰"
  },
  "source": "meta"
}