{
  "id": 545070,
  "title": "Time limit estimation",
  "url": "/competitions/jane-street-real-time-market-data-forecasting/discussion/545070",
  "author_name": "Danu A.",
  "post_date": "2024-11-08T08:18:56.769000",
  "votes": 1,
  "comment_count": 8,
  "views": 0,
  "content": "<p>The call limit of 1 min is non-sense if we take into account the max running time of the notebook of 8h + additional 1h:</p>\n<p>4 500 000 total rows / 39 rows in a call = ~115 400 calls<br>\n9h * 60min * 60 sec / 115 400 calls = ~280ms by call is the max we are allowed</p>\n<p>And this is just for the public test and then we will have additional 4.5 mln rows, so 140ms per call will be the max.</p>\n<p>All of this with the response lags of the server included.</p>\n<p>Do I miss something?</p>",
  "messages": [
    {
      "id": 3040114,
      "postDate": "2024-11-08T17:46:47.817Z",
      "content": "<p>You are assuming all calls using the same amout of time. What if one choose to do a model recalibration between two calls? For example, let's assume a normal calls takes e.g. 10ms to run, and the recalibration may take longer time, say 2 min. Here it makes a difference with or without the 1 min limit. It can also happen if one has some heavy feature engineering on the lags data, which happens only once per day. This also makes a difference if there is a 1 minute limite.</p>",
      "rawMarkdown": "You are assuming all calls using the same amout of time. What if one choose to do a model recalibration between two calls? For example, let's assume a normal calls takes e.g. 10ms to run, and the recalibration may take longer time, say 2 min. Here it makes a difference with or without the 1 min limit. It can also happen if one has some heavy feature engineering on the lags data, which happens only once per day. This also makes a difference if there is a 1 minute limite.",
      "votes": 1,
      "replies": [
        {
          "id": 3040124,
          "postDate": "2024-11-08T18:03:05.287Z",
          "rawMarkdown": "",
          "isDeleted": true
        },
        {
          "id": 3040128,
          "postDate": "2024-11-08T18:11:16.323Z",
          "content": "<p>Yes, it's just a rough aproximation and in examples like yours it gets really tricky</p>",
          "rawMarkdown": "Yes, it's just a rough aproximation and in examples like yours it gets really tricky"
        }
      ]
    },
    {
      "id": 3039721,
      "postDate": "2024-11-08T10:01:18.693Z",
      "content": "<p>I suppose you won't need to run your code against both public and private with the help of \"is_scored\" column</p>",
      "rawMarkdown": "I suppose you won't need to run your code against both public and private with the help of \"is_scored\" column",
      "votes": 1,
      "replies": [
        {
          "id": 3039765,
          "postDate": "2024-11-08T11:13:24.383Z",
          "content": "<p>I agree, but we will still have to predict ~4.5 mln rows and if my calculations are correct, their 1 min limit by call (that already scares people) is just max ~280ms by call</p>",
          "rawMarkdown": "I agree, but we will still have to predict ~4.5 mln rows and if my calculations are correct, their 1 min limit by call (that already scares people) is just max ~280ms by call"
        }
      ]
    },
    {
      "id": 3039643,
      "postDate": "2024-11-08T08:18:56.770Z",
      "content": "<p>The call limit of 1 min is non-sense if we take into account the max running time of the notebook of 8h + additional 1h:</p>\n<p>4 500 000 total rows / 39 rows in a call = ~115 400 calls<br>\n9h * 60min * 60 sec / 115 400 calls = ~280ms by call is the max we are allowed</p>\n<p>And this is just for the public test and then we will have additional 4.5 mln rows, so 140ms per call will be the max.</p>\n<p>All of this with the response lags of the server included.</p>\n<p>Do I miss something?</p>",
      "rawMarkdown": "The call limit of 1 min is non-sense if we take into account the max running time of the notebook of 8h + additional 1h:\n\n4 500 000 total rows / 39 rows in a call = ~115 400 calls\n9h * 60min * 60 sec / 115 400 calls = ~280ms by call is the max we are allowed\n\nAnd this is just for the public test and then we will have additional 4.5 mln rows, so 140ms per call will be the max.\n\nAll of this with the response lags of the server included.\n\nDo I miss something?\n",
      "votes": 1
    },
    {
      "id": 3044783,
      "postDate": "2024-11-13T19:50:22.377Z",
      "content": "<p>What's the difference between the public test and the additional 4.5 mln rows? Where is mentioned?</p>",
      "rawMarkdown": "What's the difference between the public test and the additional 4.5 mln rows? Where is mentioned?",
      "replies": [
        {
          "id": 3044789,
          "postDate": "2024-11-13T19:59:32.533Z",
          "content": "<blockquote>\n  <p>the competition will proceed in two phases:</p>\n  <p>A model training phase with a test set of historical data. This test set has about 4.5 million rows.<br>\n  A forecasting phase with a test set to be collected after submissions close. You should expect this test set to be about the same size as the test set in the first phase.</p>\n</blockquote>",
          "rawMarkdown": ">the competition will proceed in two phases:\n\n>A model training phase with a test set of historical data. This test set has about 4.5 million rows.\n>A forecasting phase with a test set to be collected after submissions close. You should expect this test set to be about the same size as the test set in the first phase.",
          "replies": [
            {
              "id": 3044796,
              "postDate": "2024-11-13T20:17:15.327Z",
              "content": "<p>And for the first phase the evaluation in total is supposed to be up to 8 hours based on which we’re guessing less than 1 min per call to predict?</p>",
              "rawMarkdown": "And for the first phase the evaluation in total is supposed to be up to 8 hours based on which we’re guessing less than 1 min per call to predict?"
            }
          ]
        }
      ]
    }
  ],
  "comments": [
    {
      "id": 3040114,
      "author_name": "SLi",
      "author_url": "",
      "post_date": "2024-11-08T17:46:47.817000",
      "content": "<p>You are assuming all calls using the same amout of time. What if one choose to do a model recalibration between two calls? For example, let's assume a normal calls takes e.g. 10ms to run, and the recalibration may take longer time, say 2 min. Here it makes a difference with or without the 1 min limit. It can also happen if one has some heavy feature engineering on the lags data, which happens only once per day. This also makes a difference if there is a 1 minute limite.</p>",
      "votes": 1,
      "replies": [
        {
          "id": 3040124,
          "author_name": "",
          "author_url": "",
          "post_date": "2024-11-08T18:03:05.287000",
          "content": "",
          "votes": 0,
          "replies": []
        },
        {
          "id": 3040128,
          "author_name": "Danu A.",
          "author_url": "",
          "post_date": "2024-11-08T18:11:16.323000",
          "content": "<p>Yes, it's just a rough aproximation and in examples like yours it gets really tricky</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 3039721,
      "author_name": "Evgeniia Grigoreva",
      "author_url": "",
      "post_date": "2024-11-08T10:01:18.693000",
      "content": "<p>I suppose you won't need to run your code against both public and private with the help of \"is_scored\" column</p>",
      "votes": 1,
      "replies": [
        {
          "id": 3039765,
          "author_name": "Danu A.",
          "author_url": "",
          "post_date": "2024-11-08T11:13:24.383000",
          "content": "<p>I agree, but we will still have to predict ~4.5 mln rows and if my calculations are correct, their 1 min limit by call (that already scares people) is just max ~280ms by call</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 3044783,
      "author_name": "sophia",
      "author_url": "",
      "post_date": "2024-11-13T19:50:22.377000",
      "content": "<p>What's the difference between the public test and the additional 4.5 mln rows? Where is mentioned?</p>",
      "votes": 0,
      "replies": [
        {
          "id": 3044789,
          "author_name": "Danu A.",
          "author_url": "",
          "post_date": "2024-11-13T19:59:32.533000",
          "content": "<blockquote>\n  <p>the competition will proceed in two phases:</p>\n  <p>A model training phase with a test set of historical data. This test set has about 4.5 million rows.<br>\n  A forecasting phase with a test set to be collected after submissions close. You should expect this test set to be about the same size as the test set in the first phase.</p>\n</blockquote>",
          "votes": 0,
          "replies": [
            {
              "id": 3044796,
              "author_name": "sophia",
              "author_url": "",
              "post_date": "2024-11-13T20:17:15.327000",
              "content": "<p>And for the first phase the evaluation in total is supposed to be up to 8 hours based on which we’re guessing less than 1 min per call to predict?</p>",
              "votes": 0,
              "replies": []
            }
          ]
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "3040114": "You are assuming all calls using the same amout of time. What if one choose to do a model recalibration between two calls? For example, let's assume a normal calls takes e.g. 10ms to run, and the recalibration may take longer time, say 2 min. Here it makes a difference with or without the 1 min limit. It can also happen if one has some heavy feature engineering on the lags data, which happens only once per day. This also makes a difference if there is a 1 minute limite.",
    "3039721": "I suppose you won't need to run your code against both public and private with the help of \"is_scored\" column",
    "3039643": "The call limit of 1 min is non-sense if we take into account the max running time of the notebook of 8h + additional 1h:\n\n4 500 000 total rows / 39 rows in a call = ~115 400 calls\n9h * 60min * 60 sec / 115 400 calls = ~280ms by call is the max we are allowed\n\nAnd this is just for the public test and then we will have additional 4.5 mln rows, so 140ms per call will be the max.\n\nAll of this with the response lags of the server included.\n\nDo I miss something?\n",
    "3044783": "What's the difference between the public test and the additional 4.5 mln rows? Where is mentioned?"
  }
}