{
  "id": 551724,
  "title": "The evaluation mechanics",
  "url": "/competitions/jane-street-real-time-market-data-forecasting/discussion/551724",
  "author_name": "",
  "post_date": "2024-12-15T03:31:10.479433700Z",
  "votes": 1,
  "comment_count": 3,
  "views": 0,
  "content": "<p>Am I correct in understanding that the evaluation server sends the samples for one date at a time to the predict function, and the average number of samples on one date is always below 40,000? </p>",
  "messages": [
    {
      "id": "3072328",
      "postDate": "12/15/2024 03:31:10",
      "content": "<p>Am I correct in understanding that the evaluation server sends the samples for one date at a time to the predict function, and the average number of samples on one date is always below 40,000? </p>",
      "rawMarkdown": "Am I correct in understanding that the evaluation server sends the samples for one date at a time to the predict function, and the average number of samples on one date is always below 40,000?",
      "votes": null
    },
    {
      "id": "3073015",
      "postDate": "12/15/2024 23:18:16",
      "content": "<p>It sends each time_id individually I believe. So you will get roughly one row per symbol on every invocation and roughly 968 or whatever it is invocations per date_id</p>",
      "rawMarkdown": "It sends each time_id individually I believe. So you will get roughly one row per symbol on every invocation and roughly 968 or whatever it is invocations per date_id",
      "votes": null
    },
    {
      "id": "3073020",
      "postDate": "12/15/2024 23:52:58",
      "content": "<p>Thanks, just to clarify, the dataframe that is provided by the server at each call to the predict function includes all the time_id and symbols for one date_id? In the validation set there are about 40,000 rows for each date_id file, so I assumed that for the hidden set should be about the same. Is that correct?</p>",
      "rawMarkdown": "Thanks, just to clarify, the dataframe that is provided by the server at each call to the predict function includes all the time_id and symbols for one date_id? In the validation set there are about 40,000 rows for each date_id file, so I assumed that for the hidden set should be about the same. Is that correct?",
      "votes": null
    },
    {
      "id": "3073035",
      "postDate": "12/16/2024 00:36:46",
      "content": "<p>No. It is partitioned by time_id.  You will never get multiple time_id at the same time other than in lags. </p>\n<p>e.g</p>\n<p>---- first predict call -----<br>\ndate_id, time_id, symbol_id<br>\n1, 0, 1<br>\n1, 0, 2<br>\n1, 0, 3</p>\n<h2>…..</h2>\n<p>---- second predict call -----<br>\ndate_id, time_id, symbol_id<br>\n1, 1, 1<br>\n1, 1, 2<br>\n1, 1, 3</p>\n<h2>…..</h2>\n<p>IF you want to replicate this locally you can do</p>\n<pre><code>time_partitions = df.collect().partition_by(\n    , maintain_order=\n)\n\n timestep_data  time_partitions:\n    timestep_preds = predict(\n        timestep_data\n    )\n</code></pre>",
      "rawMarkdown": "No. It is partitioned by time_id.  You will never get multiple time_id at the same time other than in lags. \n\ne.g\n\n---- first predict call -----\ndate_id, time_id, symbol_id\n1, 0, 1\n1, 0, 2\n1, 0, 3\n.....\n------\n---- second predict call -----\ndate_id, time_id, symbol_id\n1, 1, 1\n1, 1, 2\n1, 1, 3\n.....\n------\n\nIF you want to replicate this locally you can do\n\n```python\ntime_partitions = df.collect().partition_by(\n    \"time_id\", maintain_order=True\n)\n\nfor timestep_data in time_partitions:\n    timestep_preds = predict(\n        timestep_data\n    )\n```",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 3073015,
      "author_name": "michaeltimbs",
      "author_url": "",
      "post_date": "12/15/2024 23:18:16",
      "content": "<p>It sends each time_id individually I believe. So you will get roughly one row per symbol on every invocation and roughly 968 or whatever it is invocations per date_id</p>",
      "votes": null,
      "replies": [
        {
          "id": 3073020,
          "author_name": "izaznov",
          "author_url": "",
          "post_date": "12/15/2024 23:52:58",
          "content": "<p>Thanks, just to clarify, the dataframe that is provided by the server at each call to the predict function includes all the time_id and symbols for one date_id? In the validation set there are about 40,000 rows for each date_id file, so I assumed that for the hidden set should be about the same. Is that correct?</p>",
          "votes": null,
          "replies": [
            {
              "id": 3073035,
              "author_name": "michaeltimbs",
              "author_url": "",
              "post_date": "12/16/2024 00:36:46",
              "content": "<p>No. It is partitioned by time_id.  You will never get multiple time_id at the same time other than in lags. </p>\n<p>e.g</p>\n<p>---- first predict call -----<br>\ndate_id, time_id, symbol_id<br>\n1, 0, 1<br>\n1, 0, 2<br>\n1, 0, 3</p>\n<h2>…..</h2>\n<p>---- second predict call -----<br>\ndate_id, time_id, symbol_id<br>\n1, 1, 1<br>\n1, 1, 2<br>\n1, 1, 3</p>\n<h2>…..</h2>\n<p>IF you want to replicate this locally you can do</p>\n<pre><code>time_partitions = df.collect().partition_by(\n    , maintain_order=\n)\n\n timestep_data  time_partitions:\n    timestep_preds = predict(\n        timestep_data\n    )\n</code></pre>",
              "votes": null,
              "replies": []
            }
          ]
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "3072328": "Am I correct in understanding that the evaluation server sends the samples for one date at a time to the predict function, and the average number of samples on one date is always below 40,000?",
    "3073015": "It sends each time_id individually I believe. So you will get roughly one row per symbol on every invocation and roughly 968 or whatever it is invocations per date_id",
    "3073020": "Thanks, just to clarify, the dataframe that is provided by the server at each call to the predict function includes all the time_id and symbols for one date_id? In the validation set there are about 40,000 rows for each date_id file, so I assumed that for the hidden set should be about the same. Is that correct?",
    "3073035": "No. It is partitioned by time_id.  You will never get multiple time_id at the same time other than in lags. \n\ne.g\n\n---- first predict call -----\ndate_id, time_id, symbol_id\n1, 0, 1\n1, 0, 2\n1, 0, 3\n.....\n------\n---- second predict call -----\ndate_id, time_id, symbol_id\n1, 1, 1\n1, 1, 2\n1, 1, 3\n.....\n------\n\nIF you want to replicate this locally you can do\n\n```python\ntime_partitions = df.collect().partition_by(\n    \"time_id\", maintain_order=True\n)\n\nfor timestep_data in time_partitions:\n    timestep_preds = predict(\n        timestep_data\n    )\n```"
  },
  "source": "meta"
}