{
  "id": 549461,
  "title": "Can we save provided test data and prediction result during API call to use them in next batch?",
  "url": "/competitions/jane-street-real-time-market-data-forecasting/discussion/549461",
  "author_name": "",
  "post_date": "2024-12-02T13:31:50.447381100Z",
  "votes": null,
  "comment_count": 6,
  "views": 0,
  "content": "<p>Hello all, </p>\n<p>I would like to ask your support about data handling during API call.</p>\n<p>From the notebook <a href=\"https://www.kaggle.com/code/ryanholbrook/jane-street-rmf-demo-submission\" target=\"_blank\">Jane Street RMF Demo Submission</a> by <a href=\"https://www.kaggle.com/ryanholbrook\" target=\"_blank\">@ryanholbrook</a> &amp; <a href=\"https://www.kaggle.com/sohier\" target=\"_blank\">@sohier</a> , I understand lag data is provided at time_id=0 in each date and that is stored in global variable, lags_.</p>\n<p>My question here is can I save Test data and prediction data in batch as same as lag data and can I use it as additional data in next provided batch?</p>\n<p>I checked some discussions and notebooks but could not find the answer for that. If you would give me any hint, I really appreciate it 🙏</p>",
  "messages": [
    {
      "id": "3061230",
      "postDate": "12/02/2024 13:31:50",
      "content": "<p>Hello all, </p>\n<p>I would like to ask your support about data handling during API call.</p>\n<p>From the notebook <a href=\"https://www.kaggle.com/code/ryanholbrook/jane-street-rmf-demo-submission\" target=\"_blank\">Jane Street RMF Demo Submission</a> by <a href=\"https://www.kaggle.com/ryanholbrook\" target=\"_blank\">@ryanholbrook</a> &amp; <a href=\"https://www.kaggle.com/sohier\" target=\"_blank\">@sohier</a> , I understand lag data is provided at time_id=0 in each date and that is stored in global variable, lags_.</p>\n<p>My question here is can I save Test data and prediction data in batch as same as lag data and can I use it as additional data in next provided batch?</p>\n<p>I checked some discussions and notebooks but could not find the answer for that. If you would give me any hint, I really appreciate it 🙏</p>",
      "rawMarkdown": "Hello all, \n\nI would like to ask your support about data handling during API call.\n\nFrom the notebook [Jane Street RMF Demo Submission](https://www.kaggle.com/code/ryanholbrook/jane-street-rmf-demo-submission) by @ryanholbrook & @sohier , I understand lag data is provided at time_id=0 in each date and that is stored in global variable, lags_.\n\nMy question here is can I save Test data and prediction data in batch as same as lag data and can I use it as additional data in next provided batch?\n\nI checked some discussions and notebooks but could not find the answer for that. If you would give me any hint, I really appreciate it 🙏",
      "votes": null
    },
    {
      "id": "3061617",
      "postDate": "12/02/2024 21:04:34",
      "content": "<p>Currently, in this phase of the competition, the answer is YES.  The test set evaluation is run entirely in a single process with many invocations of the predict() method -- one for each time_id.  This allows you to store any information you want in \"global\" variables.  That said, I assume that we will need to make some changes once we get to the next phase of the competition since I would assume that each day of the 6 month evaluation period is going to launch a new process instance and that process instance will be used for all the time_id specific calls to the predict() method for the that day.  This would imply that to store data between process invocations (between days), we'll have to write to disk.  However, given that this is my first kaggle competition, I don't know the details of how this will run in the next phase.</p>",
      "rawMarkdown": "Currently, in this phase of the competition, the answer is YES.  The test set evaluation is run entirely in a single process with many invocations of the predict() method -- one for each time_id.  This allows you to store any information you want in \"global\" variables.  That said, I assume that we will need to make some changes once we get to the next phase of the competition since I would assume that each day of the 6 month evaluation period is going to launch a new process instance and that process instance will be used for all the time_id specific calls to the predict() method for the that day.  This would imply that to store data between process invocations (between days), we'll have to write to disk.  However, given that this is my first kaggle competition, I don't know the details of how this will run in the next phase.",
      "votes": null
    },
    {
      "id": "3061723",
      "postDate": "12/03/2024 00:18:55",
      "content": "<p>Hello <a href=\"https://www.kaggle.com/maciejzawadzki\" target=\"_blank\">@maciejzawadzki</a> , thank you for your prompt reply. That is helpful for me! OK I will try to save test data in global variable during API call. On the other hand I am bit concern about possible change of way of providing data in future prediction phase. Could you kindly elaborate from which information do you think that change happen? Is there announcement from competition owner? Thanks for your advice in advance 🙏</p>",
      "rawMarkdown": "Hello @maciejzawadzki , thank you for your prompt reply. That is helpful for me! OK I will try to save test data in global variable during API call. On the other hand I am bit concern about possible change of way of providing data in future prediction phase. Could you kindly elaborate from which information do you think that change happen? Is there announcement from competition owner? Thanks for your advice in advance 🙏",
      "votes": null
    },
    {
      "id": "3061878",
      "postDate": "12/03/2024 04:01:19",
      "content": "<p>Looks like I'm wrong -- \"During the forecasting phase, the evaluation API will serve test data from the beginning of the public set to the end of the private set. You must make predictions at every timestep, but, in this phase, only predictions on the private set are scored. (You may predict 0.0 on the unscored segments, if you like.)\" from <a href=\"https://www.kaggle.com/competitions/jane-street-real-time-market-data-forecasting/data\" target=\"_blank\">https://www.kaggle.com/competitions/jane-street-real-time-market-data-forecasting/data</a></p>\n<p>This would indicate that during the forecasting phase, they will call the predict() method each day with not just the latest day data, but with data from the beginning of evaluation time.  So, storing data in global variables should be fine.</p>\n<p>All my statements are IMHO only as I've not participated in any prior Jane Street competitions or even any kaggle competitions.</p>\n<p>Good luck.</p>",
      "rawMarkdown": "Looks like I'm wrong -- \"During the forecasting phase, the evaluation API will serve test data from the beginning of the public set to the end of the private set. You must make predictions at every timestep, but, in this phase, only predictions on the private set are scored. (You may predict 0.0 on the unscored segments, if you like.)\" from https://www.kaggle.com/competitions/jane-street-real-time-market-data-forecasting/data\n\nThis would indicate that during the forecasting phase, they will call the predict() method each day with not just the latest day data, but with data from the beginning of evaluation time.  So, storing data in global variables should be fine.\n\nAll my statements are IMHO only as I've not participated in any prior Jane Street competitions or even any kaggle competitions.\n\nGood luck.",
      "votes": null
    },
    {
      "id": "3062418",
      "postDate": "12/03/2024 15:13:15",
      "content": "<p><a href=\"https://www.kaggle.com/maciejzawadzki\" target=\"_blank\">@maciejzawadzki</a> , thanks again for your update! Looks like there will be the way to keep data in global variable even during forecasting phase.  Will try it out. </p>\n<p>Thank you and good luck to you 😉 </p>",
      "rawMarkdown": "maciejzawadzki , thanks again for your update! Looks like there will be the way to keep data in global variable even during forecasting phase.  Will try it out. \n\nThank you and good luck to you 😉",
      "votes": null
    },
    {
      "id": "3062600",
      "postDate": "12/03/2024 17:49:35",
      "content": "<p>yes you can.</p>\n<p>I wonder why you could not find relevant information in the discussion/code.</p>\n<p>For example:</p>\n<p><a href=\"https://www.kaggle.com/code/shiyili/js2024-rmf-gru-inference\" target=\"_blank\">https://www.kaggle.com/code/shiyili/js2024-rmf-gru-inference</a></p>",
      "rawMarkdown": "yes you can.\n\nI wonder why you could not find relevant information in the discussion/code.\n\nFor example:\n\nhttps://www.kaggle.com/code/shiyili/js2024-rmf-gru-inference",
      "votes": null
    },
    {
      "id": "3063052",
      "postDate": "12/04/2024 05:53:13",
      "content": "<p>Hello <a href=\"https://www.kaggle.com/shiyili\" target=\"_blank\">@shiyili</a> , thank you very much for your comment and sharing that great example how to save provided data! It is really helpful for my understanding 🙏</p>",
      "rawMarkdown": "Hello @shiyili , thank you very much for your comment and sharing that great example how to save provided data! It is really helpful for my understanding 🙏",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 3061617,
      "author_name": "maciejzawadzki",
      "author_url": "",
      "post_date": "12/02/2024 21:04:34",
      "content": "<p>Currently, in this phase of the competition, the answer is YES.  The test set evaluation is run entirely in a single process with many invocations of the predict() method -- one for each time_id.  This allows you to store any information you want in \"global\" variables.  That said, I assume that we will need to make some changes once we get to the next phase of the competition since I would assume that each day of the 6 month evaluation period is going to launch a new process instance and that process instance will be used for all the time_id specific calls to the predict() method for the that day.  This would imply that to store data between process invocations (between days), we'll have to write to disk.  However, given that this is my first kaggle competition, I don't know the details of how this will run in the next phase.</p>",
      "votes": null,
      "replies": [
        {
          "id": 3061723,
          "author_name": "hechtjp",
          "author_url": "",
          "post_date": "12/03/2024 00:18:55",
          "content": "<p>Hello <a href=\"https://www.kaggle.com/maciejzawadzki\" target=\"_blank\">@maciejzawadzki</a> , thank you for your prompt reply. That is helpful for me! OK I will try to save test data in global variable during API call. On the other hand I am bit concern about possible change of way of providing data in future prediction phase. Could you kindly elaborate from which information do you think that change happen? Is there announcement from competition owner? Thanks for your advice in advance 🙏</p>",
          "votes": null,
          "replies": [
            {
              "id": 3061878,
              "author_name": "maciejzawadzki",
              "author_url": "",
              "post_date": "12/03/2024 04:01:19",
              "content": "<p>Looks like I'm wrong -- \"During the forecasting phase, the evaluation API will serve test data from the beginning of the public set to the end of the private set. You must make predictions at every timestep, but, in this phase, only predictions on the private set are scored. (You may predict 0.0 on the unscored segments, if you like.)\" from <a href=\"https://www.kaggle.com/competitions/jane-street-real-time-market-data-forecasting/data\" target=\"_blank\">https://www.kaggle.com/competitions/jane-street-real-time-market-data-forecasting/data</a></p>\n<p>This would indicate that during the forecasting phase, they will call the predict() method each day with not just the latest day data, but with data from the beginning of evaluation time.  So, storing data in global variables should be fine.</p>\n<p>All my statements are IMHO only as I've not participated in any prior Jane Street competitions or even any kaggle competitions.</p>\n<p>Good luck.</p>",
              "votes": null,
              "replies": [
                {
                  "id": 3062418,
                  "author_name": "hechtjp",
                  "author_url": "",
                  "post_date": "12/03/2024 15:13:15",
                  "content": "<p><a href=\"https://www.kaggle.com/maciejzawadzki\" target=\"_blank\">@maciejzawadzki</a> , thanks again for your update! Looks like there will be the way to keep data in global variable even during forecasting phase.  Will try it out. </p>\n<p>Thank you and good luck to you 😉 </p>",
                  "votes": null,
                  "replies": []
                }
              ]
            }
          ]
        }
      ]
    },
    {
      "id": 3062600,
      "author_name": "shiyili",
      "author_url": "",
      "post_date": "12/03/2024 17:49:35",
      "content": "<p>yes you can.</p>\n<p>I wonder why you could not find relevant information in the discussion/code.</p>\n<p>For example:</p>\n<p><a href=\"https://www.kaggle.com/code/shiyili/js2024-rmf-gru-inference\" target=\"_blank\">https://www.kaggle.com/code/shiyili/js2024-rmf-gru-inference</a></p>",
      "votes": null,
      "replies": [
        {
          "id": 3063052,
          "author_name": "hechtjp",
          "author_url": "",
          "post_date": "12/04/2024 05:53:13",
          "content": "<p>Hello <a href=\"https://www.kaggle.com/shiyili\" target=\"_blank\">@shiyili</a> , thank you very much for your comment and sharing that great example how to save provided data! It is really helpful for my understanding 🙏</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "3061230": "Hello all, \n\nI would like to ask your support about data handling during API call.\n\nFrom the notebook [Jane Street RMF Demo Submission](https://www.kaggle.com/code/ryanholbrook/jane-street-rmf-demo-submission) by @ryanholbrook & @sohier , I understand lag data is provided at time_id=0 in each date and that is stored in global variable, lags_.\n\nMy question here is can I save Test data and prediction data in batch as same as lag data and can I use it as additional data in next provided batch?\n\nI checked some discussions and notebooks but could not find the answer for that. If you would give me any hint, I really appreciate it 🙏",
    "3061617": "Currently, in this phase of the competition, the answer is YES.  The test set evaluation is run entirely in a single process with many invocations of the predict() method -- one for each time_id.  This allows you to store any information you want in \"global\" variables.  That said, I assume that we will need to make some changes once we get to the next phase of the competition since I would assume that each day of the 6 month evaluation period is going to launch a new process instance and that process instance will be used for all the time_id specific calls to the predict() method for the that day.  This would imply that to store data between process invocations (between days), we'll have to write to disk.  However, given that this is my first kaggle competition, I don't know the details of how this will run in the next phase.",
    "3061723": "Hello @maciejzawadzki , thank you for your prompt reply. That is helpful for me! OK I will try to save test data in global variable during API call. On the other hand I am bit concern about possible change of way of providing data in future prediction phase. Could you kindly elaborate from which information do you think that change happen? Is there announcement from competition owner? Thanks for your advice in advance 🙏",
    "3061878": "Looks like I'm wrong -- \"During the forecasting phase, the evaluation API will serve test data from the beginning of the public set to the end of the private set. You must make predictions at every timestep, but, in this phase, only predictions on the private set are scored. (You may predict 0.0 on the unscored segments, if you like.)\" from https://www.kaggle.com/competitions/jane-street-real-time-market-data-forecasting/data\n\nThis would indicate that during the forecasting phase, they will call the predict() method each day with not just the latest day data, but with data from the beginning of evaluation time.  So, storing data in global variables should be fine.\n\nAll my statements are IMHO only as I've not participated in any prior Jane Street competitions or even any kaggle competitions.\n\nGood luck.",
    "3062418": "maciejzawadzki , thanks again for your update! Looks like there will be the way to keep data in global variable even during forecasting phase.  Will try it out. \n\nThank you and good luck to you 😉",
    "3062600": "yes you can.\n\nI wonder why you could not find relevant information in the discussion/code.\n\nFor example:\n\nhttps://www.kaggle.com/code/shiyili/js2024-rmf-gru-inference",
    "3063052": "Hello @shiyili , thank you very much for your comment and sharing that great example how to save provided data! It is really helpful for my understanding 🙏"
  },
  "source": "meta"
}