{
  "id": 553598,
  "title": "Sequence Dataset",
  "url": "/competitions/jane-street-real-time-market-data-forecasting/discussion/553598",
  "author_name": "",
  "post_date": "2024-12-27T06:55:45.586441500Z",
  "votes": 11,
  "comment_count": 19,
  "views": 0,
  "content": "<p>Hi everyone<br>\nHere we present our notebook that generates a sequence dataset. I see most people model the problem as a time series problem. Our team plans to use sequence models to fully utilize the lag data. <br>\nWe phrasHere we present our notebook that generates a sequence dataset. I see most people model the problem as a time series problem. Our team plans to use sequence models to fully utilize the lag data. <br>\n<strong>We model the problem as follows:</strong></p>\n<ul>\n<li>Given 8 pairs of feature-responder sample from previous day, plus 8 feature sample from the current day, predict the responder at the last time_id</li>\n</ul>\n<p><strong>Here is a brief description of the dataset:</strong></p>\n<ul>\n<li>The dataset is composed of 3 parts that are index aligned: trajectory, output, weight</li>\n<li>Trajectory: each trajectory is a 16x4200 tensor. That is 16 timesteps, each with 4200 feature dimensions</li>\n<li>4200: 42-padded symbols by 100-padded features/responders</li>\n<li>For details of how data is prepared, please refer to the notebook</li>\n<li>Time_id is treated as a 4200-dim positional embedding, added to the feature dimension</li>\n<li>Integer columns are replaced with vector embeddings</li>\n</ul>\n<p>I did not leave much documentation in the notebook. If you have any questions, please leave comments. I am looking forward to hearing feedbacks. Also, if you are also interested in using sequence models, please let me know. We are looking for teammates.</p>\n<p>link:<br>\n<a href=\"https://www.kaggle.com/code/laodriverayu/seq-data\" target=\"_blank\">https://www.kaggle.com/code/laodriverayu/seq-data</a></p>",
  "messages": [
    {
      "id": "3081758",
      "postDate": "12/27/2024 06:55:45",
      "content": "<p>Hi everyone<br>\nHere we present our notebook that generates a sequence dataset. I see most people model the problem as a time series problem. Our team plans to use sequence models to fully utilize the lag data. <br>\nWe phrasHere we present our notebook that generates a sequence dataset. I see most people model the problem as a time series problem. Our team plans to use sequence models to fully utilize the lag data. <br>\n<strong>We model the problem as follows:</strong></p>\n<ul>\n<li>Given 8 pairs of feature-responder sample from previous day, plus 8 feature sample from the current day, predict the responder at the last time_id</li>\n</ul>\n<p><strong>Here is a brief description of the dataset:</strong></p>\n<ul>\n<li>The dataset is composed of 3 parts that are index aligned: trajectory, output, weight</li>\n<li>Trajectory: each trajectory is a 16x4200 tensor. That is 16 timesteps, each with 4200 feature dimensions</li>\n<li>4200: 42-padded symbols by 100-padded features/responders</li>\n<li>For details of how data is prepared, please refer to the notebook</li>\n<li>Time_id is treated as a 4200-dim positional embedding, added to the feature dimension</li>\n<li>Integer columns are replaced with vector embeddings</li>\n</ul>\n<p>I did not leave much documentation in the notebook. If you have any questions, please leave comments. I am looking forward to hearing feedbacks. Also, if you are also interested in using sequence models, please let me know. We are looking for teammates.</p>\n<p>link:<br>\n<a href=\"https://www.kaggle.com/code/laodriverayu/seq-data\" target=\"_blank\">https://www.kaggle.com/code/laodriverayu/seq-data</a></p>",
      "rawMarkdown": "Hi everyone\nHere we present our notebook that generates a sequence dataset. I see most people model the problem as a time series problem. Our team plans to use sequence models to fully utilize the lag data. \nWe phrasHere we present our notebook that generates a sequence dataset. I see most people model the problem as a time series problem. Our team plans to use sequence models to fully utilize the lag data. \n**We model the problem as follows:**\n- Given 8 pairs of feature-responder sample from previous day, plus 8 feature sample from the current day, predict the responder at the last time_id\n\n**Here is a brief description of the dataset:**\n- The dataset is composed of 3 parts that are index aligned: trajectory, output, weight\n- Trajectory: each trajectory is a 16x4200 tensor. That is 16 timesteps, each with 4200 feature dimensions\n- 4200: 42-padded symbols by 100-padded features/responders\n- For details of how data is prepared, please refer to the notebook\n- Time_id is treated as a 4200-dim positional embedding, added to the feature dimension\n- Integer columns are replaced with vector embeddings\n\nI did not leave much documentation in the notebook. If you have any questions, please leave comments. I am looking forward to hearing feedbacks. Also, if you are also interested in using sequence models, please let me know. We are looking for teammates.\n\nlink:\nhttps://www.kaggle.com/code/laodriverayu/seq-data",
      "votes": null
    },
    {
      "id": "3081845",
      "postDate": "12/27/2024 09:29:54",
      "content": "<p>Good dataset!</p>",
      "rawMarkdown": "Good dataset!",
      "votes": null
    },
    {
      "id": "3081853",
      "postDate": "12/27/2024 09:37:02",
      "content": "<p>May I ask how many your model's parameter size is and how long it take to train? It seems that the dimension of input data is large.</p>",
      "rawMarkdown": "May I ask how many your model's parameter size is and how long it take to train? It seems that the dimension of input data is large.",
      "votes": null
    },
    {
      "id": "3081856",
      "postDate": "12/27/2024 09:41:08",
      "content": "<p>For the model part, I am still testing. But I am using vanilla models built from scratch. Most of the trainings are done in less than 1 hr. But you are right, currently I am simply concatenating all info about one time step</p>",
      "rawMarkdown": "For the model part, I am still testing. But I am using vanilla models built from scratch. Most of the trainings are done in less than 1 hr. But you are right, currently I am simply concatenating all info about one time step",
      "votes": null
    },
    {
      "id": "3081912",
      "postDate": "12/27/2024 12:04:09",
      "content": "<p>What is the difference between modeling as sequence and modeling as time series? I think that time series can be considered a sequence.</p>",
      "rawMarkdown": "What is the difference between modeling as sequence and modeling as time series? I think that time series can be considered a sequence.",
      "votes": null
    },
    {
      "id": "3081920",
      "postDate": "12/27/2024 12:16:15",
      "content": "<p>I am trying to model the problem as a time series using the darts library. The main drawback is the converstion between Dataframe and darts.TimeSeries data types. </p>\n<p>The time to execute one predict call is around 0.3 seconds (which is apparently insufficient to predict all private data in time)<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1342094%2F4aed07746a8cac27a8c5a69ec94b3212%2Ftime.png?generation=1735301619169569&amp;alt=media\" alt=\"\"></p>\n<p>I am curious to know about your time in the inference step.</p>",
      "rawMarkdown": "I am trying to model the problem as a time series using the darts library. The main drawback is the converstion between Dataframe and darts.TimeSeries data types. \n\nThe time to execute one predict call is around 0.3 seconds (which is apparently insufficient to predict all private data in time)\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1342094%2F4aed07746a8cac27a8c5a69ec94b3212%2Ftime.png?generation=1735301619169569&alt=media)\n\nI am curious to know about your time in the inference step.",
      "votes": null
    },
    {
      "id": "3082172",
      "postDate": "12/27/2024 18:41:42",
      "content": "<p>For the time series, I mean markovian model, or 1-lag model</p>",
      "rawMarkdown": "For the time series, I mean markovian model, or 1-lag model",
      "votes": null
    },
    {
      "id": "3082173",
      "postDate": "12/27/2024 18:43:13",
      "content": "<p>How do you know how many predictions we need to make in the private data?</p>",
      "rawMarkdown": "How do you know how many predictions we need to make in the private data?",
      "votes": null
    },
    {
      "id": "3082182",
      "postDate": "12/27/2024 19:07:57",
      "content": "<p>hmmm let me see if I understand. By time series modeling you mean a sequence with a lookback window of 1 and by a sequence modeling you mean a sequence with lookback window that can assume N values, right?</p>",
      "rawMarkdown": "hmmm let me see if I understand. By time series modeling you mean a sequence with a lookback window of 1 and by a sequence modeling you mean a sequence with lookback window that can assume N values, right?",
      "votes": null
    },
    {
      "id": "3082188",
      "postDate": "12/27/2024 19:18:49",
      "content": "<p>I don't know exactly. It is a gross estimate. </p>\n<p>I am assuming 120 date_ids with 968 time_ids. The predict mehtod is called for every time id.</p>\n<p>We have 8 hours to complete the submission. So:<br>\n(8 hours * 60 minutes * 60 seconds) / (120 days * 968 time units) = 0,24.</p>\n<p>If I'm not mistaken, each prediction call should take 0.24 seconds or less to complete.</p>",
      "rawMarkdown": "I don't know exactly. It is a gross estimate. \n\nI am assuming 120 date_ids with 968 time_ids. The predict mehtod is called for every time id.\n\nWe have 8 hours to complete the submission. So:\n(8 hours * 60 minutes * 60 seconds) / (120 days * 968 time units) = 0,24.\n\nIf I'm not mistaken, each prediction call should take 0.24 seconds or less to complete.",
      "votes": null
    },
    {
      "id": "3082212",
      "postDate": "12/27/2024 19:54:29",
      "content": "<p>right. so with sequence modeling we can use a bunch of sequence models like rnn, lstm, transformer</p>",
      "rawMarkdown": "right. so with sequence modeling we can use a bunch of sequence models like rnn, lstm, transformer",
      "votes": null
    },
    {
      "id": "3082219",
      "postDate": "12/27/2024 20:02:32",
      "content": "<p>Did not really get the following two points:<br>\n1) what are the 8-pairs sample from the previous day and the current day?<br>\n2) why 16 time steps?</p>\n<p>Can you elaborate? 🙏 thx</p>",
      "rawMarkdown": "Did not really get the following two points:\n1) what are the 8-pairs sample from the previous day and the current day?\n2) why 16 time steps?\n\nCan you elaborate? 🙏 thx",
      "votes": null
    },
    {
      "id": "3082229",
      "postDate": "12/27/2024 20:25:12",
      "content": "<p>Suppose the model is deployed. We can cache the test dataframe, which will provide the 'feature columns' of the current day. We can also access the lag dataframe, which will provide the 'responder columns' of the previous day. With a little more work, we can access previous day's 'feature columns' as well. So the 8 pairs sampled from previous day actually means: </p>\n<ol>\n<li>randomly sample 8 time_id</li>\n<li>group by time_id, get the chunk of feature cols and responder cols across all symbol_id rows</li>\n<li>flatten each group as a 'token' vector, concatenate the 8 tokens as a sequence</li>\n</ol>\n<p>The 8 samples from the current works the same way. Only difference is that responder columns are left blank</p>\n<p>For your second question, 16 is kinda arbitrary. If it turns out the inference is too slow, I might shorten the trajectory length</p>",
      "rawMarkdown": "Suppose the model is deployed. We can cache the test dataframe, which will provide the 'feature columns' of the current day. We can also access the lag dataframe, which will provide the 'responder columns' of the previous day. With a little more work, we can access previous day's 'feature columns' as well. So the 8 pairs sampled from previous day actually means: \n1. randomly sample 8 time_id\n2. group by time_id, get the chunk of feature cols and responder cols across all symbol_id rows\n3. flatten each group as a 'token' vector, concatenate the 8 tokens as a sequence\n\nThe 8 samples from the current works the same way. Only difference is that responder columns are left blank\n\nFor your second question, 16 is kinda arbitrary. If it turns out the inference is too slow, I might shorten the trajectory length",
      "votes": null
    },
    {
      "id": "3082269",
      "postDate": "12/27/2024 22:15:36",
      "content": "<p>randomly sample 8 time_id --&gt; Is here the 8 time_id some kind of hyper-param? Is it possible to use all time_id?</p>",
      "rawMarkdown": "randomly sample 8 time_id --> Is here the 8 time_id some kind of hyper-param? Is it possible to use all time_id?",
      "votes": null
    },
    {
      "id": "3082324",
      "postDate": "12/28/2024 01:28:35",
      "content": "<p>8 is just an arbitrary number. Using all time_id would take way too long to compute</p>",
      "rawMarkdown": "8 is just an arbitrary number. Using all time_id would take way too long to compute",
      "votes": null
    },
    {
      "id": "3083793",
      "postDate": "12/30/2024 03:13:22",
      "content": "<p>For the “16 timesteps” setting, does it affect reasoning speed and score?</p>",
      "rawMarkdown": "For the “16 timesteps” setting, does it affect reasoning speed and score?",
      "votes": null
    },
    {
      "id": "3083800",
      "postDate": "12/30/2024 03:24:57",
      "content": "<p>That is exactly what I am trying to find out</p>",
      "rawMarkdown": "That is exactly what I am trying to find out",
      "votes": null
    },
    {
      "id": "3084464",
      "postDate": "12/30/2024 20:54:02",
      "content": "<p>for the sequence data as the input of deep model, the timesteps here, are they created by the same time_id on different days or on the same date but different time_id? </p>",
      "rawMarkdown": "for the sequence data as the input of deep model, the timesteps here, are they created by the same time_id on different days or on the same date but different time_id?",
      "votes": null
    },
    {
      "id": "3084465",
      "postDate": "12/30/2024 20:55:12",
      "content": "<p>Creating sequence samples by both methods should be fine. Just curious about which format will have better results? </p>",
      "rawMarkdown": "Creating sequence samples by both methods should be fine. Just curious about which format will have better results?",
      "votes": null
    },
    {
      "id": "3084565",
      "postDate": "12/31/2024 02:55:31",
      "content": "<p>Do you use LSTM or Transformer other than MLP as your vanilla model? How does the model perform on cv and lb dataset?</p>",
      "rawMarkdown": "Do you use LSTM or Transformer other than MLP as your vanilla model? How does the model perform on cv and lb dataset?",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 3081845,
      "author_name": "mrsimple07",
      "author_url": "",
      "post_date": "12/27/2024 09:29:54",
      "content": "<p>Good dataset!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 3081853,
      "author_name": "i2nfinit3y",
      "author_url": "",
      "post_date": "12/27/2024 09:37:02",
      "content": "<p>May I ask how many your model's parameter size is and how long it take to train? It seems that the dimension of input data is large.</p>",
      "votes": null,
      "replies": [
        {
          "id": 3081856,
          "author_name": "laodriverayu",
          "author_url": "",
          "post_date": "12/27/2024 09:41:08",
          "content": "<p>For the model part, I am still testing. But I am using vanilla models built from scratch. Most of the trainings are done in less than 1 hr. But you are right, currently I am simply concatenating all info about one time step</p>",
          "votes": null,
          "replies": [
            {
              "id": 3084565,
              "author_name": "shanyun",
              "author_url": "",
              "post_date": "12/31/2024 02:55:31",
              "content": "<p>Do you use LSTM or Transformer other than MLP as your vanilla model? How does the model perform on cv and lb dataset?</p>",
              "votes": null,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 3081912,
      "author_name": "serjhenrique",
      "author_url": "",
      "post_date": "12/27/2024 12:04:09",
      "content": "<p>What is the difference between modeling as sequence and modeling as time series? I think that time series can be considered a sequence.</p>",
      "votes": null,
      "replies": [
        {
          "id": 3082172,
          "author_name": "laodriverayu",
          "author_url": "",
          "post_date": "12/27/2024 18:41:42",
          "content": "<p>For the time series, I mean markovian model, or 1-lag model</p>",
          "votes": null,
          "replies": [
            {
              "id": 3082182,
              "author_name": "serjhenrique",
              "author_url": "",
              "post_date": "12/27/2024 19:07:57",
              "content": "<p>hmmm let me see if I understand. By time series modeling you mean a sequence with a lookback window of 1 and by a sequence modeling you mean a sequence with lookback window that can assume N values, right?</p>",
              "votes": null,
              "replies": [
                {
                  "id": 3082212,
                  "author_name": "laodriverayu",
                  "author_url": "",
                  "post_date": "12/27/2024 19:54:29",
                  "content": "<p>right. so with sequence modeling we can use a bunch of sequence models like rnn, lstm, transformer</p>",
                  "votes": null,
                  "replies": []
                }
              ]
            }
          ]
        }
      ]
    },
    {
      "id": 3081920,
      "author_name": "serjhenrique",
      "author_url": "",
      "post_date": "12/27/2024 12:16:15",
      "content": "<p>I am trying to model the problem as a time series using the darts library. The main drawback is the converstion between Dataframe and darts.TimeSeries data types. </p>\n<p>The time to execute one predict call is around 0.3 seconds (which is apparently insufficient to predict all private data in time)<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1342094%2F4aed07746a8cac27a8c5a69ec94b3212%2Ftime.png?generation=1735301619169569&amp;alt=media\" alt=\"\"></p>\n<p>I am curious to know about your time in the inference step.</p>",
      "votes": null,
      "replies": [
        {
          "id": 3082173,
          "author_name": "laodriverayu",
          "author_url": "",
          "post_date": "12/27/2024 18:43:13",
          "content": "<p>How do you know how many predictions we need to make in the private data?</p>",
          "votes": null,
          "replies": [
            {
              "id": 3082188,
              "author_name": "serjhenrique",
              "author_url": "",
              "post_date": "12/27/2024 19:18:49",
              "content": "<p>I don't know exactly. It is a gross estimate. </p>\n<p>I am assuming 120 date_ids with 968 time_ids. The predict mehtod is called for every time id.</p>\n<p>We have 8 hours to complete the submission. So:<br>\n(8 hours * 60 minutes * 60 seconds) / (120 days * 968 time units) = 0,24.</p>\n<p>If I'm not mistaken, each prediction call should take 0.24 seconds or less to complete.</p>",
              "votes": null,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 3082219,
      "author_name": "shiyili",
      "author_url": "",
      "post_date": "12/27/2024 20:02:32",
      "content": "<p>Did not really get the following two points:<br>\n1) what are the 8-pairs sample from the previous day and the current day?<br>\n2) why 16 time steps?</p>\n<p>Can you elaborate? 🙏 thx</p>",
      "votes": null,
      "replies": [
        {
          "id": 3082229,
          "author_name": "laodriverayu",
          "author_url": "",
          "post_date": "12/27/2024 20:25:12",
          "content": "<p>Suppose the model is deployed. We can cache the test dataframe, which will provide the 'feature columns' of the current day. We can also access the lag dataframe, which will provide the 'responder columns' of the previous day. With a little more work, we can access previous day's 'feature columns' as well. So the 8 pairs sampled from previous day actually means: </p>\n<ol>\n<li>randomly sample 8 time_id</li>\n<li>group by time_id, get the chunk of feature cols and responder cols across all symbol_id rows</li>\n<li>flatten each group as a 'token' vector, concatenate the 8 tokens as a sequence</li>\n</ol>\n<p>The 8 samples from the current works the same way. Only difference is that responder columns are left blank</p>\n<p>For your second question, 16 is kinda arbitrary. If it turns out the inference is too slow, I might shorten the trajectory length</p>",
          "votes": null,
          "replies": [
            {
              "id": 3082269,
              "author_name": "shiyili",
              "author_url": "",
              "post_date": "12/27/2024 22:15:36",
              "content": "<p>randomly sample 8 time_id --&gt; Is here the 8 time_id some kind of hyper-param? Is it possible to use all time_id?</p>",
              "votes": null,
              "replies": [
                {
                  "id": 3082324,
                  "author_name": "laodriverayu",
                  "author_url": "",
                  "post_date": "12/28/2024 01:28:35",
                  "content": "<p>8 is just an arbitrary number. Using all time_id would take way too long to compute</p>",
                  "votes": null,
                  "replies": []
                }
              ]
            }
          ]
        }
      ]
    },
    {
      "id": 3083793,
      "author_name": "linghx",
      "author_url": "",
      "post_date": "12/30/2024 03:13:22",
      "content": "<p>For the “16 timesteps” setting, does it affect reasoning speed and score?</p>",
      "votes": null,
      "replies": [
        {
          "id": 3083800,
          "author_name": "laodriverayu",
          "author_url": "",
          "post_date": "12/30/2024 03:24:57",
          "content": "<p>That is exactly what I am trying to find out</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 3084464,
      "author_name": "jiadiwang",
      "author_url": "",
      "post_date": "12/30/2024 20:54:02",
      "content": "<p>for the sequence data as the input of deep model, the timesteps here, are they created by the same time_id on different days or on the same date but different time_id? </p>",
      "votes": null,
      "replies": [
        {
          "id": 3084465,
          "author_name": "jiadiwang",
          "author_url": "",
          "post_date": "12/30/2024 20:55:12",
          "content": "<p>Creating sequence samples by both methods should be fine. Just curious about which format will have better results? </p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "3081758": "Hi everyone\nHere we present our notebook that generates a sequence dataset. I see most people model the problem as a time series problem. Our team plans to use sequence models to fully utilize the lag data. \nWe phrasHere we present our notebook that generates a sequence dataset. I see most people model the problem as a time series problem. Our team plans to use sequence models to fully utilize the lag data. \n**We model the problem as follows:**\n- Given 8 pairs of feature-responder sample from previous day, plus 8 feature sample from the current day, predict the responder at the last time_id\n\n**Here is a brief description of the dataset:**\n- The dataset is composed of 3 parts that are index aligned: trajectory, output, weight\n- Trajectory: each trajectory is a 16x4200 tensor. That is 16 timesteps, each with 4200 feature dimensions\n- 4200: 42-padded symbols by 100-padded features/responders\n- For details of how data is prepared, please refer to the notebook\n- Time_id is treated as a 4200-dim positional embedding, added to the feature dimension\n- Integer columns are replaced with vector embeddings\n\nI did not leave much documentation in the notebook. If you have any questions, please leave comments. I am looking forward to hearing feedbacks. Also, if you are also interested in using sequence models, please let me know. We are looking for teammates.\n\nlink:\nhttps://www.kaggle.com/code/laodriverayu/seq-data",
    "3081845": "Good dataset!",
    "3081853": "May I ask how many your model's parameter size is and how long it take to train? It seems that the dimension of input data is large.",
    "3081856": "For the model part, I am still testing. But I am using vanilla models built from scratch. Most of the trainings are done in less than 1 hr. But you are right, currently I am simply concatenating all info about one time step",
    "3081912": "What is the difference between modeling as sequence and modeling as time series? I think that time series can be considered a sequence.",
    "3081920": "I am trying to model the problem as a time series using the darts library. The main drawback is the converstion between Dataframe and darts.TimeSeries data types. \n\nThe time to execute one predict call is around 0.3 seconds (which is apparently insufficient to predict all private data in time)\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1342094%2F4aed07746a8cac27a8c5a69ec94b3212%2Ftime.png?generation=1735301619169569&alt=media)\n\nI am curious to know about your time in the inference step.",
    "3082172": "For the time series, I mean markovian model, or 1-lag model",
    "3082173": "How do you know how many predictions we need to make in the private data?",
    "3082182": "hmmm let me see if I understand. By time series modeling you mean a sequence with a lookback window of 1 and by a sequence modeling you mean a sequence with lookback window that can assume N values, right?",
    "3082188": "I don't know exactly. It is a gross estimate. \n\nI am assuming 120 date_ids with 968 time_ids. The predict mehtod is called for every time id.\n\nWe have 8 hours to complete the submission. So:\n(8 hours * 60 minutes * 60 seconds) / (120 days * 968 time units) = 0,24.\n\nIf I'm not mistaken, each prediction call should take 0.24 seconds or less to complete.",
    "3082212": "right. so with sequence modeling we can use a bunch of sequence models like rnn, lstm, transformer",
    "3082219": "Did not really get the following two points:\n1) what are the 8-pairs sample from the previous day and the current day?\n2) why 16 time steps?\n\nCan you elaborate? 🙏 thx",
    "3082229": "Suppose the model is deployed. We can cache the test dataframe, which will provide the 'feature columns' of the current day. We can also access the lag dataframe, which will provide the 'responder columns' of the previous day. With a little more work, we can access previous day's 'feature columns' as well. So the 8 pairs sampled from previous day actually means: \n1. randomly sample 8 time_id\n2. group by time_id, get the chunk of feature cols and responder cols across all symbol_id rows\n3. flatten each group as a 'token' vector, concatenate the 8 tokens as a sequence\n\nThe 8 samples from the current works the same way. Only difference is that responder columns are left blank\n\nFor your second question, 16 is kinda arbitrary. If it turns out the inference is too slow, I might shorten the trajectory length",
    "3082269": "randomly sample 8 time_id --> Is here the 8 time_id some kind of hyper-param? Is it possible to use all time_id?",
    "3082324": "8 is just an arbitrary number. Using all time_id would take way too long to compute",
    "3083793": "For the “16 timesteps” setting, does it affect reasoning speed and score?",
    "3083800": "That is exactly what I am trying to find out",
    "3084464": "for the sequence data as the input of deep model, the timesteps here, are they created by the same time_id on different days or on the same date but different time_id?",
    "3084465": "Creating sequence samples by both methods should be fine. Just curious about which format will have better results?",
    "3084565": "Do you use LSTM or Transformer other than MLP as your vanilla model? How does the model perform on cv and lb dataset?"
  },
  "source": "meta"
}