{
  "id": 548636,
  "title": "Some domain knowledge",
  "url": "/competitions/jane-street-real-time-market-data-forecasting/discussion/548636",
  "author_name": "Tucker Arrants",
  "post_date": "2024-11-28T00:56:31.048000",
  "votes": 77,
  "comment_count": 18,
  "views": 0,
  "content": "<p>I have some domain knowledge for this competition, as I am a full-time quant trader. I would like to share some of my thoughts. I joined this competition a few days ago, so I apologize for the later discussion post. </p>\n<p>Firstly, who is Jane Street? As per their website, they are \"a quantitative trading firm and liquidity provider with a unique focus on technology and collaborative problem solving.\" Liquidity providers use sophisticated high-speed computers and algorithms to create volume on exchanges in order to add liquidity to the markets.  </p>\n<p>It is very likely that the data we are training on is data about these high-speed algorithms and the trades they execute. Algorithmic trading bots execute on rule-based strategies. I believe the features metadata represents these particular rules, and the feature columns are certain indicators / signal values at the time of execution.</p>\n<p>Assuming this is true, we can guess why the data is so seemingly inconsistent: different <code>time_ids</code> per <code>date_id</code>, different <code>symbol_ids</code> per <code>date_id</code>, the missing values before day 500, and the increasing number of <code>time_ids</code> per <code>date_id</code>, as <code>date_id</code> increases. I do not believe that these are lagging indicators, as I have seen some suggest. This is primarily because financial markets are extremely efficient on small time frames, meaning there is no extra \"alpha\" to gain by considering lagging features. Perhaps some are, but to me, it makes more sense that these are the byproduct of several different algo trading at once. </p>\n<p>Perhaps the initial algorithms relied on different indicators and signals, or were trading less frequently. As Jane Street created more sophisticated algorithms, it seems like they \"consolidated\" them over time, deploying them on new <code>symbol_ids</code>. This is why the data is more structured as <code>date_id</code> increases. </p>\n<p>So what do they hope to learn from us Kagglers. I can assure you these algorithmic bots are already highly profitable. They are not asking us to predict something simple like the \"price of the S&amp;P500 in 3 months.\" They want us to project the results of their existing algorithmic bots into the future, to help with risk-management, which is crucial for liquidity providers. </p>\n<p>Furthermore, it seems safe to assume that <code>time_id</code> per <code>date_id</code> will be roughly constant in the public test dataset and the private test dataset, which is much different from the train dataset. It feels like the training dataset represents a period of improving their algorithms, and now that they have, they are interested in not only forecasting their results into the future, but also on deploying them on new financial instruments or <code>symbol_ids</code>. This assumption will impact modeling, especially if you are using a sequential approach. </p>\n<p>But ideally, they can learn something from their old algo trades as well, meaning an auxiliary ideal outcome of this competition is finding a way of using these first algo trades in their current algos. I think they would be \"disappointed\" if the winning solution just dropped them in their pipeline. As such, it should be a priority for us as well. Dropping, filling with 0, or forward filling, does not feel right here, no matter the impact on model performance. It may be worth using more sophisticated means of handling the missing data, like clustering. </p>\n<p>We also need to help our model generalize to new <code>symbol_ids</code>, as this is likely what Jane Street wants to do. We cannot simply train models on the raw <code>symbol_id</code>, or it will not generalize well to new ones. A simple way of teaching a model to learn <code>symbol_ids</code> without it overfitting, is with an embedding layer. </p>\n<p>I hope this helps, please let me know your thoughts below. With the above assumptions, I am focusing on validating with more recent samples, as they likely reflect the most recent trading algos, which are the ones used on the private and public test set. Furthermore, I am prioritizing <code>symbol_id</code> and assuming a fairly consistent number of <code>time_ids</code> per <code>date_id</code> in the test datasets. Below are some graphs of my current exploration into the above thesis:<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4379159%2F218f8242c901e566a0a073ba427860c7%2FScreenshot%202024-11-30%20235444.png?generation=1733028902671124&amp;alt=media\" alt=\"\"><br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4379159%2F0f1c7b71aac780d6454ce1b71a745c3c%2FScreenshot%202024-11-30%20235917.png?generation=1733029177791681&amp;alt=media\" alt=\"\"></p>",
  "messages": [
    {
      "id": 3057276,
      "postDate": "2024-11-28T00:56:31.050Z",
      "content": "<p>I have some domain knowledge for this competition, as I am a full-time quant trader. I would like to share some of my thoughts. I joined this competition a few days ago, so I apologize for the later discussion post. </p>\n<p>Firstly, who is Jane Street? As per their website, they are \"a quantitative trading firm and liquidity provider with a unique focus on technology and collaborative problem solving.\" Liquidity providers use sophisticated high-speed computers and algorithms to create volume on exchanges in order to add liquidity to the markets.  </p>\n<p>It is very likely that the data we are training on is data about these high-speed algorithms and the trades they execute. Algorithmic trading bots execute on rule-based strategies. I believe the features metadata represents these particular rules, and the feature columns are certain indicators / signal values at the time of execution.</p>\n<p>Assuming this is true, we can guess why the data is so seemingly inconsistent: different <code>time_ids</code> per <code>date_id</code>, different <code>symbol_ids</code> per <code>date_id</code>, the missing values before day 500, and the increasing number of <code>time_ids</code> per <code>date_id</code>, as <code>date_id</code> increases. I do not believe that these are lagging indicators, as I have seen some suggest. This is primarily because financial markets are extremely efficient on small time frames, meaning there is no extra \"alpha\" to gain by considering lagging features. Perhaps some are, but to me, it makes more sense that these are the byproduct of several different algo trading at once. </p>\n<p>Perhaps the initial algorithms relied on different indicators and signals, or were trading less frequently. As Jane Street created more sophisticated algorithms, it seems like they \"consolidated\" them over time, deploying them on new <code>symbol_ids</code>. This is why the data is more structured as <code>date_id</code> increases. </p>\n<p>So what do they hope to learn from us Kagglers. I can assure you these algorithmic bots are already highly profitable. They are not asking us to predict something simple like the \"price of the S&amp;P500 in 3 months.\" They want us to project the results of their existing algorithmic bots into the future, to help with risk-management, which is crucial for liquidity providers. </p>\n<p>Furthermore, it seems safe to assume that <code>time_id</code> per <code>date_id</code> will be roughly constant in the public test dataset and the private test dataset, which is much different from the train dataset. It feels like the training dataset represents a period of improving their algorithms, and now that they have, they are interested in not only forecasting their results into the future, but also on deploying them on new financial instruments or <code>symbol_ids</code>. This assumption will impact modeling, especially if you are using a sequential approach. </p>\n<p>But ideally, they can learn something from their old algo trades as well, meaning an auxiliary ideal outcome of this competition is finding a way of using these first algo trades in their current algos. I think they would be \"disappointed\" if the winning solution just dropped them in their pipeline. As such, it should be a priority for us as well. Dropping, filling with 0, or forward filling, does not feel right here, no matter the impact on model performance. It may be worth using more sophisticated means of handling the missing data, like clustering. </p>\n<p>We also need to help our model generalize to new <code>symbol_ids</code>, as this is likely what Jane Street wants to do. We cannot simply train models on the raw <code>symbol_id</code>, or it will not generalize well to new ones. A simple way of teaching a model to learn <code>symbol_ids</code> without it overfitting, is with an embedding layer. </p>\n<p>I hope this helps, please let me know your thoughts below. With the above assumptions, I am focusing on validating with more recent samples, as they likely reflect the most recent trading algos, which are the ones used on the private and public test set. Furthermore, I am prioritizing <code>symbol_id</code> and assuming a fairly consistent number of <code>time_ids</code> per <code>date_id</code> in the test datasets. Below are some graphs of my current exploration into the above thesis:<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4379159%2F218f8242c901e566a0a073ba427860c7%2FScreenshot%202024-11-30%20235444.png?generation=1733028902671124&amp;alt=media\" alt=\"\"><br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4379159%2F0f1c7b71aac780d6454ce1b71a745c3c%2FScreenshot%202024-11-30%20235917.png?generation=1733029177791681&amp;alt=media\" alt=\"\"></p>",
      "rawMarkdown": "I have some domain knowledge for this competition, as I am a full-time quant trader. I would like to share some of my thoughts. I joined this competition a few days ago, so I apologize for the later discussion post. \n\nFirstly, who is Jane Street? As per their website, they are \"a quantitative trading firm and liquidity provider with a unique focus on technology and collaborative problem solving.\" Liquidity providers use sophisticated high-speed computers and algorithms to create volume on exchanges in order to add liquidity to the markets.  \n\nIt is very likely that the data we are training on is data about these high-speed algorithms and the trades they execute. Algorithmic trading bots execute on rule-based strategies. I believe the features metadata represents these particular rules, and the feature columns are certain indicators / signal values at the time of execution.\n\nAssuming this is true, we can guess why the data is so seemingly inconsistent: different `time_ids` per `date_id`, different `symbol_ids` per `date_id`, the missing values before day 500, and the increasing number of `time_ids` per `date_id`, as `date_id` increases. I do not believe that these are lagging indicators, as I have seen some suggest. This is primarily because financial markets are extremely efficient on small time frames, meaning there is no extra \"alpha\" to gain by considering lagging features. Perhaps some are, but to me, it makes more sense that these are the byproduct of several different algo trading at once. \n\nPerhaps the initial algorithms relied on different indicators and signals, or were trading less frequently. As Jane Street created more sophisticated algorithms, it seems like they \"consolidated\" them over time, deploying them on new `symbol_ids`. This is why the data is more structured as `date_id` increases. \n\nSo what do they hope to learn from us Kagglers. I can assure you these algorithmic bots are already highly profitable. They are not asking us to predict something simple like the \"price of the S&P500 in 3 months.\" They want us to project the results of their existing algorithmic bots into the future, to help with risk-management, which is crucial for liquidity providers. \n\nFurthermore, it seems safe to assume that `time_id` per `date_id` will be roughly constant in the public test dataset and the private test dataset, which is much different from the train dataset. It feels like the training dataset represents a period of improving their algorithms, and now that they have, they are interested in not only forecasting their results into the future, but also on deploying them on new financial instruments or `symbol_ids`. This assumption will impact modeling, especially if you are using a sequential approach. \n\nBut ideally, they can learn something from their old algo trades as well, meaning an auxiliary ideal outcome of this competition is finding a way of using these first algo trades in their current algos. I think they would be \"disappointed\" if the winning solution just dropped them in their pipeline. As such, it should be a priority for us as well. Dropping, filling with 0, or forward filling, does not feel right here, no matter the impact on model performance. It may be worth using more sophisticated means of handling the missing data, like clustering. \n\nWe also need to help our model generalize to new `symbol_ids`, as this is likely what Jane Street wants to do. We cannot simply train models on the raw `symbol_id`, or it will not generalize well to new ones. A simple way of teaching a model to learn `symbol_ids` without it overfitting, is with an embedding layer. \n\nI hope this helps, please let me know your thoughts below. With the above assumptions, I am focusing on validating with more recent samples, as they likely reflect the most recent trading algos, which are the ones used on the private and public test set. Furthermore, I am prioritizing `symbol_id` and assuming a fairly consistent number of `time_ids` per `date_id` in the test datasets. Below are some graphs of my current exploration into the above thesis:\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4379159%2F218f8242c901e566a0a073ba427860c7%2FScreenshot%202024-11-30%20235444.png?generation=1733028902671124&alt=media)\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4379159%2F0f1c7b71aac780d6454ce1b71a745c3c%2FScreenshot%202024-11-30%20235917.png?generation=1733029177791681&alt=media)",
      "votes": 77
    },
    {
      "id": 3058310,
      "postDate": "2024-11-29T10:14:59.813Z",
      "content": "<p>Great thoughts. The responders have time and symbol correlations. That is, the responders are correlated to themselves in the past, and correlated between symbols. Whatever implementation that considers only the features to get to the predictions is missing on important information (all public codes shared have this issue).<br>\nI'm not so sure about using the first trades, however. As another kaggler pointed out in a great thread, whatever relations we find tend to be very time-specific. This is something I'm currently struggling with - the more data you use, the more your model can generalize, but also, the less it can fit to the patterns that pop up and disappear in specific time frames. Is it better to use a generalist model, that fits to all data, but not so well, or a short-term model with frequent retrains, that fits very well to specific time-frames, but tends to overfit? Maybe the answer is a combination of these two.</p>",
      "rawMarkdown": "Great thoughts. The responders have time and symbol correlations. That is, the responders are correlated to themselves in the past, and correlated between symbols. Whatever implementation that considers only the features to get to the predictions is missing on important information (all public codes shared have this issue).\nI'm not so sure about using the first trades, however. As another kaggler pointed out in a great thread, whatever relations we find tend to be very time-specific. This is something I'm currently struggling with - the more data you use, the more your model can generalize, but also, the less it can fit to the patterns that pop up and disappear in specific time frames. Is it better to use a generalist model, that fits to all data, but not so well, or a short-term model with frequent retrains, that fits very well to specific time-frames, but tends to overfit? Maybe the answer is a combination of these two.",
      "votes": 10,
      "replies": [
        {
          "id": 3058499,
          "postDate": "2024-11-29T13:59:57.483Z",
          "content": "<p>Have you tried time decaying the datapoints, so that the model can both see more data but learn to prioritize more recent samples? </p>",
          "rawMarkdown": "Have you tried time decaying the datapoints, so that the model can both see more data but learn to prioritize more recent samples? ",
          "replies": [
            {
              "id": 3058511,
              "postDate": "2024-11-29T14:29:01.027Z",
              "content": "<p>I have in fact, yes. Not with weights but with probability sampling. And it didn't quite work. This has been some time ago so maybe I should try it again but I'm not very optimistic about it.</p>",
              "rawMarkdown": "I have in fact, yes. Not with weights but with probability sampling. And it didn't quite work. This has been some time ago so maybe I should try it again but I'm not very optimistic about it."
            },
            {
              "id": 3058513,
              "postDate": "2024-11-29T14:31:25.880Z",
              "content": "<p>I came to the same conclusion, didn't quite work for me either. Might need to play around with a self supervised approach here. There is too much to try and not enough time. </p>",
              "rawMarkdown": "I came to the same conclusion, didn't quite work for me either. Might need to play around with a self supervised approach here. There is too much to try and not enough time. ",
              "votes": 2
            }
          ]
        }
      ]
    },
    {
      "id": 3057279,
      "postDate": "2024-11-28T01:04:26.650Z",
      "content": "<p>Great insight! Literally just created an embedding layer this morning 😆 When new symbols appear, they will start with the random initialization. And if I can figure out how to do online learning, then those embeddings will evolve. I set the number of symbols to 200 to play it safe, but suspect we will see much less from other comments I saw.</p>",
      "rawMarkdown": "Great insight! Literally just created an embedding layer this morning 😆 When new symbols appear, they will start with the random initialization. And if I can figure out how to do online learning, then those embeddings will evolve. I set the number of symbols to 200 to play it safe, but suspect we will see much less from other comments I saw.",
      "votes": 5,
      "replies": [
        {
          "id": 3057295,
          "postDate": "2024-11-28T01:52:51.093Z",
          "content": "<p>Nicely done, I think <code>symbol_id</code> is a very underestimated feature, at least from the public notebooks/discussions I have seen thus far. I can think of at least 5 different ways to use it in training, without directly training on the column itself. An embedding layer is a great start. </p>",
          "rawMarkdown": "Nicely done, I think `symbol_id` is a very underestimated feature, at least from the public notebooks/discussions I have seen thus far. I can think of at least 5 different ways to use it in training, without directly training on the column itself. An embedding layer is a great start. ",
          "votes": 3,
          "replies": [
            {
              "id": 3057324,
              "postDate": "2024-11-28T03:10:58.363Z",
              "content": "<p>I thought so too but Im a newbie and this is my first finance dataset. In fact, Im trying to use this comp to learn some basic stuff and how finance data and models are.</p>",
              "rawMarkdown": "I thought so too but Im a newbie and this is my first finance dataset. In fact, Im trying to use this comp to learn some basic stuff and how finance data and models are."
            },
            {
              "id": 3057339,
              "postDate": "2024-11-28T03:30:28.830Z",
              "content": "<p>This is one of the more advanced competitions I have joined on Kaggle, but I took a 4 year break so perhaps I am just behind. It seems you are doing well based on current LB score. If you can \"hold your own\" in this competition, I'd say you are a fast learner and will not be a newbie for long. Good luck. </p>",
              "rawMarkdown": "This is one of the more advanced competitions I have joined on Kaggle, but I took a 4 year break so perhaps I am just behind. It seems you are doing well based on current LB score. If you can \"hold your own\" in this competition, I'd say you are a fast learner and will not be a newbie for long. Good luck. ",
              "votes": 2
            }
          ]
        }
      ]
    },
    {
      "id": 3083835,
      "postDate": "2024-12-30T05:02:13.547Z",
      "content": "<p>It seems the most voted kernels are based on forward filling and zero filling so far. Do you think other imputation methods are better?</p>",
      "rawMarkdown": "It seems the most voted kernels are based on forward filling and zero filling so far. Do you think other imputation methods are better?"
    },
    {
      "id": 3068327,
      "postDate": "2024-12-10T06:43:02.843Z",
      "content": "<p>This is really helpful</p>",
      "rawMarkdown": "This is really helpful"
    },
    {
      "id": 3062732,
      "postDate": "2024-12-03T21:21:45.883Z",
      "content": "<p><a href=\"https://www.kaggle.com/tuckerarrants\" target=\"_blank\">@tuckerarrants</a> thanks for sharing your great thoughts in the post ! <br>\nJust one comment: how will hypothetical HFT nature of train_data logic match with  \"The goal of the competition is to forecast one of these responders, i.e., responder_6,** for up to six months **in the future\"? 6M are very far from being HF.</p>\n<p>And generally  guys, what is your understanding of how on inference to forecast  for up to 6M time-period given that timestamp granularity of training data is not fully disclosed ? ..First, making an aggregation/symbol by symbol etc..(Maybe I am missing smth)</p>",
      "rawMarkdown": "@tuckerarrants thanks for sharing your great thoughts in the post ! \nJust one comment: how will hypothetical HFT nature of train_data logic match with  \"The goal of the competition is to forecast one of these responders, i.e., responder_6,** for up to six months **in the future\"? 6M are very far from being HF.\n\nAnd generally  guys, what is your understanding of how on inference to forecast  for up to 6M time-period given that timestamp granularity of training data is not fully disclosed ? ..First, making an aggregation/symbol by symbol etc..(Maybe I am missing smth)",
      "replies": [
        {
          "id": 3062745,
          "postDate": "2024-12-03T21:44:35.180Z",
          "content": "<p>Just like any other trading strategy. HFT trading just means the trade is entered and exited quickly. That can go on for as long into the future as you like. (E.g. 1000 trades a day, for 6 months into the future.) </p>\n<p>It may be counterintuitive, but it is easier to forecast a HFT strategy into the future than a LFT one, since HFT strategies do not rely as heavily on lagged signals whereas LFT strategies solely rely on lagged signals.</p>",
          "rawMarkdown": "Just like any other trading strategy. HFT trading just means the trade is entered and exited quickly. That can go on for as long into the future as you like. (E.g. 1000 trades a day, for 6 months into the future.) \n\nIt may be counterintuitive, but it is easier to forecast a HFT strategy into the future than a LFT one, since HFT strategies do not rely as heavily on lagged signals whereas LFT strategies solely rely on lagged signals.",
          "votes": 3
        }
      ]
    },
    {
      "id": 3059667,
      "postDate": "2024-12-01T00:55:22.283Z",
      "content": "<p>Do you have any hypothesis on what market this data could be from?</p>",
      "rawMarkdown": "Do you have any hypothesis on what market this data could be from?",
      "replies": [
        {
          "id": 3059674,
          "postDate": "2024-12-01T01:06:44.843Z",
          "content": "<p>If I had to guess, they are from futures markets, not spot markets. Likely the most liquid futures tickers, so US treasury futures, US equity futures, crude oil/natural gas futures, gold futures, and maybe some forex futures, like the Euro and the British Pound.</p>",
          "rawMarkdown": "If I had to guess, they are from futures markets, not spot markets. Likely the most liquid futures tickers, so US treasury futures, US equity futures, crude oil/natural gas futures, gold futures, and maybe some forex futures, like the Euro and the British Pound.",
          "votes": 4
        }
      ]
    },
    {
      "id": 3057326,
      "postDate": "2024-11-28T03:16:33.357Z",
      "content": "<blockquote>\n  <p>We also need to help our model generalize to new symbol_ids, as this is likely what Jane Street wants to do. We cannot simply train models on the raw symbol_id, or it will not generalize well to new ones. A simple way of teaching a model to learn symbol_ids without it overfitting, is with an embedding layer.</p>\n</blockquote>\n<p>How would the embedding layer deal with unseen data?</p>",
      "rawMarkdown": ">We also need to help our model generalize to new symbol_ids, as this is likely what Jane Street wants to do. We cannot simply train models on the raw symbol_id, or it will not generalize well to new ones. A simple way of teaching a model to learn symbol_ids without it overfitting, is with an embedding layer.\n\n\nHow would the embedding layer deal with unseen data?",
      "replies": [
        {
          "id": 3057333,
          "postDate": "2024-11-28T03:27:27.353Z",
          "content": "<p>A crude implementation would be to use the mean of seen embedding values as a fallback for unseen <code>symbol id</code> embeddings. This would allow the model to perform well on seen <code>symbol_ids</code>, while reducing the risk of overfitting. There will still be some overfitting, but it is worth shot. You can \"probe\" the leaderboard to see how this compares to no <code>symbol_ids</code> in training at all. </p>\n<p>A more refined implementation would be to occasionally retrain the embeddings with new <code>symbol_ids</code>, also called online learning. </p>\n<p>An intermediate implementation would be to group unseen <code>symbol_ids</code> based on their similarity (<code>features</code> columns) to known ones and assigning them existing embeddings or surrogate representations.</p>",
          "rawMarkdown": "A crude implementation would be to use the mean of seen embedding values as a fallback for unseen `symbol id` embeddings. This would allow the model to perform well on seen `symbol_ids`, while reducing the risk of overfitting. There will still be some overfitting, but it is worth shot. You can \"probe\" the leaderboard to see how this compares to no `symbol_ids` in training at all. \n\nA more refined implementation would be to occasionally retrain the embeddings with new `symbol_ids`, also called online learning. \n\nAn intermediate implementation would be to group unseen `symbol_ids` based on their similarity (`features` columns) to known ones and assigning them existing embeddings or surrogate representations.",
          "votes": 5,
          "replies": [
            {
              "id": 3057358,
              "postDate": "2024-11-28T04:02:08.893Z",
              "content": "<p>Great. Thanks</p>",
              "rawMarkdown": "Great. Thanks"
            }
          ]
        }
      ]
    },
    {
      "id": 3057435,
      "postDate": "2024-11-28T06:26:20.610Z",
      "rawMarkdown": "",
      "isDeleted": true
    }
  ],
  "comments": [
    {
      "id": 3058310,
      "author_name": "Natan Labarrère",
      "author_url": "",
      "post_date": "2024-11-29T10:14:59.813000",
      "content": "<p>Great thoughts. The responders have time and symbol correlations. That is, the responders are correlated to themselves in the past, and correlated between symbols. Whatever implementation that considers only the features to get to the predictions is missing on important information (all public codes shared have this issue).<br>\nI'm not so sure about using the first trades, however. As another kaggler pointed out in a great thread, whatever relations we find tend to be very time-specific. This is something I'm currently struggling with - the more data you use, the more your model can generalize, but also, the less it can fit to the patterns that pop up and disappear in specific time frames. Is it better to use a generalist model, that fits to all data, but not so well, or a short-term model with frequent retrains, that fits very well to specific time-frames, but tends to overfit? Maybe the answer is a combination of these two.</p>",
      "votes": 10,
      "replies": [
        {
          "id": 3058499,
          "author_name": "Tucker Arrants",
          "author_url": "",
          "post_date": "2024-11-29T13:59:57.483000",
          "content": "<p>Have you tried time decaying the datapoints, so that the model can both see more data but learn to prioritize more recent samples? </p>",
          "votes": 0,
          "replies": [
            {
              "id": 3058511,
              "author_name": "Natan Labarrère",
              "author_url": "",
              "post_date": "2024-11-29T14:29:01.027000",
              "content": "<p>I have in fact, yes. Not with weights but with probability sampling. And it didn't quite work. This has been some time ago so maybe I should try it again but I'm not very optimistic about it.</p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 3058513,
              "author_name": "Tucker Arrants",
              "author_url": "",
              "post_date": "2024-11-29T14:31:25.880000",
              "content": "<p>I came to the same conclusion, didn't quite work for me either. Might need to play around with a self supervised approach here. There is too much to try and not enough time. </p>",
              "votes": 2,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 3057279,
      "author_name": "Jim Beno",
      "author_url": "",
      "post_date": "2024-11-28T01:04:26.650000",
      "content": "<p>Great insight! Literally just created an embedding layer this morning 😆 When new symbols appear, they will start with the random initialization. And if I can figure out how to do online learning, then those embeddings will evolve. I set the number of symbols to 200 to play it safe, but suspect we will see much less from other comments I saw.</p>",
      "votes": 5,
      "replies": [
        {
          "id": 3057295,
          "author_name": "Tucker Arrants",
          "author_url": "",
          "post_date": "2024-11-28T01:52:51.093000",
          "content": "<p>Nicely done, I think <code>symbol_id</code> is a very underestimated feature, at least from the public notebooks/discussions I have seen thus far. I can think of at least 5 different ways to use it in training, without directly training on the column itself. An embedding layer is a great start. </p>",
          "votes": 3,
          "replies": [
            {
              "id": 3057324,
              "author_name": "Seqaeon",
              "author_url": "",
              "post_date": "2024-11-28T03:10:58.363000",
              "content": "<p>I thought so too but Im a newbie and this is my first finance dataset. In fact, Im trying to use this comp to learn some basic stuff and how finance data and models are.</p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 3057339,
              "author_name": "Tucker Arrants",
              "author_url": "",
              "post_date": "2024-11-28T03:30:28.830000",
              "content": "<p>This is one of the more advanced competitions I have joined on Kaggle, but I took a 4 year break so perhaps I am just behind. It seems you are doing well based on current LB score. If you can \"hold your own\" in this competition, I'd say you are a fast learner and will not be a newbie for long. Good luck. </p>",
              "votes": 2,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 3083835,
      "author_name": "Ak11",
      "author_url": "",
      "post_date": "2024-12-30T05:02:13.547000",
      "content": "<p>It seems the most voted kernels are based on forward filling and zero filling so far. Do you think other imputation methods are better?</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 3068327,
      "author_name": "Anna Sai Nikhil",
      "author_url": "",
      "post_date": "2024-12-10T06:43:02.843000",
      "content": "<p>This is really helpful</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 3062732,
      "author_name": "denis lukyanov",
      "author_url": "",
      "post_date": "2024-12-03T21:21:45.883000",
      "content": "<p><a href=\"https://www.kaggle.com/tuckerarrants\" target=\"_blank\">@tuckerarrants</a> thanks for sharing your great thoughts in the post ! <br>\nJust one comment: how will hypothetical HFT nature of train_data logic match with  \"The goal of the competition is to forecast one of these responders, i.e., responder_6,** for up to six months **in the future\"? 6M are very far from being HF.</p>\n<p>And generally  guys, what is your understanding of how on inference to forecast  for up to 6M time-period given that timestamp granularity of training data is not fully disclosed ? ..First, making an aggregation/symbol by symbol etc..(Maybe I am missing smth)</p>",
      "votes": 0,
      "replies": [
        {
          "id": 3062745,
          "author_name": "Tucker Arrants",
          "author_url": "",
          "post_date": "2024-12-03T21:44:35.180000",
          "content": "<p>Just like any other trading strategy. HFT trading just means the trade is entered and exited quickly. That can go on for as long into the future as you like. (E.g. 1000 trades a day, for 6 months into the future.) </p>\n<p>It may be counterintuitive, but it is easier to forecast a HFT strategy into the future than a LFT one, since HFT strategies do not rely as heavily on lagged signals whereas LFT strategies solely rely on lagged signals.</p>",
          "votes": 3,
          "replies": []
        }
      ]
    },
    {
      "id": 3059667,
      "author_name": "Lu Bin Liu",
      "author_url": "",
      "post_date": "2024-12-01T00:55:22.283000",
      "content": "<p>Do you have any hypothesis on what market this data could be from?</p>",
      "votes": 0,
      "replies": [
        {
          "id": 3059674,
          "author_name": "Tucker Arrants",
          "author_url": "",
          "post_date": "2024-12-01T01:06:44.843000",
          "content": "<p>If I had to guess, they are from futures markets, not spot markets. Likely the most liquid futures tickers, so US treasury futures, US equity futures, crude oil/natural gas futures, gold futures, and maybe some forex futures, like the Euro and the British Pound.</p>",
          "votes": 4,
          "replies": []
        }
      ]
    },
    {
      "id": 3057326,
      "author_name": "Seqaeon",
      "author_url": "",
      "post_date": "2024-11-28T03:16:33.357000",
      "content": "<blockquote>\n  <p>We also need to help our model generalize to new symbol_ids, as this is likely what Jane Street wants to do. We cannot simply train models on the raw symbol_id, or it will not generalize well to new ones. A simple way of teaching a model to learn symbol_ids without it overfitting, is with an embedding layer.</p>\n</blockquote>\n<p>How would the embedding layer deal with unseen data?</p>",
      "votes": 0,
      "replies": [
        {
          "id": 3057333,
          "author_name": "Tucker Arrants",
          "author_url": "",
          "post_date": "2024-11-28T03:27:27.353000",
          "content": "<p>A crude implementation would be to use the mean of seen embedding values as a fallback for unseen <code>symbol id</code> embeddings. This would allow the model to perform well on seen <code>symbol_ids</code>, while reducing the risk of overfitting. There will still be some overfitting, but it is worth shot. You can \"probe\" the leaderboard to see how this compares to no <code>symbol_ids</code> in training at all. </p>\n<p>A more refined implementation would be to occasionally retrain the embeddings with new <code>symbol_ids</code>, also called online learning. </p>\n<p>An intermediate implementation would be to group unseen <code>symbol_ids</code> based on their similarity (<code>features</code> columns) to known ones and assigning them existing embeddings or surrogate representations.</p>",
          "votes": 5,
          "replies": [
            {
              "id": 3057358,
              "author_name": "Seqaeon",
              "author_url": "",
              "post_date": "2024-11-28T04:02:08.893000",
              "content": "<p>Great. Thanks</p>",
              "votes": 0,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 3057435,
      "author_name": "",
      "author_url": "",
      "post_date": "2024-11-28T06:26:20.610000",
      "content": "",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "3057276": "I have some domain knowledge for this competition, as I am a full-time quant trader. I would like to share some of my thoughts. I joined this competition a few days ago, so I apologize for the later discussion post. \n\nFirstly, who is Jane Street? As per their website, they are \"a quantitative trading firm and liquidity provider with a unique focus on technology and collaborative problem solving.\" Liquidity providers use sophisticated high-speed computers and algorithms to create volume on exchanges in order to add liquidity to the markets.  \n\nIt is very likely that the data we are training on is data about these high-speed algorithms and the trades they execute. Algorithmic trading bots execute on rule-based strategies. I believe the features metadata represents these particular rules, and the feature columns are certain indicators / signal values at the time of execution.\n\nAssuming this is true, we can guess why the data is so seemingly inconsistent: different `time_ids` per `date_id`, different `symbol_ids` per `date_id`, the missing values before day 500, and the increasing number of `time_ids` per `date_id`, as `date_id` increases. I do not believe that these are lagging indicators, as I have seen some suggest. This is primarily because financial markets are extremely efficient on small time frames, meaning there is no extra \"alpha\" to gain by considering lagging features. Perhaps some are, but to me, it makes more sense that these are the byproduct of several different algo trading at once. \n\nPerhaps the initial algorithms relied on different indicators and signals, or were trading less frequently. As Jane Street created more sophisticated algorithms, it seems like they \"consolidated\" them over time, deploying them on new `symbol_ids`. This is why the data is more structured as `date_id` increases. \n\nSo what do they hope to learn from us Kagglers. I can assure you these algorithmic bots are already highly profitable. They are not asking us to predict something simple like the \"price of the S&P500 in 3 months.\" They want us to project the results of their existing algorithmic bots into the future, to help with risk-management, which is crucial for liquidity providers. \n\nFurthermore, it seems safe to assume that `time_id` per `date_id` will be roughly constant in the public test dataset and the private test dataset, which is much different from the train dataset. It feels like the training dataset represents a period of improving their algorithms, and now that they have, they are interested in not only forecasting their results into the future, but also on deploying them on new financial instruments or `symbol_ids`. This assumption will impact modeling, especially if you are using a sequential approach. \n\nBut ideally, they can learn something from their old algo trades as well, meaning an auxiliary ideal outcome of this competition is finding a way of using these first algo trades in their current algos. I think they would be \"disappointed\" if the winning solution just dropped them in their pipeline. As such, it should be a priority for us as well. Dropping, filling with 0, or forward filling, does not feel right here, no matter the impact on model performance. It may be worth using more sophisticated means of handling the missing data, like clustering. \n\nWe also need to help our model generalize to new `symbol_ids`, as this is likely what Jane Street wants to do. We cannot simply train models on the raw `symbol_id`, or it will not generalize well to new ones. A simple way of teaching a model to learn `symbol_ids` without it overfitting, is with an embedding layer. \n\nI hope this helps, please let me know your thoughts below. With the above assumptions, I am focusing on validating with more recent samples, as they likely reflect the most recent trading algos, which are the ones used on the private and public test set. Furthermore, I am prioritizing `symbol_id` and assuming a fairly consistent number of `time_ids` per `date_id` in the test datasets. Below are some graphs of my current exploration into the above thesis:\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4379159%2F218f8242c901e566a0a073ba427860c7%2FScreenshot%202024-11-30%20235444.png?generation=1733028902671124&alt=media)\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4379159%2F0f1c7b71aac780d6454ce1b71a745c3c%2FScreenshot%202024-11-30%20235917.png?generation=1733029177791681&alt=media)",
    "3058310": "Great thoughts. The responders have time and symbol correlations. That is, the responders are correlated to themselves in the past, and correlated between symbols. Whatever implementation that considers only the features to get to the predictions is missing on important information (all public codes shared have this issue).\nI'm not so sure about using the first trades, however. As another kaggler pointed out in a great thread, whatever relations we find tend to be very time-specific. This is something I'm currently struggling with - the more data you use, the more your model can generalize, but also, the less it can fit to the patterns that pop up and disappear in specific time frames. Is it better to use a generalist model, that fits to all data, but not so well, or a short-term model with frequent retrains, that fits very well to specific time-frames, but tends to overfit? Maybe the answer is a combination of these two.",
    "3057279": "Great insight! Literally just created an embedding layer this morning 😆 When new symbols appear, they will start with the random initialization. And if I can figure out how to do online learning, then those embeddings will evolve. I set the number of symbols to 200 to play it safe, but suspect we will see much less from other comments I saw.",
    "3083835": "It seems the most voted kernels are based on forward filling and zero filling so far. Do you think other imputation methods are better?",
    "3068327": "This is really helpful",
    "3062732": "@tuckerarrants thanks for sharing your great thoughts in the post ! \nJust one comment: how will hypothetical HFT nature of train_data logic match with  \"The goal of the competition is to forecast one of these responders, i.e., responder_6,** for up to six months **in the future\"? 6M are very far from being HF.\n\nAnd generally  guys, what is your understanding of how on inference to forecast  for up to 6M time-period given that timestamp granularity of training data is not fully disclosed ? ..First, making an aggregation/symbol by symbol etc..(Maybe I am missing smth)",
    "3059667": "Do you have any hypothesis on what market this data could be from?",
    "3057326": ">We also need to help our model generalize to new symbol_ids, as this is likely what Jane Street wants to do. We cannot simply train models on the raw symbol_id, or it will not generalize well to new ones. A simple way of teaching a model to learn symbol_ids without it overfitting, is with an embedding layer.\n\n\nHow would the embedding layer deal with unseen data?",
    "3057435": ""
  }
}