{
  "id": 555562,
  "title": "Reverse Engineering the Responders",
  "url": "/competitions/jane-street-real-time-market-data-forecasting/discussion/555562",
  "author_name": "John Payne",
  "post_date": "2025-01-08T05:26:56.382000",
  "votes": 165,
  "comment_count": 42,
  "views": 0,
  "content": "<p>Hi everyone, I joined the competition a couple weeks ago and since then I have spent a good portion of that time studying the responders. In hindsight, I definitely spent too much time studying the data instead of developing my model, but regardless I've decided to reveal the insights I've found and a potential explanation of what the responders could be. This post will walk through roughly the process I went through to get to my findings.</p>\n<p><strong>1.  Autocorrelation Function (ACF)</strong></p>\n<p>One of the first things I noticed with the responders is their very peculiar autocorrelation functions. For example, here's a graph of responder 6's ACF for symbol 1 across the entire dataset:</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F20686390%2F6d2945b91023f6a7a453e6aa1cbd5954%2Foutput.png?generation=1736304198418982&amp;alt=media\" alt=\"\"></p>\n<p>Notably, the autocorrelation appears to decrease linearly with each lag but have a sharp cutoff in autocorrelation at exactly lag 20. If we look at the ACF of responders 7 and 8, they behave similarly but cutoff at lags 120 and 4, respectively:</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F20686390%2F077216e8d8fc58407ffeddd1ef5e132a%2Foutput1.png?generation=1736305724902604&amp;alt=media\" alt=\"\"></p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F20686390%2F91b61be85c11463de1c07471dbfe828b%2Foutput2.png?generation=1736305733028007&amp;alt=media\" alt=\"\"></p>\n<p>The ACFs of the rest of the responders are slightly different, but they all share the same property of having a sharp change in autocorrelation at specific lags. Interestingly, there are similarities of the ACF between different responders, with responders 0,3,6 having the same ACF cutoff at lag 20, responders 1,4,7 having it at lag 120, and responders 2,5,8 having it lag 4.</p>\n<p><strong>2. Responder Relationships</strong></p>\n<p>To make matters more interesting, I found that the features at time t are weirdly good at predicting responders 0,3,6 at time t-20. (You can try this yourself, using linear regression on the features I was able to get a consistent R^2 of around 0.5 on responder 6 lag 20). This property holds on the other responder groups too with their respective lags. This led me to hypothesize that the responders we're given are actually shifted in time, i.e, we're actually trying to predict responder 6 at time t+20 given the features at time t. We can see evidence of this when we graph the lagged responders together:</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F20686390%2Fd3d1f158b2986dc02bde34730452954e%2Foutput3.png?generation=1736308270543101&amp;alt=media\" alt=\"\"></p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F20686390%2F4dfec682d93e8a655ab6078b8dc079b7%2Foutput4.png?generation=1736308288491253&amp;alt=media\" alt=\"\"></p>\n<p>Responders 6,7,8 correlate <em>much</em> more nicely at their lags as compared to the normal values. In fact, if you look a little bit closer at these graphs, it almost seems like responder 6 lag 20 is a \"derivative\" of responder 7 lag 120, and responder 8 lag 4 is a \"derivative\" of responder 6 lag 20. Whenever responder 6 is positive, it seems responder 7 moves upwards, and whenever responder 8 is positive, responder 6 moves upwards. Evidence of this is very strong when we graph moving averages of the lagged responders:</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F20686390%2F033442ee8ccbaa69a134f72e62ee73f8%2Foutput5.png?generation=1736309168179946&amp;alt=media\" alt=\"\"></p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F20686390%2Fae2402a655b00fe270c7e9bf69762102%2Foutput6.png?generation=1736309176377111&amp;alt=media\" alt=\"\"></p>\n<p>Clearly there's a strong relationship between responders 6,7,8.  Additionally, responders 3,4,5 seem to be very similar to 6,7,8 and therefore also share this same relationship. The last group of responders, 0,1,2 are quite different from the rest of the responders, but after studying some of the graphs it appears that responders 0,1,2 are roughly the difference between responders 3,4,5 and 6,7,8:</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F20686390%2F7c62fafd2d9986d3f42a9824570df601%2Foutput7.png?generation=1736310301191748&amp;alt=media\" alt=\"\"></p>\n<p><strong>3. Fourier Analysis</strong></p>\n<p>The last bit of analysis I tried was Fourier analysis on responder 6. Here's what the Fourier transform looks like:</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F20686390%2F18494492bc06afe8e70c38a31eef0f2d%2Foutput8.png?generation=1736310892074227&amp;alt=media\" alt=\"\"></p>\n<p>Notably, there's a distinct pattern in the transform with interval 1/20. In fact it's pretty clear to see that a curve of sin(20pi x)/(20pi x) roughly fits the transform:</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F20686390%2F3a081220abbbb8ff6e629afe61f56c21%2Foutput9.png?generation=1736311350550145&amp;alt=media\" alt=\"\"></p>\n<p>It looks as though the signal was originally white noise, but had been multiplied by a sin(x)/x curve in the frequency domain (which translates to a convolution in the time domain). Since the inverse Fourier transform of sin(x)/x is a box function, that implies that the original signal was convolved with a uniformly weighted kernel, i.e. a moving average. To test this idea, I generated a white noise signal, applied a moving average of window size 20 to it, and plotted the Fourier transform. And sure enough, we find a very similar graph:</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F20686390%2F81e9a3fc373b8bb6d2b4ec23fa259360%2Foutput10.png?generation=1736312764974135&amp;alt=media\" alt=\"\"></p>\n<p>The main difference being that it has sharper troughs than responder 6's transform. However, by adding another gaussian white noise variable to our generated white noise, we see that it matches almost perfectly to responder 6's transform:</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F20686390%2Fdabece1d83100cd18e42e42c4ce1c554%2Foutput12.png?generation=1736312811480746&amp;alt=media\" alt=\"\"></p>\n<p>Which implies that responder 6 is a moving average of width 20 over gaussian white noise.</p>\n<p><strong>What does this mean?</strong></p>\n<p>Given all of these findings, my best theory is that responders 6, 7, 8 are moving averages with widths 20, 120, 4 of the same underlying signal, but have a slight added noise to them (either intentionally for obfuscation or some sort of measurement error). The underlying signal is gaussian (but also has fat tails), which leads me to believe they are minute-by-minute returns of the symbol. Since responders 3, 4, 5 are so similar to 6, 7, 8, I think they are returns on a separate exchange, and responders 0, 1, 2 are the difference in returns. Since the responders given appear shifted in time, it seems the goal of this competition is to come up with a model that predicts the average returns over the next 20 minutes (which would explain why everyone's score is so \"low\", it is notoriously difficult to predict future returns from present data).</p>\n<p>I will admit that I lack a lot of domain knowledge and this is just a guess, so if anyone has different ideas, please feel free to propose.</p>",
  "messages": [
    {
      "id": 3091145,
      "postDate": "2025-01-08T05:26:56.383Z",
      "content": "<p>Hi everyone, I joined the competition a couple weeks ago and since then I have spent a good portion of that time studying the responders. In hindsight, I definitely spent too much time studying the data instead of developing my model, but regardless I've decided to reveal the insights I've found and a potential explanation of what the responders could be. This post will walk through roughly the process I went through to get to my findings.</p>\n<p><strong>1.  Autocorrelation Function (ACF)</strong></p>\n<p>One of the first things I noticed with the responders is their very peculiar autocorrelation functions. For example, here's a graph of responder 6's ACF for symbol 1 across the entire dataset:</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F20686390%2F6d2945b91023f6a7a453e6aa1cbd5954%2Foutput.png?generation=1736304198418982&amp;alt=media\" alt=\"\"></p>\n<p>Notably, the autocorrelation appears to decrease linearly with each lag but have a sharp cutoff in autocorrelation at exactly lag 20. If we look at the ACF of responders 7 and 8, they behave similarly but cutoff at lags 120 and 4, respectively:</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F20686390%2F077216e8d8fc58407ffeddd1ef5e132a%2Foutput1.png?generation=1736305724902604&amp;alt=media\" alt=\"\"></p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F20686390%2F91b61be85c11463de1c07471dbfe828b%2Foutput2.png?generation=1736305733028007&amp;alt=media\" alt=\"\"></p>\n<p>The ACFs of the rest of the responders are slightly different, but they all share the same property of having a sharp change in autocorrelation at specific lags. Interestingly, there are similarities of the ACF between different responders, with responders 0,3,6 having the same ACF cutoff at lag 20, responders 1,4,7 having it at lag 120, and responders 2,5,8 having it lag 4.</p>\n<p><strong>2. Responder Relationships</strong></p>\n<p>To make matters more interesting, I found that the features at time t are weirdly good at predicting responders 0,3,6 at time t-20. (You can try this yourself, using linear regression on the features I was able to get a consistent R^2 of around 0.5 on responder 6 lag 20). This property holds on the other responder groups too with their respective lags. This led me to hypothesize that the responders we're given are actually shifted in time, i.e, we're actually trying to predict responder 6 at time t+20 given the features at time t. We can see evidence of this when we graph the lagged responders together:</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F20686390%2Fd3d1f158b2986dc02bde34730452954e%2Foutput3.png?generation=1736308270543101&amp;alt=media\" alt=\"\"></p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F20686390%2F4dfec682d93e8a655ab6078b8dc079b7%2Foutput4.png?generation=1736308288491253&amp;alt=media\" alt=\"\"></p>\n<p>Responders 6,7,8 correlate <em>much</em> more nicely at their lags as compared to the normal values. In fact, if you look a little bit closer at these graphs, it almost seems like responder 6 lag 20 is a \"derivative\" of responder 7 lag 120, and responder 8 lag 4 is a \"derivative\" of responder 6 lag 20. Whenever responder 6 is positive, it seems responder 7 moves upwards, and whenever responder 8 is positive, responder 6 moves upwards. Evidence of this is very strong when we graph moving averages of the lagged responders:</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F20686390%2F033442ee8ccbaa69a134f72e62ee73f8%2Foutput5.png?generation=1736309168179946&amp;alt=media\" alt=\"\"></p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F20686390%2Fae2402a655b00fe270c7e9bf69762102%2Foutput6.png?generation=1736309176377111&amp;alt=media\" alt=\"\"></p>\n<p>Clearly there's a strong relationship between responders 6,7,8.  Additionally, responders 3,4,5 seem to be very similar to 6,7,8 and therefore also share this same relationship. The last group of responders, 0,1,2 are quite different from the rest of the responders, but after studying some of the graphs it appears that responders 0,1,2 are roughly the difference between responders 3,4,5 and 6,7,8:</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F20686390%2F7c62fafd2d9986d3f42a9824570df601%2Foutput7.png?generation=1736310301191748&amp;alt=media\" alt=\"\"></p>\n<p><strong>3. Fourier Analysis</strong></p>\n<p>The last bit of analysis I tried was Fourier analysis on responder 6. Here's what the Fourier transform looks like:</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F20686390%2F18494492bc06afe8e70c38a31eef0f2d%2Foutput8.png?generation=1736310892074227&amp;alt=media\" alt=\"\"></p>\n<p>Notably, there's a distinct pattern in the transform with interval 1/20. In fact it's pretty clear to see that a curve of sin(20pi x)/(20pi x) roughly fits the transform:</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F20686390%2F3a081220abbbb8ff6e629afe61f56c21%2Foutput9.png?generation=1736311350550145&amp;alt=media\" alt=\"\"></p>\n<p>It looks as though the signal was originally white noise, but had been multiplied by a sin(x)/x curve in the frequency domain (which translates to a convolution in the time domain). Since the inverse Fourier transform of sin(x)/x is a box function, that implies that the original signal was convolved with a uniformly weighted kernel, i.e. a moving average. To test this idea, I generated a white noise signal, applied a moving average of window size 20 to it, and plotted the Fourier transform. And sure enough, we find a very similar graph:</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F20686390%2F81e9a3fc373b8bb6d2b4ec23fa259360%2Foutput10.png?generation=1736312764974135&amp;alt=media\" alt=\"\"></p>\n<p>The main difference being that it has sharper troughs than responder 6's transform. However, by adding another gaussian white noise variable to our generated white noise, we see that it matches almost perfectly to responder 6's transform:</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F20686390%2Fdabece1d83100cd18e42e42c4ce1c554%2Foutput12.png?generation=1736312811480746&amp;alt=media\" alt=\"\"></p>\n<p>Which implies that responder 6 is a moving average of width 20 over gaussian white noise.</p>\n<p><strong>What does this mean?</strong></p>\n<p>Given all of these findings, my best theory is that responders 6, 7, 8 are moving averages with widths 20, 120, 4 of the same underlying signal, but have a slight added noise to them (either intentionally for obfuscation or some sort of measurement error). The underlying signal is gaussian (but also has fat tails), which leads me to believe they are minute-by-minute returns of the symbol. Since responders 3, 4, 5 are so similar to 6, 7, 8, I think they are returns on a separate exchange, and responders 0, 1, 2 are the difference in returns. Since the responders given appear shifted in time, it seems the goal of this competition is to come up with a model that predicts the average returns over the next 20 minutes (which would explain why everyone's score is so \"low\", it is notoriously difficult to predict future returns from present data).</p>\n<p>I will admit that I lack a lot of domain knowledge and this is just a guess, so if anyone has different ideas, please feel free to propose.</p>",
      "rawMarkdown": "Hi everyone, I joined the competition a couple weeks ago and since then I have spent a good portion of that time studying the responders. In hindsight, I definitely spent too much time studying the data instead of developing my model, but regardless I've decided to reveal the insights I've found and a potential explanation of what the responders could be. This post will walk through roughly the process I went through to get to my findings.\n\n**1.  Autocorrelation Function (ACF)**\n\nOne of the first things I noticed with the responders is their very peculiar autocorrelation functions. For example, here's a graph of responder 6's ACF for symbol 1 across the entire dataset:\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F20686390%2F6d2945b91023f6a7a453e6aa1cbd5954%2Foutput.png?generation=1736304198418982&alt=media)\n\nNotably, the autocorrelation appears to decrease linearly with each lag but have a sharp cutoff in autocorrelation at exactly lag 20. If we look at the ACF of responders 7 and 8, they behave similarly but cutoff at lags 120 and 4, respectively:\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F20686390%2F077216e8d8fc58407ffeddd1ef5e132a%2Foutput1.png?generation=1736305724902604&alt=media)\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F20686390%2F91b61be85c11463de1c07471dbfe828b%2Foutput2.png?generation=1736305733028007&alt=media)\n\nThe ACFs of the rest of the responders are slightly different, but they all share the same property of having a sharp change in autocorrelation at specific lags. Interestingly, there are similarities of the ACF between different responders, with responders 0,3,6 having the same ACF cutoff at lag 20, responders 1,4,7 having it at lag 120, and responders 2,5,8 having it lag 4.\n\n**2. Responder Relationships**\n\nTo make matters more interesting, I found that the features at time t are weirdly good at predicting responders 0,3,6 at time t-20. (You can try this yourself, using linear regression on the features I was able to get a consistent R^2 of around 0.5 on responder 6 lag 20). This property holds on the other responder groups too with their respective lags. This led me to hypothesize that the responders we're given are actually shifted in time, i.e, we're actually trying to predict responder 6 at time t+20 given the features at time t. We can see evidence of this when we graph the lagged responders together:\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F20686390%2Fd3d1f158b2986dc02bde34730452954e%2Foutput3.png?generation=1736308270543101&alt=media)\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F20686390%2F4dfec682d93e8a655ab6078b8dc079b7%2Foutput4.png?generation=1736308288491253&alt=media)\n\nResponders 6,7,8 correlate *much* more nicely at their lags as compared to the normal values. In fact, if you look a little bit closer at these graphs, it almost seems like responder 6 lag 20 is a \"derivative\" of responder 7 lag 120, and responder 8 lag 4 is a \"derivative\" of responder 6 lag 20. Whenever responder 6 is positive, it seems responder 7 moves upwards, and whenever responder 8 is positive, responder 6 moves upwards. Evidence of this is very strong when we graph moving averages of the lagged responders:\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F20686390%2F033442ee8ccbaa69a134f72e62ee73f8%2Foutput5.png?generation=1736309168179946&alt=media)\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F20686390%2Fae2402a655b00fe270c7e9bf69762102%2Foutput6.png?generation=1736309176377111&alt=media)\n\nClearly there's a strong relationship between responders 6,7,8.  Additionally, responders 3,4,5 seem to be very similar to 6,7,8 and therefore also share this same relationship. The last group of responders, 0,1,2 are quite different from the rest of the responders, but after studying some of the graphs it appears that responders 0,1,2 are roughly the difference between responders 3,4,5 and 6,7,8:\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F20686390%2F7c62fafd2d9986d3f42a9824570df601%2Foutput7.png?generation=1736310301191748&alt=media)\n\n**3. Fourier Analysis**\n\nThe last bit of analysis I tried was Fourier analysis on responder 6. Here's what the Fourier transform looks like:\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F20686390%2F18494492bc06afe8e70c38a31eef0f2d%2Foutput8.png?generation=1736310892074227&alt=media)\n\nNotably, there's a distinct pattern in the transform with interval 1/20. In fact it's pretty clear to see that a curve of sin(20pi x)/(20pi x) roughly fits the transform:\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F20686390%2F3a081220abbbb8ff6e629afe61f56c21%2Foutput9.png?generation=1736311350550145&alt=media)\n\nIt looks as though the signal was originally white noise, but had been multiplied by a sin(x)/x curve in the frequency domain (which translates to a convolution in the time domain). Since the inverse Fourier transform of sin(x)/x is a box function, that implies that the original signal was convolved with a uniformly weighted kernel, i.e. a moving average. To test this idea, I generated a white noise signal, applied a moving average of window size 20 to it, and plotted the Fourier transform. And sure enough, we find a very similar graph:\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F20686390%2F81e9a3fc373b8bb6d2b4ec23fa259360%2Foutput10.png?generation=1736312764974135&alt=media)\n\nThe main difference being that it has sharper troughs than responder 6's transform. However, by adding another gaussian white noise variable to our generated white noise, we see that it matches almost perfectly to responder 6's transform:\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F20686390%2Fdabece1d83100cd18e42e42c4ce1c554%2Foutput12.png?generation=1736312811480746&alt=media)\n\nWhich implies that responder 6 is a moving average of width 20 over gaussian white noise.\n\n**What does this mean?**\n\nGiven all of these findings, my best theory is that responders 6, 7, 8 are moving averages with widths 20, 120, 4 of the same underlying signal, but have a slight added noise to them (either intentionally for obfuscation or some sort of measurement error). The underlying signal is gaussian (but also has fat tails), which leads me to believe they are minute-by-minute returns of the symbol. Since responders 3, 4, 5 are so similar to 6, 7, 8, I think they are returns on a separate exchange, and responders 0, 1, 2 are the difference in returns. Since the responders given appear shifted in time, it seems the goal of this competition is to come up with a model that predicts the average returns over the next 20 minutes (which would explain why everyone's score is so \"low\", it is notoriously difficult to predict future returns from present data).\n\nI will admit that I lack a lot of domain knowledge and this is just a guess, so if anyone has different ideas, please feel free to propose.",
      "votes": 165
    },
    {
      "id": 3091148,
      "postDate": "2025-01-08T05:37:30.210Z",
      "content": "<p>I forgot to mention, but responders.csv seems to have further evidence of my findings:</p>\n<ul>\n<li>responders 0,1,2 have tag_0 (they're all differences between the other responders)</li>\n<li>responders 2,5,8 have tag_1 (SMA 4)</li>\n<li>responders 0,3,6 have tag_2 (SMA 20)</li>\n<li>responders 1,4,7 have tag_3 (SMA 120)</li>\n<li>responders 3,4,5 have tag_4 (they're all moving averages of the same signal)</li>\n</ul>\n<p>Interestingly though there's no tag_5 to group responders 6,7,8 together, not sure why.</p>",
      "rawMarkdown": "I forgot to mention, but responders.csv seems to have further evidence of my findings:\n- responders 0,1,2 have tag_0 (they're all differences between the other responders)\n- responders 2,5,8 have tag_1 (SMA 4)\n- responders 0,3,6 have tag_2 (SMA 20)\n- responders 1,4,7 have tag_3 (SMA 120)\n- responders 3,4,5 have tag_4 (they're all moving averages of the same signal)\n\nInterestingly though there's no tag_5 to group responders 6,7,8 together, not sure why.",
      "votes": 18,
      "replies": [
        {
          "id": 3091165,
          "postDate": "2025-01-08T06:17:35.007Z",
          "content": "<p>great analysis!</p>",
          "rawMarkdown": "great analysis!"
        },
        {
          "id": 3091289,
          "postDate": "2025-01-08T09:18:06.603Z",
          "content": "<p>Maybe there's no need for <code>tag_5</code> because that's the baseline/target, and all the others are either secondary source <code>tag_4</code> or differences between primary source and secondary source <code>tag_0</code>?</p>",
          "rawMarkdown": "Maybe there's no need for `tag_5` because that's the baseline/target, and all the others are either secondary source `tag_4` or differences between primary source and secondary source `tag_0`?"
        }
      ]
    },
    {
      "id": 3092013,
      "postDate": "2025-01-09T04:44:30.577Z",
      "content": "<p>Outstanding, very impressive analysis. Can you elaborate on this? \"The underlying signal is gaussian (but also has fat tails), which leads me to believe they are minute-by-minute returns of the symbol.\" I wasn't clear on the connection there.</p>",
      "rawMarkdown": "Outstanding, very impressive analysis. Can you elaborate on this? \"The underlying signal is gaussian (but also has fat tails), which leads me to believe they are minute-by-minute returns of the symbol.\" I wasn't clear on the connection there.",
      "votes": 3
    },
    {
      "id": 3091394,
      "postDate": "2025-01-08T11:29:38.723Z",
      "content": "<p>Your findings make we think we are looking at cross-exchange arbitrage opportunities/signals. </p>",
      "rawMarkdown": "Your findings make we think we are looking at cross-exchange arbitrage opportunities/signals. ",
      "votes": 3,
      "replies": [
        {
          "id": 3093450,
          "postDate": "2025-01-10T23:25:04.557Z",
          "content": "<p>Yepp i was Thinking the same, also the way, the Difference of Responder 0,1 and 2 are correlated makes one wonder.</p>",
          "rawMarkdown": "Yepp i was Thinking the same, also the way, the Difference of Responder 0,1 and 2 are correlated makes one wonder."
        }
      ]
    },
    {
      "id": 3092465,
      "postDate": "2025-01-09T15:20:06.740Z",
      "content": "<p>Beyond returns from another exchange, responders 3, 4, 5 might show ask price returns while 0, 3, 6 show bid price returns (or vice versa).</p>",
      "rawMarkdown": "Beyond returns from another exchange, responders 3, 4, 5 might show ask price returns while 0, 3, 6 show bid price returns (or vice versa).",
      "votes": 2,
      "replies": [
        {
          "id": 3092492,
          "postDate": "2025-01-09T15:50:46.410Z",
          "content": "<p>But then bid should never be higher than ask, while the responders' differences can be both positive or negative.</p>",
          "rawMarkdown": "But then bid should never be higher than ask, while the responders' differences can be both positive or negative.",
          "votes": 1
        }
      ]
    },
    {
      "id": 3091393,
      "postDate": "2025-01-08T11:28:02.303Z",
      "content": "<p>Curiously, sometimes the differences can move in an opposite direction from responders 0, 1, 2, or ever go haywire entirely.</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F10943849%2Ffcec62520367085678e28ba30a59904d%2Fvisualization(1).png?generation=1736335632042318&amp;alt=media\" alt=\"Difference between responder 3 and responder 6\"></p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F10943849%2Fa17e362bdd3c6d9f4198143714bc3b43%2Fvisualization(2).png?generation=1736335641095285&amp;alt=media\" alt=\"Difference between responder 4 and responder 7\"></p>",
      "rawMarkdown": "Curiously, sometimes the differences can move in an opposite direction from responders 0, 1, 2, or ever go haywire entirely.\n\n![Difference between responder 3 and responder 6](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F10943849%2Ffcec62520367085678e28ba30a59904d%2Fvisualization(1).png?generation=1736335632042318&alt=media)\n\n![Difference between responder 4 and responder 7](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F10943849%2Fa17e362bdd3c6d9f4198143714bc3b43%2Fvisualization(2).png?generation=1736335641095285&alt=media)",
      "votes": 2
    },
    {
      "id": 3092591,
      "postDate": "2025-01-09T18:24:50.763Z",
      "content": "<blockquote>\n  <p>which would explain why everyone's score is so \"low\", it is notoriously difficult to predict future returns from present data</p>\n</blockquote>\n<p>In the limit where it's infinitely difficult to predict future results from present data, the problem essentially becomes one where present data are no more useful than random coin flips for the purpose of prediction.  In this hypothetical scenario, big LB shakeups are fundamentally unavoidable.</p>\n<p>Even though \"notoriously difficult\" is not as bad as \"infinitely difficult\", I wonder if it's bad enough to cause (big) LB shakeups, esp since this competition will test on months of future data no one has seen yet.</p>",
      "rawMarkdown": ">which would explain why everyone's score is so \"low\", it is notoriously difficult to predict future returns from present data\n\nIn the limit where it's infinitely difficult to predict future results from present data, the problem essentially becomes one where present data are no more useful than random coin flips for the purpose of prediction.  In this hypothetical scenario, big LB shakeups are fundamentally unavoidable.\n\nEven though \"notoriously difficult\" is not as bad as \"infinitely difficult\", I wonder if it's bad enough to cause (big) LB shakeups, esp since this competition will test on months of future data no one has seen yet.",
      "votes": -1,
      "replies": [
        {
          "id": 3092622,
          "postDate": "2025-01-09T19:17:29.347Z",
          "content": "<p>Eventually Trump will dominate the results LOL</p>",
          "rawMarkdown": "Eventually Trump will dominate the results LOL",
          "votes": 4
        }
      ]
    },
    {
      "id": 3095673,
      "postDate": "2025-01-13T15:18:08.590Z",
      "content": "<p>This idea is perfect, but API only provide with only one day's data to predict, it's difficult to progress time series(although the idea perform good in training). Of course, it may be my thought mistake.</p>",
      "rawMarkdown": "This idea is perfect, but API only provide with only one day's data to predict, it's difficult to progress time series(although the idea perform good in training). Of course, it may be my thought mistake."
    },
    {
      "id": 3095299,
      "postDate": "2025-01-13T07:24:23.220Z",
      "content": "<p>Interesting. But I used to hear that technical analysis works for high frequency trading. Do they use high frequency features then？</p>",
      "rawMarkdown": "Interesting. But I used to hear that technical analysis works for high frequency trading. Do they use high frequency features then？"
    },
    {
      "id": 3093605,
      "postDate": "2025-01-11T07:40:51.817Z",
      "content": "<p>OMG if this is the case, what R2 score can you make off the 4min moving average? like, if thats predictable, it shows how predictable are the markets and potentially it would be immensely interesting</p>",
      "rawMarkdown": "OMG if this is the case, what R2 score can you make off the 4min moving average? like, if thats predictable, it shows how predictable are the markets and potentially it would be immensely interesting",
      "replies": [
        {
          "id": 3093894,
          "postDate": "2025-01-11T13:05:46.260Z",
          "content": "<p>I mean we could look at the 4min bar and see how easy it is to predict the market 😆</p>",
          "rawMarkdown": "I mean we could look at the 4min bar and see how easy it is to predict the market 😆"
        }
      ]
    },
    {
      "id": 3091700,
      "postDate": "2025-01-08T17:21:54.783Z",
      "content": "<p>Great analysis 🤩, do you think that by adding more lagged (-1, -5 etc…) responders as input features could allow to improve the score of a model? I used this approach some weeks ago but I ended up by strongly overfitting the train data.</p>",
      "rawMarkdown": "Great analysis 🤩, do you think that by adding more lagged (-1, -5 etc...) responders as input features could allow to improve the score of a model? I used this approach some weeks ago but I ended up by strongly overfitting the train data.\n",
      "replies": [
        {
          "id": 3091705,
          "postDate": "2025-01-08T17:29:39.297Z",
          "content": "<p>In the inference API, we don’t have responder labels per time_id, only the daily lags. Directly adding responder labels lagged by time_ids is therefore not feasible, is it?</p>",
          "rawMarkdown": "In the inference API, we don’t have responder labels per time_id, only the daily lags. Directly adding responder labels lagged by time_ids is therefore not feasible, is it?",
          "replies": [
            {
              "id": 3091712,
              "postDate": "2025-01-08T17:35:58.297Z",
              "content": "<p>Maybe I'm wrong but the lags passed at the start of each day are the responders for each time_id of the previous day.</p>",
              "rawMarkdown": "Maybe I'm wrong but the lags passed at the start of each day are the responders for each time_id of the previous day.",
              "votes": 2
            },
            {
              "id": 3091719,
              "postDate": "2025-01-08T17:41:46.580Z",
              "content": "<p>if you use time_id from the lags.parquet and shift it by n steps in the api, you are actually using a time_id shifted by 968+n time steps, right?</p>",
              "rawMarkdown": "if you use time_id from the lags.parquet and shift it by n steps in the api, you are actually using a time_id shifted by 968+n time steps, right?"
            },
            {
              "id": 3091741,
              "postDate": "2025-01-08T17:58:39.240Z",
              "content": "<p>I'm not following you here I'm sorry.<br>\nMy point basically was to use to do a prediction, given the responder value for a time Tx of day D,<br>\nThe responder for T_[last] of day D-1 and T_[last-1] of day D-1 ideally to help to capture some trend.</p>",
              "rawMarkdown": "I'm not following you here I'm sorry.\nMy point basically was to use to do a prediction, given the responder value for a time Tx of day D,\nThe responder for T_[last] of day D-1 and T_[last-1] of day D-1 ideally to help to capture some trend."
            },
            {
              "id": 3091760,
              "postDate": "2025-01-08T18:10:41.070Z",
              "content": "<p>Consider this: say we are at T=100 at day D. There is a gap of 100 time steps between T=100 and T_last of D-1. </p>\n<p>If we proceed to T=968, we would have a temporal gap of 968 steps between T and T_last of D-1.</p>\n<p>I don’t think any trend can pass over such gap…</p>",
              "rawMarkdown": "Consider this: say we are at T=100 at day D. There is a gap of 100 time steps between T=100 and T_last of D-1. \n\nIf we proceed to T=968, we would have a temporal gap of 968 steps between T and T_last of D-1.\n\nI don’t think any trend can pass over such gap…",
              "votes": 3
            },
            {
              "id": 3091770,
              "postDate": "2025-01-08T18:27:02.813Z",
              "content": "<p>Yes I know what you're saying  in the sense the \"weight\" of the lagged features will decrease with the time id. But if we go in this direction we should not use the lags at all. <br>\nMaybe more lags can help to capture cyclic trends.</p>",
              "rawMarkdown": "Yes I know what you're saying  in the sense the \"weight\" of the lagged features will decrease with the time id. But if we go in this direction we should not use the lags at all. \nMaybe more lags can help to capture cyclic trends."
            }
          ]
        }
      ]
    },
    {
      "id": 3091651,
      "postDate": "2025-01-08T16:02:19.630Z",
      "content": "<p>I’m a bit confused about a point you mentioned in the Responder Relationships section. You stated that the features at time t predict responder 6 at time t-20 quite well. I don’t fully understand this. For example, if t = 150, where X represents the features and y represents responder 6, does this mean X_{150} predicts y_{130} quite well? This seems strange to me because it appears to use future data to predict y, which would already be known at that time.</p>\n<p>By the way, great analysis and thanks for sharing!</p>",
      "rawMarkdown": "I’m a bit confused about a point you mentioned in the Responder Relationships section. You stated that the features at time t predict responder 6 at time t-20 quite well. I don’t fully understand this. For example, if t = 150, where X represents the features and y represents responder 6, does this mean X_{150} predicts y_{130} quite well? This seems strange to me because it appears to use future data to predict y, which would already be known at that time.\n\nBy the way, great analysis and thanks for sharing!",
      "replies": [
        {
          "id": 3091653,
          "postDate": "2025-01-08T16:07:00.243Z",
          "content": "<p>y_{130} is the actual target value corresponding to X_150; the r6 value at X_150 given in the dataset is acutally corresponds to y_170. In other words, the host may have done <code>.shift(-20)</code> when preparing the data.</p>",
          "rawMarkdown": "y_{130} is the actual target value corresponding to X_150; the r6 value at X_150 given in the dataset is acutally corresponds to y_170. In other words, the host may have done `.shift(-20)` when preparing the data.",
          "votes": 3,
          "replies": [
            {
              "id": 3091673,
              "postDate": "2025-01-08T16:41:48.463Z",
              "content": "<p>Thanks for the clarification! That makes sense now. Appreciate your help!</p>",
              "rawMarkdown": "Thanks for the clarification! That makes sense now. Appreciate your help!"
            }
          ]
        }
      ]
    },
    {
      "id": 3091363,
      "postDate": "2025-01-08T10:36:46.647Z",
      "content": "<p>Good analysis. Now to try and think how I can use this in my models </p>",
      "rawMarkdown": "Good analysis. Now to try and think how I can use this in my models "
    },
    {
      "id": 3091331,
      "postDate": "2025-01-08T10:01:42.457Z",
      "content": "<p>Amazing analysis! Thank you for sharing!</p>",
      "rawMarkdown": "Amazing analysis! Thank you for sharing!"
    },
    {
      "id": 3091264,
      "postDate": "2025-01-08T08:38:12.910Z",
      "content": "<p>Incrediable analysis!</p>",
      "rawMarkdown": "Incrediable analysis!"
    },
    {
      "id": 3091230,
      "postDate": "2025-01-08T07:38:01.260Z",
      "content": "<p>Thanks for sharing.</p>\n<p>If the target is in the future, then the \"pre-market\" / \"after-hours\" shared here \"https://www.kaggle.com/competitions/jane-street-real-time-market-data-forecasting/discussion/549223\" might not be correct? or the final 20 time_ids are removed.</p>",
      "rawMarkdown": "Thanks for sharing.\n\nIf the target is in the future, then the \"pre-market\" / \"after-hours\" shared here \"https://www.kaggle.com/competitions/jane-street-real-time-market-data-forecasting/discussion/549223\" might not be correct? or the final 20 time_ids are removed."
    },
    {
      "id": 3091823,
      "postDate": "2025-01-08T19:57:06.380Z",
      "rawMarkdown": "",
      "isDeleted": true
    },
    {
      "id": 3091813,
      "postDate": "2025-01-08T19:31:10.257Z",
      "content": "<p>I actually tried to leverage the oscillatory nature of responder 6 a few weeks before. But the problem is that we do not have the initial conditions. Lagged repsonder from previous day don't give us a good prediction for the next day. I have tried using the features to predict differenced responder 6, and the R squared can reach 0.17. The features are better at predicting the change , and we need the current value for predicting responder 6. I still got some ideas, but I don't think I have enough time to try those…</p>",
      "rawMarkdown": "I actually tried to leverage the oscillatory nature of responder 6 a few weeks before. But the problem is that we do not have the initial conditions. Lagged repsonder from previous day don't give us a good prediction for the next day. I have tried using the features to predict differenced responder 6, and the R squared can reach 0.17. The features are better at predicting the change , and we need the current value for predicting responder 6. I still got some ideas, but I don't think I have enough time to try those...",
      "votes": 1,
      "isDeleted": true,
      "replies": [
        {
          "id": 3091814,
          "postDate": "2025-01-08T19:33:28.243Z",
          "content": "<p>I have tried LSTM and CNN on time series, but they are not working yet.</p>",
          "rawMarkdown": "I have tried LSTM and CNN on time series, but they are not working yet.",
          "votes": 1,
          "isDeleted": true
        },
        {
          "id": 3091817,
          "postDate": "2025-01-08T19:41:04.083Z",
          "content": "<p>Even if you had a perfect estimation of the responder at lag 20, the signal is a moving average so the expected value at time t given the value at time t-20 is 0, so any sort of forecasting is pretty much a dead end.</p>",
          "rawMarkdown": "Even if you had a perfect estimation of the responder at lag 20, the signal is a moving average so the expected value at time t given the value at time t-20 is 0, so any sort of forecasting is pretty much a dead end.",
          "votes": 4,
          "replies": [
            {
              "id": 3091829,
              "postDate": "2025-01-08T20:03:15.123Z",
              "content": "<blockquote>\n  <p>the expected value at time t given the value at time t-20 is 0</p>\n</blockquote>\n<p>That’s exactly the point. I have experimented to predict r6 from T-20 towards T using features at T. The result confirms this. I got pretty good prediction for T-20, and at T the result has lost most of the variation and got “converged “ to zero.</p>\n<p>How would you suggest to model the series differently? </p>",
              "rawMarkdown": ">the expected value at time t given the value at time t-20 is 0\n\nThat’s exactly the point. I have experimented to predict r6 from T-20 towards T using features at T. The result confirms this. I got pretty good prediction for T-20, and at T the result has lost most of the variation and got “converged “ to zero.\n\nHow would you suggest to model the series differently? ",
              "votes": 2
            },
            {
              "id": 3091862,
              "postDate": "2025-01-08T20:55:40.607Z",
              "content": "<p>I'm not sure there's a better way to model the series other than:</p>\n<p>$$Y^6_t=\\alpha\\sum_{i=0}^{19}Z_{t-i}+\\epsilon_t$$</p>\n<p>Where Z_t and epsilon_t are i.i.d normal variables. With this model, knowing Y6 at time t-20 gives you zero information about Y6 at time t. I think this competition was deliberately designed so that we must predict Y6 by finding some sort of hidden dependence between Z_t and the features.</p>",
              "rawMarkdown": "I'm not sure there's a better way to model the series other than:\n\n$$Y^6_t=\\alpha\\sum_{i=0}^{19}Z_{t-i}+\\epsilon_t$$\n\nWhere Z_t and epsilon_t are i.i.d normal variables. With this model, knowing Y6 at time t-20 gives you zero information about Y6 at time t. I think this competition was deliberately designed so that we must predict Y6 by finding some sort of hidden dependence between Z_t and the features.",
              "votes": 1
            },
            {
              "id": 3092505,
              "postDate": "2025-01-09T16:14:44.773Z",
              "content": "<p>What about the following approach: first we predict forward features for <code>T~T+20</code> given features before T, next we predict <code>Y^6</code>  using the predicted forward features. In this way, we can design a model to predict forward features by leveraging both the temporal dependency of a single feature as well as the cross-channel dependency between features, which would be easier than predicting <code>Y^6</code> using 20 steps of lagged <code>Y^6</code> values (since returns are random walk).<br>\nI dont't think I have time to implement this, but I am curious if this could work, even just theoretically. </p>",
              "rawMarkdown": "What about the following approach: first we predict forward features for `T~T+20` given features before T, next we predict `Y^6`  using the predicted forward features. In this way, we can design a model to predict forward features by leveraging both the temporal dependency of a single feature as well as the cross-channel dependency between features, which would be easier than predicting `Y^6` using 20 steps of lagged `Y^6` values (since returns are random walk).\nI dont't think I have time to implement this, but I am curious if this could work, even just theoretically. ",
              "votes": 2
            },
            {
              "id": 3092571,
              "postDate": "2025-01-09T17:45:32.087Z",
              "content": "<p>I did think of doing this a while back, but I figured that predicting the relevant future features would end up being just as difficult as predicting Y6 itself, but I also don't have much time to test this idea.</p>\n<p>I can tell you though that I've spent hours analyzing and testing ideas using this information I found, and so far nothing has worked better than a simple NN from features to responders.</p>",
              "rawMarkdown": "I did think of doing this a while back, but I figured that predicting the relevant future features would end up being just as difficult as predicting Y6 itself, but I also don't have much time to test this idea.\n\nI can tell you though that I've spent hours analyzing and testing ideas using this information I found, and so far nothing has worked better than a simple NN from features to responders.\n\n",
              "votes": 1
            },
            {
              "id": 3092650,
              "postDate": "2025-01-09T20:02:22.207Z",
              "content": "<p>I don't think the key is about the temporal dependency of the features. It's the temporal dependency of the responder. I have experimented that the change in responder 6 depends on both its current value and also its momentum. </p>\n<p>Using differential equation as an analogy, the features only gives us the equation. We still need the initial value and first derivative of responder 6, which we have no way to get those.</p>\n<p>I have think of using the features to identify whether responder 6 has return to zero, since I think if the market doesn't have trading opportunity for a while, we shouldn't take any position and responder 6 should reflect that. If we can identify regions that are near 0, then we can use those as initial conditions. But nothing worked yet.</p>",
              "rawMarkdown": "I don't think the key is about the temporal dependency of the features. It's the temporal dependency of the responder. I have experimented that the change in responder 6 depends on both its current value and also its momentum. \n\nUsing differential equation as an analogy, the features only gives us the equation. We still need the initial value and first derivative of responder 6, which we have no way to get those.\n\nI have think of using the features to identify whether responder 6 has return to zero, since I think if the market doesn't have trading opportunity for a while, we shouldn't take any position and responder 6 should reflect that. If we can identify regions that are near 0, then we can use those as initial conditions. But nothing worked yet.",
              "isDeleted": true
            },
            {
              "id": 3092655,
              "postDate": "2025-01-09T20:07:37.493Z",
              "content": "<p>Let me give you an example. Let say now responder 6 equals 4 right now and it's in a down trend, even if the features consistently showing an up-signal, since the downward momentum is so strong, reponsder 6 still decreases.</p>",
              "rawMarkdown": "Let me give you an example. Let say now responder 6 equals 4 right now and it's in a down trend, even if the features consistently showing an up-signal, since the downward momentum is so strong, reponsder 6 still decreases.",
              "isDeleted": true
            },
            {
              "id": 3092711,
              "postDate": "2025-01-09T22:18:57.753Z",
              "content": "<p>Here's the conditional expectation of responder 6:</p>\n<p>$$E\\left[ Y^6_{t+s} \\mid Y^6_t \\right] =\\frac{20 - s}{20}Y^6_t \\,\\,\\,\\,\\,\\, \\text{for } s = 0, 1, \\dots, 19$$<br>\n$$E\\left[ Y^6_{t+s} \\mid Y^6_t \\right] =0 \\,\\,\\,\\,\\,\\, \\text{for } s \\geq 20$$</p>\n<p>This tells us that responder 6 is always expected to decrease towards 0, and if you're trying to forecast any more than 19 time ticks forward, then there is really no telling what responder 6 will be, all you know is that it will on average be 0 (with high variance).</p>\n<p>If you really believe that Y6 has a dependence on its momentum, then try and come up with a mathematical model for it. I personally think though that even if there is something there, NNs on the features would outperform anyway.</p>",
              "rawMarkdown": "Here's the conditional expectation of responder 6:\n\n$$E\\left[ Y^6_{t+s} \\mid Y^6_t \\right] =\\frac{20 - s}{20}Y^6_t \\,\\,\\,\\,\\,\\, \\text{for } s = 0, 1, \\dots, 19$$\n$$E\\left[ Y^6_{t+s} \\mid Y^6_t \\right] =0 \\,\\,\\,\\,\\,\\, \\text{for } s \\geq 20$$\n\nThis tells us that responder 6 is always expected to decrease towards 0, and if you're trying to forecast any more than 19 time ticks forward, then there is really no telling what responder 6 will be, all you know is that it will on average be 0 (with high variance).\n\nIf you really believe that Y6 has a dependence on its momentum, then try and come up with a mathematical model for it. I personally think though that even if there is something there, NNs on the features would outperform anyway.",
              "votes": 2
            }
          ]
        }
      ]
    },
    {
      "id": 3091210,
      "postDate": "2025-01-08T07:16:37.387Z",
      "rawMarkdown": "",
      "isDeleted": true
    },
    {
      "id": 3091164,
      "postDate": "2025-01-08T06:17:15.103Z",
      "content": "<p>Great work! Thanks for sharing.</p>",
      "rawMarkdown": "Great work! Thanks for sharing.",
      "votes": 2
    }
  ],
  "comments": [
    {
      "id": 3091148,
      "author_name": "John Payne",
      "author_url": "",
      "post_date": "2025-01-08T05:37:30.210000",
      "content": "<p>I forgot to mention, but responders.csv seems to have further evidence of my findings:</p>\n<ul>\n<li>responders 0,1,2 have tag_0 (they're all differences between the other responders)</li>\n<li>responders 2,5,8 have tag_1 (SMA 4)</li>\n<li>responders 0,3,6 have tag_2 (SMA 20)</li>\n<li>responders 1,4,7 have tag_3 (SMA 120)</li>\n<li>responders 3,4,5 have tag_4 (they're all moving averages of the same signal)</li>\n</ul>\n<p>Interestingly though there's no tag_5 to group responders 6,7,8 together, not sure why.</p>",
      "votes": 18,
      "replies": [
        {
          "id": 3091165,
          "author_name": "ZT",
          "author_url": "",
          "post_date": "2025-01-08T06:17:35.007000",
          "content": "<p>great analysis!</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 3091289,
          "author_name": "Abstr Phil",
          "author_url": "",
          "post_date": "2025-01-08T09:18:06.603000",
          "content": "<p>Maybe there's no need for <code>tag_5</code> because that's the baseline/target, and all the others are either secondary source <code>tag_4</code> or differences between primary source and secondary source <code>tag_0</code>?</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 3092013,
      "author_name": "Jim Beno",
      "author_url": "",
      "post_date": "2025-01-09T04:44:30.577000",
      "content": "<p>Outstanding, very impressive analysis. Can you elaborate on this? \"The underlying signal is gaussian (but also has fat tails), which leads me to believe they are minute-by-minute returns of the symbol.\" I wasn't clear on the connection there.</p>",
      "votes": 3,
      "replies": []
    },
    {
      "id": 3091394,
      "author_name": "Michael Timbs",
      "author_url": "",
      "post_date": "2025-01-08T11:29:38.723000",
      "content": "<p>Your findings make we think we are looking at cross-exchange arbitrage opportunities/signals. </p>",
      "votes": 3,
      "replies": [
        {
          "id": 3093450,
          "author_name": "YMJA",
          "author_url": "",
          "post_date": "2025-01-10T23:25:04.557000",
          "content": "<p>Yepp i was Thinking the same, also the way, the Difference of Responder 0,1 and 2 are correlated makes one wonder.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 3092465,
      "author_name": "Maaax",
      "author_url": "",
      "post_date": "2025-01-09T15:20:06.740000",
      "content": "<p>Beyond returns from another exchange, responders 3, 4, 5 might show ask price returns while 0, 3, 6 show bid price returns (or vice versa).</p>",
      "votes": 2,
      "replies": [
        {
          "id": 3092492,
          "author_name": "tinkei",
          "author_url": "",
          "post_date": "2025-01-09T15:50:46.410000",
          "content": "<p>But then bid should never be higher than ask, while the responders' differences can be both positive or negative.</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 3091393,
      "author_name": "tinkei",
      "author_url": "",
      "post_date": "2025-01-08T11:28:02.303000",
      "content": "<p>Curiously, sometimes the differences can move in an opposite direction from responders 0, 1, 2, or ever go haywire entirely.</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F10943849%2Ffcec62520367085678e28ba30a59904d%2Fvisualization(1).png?generation=1736335632042318&amp;alt=media\" alt=\"Difference between responder 3 and responder 6\"></p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F10943849%2Fa17e362bdd3c6d9f4198143714bc3b43%2Fvisualization(2).png?generation=1736335641095285&amp;alt=media\" alt=\"Difference between responder 4 and responder 7\"></p>",
      "votes": 2,
      "replies": []
    },
    {
      "id": 3092591,
      "author_name": "Truth Seeker",
      "author_url": "",
      "post_date": "2025-01-09T18:24:50.763000",
      "content": "<blockquote>\n  <p>which would explain why everyone's score is so \"low\", it is notoriously difficult to predict future returns from present data</p>\n</blockquote>\n<p>In the limit where it's infinitely difficult to predict future results from present data, the problem essentially becomes one where present data are no more useful than random coin flips for the purpose of prediction.  In this hypothetical scenario, big LB shakeups are fundamentally unavoidable.</p>\n<p>Even though \"notoriously difficult\" is not as bad as \"infinitely difficult\", I wonder if it's bad enough to cause (big) LB shakeups, esp since this competition will test on months of future data no one has seen yet.</p>",
      "votes": -1,
      "replies": [
        {
          "id": 3092622,
          "author_name": "SLi",
          "author_url": "",
          "post_date": "2025-01-09T19:17:29.347000",
          "content": "<p>Eventually Trump will dominate the results LOL</p>",
          "votes": 4,
          "replies": []
        }
      ]
    },
    {
      "id": 3095673,
      "author_name": "llssss1021",
      "author_url": "",
      "post_date": "2025-01-13T15:18:08.590000",
      "content": "<p>This idea is perfect, but API only provide with only one day's data to predict, it's difficult to progress time series(although the idea perform good in training). Of course, it may be my thought mistake.</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 3095299,
      "author_name": "GodGod3",
      "author_url": "",
      "post_date": "2025-01-13T07:24:23.220000",
      "content": "<p>Interesting. But I used to hear that technical analysis works for high frequency trading. Do they use high frequency features then？</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 3093605,
      "author_name": "Simon GUhhhh",
      "author_url": "",
      "post_date": "2025-01-11T07:40:51.817000",
      "content": "<p>OMG if this is the case, what R2 score can you make off the 4min moving average? like, if thats predictable, it shows how predictable are the markets and potentially it would be immensely interesting</p>",
      "votes": 0,
      "replies": [
        {
          "id": 3093894,
          "author_name": "Simon GUhhhh",
          "author_url": "",
          "post_date": "2025-01-11T13:05:46.260000",
          "content": "<p>I mean we could look at the 4min bar and see how easy it is to predict the market 😆</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 3091700,
      "author_name": "Simone De Gasperis",
      "author_url": "",
      "post_date": "2025-01-08T17:21:54.783000",
      "content": "<p>Great analysis 🤩, do you think that by adding more lagged (-1, -5 etc…) responders as input features could allow to improve the score of a model? I used this approach some weeks ago but I ended up by strongly overfitting the train data.</p>",
      "votes": 0,
      "replies": [
        {
          "id": 3091705,
          "author_name": "SLi",
          "author_url": "",
          "post_date": "2025-01-08T17:29:39.297000",
          "content": "<p>In the inference API, we don’t have responder labels per time_id, only the daily lags. Directly adding responder labels lagged by time_ids is therefore not feasible, is it?</p>",
          "votes": 0,
          "replies": [
            {
              "id": 3091712,
              "author_name": "Simone De Gasperis",
              "author_url": "",
              "post_date": "2025-01-08T17:35:58.297000",
              "content": "<p>Maybe I'm wrong but the lags passed at the start of each day are the responders for each time_id of the previous day.</p>",
              "votes": 2,
              "replies": []
            },
            {
              "id": 3091719,
              "author_name": "SLi",
              "author_url": "",
              "post_date": "2025-01-08T17:41:46.580000",
              "content": "<p>if you use time_id from the lags.parquet and shift it by n steps in the api, you are actually using a time_id shifted by 968+n time steps, right?</p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 3091741,
              "author_name": "Simone De Gasperis",
              "author_url": "",
              "post_date": "2025-01-08T17:58:39.240000",
              "content": "<p>I'm not following you here I'm sorry.<br>\nMy point basically was to use to do a prediction, given the responder value for a time Tx of day D,<br>\nThe responder for T_[last] of day D-1 and T_[last-1] of day D-1 ideally to help to capture some trend.</p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 3091760,
              "author_name": "SLi",
              "author_url": "",
              "post_date": "2025-01-08T18:10:41.070000",
              "content": "<p>Consider this: say we are at T=100 at day D. There is a gap of 100 time steps between T=100 and T_last of D-1. </p>\n<p>If we proceed to T=968, we would have a temporal gap of 968 steps between T and T_last of D-1.</p>\n<p>I don’t think any trend can pass over such gap…</p>",
              "votes": 3,
              "replies": []
            },
            {
              "id": 3091770,
              "author_name": "Simone De Gasperis",
              "author_url": "",
              "post_date": "2025-01-08T18:27:02.813000",
              "content": "<p>Yes I know what you're saying  in the sense the \"weight\" of the lagged features will decrease with the time id. But if we go in this direction we should not use the lags at all. <br>\nMaybe more lags can help to capture cyclic trends.</p>",
              "votes": 0,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 3091651,
      "author_name": "Calvin Hui",
      "author_url": "",
      "post_date": "2025-01-08T16:02:19.630000",
      "content": "<p>I’m a bit confused about a point you mentioned in the Responder Relationships section. You stated that the features at time t predict responder 6 at time t-20 quite well. I don’t fully understand this. For example, if t = 150, where X represents the features and y represents responder 6, does this mean X_{150} predicts y_{130} quite well? This seems strange to me because it appears to use future data to predict y, which would already be known at that time.</p>\n<p>By the way, great analysis and thanks for sharing!</p>",
      "votes": 0,
      "replies": [
        {
          "id": 3091653,
          "author_name": "SLi",
          "author_url": "",
          "post_date": "2025-01-08T16:07:00.243000",
          "content": "<p>y_{130} is the actual target value corresponding to X_150; the r6 value at X_150 given in the dataset is acutally corresponds to y_170. In other words, the host may have done <code>.shift(-20)</code> when preparing the data.</p>",
          "votes": 3,
          "replies": [
            {
              "id": 3091673,
              "author_name": "Calvin Hui",
              "author_url": "",
              "post_date": "2025-01-08T16:41:48.463000",
              "content": "<p>Thanks for the clarification! That makes sense now. Appreciate your help!</p>",
              "votes": 0,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 3091363,
      "author_name": "Michael Timbs",
      "author_url": "",
      "post_date": "2025-01-08T10:36:46.647000",
      "content": "<p>Good analysis. Now to try and think how I can use this in my models </p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 3091331,
      "author_name": "WindClimber",
      "author_url": "",
      "post_date": "2025-01-08T10:01:42.457000",
      "content": "<p>Amazing analysis! Thank you for sharing!</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 3091264,
      "author_name": "Chinnn",
      "author_url": "",
      "post_date": "2025-01-08T08:38:12.910000",
      "content": "<p>Incrediable analysis!</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 3091230,
      "author_name": "Anthony Chiu",
      "author_url": "",
      "post_date": "2025-01-08T07:38:01.260000",
      "content": "<p>Thanks for sharing.</p>\n<p>If the target is in the future, then the \"pre-market\" / \"after-hours\" shared here \"https://www.kaggle.com/competitions/jane-street-real-time-market-data-forecasting/discussion/549223\" might not be correct? or the final 20 time_ids are removed.</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 3091823,
      "author_name": "",
      "author_url": "",
      "post_date": "2025-01-08T19:57:06.380000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 3091813,
      "author_name": "",
      "author_url": "",
      "post_date": "2025-01-08T19:31:10.257000",
      "content": "<p>I actually tried to leverage the oscillatory nature of responder 6 a few weeks before. But the problem is that we do not have the initial conditions. Lagged repsonder from previous day don't give us a good prediction for the next day. I have tried using the features to predict differenced responder 6, and the R squared can reach 0.17. The features are better at predicting the change , and we need the current value for predicting responder 6. I still got some ideas, but I don't think I have enough time to try those…</p>",
      "votes": 1,
      "replies": [
        {
          "id": 3091814,
          "author_name": "",
          "author_url": "",
          "post_date": "2025-01-08T19:33:28.243000",
          "content": "<p>I have tried LSTM and CNN on time series, but they are not working yet.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 3091817,
          "author_name": "John Payne",
          "author_url": "",
          "post_date": "2025-01-08T19:41:04.083000",
          "content": "<p>Even if you had a perfect estimation of the responder at lag 20, the signal is a moving average so the expected value at time t given the value at time t-20 is 0, so any sort of forecasting is pretty much a dead end.</p>",
          "votes": 4,
          "replies": [
            {
              "id": 3091829,
              "author_name": "SLi",
              "author_url": "",
              "post_date": "2025-01-08T20:03:15.123000",
              "content": "<blockquote>\n  <p>the expected value at time t given the value at time t-20 is 0</p>\n</blockquote>\n<p>That’s exactly the point. I have experimented to predict r6 from T-20 towards T using features at T. The result confirms this. I got pretty good prediction for T-20, and at T the result has lost most of the variation and got “converged “ to zero.</p>\n<p>How would you suggest to model the series differently? </p>",
              "votes": 2,
              "replies": []
            },
            {
              "id": 3091862,
              "author_name": "John Payne",
              "author_url": "",
              "post_date": "2025-01-08T20:55:40.607000",
              "content": "<p>I'm not sure there's a better way to model the series other than:</p>\n<p>$$Y^6_t=\\alpha\\sum_{i=0}^{19}Z_{t-i}+\\epsilon_t$$</p>\n<p>Where Z_t and epsilon_t are i.i.d normal variables. With this model, knowing Y6 at time t-20 gives you zero information about Y6 at time t. I think this competition was deliberately designed so that we must predict Y6 by finding some sort of hidden dependence between Z_t and the features.</p>",
              "votes": 1,
              "replies": []
            },
            {
              "id": 3092505,
              "author_name": "SLi",
              "author_url": "",
              "post_date": "2025-01-09T16:14:44.773000",
              "content": "<p>What about the following approach: first we predict forward features for <code>T~T+20</code> given features before T, next we predict <code>Y^6</code>  using the predicted forward features. In this way, we can design a model to predict forward features by leveraging both the temporal dependency of a single feature as well as the cross-channel dependency between features, which would be easier than predicting <code>Y^6</code> using 20 steps of lagged <code>Y^6</code> values (since returns are random walk).<br>\nI dont't think I have time to implement this, but I am curious if this could work, even just theoretically. </p>",
              "votes": 2,
              "replies": []
            },
            {
              "id": 3092571,
              "author_name": "John Payne",
              "author_url": "",
              "post_date": "2025-01-09T17:45:32.087000",
              "content": "<p>I did think of doing this a while back, but I figured that predicting the relevant future features would end up being just as difficult as predicting Y6 itself, but I also don't have much time to test this idea.</p>\n<p>I can tell you though that I've spent hours analyzing and testing ideas using this information I found, and so far nothing has worked better than a simple NN from features to responders.</p>",
              "votes": 1,
              "replies": []
            },
            {
              "id": 3092650,
              "author_name": "",
              "author_url": "",
              "post_date": "2025-01-09T20:02:22.207000",
              "content": "<p>I don't think the key is about the temporal dependency of the features. It's the temporal dependency of the responder. I have experimented that the change in responder 6 depends on both its current value and also its momentum. </p>\n<p>Using differential equation as an analogy, the features only gives us the equation. We still need the initial value and first derivative of responder 6, which we have no way to get those.</p>\n<p>I have think of using the features to identify whether responder 6 has return to zero, since I think if the market doesn't have trading opportunity for a while, we shouldn't take any position and responder 6 should reflect that. If we can identify regions that are near 0, then we can use those as initial conditions. But nothing worked yet.</p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 3092655,
              "author_name": "",
              "author_url": "",
              "post_date": "2025-01-09T20:07:37.493000",
              "content": "<p>Let me give you an example. Let say now responder 6 equals 4 right now and it's in a down trend, even if the features consistently showing an up-signal, since the downward momentum is so strong, reponsder 6 still decreases.</p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 3092711,
              "author_name": "John Payne",
              "author_url": "",
              "post_date": "2025-01-09T22:18:57.753000",
              "content": "<p>Here's the conditional expectation of responder 6:</p>\n<p>$$E\\left[ Y^6_{t+s} \\mid Y^6_t \\right] =\\frac{20 - s}{20}Y^6_t \\,\\,\\,\\,\\,\\, \\text{for } s = 0, 1, \\dots, 19$$<br>\n$$E\\left[ Y^6_{t+s} \\mid Y^6_t \\right] =0 \\,\\,\\,\\,\\,\\, \\text{for } s \\geq 20$$</p>\n<p>This tells us that responder 6 is always expected to decrease towards 0, and if you're trying to forecast any more than 19 time ticks forward, then there is really no telling what responder 6 will be, all you know is that it will on average be 0 (with high variance).</p>\n<p>If you really believe that Y6 has a dependence on its momentum, then try and come up with a mathematical model for it. I personally think though that even if there is something there, NNs on the features would outperform anyway.</p>",
              "votes": 2,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 3091210,
      "author_name": "",
      "author_url": "",
      "post_date": "2025-01-08T07:16:37.387000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 3091164,
      "author_name": "Ahmet Erdem",
      "author_url": "",
      "post_date": "2025-01-08T06:17:15.103000",
      "content": "<p>Great work! Thanks for sharing.</p>",
      "votes": 2,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "3091145": "Hi everyone, I joined the competition a couple weeks ago and since then I have spent a good portion of that time studying the responders. In hindsight, I definitely spent too much time studying the data instead of developing my model, but regardless I've decided to reveal the insights I've found and a potential explanation of what the responders could be. This post will walk through roughly the process I went through to get to my findings.\n\n**1.  Autocorrelation Function (ACF)**\n\nOne of the first things I noticed with the responders is their very peculiar autocorrelation functions. For example, here's a graph of responder 6's ACF for symbol 1 across the entire dataset:\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F20686390%2F6d2945b91023f6a7a453e6aa1cbd5954%2Foutput.png?generation=1736304198418982&alt=media)\n\nNotably, the autocorrelation appears to decrease linearly with each lag but have a sharp cutoff in autocorrelation at exactly lag 20. If we look at the ACF of responders 7 and 8, they behave similarly but cutoff at lags 120 and 4, respectively:\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F20686390%2F077216e8d8fc58407ffeddd1ef5e132a%2Foutput1.png?generation=1736305724902604&alt=media)\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F20686390%2F91b61be85c11463de1c07471dbfe828b%2Foutput2.png?generation=1736305733028007&alt=media)\n\nThe ACFs of the rest of the responders are slightly different, but they all share the same property of having a sharp change in autocorrelation at specific lags. Interestingly, there are similarities of the ACF between different responders, with responders 0,3,6 having the same ACF cutoff at lag 20, responders 1,4,7 having it at lag 120, and responders 2,5,8 having it lag 4.\n\n**2. Responder Relationships**\n\nTo make matters more interesting, I found that the features at time t are weirdly good at predicting responders 0,3,6 at time t-20. (You can try this yourself, using linear regression on the features I was able to get a consistent R^2 of around 0.5 on responder 6 lag 20). This property holds on the other responder groups too with their respective lags. This led me to hypothesize that the responders we're given are actually shifted in time, i.e, we're actually trying to predict responder 6 at time t+20 given the features at time t. We can see evidence of this when we graph the lagged responders together:\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F20686390%2Fd3d1f158b2986dc02bde34730452954e%2Foutput3.png?generation=1736308270543101&alt=media)\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F20686390%2F4dfec682d93e8a655ab6078b8dc079b7%2Foutput4.png?generation=1736308288491253&alt=media)\n\nResponders 6,7,8 correlate *much* more nicely at their lags as compared to the normal values. In fact, if you look a little bit closer at these graphs, it almost seems like responder 6 lag 20 is a \"derivative\" of responder 7 lag 120, and responder 8 lag 4 is a \"derivative\" of responder 6 lag 20. Whenever responder 6 is positive, it seems responder 7 moves upwards, and whenever responder 8 is positive, responder 6 moves upwards. Evidence of this is very strong when we graph moving averages of the lagged responders:\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F20686390%2F033442ee8ccbaa69a134f72e62ee73f8%2Foutput5.png?generation=1736309168179946&alt=media)\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F20686390%2Fae2402a655b00fe270c7e9bf69762102%2Foutput6.png?generation=1736309176377111&alt=media)\n\nClearly there's a strong relationship between responders 6,7,8.  Additionally, responders 3,4,5 seem to be very similar to 6,7,8 and therefore also share this same relationship. The last group of responders, 0,1,2 are quite different from the rest of the responders, but after studying some of the graphs it appears that responders 0,1,2 are roughly the difference between responders 3,4,5 and 6,7,8:\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F20686390%2F7c62fafd2d9986d3f42a9824570df601%2Foutput7.png?generation=1736310301191748&alt=media)\n\n**3. Fourier Analysis**\n\nThe last bit of analysis I tried was Fourier analysis on responder 6. Here's what the Fourier transform looks like:\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F20686390%2F18494492bc06afe8e70c38a31eef0f2d%2Foutput8.png?generation=1736310892074227&alt=media)\n\nNotably, there's a distinct pattern in the transform with interval 1/20. In fact it's pretty clear to see that a curve of sin(20pi x)/(20pi x) roughly fits the transform:\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F20686390%2F3a081220abbbb8ff6e629afe61f56c21%2Foutput9.png?generation=1736311350550145&alt=media)\n\nIt looks as though the signal was originally white noise, but had been multiplied by a sin(x)/x curve in the frequency domain (which translates to a convolution in the time domain). Since the inverse Fourier transform of sin(x)/x is a box function, that implies that the original signal was convolved with a uniformly weighted kernel, i.e. a moving average. To test this idea, I generated a white noise signal, applied a moving average of window size 20 to it, and plotted the Fourier transform. And sure enough, we find a very similar graph:\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F20686390%2F81e9a3fc373b8bb6d2b4ec23fa259360%2Foutput10.png?generation=1736312764974135&alt=media)\n\nThe main difference being that it has sharper troughs than responder 6's transform. However, by adding another gaussian white noise variable to our generated white noise, we see that it matches almost perfectly to responder 6's transform:\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F20686390%2Fdabece1d83100cd18e42e42c4ce1c554%2Foutput12.png?generation=1736312811480746&alt=media)\n\nWhich implies that responder 6 is a moving average of width 20 over gaussian white noise.\n\n**What does this mean?**\n\nGiven all of these findings, my best theory is that responders 6, 7, 8 are moving averages with widths 20, 120, 4 of the same underlying signal, but have a slight added noise to them (either intentionally for obfuscation or some sort of measurement error). The underlying signal is gaussian (but also has fat tails), which leads me to believe they are minute-by-minute returns of the symbol. Since responders 3, 4, 5 are so similar to 6, 7, 8, I think they are returns on a separate exchange, and responders 0, 1, 2 are the difference in returns. Since the responders given appear shifted in time, it seems the goal of this competition is to come up with a model that predicts the average returns over the next 20 minutes (which would explain why everyone's score is so \"low\", it is notoriously difficult to predict future returns from present data).\n\nI will admit that I lack a lot of domain knowledge and this is just a guess, so if anyone has different ideas, please feel free to propose.",
    "3091148": "I forgot to mention, but responders.csv seems to have further evidence of my findings:\n- responders 0,1,2 have tag_0 (they're all differences between the other responders)\n- responders 2,5,8 have tag_1 (SMA 4)\n- responders 0,3,6 have tag_2 (SMA 20)\n- responders 1,4,7 have tag_3 (SMA 120)\n- responders 3,4,5 have tag_4 (they're all moving averages of the same signal)\n\nInterestingly though there's no tag_5 to group responders 6,7,8 together, not sure why.",
    "3092013": "Outstanding, very impressive analysis. Can you elaborate on this? \"The underlying signal is gaussian (but also has fat tails), which leads me to believe they are minute-by-minute returns of the symbol.\" I wasn't clear on the connection there.",
    "3091394": "Your findings make we think we are looking at cross-exchange arbitrage opportunities/signals. ",
    "3092465": "Beyond returns from another exchange, responders 3, 4, 5 might show ask price returns while 0, 3, 6 show bid price returns (or vice versa).",
    "3091393": "Curiously, sometimes the differences can move in an opposite direction from responders 0, 1, 2, or ever go haywire entirely.\n\n![Difference between responder 3 and responder 6](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F10943849%2Ffcec62520367085678e28ba30a59904d%2Fvisualization(1).png?generation=1736335632042318&alt=media)\n\n![Difference between responder 4 and responder 7](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F10943849%2Fa17e362bdd3c6d9f4198143714bc3b43%2Fvisualization(2).png?generation=1736335641095285&alt=media)",
    "3092591": ">which would explain why everyone's score is so \"low\", it is notoriously difficult to predict future returns from present data\n\nIn the limit where it's infinitely difficult to predict future results from present data, the problem essentially becomes one where present data are no more useful than random coin flips for the purpose of prediction.  In this hypothetical scenario, big LB shakeups are fundamentally unavoidable.\n\nEven though \"notoriously difficult\" is not as bad as \"infinitely difficult\", I wonder if it's bad enough to cause (big) LB shakeups, esp since this competition will test on months of future data no one has seen yet.",
    "3095673": "This idea is perfect, but API only provide with only one day's data to predict, it's difficult to progress time series(although the idea perform good in training). Of course, it may be my thought mistake.",
    "3095299": "Interesting. But I used to hear that technical analysis works for high frequency trading. Do they use high frequency features then？",
    "3093605": "OMG if this is the case, what R2 score can you make off the 4min moving average? like, if thats predictable, it shows how predictable are the markets and potentially it would be immensely interesting",
    "3091700": "Great analysis 🤩, do you think that by adding more lagged (-1, -5 etc...) responders as input features could allow to improve the score of a model? I used this approach some weeks ago but I ended up by strongly overfitting the train data.\n",
    "3091651": "I’m a bit confused about a point you mentioned in the Responder Relationships section. You stated that the features at time t predict responder 6 at time t-20 quite well. I don’t fully understand this. For example, if t = 150, where X represents the features and y represents responder 6, does this mean X_{150} predicts y_{130} quite well? This seems strange to me because it appears to use future data to predict y, which would already be known at that time.\n\nBy the way, great analysis and thanks for sharing!",
    "3091363": "Good analysis. Now to try and think how I can use this in my models ",
    "3091331": "Amazing analysis! Thank you for sharing!",
    "3091264": "Incrediable analysis!",
    "3091230": "Thanks for sharing.\n\nIf the target is in the future, then the \"pre-market\" / \"after-hours\" shared here \"https://www.kaggle.com/competitions/jane-street-real-time-market-data-forecasting/discussion/549223\" might not be correct? or the final 20 time_ids are removed.",
    "3091823": "",
    "3091813": "I actually tried to leverage the oscillatory nature of responder 6 a few weeks before. But the problem is that we do not have the initial conditions. Lagged repsonder from previous day don't give us a good prediction for the next day. I have tried using the features to predict differenced responder 6, and the R squared can reach 0.17. The features are better at predicting the change , and we need the current value for predicting responder 6. I still got some ideas, but I don't think I have enough time to try those...",
    "3091210": "",
    "3091164": "Great work! Thanks for sharing."
  }
}