{
  "id": 556553,
  "title": "What are your brightest ideas (that didn't pan out)?",
  "url": "/competitions/jane-street-real-time-market-data-forecasting/discussion/556553",
  "author_name": "",
  "post_date": "2025-01-14T02:07:08.345203Z",
  "votes": 2,
  "comment_count": 3,
  "views": 0,
  "content": "<p>When you look at my ranking, you can already see that my brightest ideas aren't really that bright. But here goes nothing:</p>\n<ul>\n<li>I wasn't entirely sure if the R-squared score was summed over all date-time cross sections, or simple-averaged over each date-time cross section. The R-squared score ranges from (-inf, 1], because the denominator scales with -1/y^2, and responder 6 is distributed around zero as well. An unfortunate misstep could blow up the whole result. Since I was getting really negative scores, I assumed it was the latter (simple average over each <code>predict()</code> call). Therefore, I devised a devilish quantile model that would forecast when the weighted sum of ground truth responders would be close to zero, and depending on this model's confidence, I scaled down the magnitude of the  model predictions.<ul>\n<li>This was worse than a simple MLP regressing the responders directly.</li></ul></li>\n<li>Thanks to <a href=\"https://www.kaggle.com/johnpayne0\" target=\"_blank\">@johnpayne0</a>'s findings that these responders were related groups of Simple Moving Averages, shifting SMA4 (responder 8) by steps of 4 <code>time_id</code>s 4 times, and then taking the mean of the five, should be mathematically equivalent to SMA20 (responder 6). I made this an auxiliary training target, and suffice to say…<ul>\n<li>This was worse than a simple MLP regressing the responders directly.</li></ul></li>\n<li>It was rather annoying that we were only getting new ground truths the day after, instead of 20 minutes later. What if I have a model that \"cheats\", by being able to bootstrap its own predictions as previous steps' responders? In training, having access to previous ground truth responders 20 minutes ago, it had a ridiculous R-squared score of 0.7. In practice, having to bootstrap 900 time steps…<ul>\n<li>This was worse than a simple MLP regressing the responders directly.</li>\n<li>This was also my most negatively-scored model.</li></ul></li>\n</ul>\n<p>After weeks of staring at R^2 0.00XX plots, feast your eyes on this 0.7 ovefit:<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F10943849%2Faa10d400d92e13751a296f2da406208b%2FV11%20bootstrap%20model.png?generation=1736819987190150&amp;alt=media\" alt=\"Overfit is beautiful\"></p>\n<blockquote>\n  <p>\"When one overfits the population, you are become the population.\" - William Shakespeare</p>\n</blockquote>\n<p>Your turn. What great ideas did you have, that didn't work out as you had hoped?</p>",
  "messages": [
    {
      "id": "3096012",
      "postDate": "01/14/2025 02:07:08",
      "content": "<p>When you look at my ranking, you can already see that my brightest ideas aren't really that bright. But here goes nothing:</p>\n<ul>\n<li>I wasn't entirely sure if the R-squared score was summed over all date-time cross sections, or simple-averaged over each date-time cross section. The R-squared score ranges from (-inf, 1], because the denominator scales with -1/y^2, and responder 6 is distributed around zero as well. An unfortunate misstep could blow up the whole result. Since I was getting really negative scores, I assumed it was the latter (simple average over each <code>predict()</code> call). Therefore, I devised a devilish quantile model that would forecast when the weighted sum of ground truth responders would be close to zero, and depending on this model's confidence, I scaled down the magnitude of the  model predictions.<ul>\n<li>This was worse than a simple MLP regressing the responders directly.</li></ul></li>\n<li>Thanks to <a href=\"https://www.kaggle.com/johnpayne0\" target=\"_blank\">@johnpayne0</a>'s findings that these responders were related groups of Simple Moving Averages, shifting SMA4 (responder 8) by steps of 4 <code>time_id</code>s 4 times, and then taking the mean of the five, should be mathematically equivalent to SMA20 (responder 6). I made this an auxiliary training target, and suffice to say…<ul>\n<li>This was worse than a simple MLP regressing the responders directly.</li></ul></li>\n<li>It was rather annoying that we were only getting new ground truths the day after, instead of 20 minutes later. What if I have a model that \"cheats\", by being able to bootstrap its own predictions as previous steps' responders? In training, having access to previous ground truth responders 20 minutes ago, it had a ridiculous R-squared score of 0.7. In practice, having to bootstrap 900 time steps…<ul>\n<li>This was worse than a simple MLP regressing the responders directly.</li>\n<li>This was also my most negatively-scored model.</li></ul></li>\n</ul>\n<p>After weeks of staring at R^2 0.00XX plots, feast your eyes on this 0.7 ovefit:<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F10943849%2Faa10d400d92e13751a296f2da406208b%2FV11%20bootstrap%20model.png?generation=1736819987190150&amp;alt=media\" alt=\"Overfit is beautiful\"></p>\n<blockquote>\n  <p>\"When one overfits the population, you are become the population.\" - William Shakespeare</p>\n</blockquote>\n<p>Your turn. What great ideas did you have, that didn't work out as you had hoped?</p>",
      "rawMarkdown": "When you look at my ranking, you can already see that my brightest ideas aren't really that bright. But here goes nothing:\n\n- I wasn't entirely sure if the R-squared score was summed over all date-time cross sections, or simple-averaged over each date-time cross section. The R-squared score ranges from (-inf, 1], because the denominator scales with -1/y^2, and responder 6 is distributed around zero as well. An unfortunate misstep could blow up the whole result. Since I was getting really negative scores, I assumed it was the latter (simple average over each `predict()` call). Therefore, I devised a devilish quantile model that would forecast when the weighted sum of ground truth responders would be close to zero, and depending on this model's confidence, I scaled down the magnitude of the ~~bets~~ model predictions.\n  - This was worse than a simple MLP regressing the responders directly.\n- Thanks to @johnpayne0's findings that these responders were related groups of Simple Moving Averages, shifting SMA4 (responder 8) by steps of 4 `time_id`s 4 times, and then taking the mean of the five, should be mathematically equivalent to SMA20 (responder 6). I made this an auxiliary training target, and suffice to say...\n  - This was worse than a simple MLP regressing the responders directly.\n- It was rather annoying that we were only getting new ground truths the day after, instead of 20 minutes later. What if I have a model that \"cheats\", by being able to bootstrap its own predictions as previous steps' responders? In training, having access to previous ground truth responders 20 minutes ago, it had a ridiculous R-squared score of 0.7. In practice, having to bootstrap 900 time steps...\n  - This was worse than a simple MLP regressing the responders directly.\n  - This was also my most negatively-scored model.\n\nAfter weeks of staring at R^2 0.00XX plots, feast your eyes on this 0.7 ovefit:\n![Overfit is beautiful](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F10943849%2Faa10d400d92e13751a296f2da406208b%2FV11%20bootstrap%20model.png?generation=1736819987190150&alt=media)\n\n>\"When one overfits the population, you are become the population.\" - William Shakespeare\n\nYour turn. What great ideas did you have, that didn't work out as you had hoped?",
      "votes": null
    },
    {
      "id": "3096055",
      "postDate": "01/14/2025 03:38:59",
      "content": "<p>In the last few days, we tried to use lag to calculate the predicted value of the model and the real error, and then integrated it into the future prediction through certain conditions. Unfortunately, there was not enough time to explore.</p>",
      "rawMarkdown": "In the last few days, we tried to use lag to calculate the predicted value of the model and the real error, and then integrated it into the future prediction through certain conditions. Unfortunately, there was not enough time to explore.",
      "votes": null
    },
    {
      "id": "3096154",
      "postDate": "01/14/2025 06:05:56",
      "content": "<p>I tried to use multi-task learning on the 9 responders by training a separate MLP and using the output of the shared layers as added features to my main model. I thought it was a really good idea, but it only gave me a 0.0001 boost on the LB lol.</p>",
      "rawMarkdown": "I tried to use multi-task learning on the 9 responders by training a separate MLP and using the output of the shared layers as added features to my main model. I thought it was a really good idea, but it only gave me a 0.0001 boost on the LB lol.",
      "votes": null
    },
    {
      "id": "3096394",
      "postDate": "01/14/2025 10:31:21",
      "content": "<p>I did the opposite in the beginning, trying to apply a single gradient boosting model to all responders (which normally can only predict a single output): I added a <code>responder_num</code> as input feature, duplicated the other features, then trained the model with a single flattened responders target. Didn't work well. But then I tried training ensembles per feature tag. That wasn't too bad!</p>",
      "rawMarkdown": "I did the opposite in the beginning, trying to apply a single gradient boosting model to all responders (which normally can only predict a single output): I added a `responder_num` as input feature, duplicated the other features, then trained the model with a single flattened responders target. Didn't work well. But then I tried training ensembles per feature tag. That wasn't too bad!",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 3096055,
      "author_name": "lechengyan",
      "author_url": "",
      "post_date": "01/14/2025 03:38:59",
      "content": "<p>In the last few days, we tried to use lag to calculate the predicted value of the model and the real error, and then integrated it into the future prediction through certain conditions. Unfortunately, there was not enough time to explore.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 3096154,
      "author_name": "johnpayne0",
      "author_url": "",
      "post_date": "01/14/2025 06:05:56",
      "content": "<p>I tried to use multi-task learning on the 9 responders by training a separate MLP and using the output of the shared layers as added features to my main model. I thought it was a really good idea, but it only gave me a 0.0001 boost on the LB lol.</p>",
      "votes": null,
      "replies": [
        {
          "id": 3096394,
          "author_name": "tinkei",
          "author_url": "",
          "post_date": "01/14/2025 10:31:21",
          "content": "<p>I did the opposite in the beginning, trying to apply a single gradient boosting model to all responders (which normally can only predict a single output): I added a <code>responder_num</code> as input feature, duplicated the other features, then trained the model with a single flattened responders target. Didn't work well. But then I tried training ensembles per feature tag. That wasn't too bad!</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "3096012": "When you look at my ranking, you can already see that my brightest ideas aren't really that bright. But here goes nothing:\n\n- I wasn't entirely sure if the R-squared score was summed over all date-time cross sections, or simple-averaged over each date-time cross section. The R-squared score ranges from (-inf, 1], because the denominator scales with -1/y^2, and responder 6 is distributed around zero as well. An unfortunate misstep could blow up the whole result. Since I was getting really negative scores, I assumed it was the latter (simple average over each `predict()` call). Therefore, I devised a devilish quantile model that would forecast when the weighted sum of ground truth responders would be close to zero, and depending on this model's confidence, I scaled down the magnitude of the ~~bets~~ model predictions.\n  - This was worse than a simple MLP regressing the responders directly.\n- Thanks to @johnpayne0's findings that these responders were related groups of Simple Moving Averages, shifting SMA4 (responder 8) by steps of 4 `time_id`s 4 times, and then taking the mean of the five, should be mathematically equivalent to SMA20 (responder 6). I made this an auxiliary training target, and suffice to say...\n  - This was worse than a simple MLP regressing the responders directly.\n- It was rather annoying that we were only getting new ground truths the day after, instead of 20 minutes later. What if I have a model that \"cheats\", by being able to bootstrap its own predictions as previous steps' responders? In training, having access to previous ground truth responders 20 minutes ago, it had a ridiculous R-squared score of 0.7. In practice, having to bootstrap 900 time steps...\n  - This was worse than a simple MLP regressing the responders directly.\n  - This was also my most negatively-scored model.\n\nAfter weeks of staring at R^2 0.00XX plots, feast your eyes on this 0.7 ovefit:\n![Overfit is beautiful](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F10943849%2Faa10d400d92e13751a296f2da406208b%2FV11%20bootstrap%20model.png?generation=1736819987190150&alt=media)\n\n>\"When one overfits the population, you are become the population.\" - William Shakespeare\n\nYour turn. What great ideas did you have, that didn't work out as you had hoped?",
    "3096055": "In the last few days, we tried to use lag to calculate the predicted value of the model and the real error, and then integrated it into the future prediction through certain conditions. Unfortunately, there was not enough time to explore.",
    "3096154": "I tried to use multi-task learning on the 9 responders by training a separate MLP and using the output of the shared layers as added features to my main model. I thought it was a really good idea, but it only gave me a 0.0001 boost on the LB lol.",
    "3096394": "I did the opposite in the beginning, trying to apply a single gradient boosting model to all responders (which normally can only predict a single output): I added a `responder_num` as input feature, duplicated the other features, then trained the model with a single flattened responders target. Didn't work well. But then I tried training ensembles per feature tag. That wasn't too bad!"
  },
  "source": "meta"
}