{
  "id": 550920,
  "title": "Speculation on the Problem Setting",
  "url": "/competitions/jane-street-real-time-market-data-forecasting/discussion/550920",
  "author_name": "",
  "post_date": "2024-12-10T09:05:58.418612600Z",
  "votes": 11,
  "comment_count": 5,
  "views": 0,
  "content": "<p>Just wanna take a break here.</p>\n<p><strong>Time_id</strong><br>\nOne time step is one minute. This is deduced from 2 observations :<br>\n(1) Max time_id from the train data is 967. Time_id resets to 0 every day. You got 1440 minutes everyday in the real world. Factoring some off-market hours into the calculation, sounds right to me. <br>\n(2) The predict request at the API timeouts in 1 minute. You want to make a prediction before it’s too late, that’s what the limit is for. </p>\n<p><strong>Why “real-time”?</strong><br>\nFrom the time-series forecasting papers I read, the academic usually characterises the task as either “short-term” or “long-term” forecast. <br>\nDepending on the dataset they experiment with, Short term may be referring to forecasting up to 48 time steps, while long term maybe forecasting up to 720 time steps.<br>\nBut this may sound confusing in our problem here: if you think of it as a long term forecasting task, as you predict 968 time steps before you get the true target value to verify, it’s actually just one day, not that long. <br>\nIf you think of it as a short term forecasting task, and just keep predicting the next target, you will notice the prediction error grows intraday alongside with the growing gap of missing observations for the true target. <br>\nOur problem here is special in the sense that it wants you to make a prediction at “real-time” (per-minute) intraday, but it does not give you the immediately previous target values (“responder_6”). Instead you got something else at current time (79 features. Maybe more). The target values only got rebased per day. <br>\nThe timeliness here is crucial in the engineering side: It is not asking you to make one big prediction for the day with a big data-frame input. I missed this point (for not reading the instructions) and kept hitting the timeout limit wondering what went wrong. </p>\n<p><strong>Why lag 1 day?</strong><br>\nThis is a more important question in my opinion. If there is no lag, you can make a pretty accurate prediction, and there is perhaps no need for a competition. <br>\nThe setting that it can create features per minute, but only observe targets per day, has been bothering me. It is likely something that is not in control. An extremely wild guess can be some 3rd-party accountant in clearing house does his math for the P&amp;L in MS Excel and only provides the results at end-of-day. <br>\nAnyway, it is interesting to observe how the model performance varies if you try tweaking this 1 day lag to 1 hour lag or 10-minute lag. It would certainly have made things much easier. </p>\n<p><strong>What is a good predictor?</strong><br>\nThe million-dollar question to ask. No serious math here from my mobile, sorry. Let’s say responder_6 is the endogenous variable, y:<br>\n(1) By far, the immediately previous y. If plotting out y, yes it looks quite jumpy and fluctuates a lot, but it is still quite continuous. However by the problem setting you don’t have the immediately previous y except for time_id=0. Even that is questionable because it assumes nothing happens after time_id=967 and before the next day’s time_id=0<br>\n(2) The next candidate naturally, the current features X (exogenous variables). However, the mapping of x_current to y_current seems very noisy and non-stationary. If you try fitting a simple XGBoost model (the first thing I tried), it’s hard to say it predicts well for the out-of-sample test data (as evidenced by the low metric score). <br>\n(3) Lagged y_history and x_history: Sounds promising to discover more temporal relationships in theory, certainly worths more digging here; however, bigger data size means heavier engineering efforts. Yes there is a 8-hour time limit at submission to work with. <br>\n(4) Symbol_id, feature engineering etc: Haven’t looked into these seriously yet. Happy to know if it has worked for anyone. </p>\n<p><strong>Why forecast, at all?</strong><br>\nThe other threads in the discussion have been insightful. </p>\n<p>Some thoughts before returning to work: <br>\nIt is surprising to me that now approaching the competition deadline in one month and the leaderboard is still giving &lt;0.01 weighted R2. It shows how difficult this problem is. For me it has been difficult to do well at 3 sides at the same time: accurate, fast, robust. There must be something better in the end, right?</p>",
  "messages": [
    {
      "id": "3068441",
      "postDate": "12/10/2024 09:05:58",
      "content": "<p>Just wanna take a break here.</p>\n<p><strong>Time_id</strong><br>\nOne time step is one minute. This is deduced from 2 observations :<br>\n(1) Max time_id from the train data is 967. Time_id resets to 0 every day. You got 1440 minutes everyday in the real world. Factoring some off-market hours into the calculation, sounds right to me. <br>\n(2) The predict request at the API timeouts in 1 minute. You want to make a prediction before it’s too late, that’s what the limit is for. </p>\n<p><strong>Why “real-time”?</strong><br>\nFrom the time-series forecasting papers I read, the academic usually characterises the task as either “short-term” or “long-term” forecast. <br>\nDepending on the dataset they experiment with, Short term may be referring to forecasting up to 48 time steps, while long term maybe forecasting up to 720 time steps.<br>\nBut this may sound confusing in our problem here: if you think of it as a long term forecasting task, as you predict 968 time steps before you get the true target value to verify, it’s actually just one day, not that long. <br>\nIf you think of it as a short term forecasting task, and just keep predicting the next target, you will notice the prediction error grows intraday alongside with the growing gap of missing observations for the true target. <br>\nOur problem here is special in the sense that it wants you to make a prediction at “real-time” (per-minute) intraday, but it does not give you the immediately previous target values (“responder_6”). Instead you got something else at current time (79 features. Maybe more). The target values only got rebased per day. <br>\nThe timeliness here is crucial in the engineering side: It is not asking you to make one big prediction for the day with a big data-frame input. I missed this point (for not reading the instructions) and kept hitting the timeout limit wondering what went wrong. </p>\n<p><strong>Why lag 1 day?</strong><br>\nThis is a more important question in my opinion. If there is no lag, you can make a pretty accurate prediction, and there is perhaps no need for a competition. <br>\nThe setting that it can create features per minute, but only observe targets per day, has been bothering me. It is likely something that is not in control. An extremely wild guess can be some 3rd-party accountant in clearing house does his math for the P&amp;L in MS Excel and only provides the results at end-of-day. <br>\nAnyway, it is interesting to observe how the model performance varies if you try tweaking this 1 day lag to 1 hour lag or 10-minute lag. It would certainly have made things much easier. </p>\n<p><strong>What is a good predictor?</strong><br>\nThe million-dollar question to ask. No serious math here from my mobile, sorry. Let’s say responder_6 is the endogenous variable, y:<br>\n(1) By far, the immediately previous y. If plotting out y, yes it looks quite jumpy and fluctuates a lot, but it is still quite continuous. However by the problem setting you don’t have the immediately previous y except for time_id=0. Even that is questionable because it assumes nothing happens after time_id=967 and before the next day’s time_id=0<br>\n(2) The next candidate naturally, the current features X (exogenous variables). However, the mapping of x_current to y_current seems very noisy and non-stationary. If you try fitting a simple XGBoost model (the first thing I tried), it’s hard to say it predicts well for the out-of-sample test data (as evidenced by the low metric score). <br>\n(3) Lagged y_history and x_history: Sounds promising to discover more temporal relationships in theory, certainly worths more digging here; however, bigger data size means heavier engineering efforts. Yes there is a 8-hour time limit at submission to work with. <br>\n(4) Symbol_id, feature engineering etc: Haven’t looked into these seriously yet. Happy to know if it has worked for anyone. </p>\n<p><strong>Why forecast, at all?</strong><br>\nThe other threads in the discussion have been insightful. </p>\n<p>Some thoughts before returning to work: <br>\nIt is surprising to me that now approaching the competition deadline in one month and the leaderboard is still giving &lt;0.01 weighted R2. It shows how difficult this problem is. For me it has been difficult to do well at 3 sides at the same time: accurate, fast, robust. There must be something better in the end, right?</p>",
      "rawMarkdown": "Just wanna take a break here.\n\n**Time_id**\nOne time step is one minute. This is deduced from 2 observations :\n(1) Max time_id from the train data is 967. Time_id resets to 0 every day. You got 1440 minutes everyday in the real world. Factoring some off-market hours into the calculation, sounds right to me. \n(2) The predict request at the API timeouts in 1 minute. You want to make a prediction before it’s too late, that’s what the limit is for. \n\n**Why “real-time”?**\nFrom the time-series forecasting papers I read, the academic usually characterises the task as either “short-term” or “long-term” forecast. \nDepending on the dataset they experiment with, Short term may be referring to forecasting up to 48 time steps, while long term maybe forecasting up to 720 time steps.\nBut this may sound confusing in our problem here: if you think of it as a long term forecasting task, as you predict 968 time steps before you get the true target value to verify, it’s actually just one day, not that long. \nIf you think of it as a short term forecasting task, and just keep predicting the next target, you will notice the prediction error grows intraday alongside with the growing gap of missing observations for the true target. \nOur problem here is special in the sense that it wants you to make a prediction at “real-time” (per-minute) intraday, but it does not give you the immediately previous target values (“responder_6”). Instead you got something else at current time (79 features. Maybe more). The target values only got rebased per day. \nThe timeliness here is crucial in the engineering side: It is not asking you to make one big prediction for the day with a big data-frame input. I missed this point (for not reading the instructions) and kept hitting the timeout limit wondering what went wrong. \n\n**Why lag 1 day?**\nThis is a more important question in my opinion. If there is no lag, you can make a pretty accurate prediction, and there is perhaps no need for a competition. \nThe setting that it can create features per minute, but only observe targets per day, has been bothering me. It is likely something that is not in control. An extremely wild guess can be some 3rd-party accountant in clearing house does his math for the P&L in MS Excel and only provides the results at end-of-day. \nAnyway, it is interesting to observe how the model performance varies if you try tweaking this 1 day lag to 1 hour lag or 10-minute lag. It would certainly have made things much easier. \n\n**What is a good predictor?**\nThe million-dollar question to ask. No serious math here from my mobile, sorry. Let’s say responder_6 is the endogenous variable, y:\n(1) By far, the immediately previous y. If plotting out y, yes it looks quite jumpy and fluctuates a lot, but it is still quite continuous. However by the problem setting you don’t have the immediately previous y except for time_id=0. Even that is questionable because it assumes nothing happens after time_id=967 and before the next day’s time_id=0\n(2) The next candidate naturally, the current features X (exogenous variables). However, the mapping of x_current to y_current seems very noisy and non-stationary. If you try fitting a simple XGBoost model (the first thing I tried), it’s hard to say it predicts well for the out-of-sample test data (as evidenced by the low metric score). \n(3) Lagged y_history and x_history: Sounds promising to discover more temporal relationships in theory, certainly worths more digging here; however, bigger data size means heavier engineering efforts. Yes there is a 8-hour time limit at submission to work with. \n(4) Symbol_id, feature engineering etc: Haven’t looked into these seriously yet. Happy to know if it has worked for anyone. \n\n**Why forecast, at all?**\nThe other threads in the discussion have been insightful. \n\nSome thoughts before returning to work: \nIt is surprising to me that now approaching the competition deadline in one month and the leaderboard is still giving <0.01 weighted R2. It shows how difficult this problem is. For me it has been difficult to do well at 3 sides at the same time: accurate, fast, robust. There must be something better in the end, right?",
      "votes": null
    },
    {
      "id": "3092630",
      "postDate": "01/09/2025 19:34:05",
      "content": "<p>If we were given the previous time_id responder_6, r^2 scores would be closer to 0.7. There's strong autocorrelation. Unfortunately, as you realized, we work with prior day response variables.</p>",
      "rawMarkdown": "If we were given the previous time_id responder_6, r^2 scores would be closer to 0.7. There's strong autocorrelation. Unfortunately, as you realized, we work with prior day response variables.",
      "votes": null
    },
    {
      "id": "3092779",
      "postDate": "01/10/2025 03:03:49",
      "content": "<p>I am also worried that the R2 is more dictated by how closely the weighted truths are to 0 than any actual change to the models. I am quite far from the top of the leaderboard so maybe they have uncovered something I have missed, but no matter what kind of model I train or how big/small or ensemble etc, the \"shape\" of the R2 value over time is always more or less in sync.  E.g there are very clear periods of low predictive power even for online learners.</p>\n<p>Conversely, there always seem to be regions of high R2 regardless of model type/size, whether it is online etc. I haven't looked but I suspect the weighted average of these days is just closer to 0 than other days and not that the model is doing any better a job of capturing signal. e.g here is a plot of 20 day moving average of R2 during online training for a NN ensemble. You can see that even despite great diversity in models (convultions, GRU, attention, dropout, layers, lr etc) that they more or less track each other directionly. Some periods have great R2 despite this being the first time the data was ever seen (purely online).</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F21287471%2F59fe32b446e2681c22111f76c74d0cab%2Fr2ens.png?generation=1736478890800767&amp;alt=media\" alt=\"\"></p>\n<p>I have even tried ensembling models training on these different regimes to try and have robustness but it has made no impact.</p>\n<p>Another thing that supports this is simply resubmitting the same model when they updated the holdout data size made massive differences to r2 on LB scores.</p>\n<p>Also very early on I found that doing measuring training loss against full prediction and then actually predicting with a tanhh of the predicted value improved scores. Which means that being closer to 0 - with some underlying signal capture seems to be more value than say hyper parameter tuning or tweaking extra layers or ensembling!</p>",
      "rawMarkdown": "I am also worried that the R2 is more dictated by how closely the weighted truths are to 0 than any actual change to the models. I am quite far from the top of the leaderboard so maybe they have uncovered something I have missed, but no matter what kind of model I train or how big/small or ensemble etc, the \"shape\" of the R2 value over time is always more or less in sync.  E.g there are very clear periods of low predictive power even for online learners.\n\nConversely, there always seem to be regions of high R2 regardless of model type/size, whether it is online etc. I haven't looked but I suspect the weighted average of these days is just closer to 0 than other days and not that the model is doing any better a job of capturing signal. e.g here is a plot of 20 day moving average of R2 during online training for a NN ensemble. You can see that even despite great diversity in models (convultions, GRU, attention, dropout, layers, lr etc) that they more or less track each other directionly. Some periods have great R2 despite this being the first time the data was ever seen (purely online).\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F21287471%2F59fe32b446e2681c22111f76c74d0cab%2Fr2ens.png?generation=1736478890800767&alt=media)\n\nI have even tried ensembling models training on these different regimes to try and have robustness but it has made no impact.\n\nAnother thing that supports this is simply resubmitting the same model when they updated the holdout data size made massive differences to r2 on LB scores.\n\nAlso very early on I found that doing measuring training loss against full prediction and then actually predicting with a tanhh of the predicted value improved scores. Which means that being closer to 0 - with some underlying signal capture seems to be more value than say hyper parameter tuning or tweaking extra layers or ensembling!",
      "votes": null
    },
    {
      "id": "3092785",
      "postDate": "01/10/2025 03:20:34",
      "content": "<p>I think the secret lies in feature engineering because all these models use the same features. So while some perform better maybe its diminishing returns. The runtime limitations make it hard to have diversity in features though as the most expensive part is actually preparing features and if we have to hold 6-7 representations of features in memory and construct multiple different kinds then i'm not sure we will meet runtime performance constraints.</p>",
      "rawMarkdown": "I think the secret lies in feature engineering because all these models use the same features. So while some perform better maybe its diminishing returns. The runtime limitations make it hard to have diversity in features though as the most expensive part is actually preparing features and if we have to hold 6-7 representations of features in memory and construct multiple different kinds then i'm not sure we will meet runtime performance constraints.",
      "votes": null
    },
    {
      "id": "3092995",
      "postDate": "01/10/2025 10:43:52",
      "content": "<p>Do you have a similar graph plotted against time_id where the R2 is averaged across days? Wanna see if each day has a consistent shape (say, time_id=0 has a nice positive R2, time_id=967 has a negative R2)</p>",
      "rawMarkdown": "Do you have a similar graph plotted against time_id where the R2 is averaged across days? Wanna see if each day has a consistent shape (say, time_id=0 has a nice positive R2, time_id=967 has a negative R2)",
      "votes": null
    },
    {
      "id": "3093000",
      "postDate": "01/10/2025 10:50:29",
      "content": "<p>No but my next and final experiment is having different ensembles for different parts of the day and swapping them out based on time_id.  So that would be a good thing to graph </p>\n<p>I’m all out of time to try anything more radical </p>",
      "rawMarkdown": "No but my next and final experiment is having different ensembles for different parts of the day and swapping them out based on time_id.  So that would be a good thing to graph \n\nI’m all out of time to try anything more radical",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 3092630,
      "author_name": "xraygoth",
      "author_url": "",
      "post_date": "01/09/2025 19:34:05",
      "content": "<p>If we were given the previous time_id responder_6, r^2 scores would be closer to 0.7. There's strong autocorrelation. Unfortunately, as you realized, we work with prior day response variables.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 3092779,
      "author_name": "michaeltimbs",
      "author_url": "",
      "post_date": "01/10/2025 03:03:49",
      "content": "<p>I am also worried that the R2 is more dictated by how closely the weighted truths are to 0 than any actual change to the models. I am quite far from the top of the leaderboard so maybe they have uncovered something I have missed, but no matter what kind of model I train or how big/small or ensemble etc, the \"shape\" of the R2 value over time is always more or less in sync.  E.g there are very clear periods of low predictive power even for online learners.</p>\n<p>Conversely, there always seem to be regions of high R2 regardless of model type/size, whether it is online etc. I haven't looked but I suspect the weighted average of these days is just closer to 0 than other days and not that the model is doing any better a job of capturing signal. e.g here is a plot of 20 day moving average of R2 during online training for a NN ensemble. You can see that even despite great diversity in models (convultions, GRU, attention, dropout, layers, lr etc) that they more or less track each other directionly. Some periods have great R2 despite this being the first time the data was ever seen (purely online).</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F21287471%2F59fe32b446e2681c22111f76c74d0cab%2Fr2ens.png?generation=1736478890800767&amp;alt=media\" alt=\"\"></p>\n<p>I have even tried ensembling models training on these different regimes to try and have robustness but it has made no impact.</p>\n<p>Another thing that supports this is simply resubmitting the same model when they updated the holdout data size made massive differences to r2 on LB scores.</p>\n<p>Also very early on I found that doing measuring training loss against full prediction and then actually predicting with a tanhh of the predicted value improved scores. Which means that being closer to 0 - with some underlying signal capture seems to be more value than say hyper parameter tuning or tweaking extra layers or ensembling!</p>",
      "votes": null,
      "replies": [
        {
          "id": 3092785,
          "author_name": "michaeltimbs",
          "author_url": "",
          "post_date": "01/10/2025 03:20:34",
          "content": "<p>I think the secret lies in feature engineering because all these models use the same features. So while some perform better maybe its diminishing returns. The runtime limitations make it hard to have diversity in features though as the most expensive part is actually preparing features and if we have to hold 6-7 representations of features in memory and construct multiple different kinds then i'm not sure we will meet runtime performance constraints.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 3092995,
          "author_name": "benlai",
          "author_url": "",
          "post_date": "01/10/2025 10:43:52",
          "content": "<p>Do you have a similar graph plotted against time_id where the R2 is averaged across days? Wanna see if each day has a consistent shape (say, time_id=0 has a nice positive R2, time_id=967 has a negative R2)</p>",
          "votes": null,
          "replies": [
            {
              "id": 3093000,
              "author_name": "michaeltimbs",
              "author_url": "",
              "post_date": "01/10/2025 10:50:29",
              "content": "<p>No but my next and final experiment is having different ensembles for different parts of the day and swapping them out based on time_id.  So that would be a good thing to graph </p>\n<p>I’m all out of time to try anything more radical </p>",
              "votes": null,
              "replies": []
            }
          ]
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "3068441": "Just wanna take a break here.\n\n**Time_id**\nOne time step is one minute. This is deduced from 2 observations :\n(1) Max time_id from the train data is 967. Time_id resets to 0 every day. You got 1440 minutes everyday in the real world. Factoring some off-market hours into the calculation, sounds right to me. \n(2) The predict request at the API timeouts in 1 minute. You want to make a prediction before it’s too late, that’s what the limit is for. \n\n**Why “real-time”?**\nFrom the time-series forecasting papers I read, the academic usually characterises the task as either “short-term” or “long-term” forecast. \nDepending on the dataset they experiment with, Short term may be referring to forecasting up to 48 time steps, while long term maybe forecasting up to 720 time steps.\nBut this may sound confusing in our problem here: if you think of it as a long term forecasting task, as you predict 968 time steps before you get the true target value to verify, it’s actually just one day, not that long. \nIf you think of it as a short term forecasting task, and just keep predicting the next target, you will notice the prediction error grows intraday alongside with the growing gap of missing observations for the true target. \nOur problem here is special in the sense that it wants you to make a prediction at “real-time” (per-minute) intraday, but it does not give you the immediately previous target values (“responder_6”). Instead you got something else at current time (79 features. Maybe more). The target values only got rebased per day. \nThe timeliness here is crucial in the engineering side: It is not asking you to make one big prediction for the day with a big data-frame input. I missed this point (for not reading the instructions) and kept hitting the timeout limit wondering what went wrong. \n\n**Why lag 1 day?**\nThis is a more important question in my opinion. If there is no lag, you can make a pretty accurate prediction, and there is perhaps no need for a competition. \nThe setting that it can create features per minute, but only observe targets per day, has been bothering me. It is likely something that is not in control. An extremely wild guess can be some 3rd-party accountant in clearing house does his math for the P&L in MS Excel and only provides the results at end-of-day. \nAnyway, it is interesting to observe how the model performance varies if you try tweaking this 1 day lag to 1 hour lag or 10-minute lag. It would certainly have made things much easier. \n\n**What is a good predictor?**\nThe million-dollar question to ask. No serious math here from my mobile, sorry. Let’s say responder_6 is the endogenous variable, y:\n(1) By far, the immediately previous y. If plotting out y, yes it looks quite jumpy and fluctuates a lot, but it is still quite continuous. However by the problem setting you don’t have the immediately previous y except for time_id=0. Even that is questionable because it assumes nothing happens after time_id=967 and before the next day’s time_id=0\n(2) The next candidate naturally, the current features X (exogenous variables). However, the mapping of x_current to y_current seems very noisy and non-stationary. If you try fitting a simple XGBoost model (the first thing I tried), it’s hard to say it predicts well for the out-of-sample test data (as evidenced by the low metric score). \n(3) Lagged y_history and x_history: Sounds promising to discover more temporal relationships in theory, certainly worths more digging here; however, bigger data size means heavier engineering efforts. Yes there is a 8-hour time limit at submission to work with. \n(4) Symbol_id, feature engineering etc: Haven’t looked into these seriously yet. Happy to know if it has worked for anyone. \n\n**Why forecast, at all?**\nThe other threads in the discussion have been insightful. \n\nSome thoughts before returning to work: \nIt is surprising to me that now approaching the competition deadline in one month and the leaderboard is still giving <0.01 weighted R2. It shows how difficult this problem is. For me it has been difficult to do well at 3 sides at the same time: accurate, fast, robust. There must be something better in the end, right?",
    "3092630": "If we were given the previous time_id responder_6, r^2 scores would be closer to 0.7. There's strong autocorrelation. Unfortunately, as you realized, we work with prior day response variables.",
    "3092779": "I am also worried that the R2 is more dictated by how closely the weighted truths are to 0 than any actual change to the models. I am quite far from the top of the leaderboard so maybe they have uncovered something I have missed, but no matter what kind of model I train or how big/small or ensemble etc, the \"shape\" of the R2 value over time is always more or less in sync.  E.g there are very clear periods of low predictive power even for online learners.\n\nConversely, there always seem to be regions of high R2 regardless of model type/size, whether it is online etc. I haven't looked but I suspect the weighted average of these days is just closer to 0 than other days and not that the model is doing any better a job of capturing signal. e.g here is a plot of 20 day moving average of R2 during online training for a NN ensemble. You can see that even despite great diversity in models (convultions, GRU, attention, dropout, layers, lr etc) that they more or less track each other directionly. Some periods have great R2 despite this being the first time the data was ever seen (purely online).\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F21287471%2F59fe32b446e2681c22111f76c74d0cab%2Fr2ens.png?generation=1736478890800767&alt=media)\n\nI have even tried ensembling models training on these different regimes to try and have robustness but it has made no impact.\n\nAnother thing that supports this is simply resubmitting the same model when they updated the holdout data size made massive differences to r2 on LB scores.\n\nAlso very early on I found that doing measuring training loss against full prediction and then actually predicting with a tanhh of the predicted value improved scores. Which means that being closer to 0 - with some underlying signal capture seems to be more value than say hyper parameter tuning or tweaking extra layers or ensembling!",
    "3092785": "I think the secret lies in feature engineering because all these models use the same features. So while some perform better maybe its diminishing returns. The runtime limitations make it hard to have diversity in features though as the most expensive part is actually preparing features and if we have to hold 6-7 representations of features in memory and construct multiple different kinds then i'm not sure we will meet runtime performance constraints.",
    "3092995": "Do you have a similar graph plotted against time_id where the R2 is averaged across days? Wanna see if each day has a consistent shape (say, time_id=0 has a nice positive R2, time_id=967 has a negative R2)",
    "3093000": "No but my next and final experiment is having different ensembles for different parts of the day and swapping them out based on time_id.  So that would be a good thing to graph \n\nI’m all out of time to try anything more radical"
  },
  "source": "meta"
}