{
  "id": 546165,
  "title": "1-month gone, what's your key takeaway?",
  "url": "/competitions/jane-street-real-time-market-data-forecasting/discussion/546165",
  "author_name": "Ravi Ramakrishnan",
  "post_date": "2024-11-14T08:16:00.008000",
  "votes": 36,
  "comment_count": 42,
  "views": 0,
  "content": "<p>Hello all,</p>\n<p>We are exactly 1 month into the competition as on date (14-11-2024). I have learnt the below from my tryst with the competition insofar-</p>\n<ol>\n<li>This competition is very different from the rest of the time series competitions recently completed on Kaggle and in my opinion, far more difficult than the ones gone by</li>\n<li>Data size, timing constraints and inability to train models on Kaggle kernels with the entire training data is a big challenge</li>\n<li>It is very easy for one to get swayed with score clusters on the leaderboard and tread the wrong path. Such approaches are often met with disappointment later and don't work well with assignments like these</li>\n<li>CV-LB relation establishment is perhaps the biggest and herculean task here. Not all approaches have yielded good CV-Lb relations</li>\n<li>Conventional CV alignment with the LB is not likely to yield good results - one needs to be creative here</li>\n<li>The metric values (both CV score and LB score) are very low and perhaps not completely reliable, some creativity and out-of-box thinking is needed to make sense of these values </li>\n<li>Certain symbols are extremely hard to predict and are worse off with a model than conjecture</li>\n</ol>\n<p>Thoughts? Comments?</p>",
  "messages": [
    {
      "id": 3045158,
      "postDate": "2024-11-14T08:16:00.007Z",
      "content": "<p>Hello all,</p>\n<p>We are exactly 1 month into the competition as on date (14-11-2024). I have learnt the below from my tryst with the competition insofar-</p>\n<ol>\n<li>This competition is very different from the rest of the time series competitions recently completed on Kaggle and in my opinion, far more difficult than the ones gone by</li>\n<li>Data size, timing constraints and inability to train models on Kaggle kernels with the entire training data is a big challenge</li>\n<li>It is very easy for one to get swayed with score clusters on the leaderboard and tread the wrong path. Such approaches are often met with disappointment later and don't work well with assignments like these</li>\n<li>CV-LB relation establishment is perhaps the biggest and herculean task here. Not all approaches have yielded good CV-Lb relations</li>\n<li>Conventional CV alignment with the LB is not likely to yield good results - one needs to be creative here</li>\n<li>The metric values (both CV score and LB score) are very low and perhaps not completely reliable, some creativity and out-of-box thinking is needed to make sense of these values </li>\n<li>Certain symbols are extremely hard to predict and are worse off with a model than conjecture</li>\n</ol>\n<p>Thoughts? Comments?</p>",
      "rawMarkdown": "Hello all,\n\nWe are exactly 1 month into the competition as on date (14-11-2024). I have learnt the below from my tryst with the competition insofar-\n\n1. This competition is very different from the rest of the time series competitions recently completed on Kaggle and in my opinion, far more difficult than the ones gone by\n2. Data size, timing constraints and inability to train models on Kaggle kernels with the entire training data is a big challenge\n3. It is very easy for one to get swayed with score clusters on the leaderboard and tread the wrong path. Such approaches are often met with disappointment later and don't work well with assignments like these\n4. CV-LB relation establishment is perhaps the biggest and herculean task here. Not all approaches have yielded good CV-Lb relations\n5. Conventional CV alignment with the LB is not likely to yield good results - one needs to be creative here\n6. The metric values (both CV score and LB score) are very low and perhaps not completely reliable, some creativity and out-of-box thinking is needed to make sense of these values \n7. Certain symbols are extremely hard to predict and are worse off with a model than conjecture\n\nThoughts? Comments?",
      "votes": 36
    },
    {
      "id": 3045411,
      "postDate": "2024-11-14T13:13:16.463Z",
      "content": "<p>Besides, labels fluctuates greatly over time, and it is quite possible that this is due to market factors rather than statistical factors in the next 6-month prediction phase after the end of submissions.</p>",
      "rawMarkdown": "Besides, labels fluctuates greatly over time, and it is quite possible that this is due to market factors rather than statistical factors in the next 6-month prediction phase after the end of submissions.",
      "votes": 3,
      "replies": [
        {
          "id": 3045416,
          "postDate": "2024-11-14T13:17:49.897Z",
          "content": "<p>Feature engineering and cv strategy are also critical due to the time limit of submission api, but as you mentioned, sometimes it's like a gamble when meeting a good cv correlating to a good public score.</p>",
          "rawMarkdown": "Feature engineering and cv strategy are also critical due to the time limit of submission api, but as you mentioned, sometimes it's like a gamble when meeting a good cv correlating to a good public score.",
          "votes": 2,
          "replies": [
            {
              "id": 3045509,
              "postDate": "2024-11-14T14:25:28.447Z",
              "content": "<p>Polars should be utilized if doing heavy feature engineering. I need to convert my pandas to polars in submission API.</p>",
              "rawMarkdown": "Polars should be utilized if doing heavy feature engineering. I need to convert my pandas to polars in submission API.",
              "votes": 1
            },
            {
              "id": 3045513,
              "postDate": "2024-11-14T14:28:35.177Z",
              "content": "<p>I know that. I mean below</p>\n<blockquote>\n  <p>In my previous trials, EDA/feature engineering can mislead me into overfitting situation</p>\n</blockquote>",
              "rawMarkdown": "I know that. I mean below\n\n>In my previous trials, EDA/feature engineering can mislead me into overfitting situation",
              "votes": 1
            },
            {
              "id": 3045537,
              "postDate": "2024-11-14T15:00:44.307Z",
              "content": "<p>Agreed - just be careful that the patterns exhibited by a feature you engineer is shown throughout most time periods, especially your validation set. Or, suffer on submission 🫠 😭</p>",
              "rawMarkdown": "Agreed - just be careful that the patterns exhibited by a feature you engineer is shown throughout most time periods, especially your validation set. Or, suffer on submission 🫠 😭",
              "votes": 1
            }
          ]
        },
        {
          "id": 3045417,
          "postDate": "2024-11-14T13:19:05.953Z",
          "content": "<p><a href=\"https://www.kaggle.com/sweetyheehee\" target=\"_blank\">@sweetyheehee</a> We are not quite sure what these labels and features mean as well, making it difficult to structure secondary features as well. Some EDA indicates certain patterns and insights, but these are indicative rather than certain. </p>",
          "rawMarkdown": "@sweetyheehee We are not quite sure what these labels and features mean as well, making it difficult to structure secondary features as well. Some EDA indicates certain patterns and insights, but these are indicative rather than certain. ",
          "replies": [
            {
              "id": 3045429,
              "postDate": "2024-11-14T13:30:28.820Z",
              "content": "<p><a href=\"https://www.kaggle.com/ravi20076\" target=\"_blank\">@ravi20076</a> sure, even difficult to train secondly and more times. In my previous trials, EDA/feature engineering can mislead me into overfitting situation  </p>",
              "rawMarkdown": "@ravi20076 sure, even difficult to train secondly and more times. In my previous trials, EDA/feature engineering can mislead me into overfitting situation  ",
              "votes": 1
            }
          ]
        },
        {
          "id": 3045526,
          "postDate": "2024-11-14T14:48:53.623Z",
          "content": "<p>Out of curiosity, is your best model so far NN or Boosting?</p>",
          "rawMarkdown": "Out of curiosity, is your best model so far NN or Boosting?",
          "votes": 1,
          "replies": [
            {
              "id": 3045540,
              "postDate": "2024-11-14T15:04:15.200Z",
              "content": "<p>Not relating to my best cv but good one</p>",
              "rawMarkdown": "Not relating to my best cv but good one",
              "votes": 1
            }
          ]
        }
      ]
    },
    {
      "id": 3045413,
      "postDate": "2024-11-14T13:15:51.260Z",
      "content": "<p>Due to the huge amount of data, it is necessary to find a balance between the amount of data and the number of features. At present, I haven't done any feature engineering yet.</p>",
      "rawMarkdown": "Due to the huge amount of data, it is necessary to find a balance between the amount of data and the number of features. At present, I haven't done any feature engineering yet.",
      "votes": 4,
      "replies": [
        {
          "id": 3046156,
          "postDate": "2024-11-15T07:03:22.003Z",
          "content": "<p><a href=\"https://www.kaggle.com/yunsuxiaozi\" target=\"_blank\">@yunsuxiaozi</a> <br>\nHi. What kind of experiments have you done so far if feature engineering is not relavent at this stage?</p>",
          "rawMarkdown": "@yunsuxiaozi \nHi. What kind of experiments have you done so far if feature engineering is not relavent at this stage?",
          "votes": 1
        }
      ]
    },
    {
      "id": 3047029,
      "postDate": "2024-11-16T07:00:08.257Z",
      "content": "<p>I'm mostly curious about what the top 2 on LB discovered to get such a high score. yuanzhe and rib both made huge leaps in LB score very quickly. Seems like they must have discovered some magic the rest of us haven't</p>",
      "rawMarkdown": "I'm mostly curious about what the top 2 on LB discovered to get such a high score. yuanzhe and rib both made huge leaps in LB score very quickly. Seems like they must have discovered some magic the rest of us haven't",
      "votes": 2,
      "replies": [
        {
          "id": 3047137,
          "postDate": "2024-11-16T09:36:19.290Z",
          "content": "<p>I noticed that too. I'm suspecting of some kind of a post processing like people were normalizing each batch on last Optiver competition.</p>",
          "rawMarkdown": "I noticed that too. I'm suspecting of some kind of a post processing like people were normalizing each batch on last Optiver competition.",
          "votes": 3,
          "replies": [
            {
              "id": 3047531,
              "postDate": "2024-11-16T19:35:34.533Z",
              "content": "<p>Can you please explain what do you mean. Do you mean like target regularisation</p>",
              "rawMarkdown": "Can you please explain what do you mean. Do you mean like target regularisation",
              "votes": 1
            },
            {
              "id": 3047543,
              "postDate": "2024-11-16T20:08:22.993Z",
              "content": "<p>In last the Optiver competition you got a small boost by zero-meaning your prediction batch on their metric, their score is quite higher than 4th onwards so any post processing trick must be very strong</p>\n<blockquote>\n  <p><code>preds = preds - preds.mean()</code></p>\n</blockquote>",
              "rawMarkdown": "In last the Optiver competition you got a small boost by zero-meaning your prediction batch on their metric, their score is quite higher than 4th onwards so any post processing trick must be very strong\n\n>`preds = preds - preds.mean()`",
              "votes": 8
            }
          ]
        }
      ]
    },
    {
      "id": 3045628,
      "postDate": "2024-11-14T16:25:22.583Z",
      "content": "<p>I'm still stuck with this from the beginning, it is not just CV-LB relationship, small changes in the local validation produce very different results.</p>",
      "rawMarkdown": "I'm still stuck with this from the beginning, it is not just CV-LB relationship, small changes in the local validation produce very different results.",
      "votes": 1
    },
    {
      "id": 3056060,
      "postDate": "2024-11-26T14:11:53.447Z",
      "content": "<ol>\n<li>Subject to GPU and memory size, is it better to use long-term training data (using part of the features) or use all features (intercept part of the training data), and how to balance the two</li>\n<li>The model’s inference time is very demanding.</li>\n<li>I think there are serious data distribution differences, which cause the huge difference in cv-lb</li>\n<li>How to reasonably construct feature engineering to capture more information (I currently see an operation similar to positional encoding in bert)</li>\n</ol>",
      "rawMarkdown": "1. Subject to GPU and memory size, is it better to use long-term training data (using part of the features) or use all features (intercept part of the training data), and how to balance the two\n2. The model’s inference time is very demanding.\n3. I think there are serious data distribution differences, which cause the huge difference in cv-lb\n4. How to reasonably construct feature engineering to capture more information (I currently see an operation similar to positional encoding in bert)",
      "votes": 2
    },
    {
      "id": 3050605,
      "postDate": "2024-11-20T11:24:11.847Z",
      "content": "<p>For me</p>\n<ol>\n<li><p>This is actually a very hard problem when accounting for all the constraints and lack of clarity on the data. In some ways this is good and it has certainly been a lot of fun. The lack of any kind of potential for domain knowledge means everything has to be validated in the data.</p></li>\n<li><p>Spent way too long training/playing without getting a working e2e submission working. Ended up spending 2 weeks trying to use the Darts library which is a non starter for the submission requirements here.</p></li>\n<li><p>Use of lags is likely to be a distraction. I'm pretending they don't exist for now as the complexity of safely storing and infilling lags and handling new symbols seems very high for the unknown reward</p></li>\n<li><p>The size of the data has proven tricky. Without GPU access it often takes me 12 hours to do a single run and any bug is a major setback. There are limits to how many features I can even include or the depth of search. Often will have my training process killed after 3-4 hours.</p></li>\n<li><p>I think LB will shift dramatically. For example I haven't even managed to get a proper model submitted yet (other than a naive LGBM) to understand the kaggle competition due to complexityies with CV scores being unstable and missing subtle requirements. Definitely possibly many people are still struggling through to get to a foundational baseline that they can even build from (feature pruning, feature engineering, ensembling, bigger training runs et)</p></li>\n<li><p>One I am confident in model structure and assumptions it seems inevitable that I will have to rent a cloud A100 (at minimum) to train my final submission on</p></li>\n</ol>",
      "rawMarkdown": "For me\n\n0. This is actually a very hard problem when accounting for all the constraints and lack of clarity on the data. In some ways this is good and it has certainly been a lot of fun. The lack of any kind of potential for domain knowledge means everything has to be validated in the data.\n\n1. Spent way too long training/playing without getting a working e2e submission working. Ended up spending 2 weeks trying to use the Darts library which is a non starter for the submission requirements here.\n\n2. Use of lags is likely to be a distraction. I'm pretending they don't exist for now as the complexity of safely storing and infilling lags and handling new symbols seems very high for the unknown reward\n\n3. The size of the data has proven tricky. Without GPU access it often takes me 12 hours to do a single run and any bug is a major setback. There are limits to how many features I can even include or the depth of search. Often will have my training process killed after 3-4 hours.\n\n4. I think LB will shift dramatically. For example I haven't even managed to get a proper model submitted yet (other than a naive LGBM) to understand the kaggle competition due to complexityies with CV scores being unstable and missing subtle requirements. Definitely possibly many people are still struggling through to get to a foundational baseline that they can even build from (feature pruning, feature engineering, ensembling, bigger training runs et)\n\n5. One I am confident in model structure and assumptions it seems inevitable that I will have to rent a cloud A100 (at minimum) to train my final submission on",
      "votes": 2,
      "replies": [
        {
          "id": 3050620,
          "postDate": "2024-11-20T11:36:43.233Z",
          "content": "<p>Some suggestions based on what I learned so far:<br>\nFor 4. You can overcome that with random sampling. It will ultimately affect model accuracy, but non linearly (33% of samples will yield maybe 95% of results)<br>\nFor 5. What most people are doing and works fine is to separate the last 100-200 days of data for testing</p>",
          "rawMarkdown": "Some suggestions based on what I learned so far:\nFor 4. You can overcome that with random sampling. It will ultimately affect model accuracy, but non linearly (33% of samples will yield maybe 95% of results)\nFor 5. What most people are doing and works fine is to separate the last 100-200 days of data for testing",
          "votes": 2,
          "replies": [
            {
              "id": 3050630,
              "postDate": "2024-11-20T11:59:51.080Z",
              "content": "<blockquote>\n  <p>For 4. You can overcome that with random sampling. It will ultimately affect model accuracy, but non linearly (33% of samples will yield maybe 95% of results)</p>\n</blockquote>\n<p>do you mean that if i used 33% of time_ids for every date_id i will finally get a model worse by only like let's say 5%<br>\ni started by training a model on the final 5 folds and the performance only decreased whenever i decrease the number of folds i.e. when i use 4 folds performance decreased  </p>",
              "rawMarkdown": ">For 4. You can overcome that with random sampling. It will ultimately affect model accuracy, but non linearly (33% of samples will yield maybe 95% of results)\n\ndo you mean that if i used 33% of time_ids for every date_id i will finally get a model worse by only like let's say 5%\ni started by training a model on the final 5 folds and the performance only decreased whenever i decrease the number of folds i.e. when i use 4 folds performance decreased  ",
              "votes": 1
            },
            {
              "id": 3050632,
              "postDate": "2024-11-20T12:02:13.010Z",
              "content": "<blockquote>\n  <p>What most people are doing and works fine is to separate the last 100-200 days of data for testing</p>\n</blockquote>\n<p>Yes I was doing that for almost 2 weeks before I realised that the models I were training could not actually be submitted for the competition and wasted a lot of time doing it.  Since I have not yet finished training a model that will work I don't know how well my holdout R2 will match submission R2 but many are saying they are quite different</p>\n<blockquote>\n  <p>For 4. You can overcome that with random sampling. It will ultimately affect model accuracy, but non linearly (33% of samples will yield maybe 95% of results)</p>\n</blockquote>\n<p>I think I will try sampling next once I have a baseline to compare against. At current pace it will take me 5-6 days per experiment which is not feasible</p>",
              "rawMarkdown": "> What most people are doing and works fine is to separate the last 100-200 days of data for testing\n\nYes I was doing that for almost 2 weeks before I realised that the models I were training could not actually be submitted for the competition and wasted a lot of time doing it.  Since I have not yet finished training a model that will work I don't know how well my holdout R2 will match submission R2 but many are saying they are quite different\n\n> For 4. You can overcome that with random sampling. It will ultimately affect model accuracy, but non linearly (33% of samples will yield maybe 95% of results)\n\nI think I will try sampling next once I have a baseline to compare against. At current pace it will take me 5-6 days per experiment which is not feasible",
              "votes": 1
            },
            {
              "id": 3050665,
              "postDate": "2024-11-20T12:50:44.547Z",
              "content": "<blockquote>\n  <p>Yes I was doing that for almost 2 weeks before I realised that the models I were training could not actually be submitted for the competition and wasted a lot of time doing it. Since I have not yet finished training a model that will work I don't know how well my holdout R2 will match submission R2 but many are saying they are quite different</p>\n</blockquote>\n<p>They do hold out quite well for me but the thing is that you have to retrain your model with that last 100-200 days for submission. The split is just for validation purposes. If you use the same model for submission you'll have a 100-200 lagged model which does not perform well in this scenario.</p>",
              "rawMarkdown": ">Yes I was doing that for almost 2 weeks before I realised that the models I were training could not actually be submitted for the competition and wasted a lot of time doing it. Since I have not yet finished training a model that will work I don't know how well my holdout R2 will match submission R2 but many are saying they are quite different\n\nThey do hold out quite well for me but the thing is that you have to retrain your model with that last 100-200 days for submission. The split is just for validation purposes. If you use the same model for submission you'll have a 100-200 lagged model which does not perform well in this scenario.",
              "votes": 1
            },
            {
              "id": 3050669,
              "postDate": "2024-11-20T12:53:54.203Z",
              "content": "<blockquote>\n  <p>do you mean that if i used 33% of time_ids for every date_id i will finally get a model worse by only like let's say 5%<br>\n  i started by training a model on the final 5 folds and the performance only decreased whenever i decrease the number of folds i.e. when i use 4 folds performance decreased</p>\n</blockquote>\n<p>No, I mean sampling rows, not sampling time_ids. And yes, if you remove partitions the performance will decrease because the partitions are from different periods with different behaviours. If you sample out of all partitions that shouldn't be an issue.</p>",
              "rawMarkdown": ">do you mean that if i used 33% of time_ids for every date_id i will finally get a model worse by only like let's say 5%\ni started by training a model on the final 5 folds and the performance only decreased whenever i decrease the number of folds i.e. when i use 4 folds performance decreased\n\nNo, I mean sampling rows, not sampling time_ids. And yes, if you remove partitions the performance will decrease because the partitions are from different periods with different behaviours. If you sample out of all partitions that shouldn't be an issue.",
              "votes": 1
            },
            {
              "id": 3051288,
              "postDate": "2024-11-21T06:00:52.017Z",
              "content": "<p>Hopefully get an answer in the next 24 hours then as my first training run just finished with an R2 of 0.02 on the last 50 days holdout test set. Be interesting to see how much that craters over the evaluation set</p>",
              "rawMarkdown": "Hopefully get an answer in the next 24 hours then as my first training run just finished with an R2 of 0.02 on the last 50 days holdout test set. Be interesting to see how much that craters over the evaluation set",
              "votes": 1
            }
          ]
        }
      ]
    },
    {
      "id": 3048113,
      "postDate": "2024-11-17T15:13:37.167Z",
      "content": "<p>I found that just adding lag_1 feature simply will decrease my LB scores, we may need other method to handle the lags features. </p>",
      "rawMarkdown": "I found that just adding lag_1 feature simply will decrease my LB scores, we may need other method to handle the lags features. ",
      "votes": 2
    },
    {
      "id": 3046113,
      "postDate": "2024-11-15T06:06:18.500Z",
      "content": "<p>I agree with the fact that you shouldn't bother too much with public notebooks. Take feature ideas and try them by yourself. I have somewhat decent validation/lb correlation using my own splits, but I still can't replicate the results of public notebooks.</p>",
      "rawMarkdown": "I agree with the fact that you shouldn't bother too much with public notebooks. Take feature ideas and try them by yourself. I have somewhat decent validation/lb correlation using my own splits, but I still can't replicate the results of public notebooks.",
      "votes": 2,
      "replies": [
        {
          "id": 3046164,
          "postDate": "2024-11-15T07:08:16.317Z",
          "content": "<p>Agreed, most of the public notebooks are either absolute starters/ blind blends <a href=\"https://www.kaggle.com/gunesevitan\" target=\"_blank\">@gunesevitan</a> </p>",
          "rawMarkdown": "Agreed, most of the public notebooks are either absolute starters/ blind blends @gunesevitan ",
          "replies": [
            {
              "id": 3060227,
              "postDate": "2024-12-01T14:07:46.897Z",
              "content": "<p>Hi <a href=\"https://www.kaggle.com/ravi20076\" target=\"_blank\">@ravi20076</a>  <br>\nI do I systematically improve the model? Can you give me some adivice?</p>",
              "rawMarkdown": "Hi @ravi20076  \nI do I systematically improve the model? Can you give me some adivice?"
            }
          ]
        }
      ]
    },
    {
      "id": 3045418,
      "postDate": "2024-11-14T13:21:43.750Z",
      "content": "<p>My big problem so far is finding a good validation approach, the cv-lb relationship is erratic.</p>",
      "rawMarkdown": "My big problem so far is finding a good validation approach, the cv-lb relationship is erratic.",
      "votes": 2,
      "replies": [
        {
          "id": 3045422,
          "postDate": "2024-11-14T13:24:33.097Z",
          "content": "<p>Same here, cv-lb relation is erratic and very unpredictable <a href=\"https://www.kaggle.com/nandodmelo\" target=\"_blank\">@nandodmelo</a> </p>",
          "rawMarkdown": "Same here, cv-lb relation is erratic and very unpredictable @nandodmelo ",
          "votes": 2
        },
        {
          "id": 3045428,
          "postDate": "2024-11-14T13:29:18.713Z",
          "content": "<p><a href=\"https://www.kaggle.com/code/yunsuxiaozi/js2024-synthetic-data-with-purgedkfold/notebook\">JS2024 synthetic data with purgedkfold</a></p>\n<p>I think this can be used. Due to Kaggle's memory limitations, the data split for each fold cross validation here is relatively close in time. If the time gap can be widened a bit for cross validation, the results obtained may be more accurate.</p>\n<p>For example, validate 4 with [0,1,2,3] data, validate 6 with [2,3,4,5] data, and validate 8 with [4,5,6,7] data.</p>\n<p>Considering the memory issue, it may be necessary to use multiple notebooks to save the training and validation sets for each fold, and then load them together for cross validation.</p>",
          "rawMarkdown": "<a href=\"https://www.kaggle.com/code/yunsuxiaozi/js2024-synthetic-data-with-purgedkfold/notebook\">JS2024 synthetic data with purgedkfold</a>\n\nI think this can be used. Due to Kaggle's memory limitations, the data split for each fold cross validation here is relatively close in time. If the time gap can be widened a bit for cross validation, the results obtained may be more accurate.\n\nFor example, validate 4 with [0,1,2,3] data, validate 6 with [2,3,4,5] data, and validate 8 with [4,5,6,7] data.\n\nConsidering the memory issue, it may be necessary to use multiple notebooks to save the training and validation sets for each fold, and then load them together for cross validation.",
          "votes": 2,
          "replies": [
            {
              "id": 3045442,
              "postDate": "2024-11-14T13:40:20.663Z",
              "content": "<p><a href=\"https://www.kaggle.com/yunsuxiaozi\" target=\"_blank\">@yunsuxiaozi</a> are you getting any correlation with the lb with this strategy?</p>",
              "rawMarkdown": "@yunsuxiaozi are you getting any correlation with the lb with this strategy?",
              "votes": 1
            },
            {
              "id": 3045502,
              "postDate": "2024-11-14T14:19:39.733Z",
              "content": "<p>Due to my ignorance, I have some prejudice against synthetic data. Are your results good with this data? What is the correlation between cross-validation (CV) and leaderboard (LB) scores?</p>",
              "rawMarkdown": "Due to my ignorance, I have some prejudice against synthetic data. Are your results good with this data? What is the correlation between cross-validation (CV) and leaderboard (LB) scores?",
              "votes": 1
            },
            {
              "id": 3046092,
              "postDate": "2024-11-15T05:29:45.897Z",
              "content": "<p><a href=\"https://www.kaggle.com/nandodmelo\" target=\"_blank\">@nandodmelo</a> I don't think this is synthetic data, but it is some form of transformation on actual data. </p>",
              "rawMarkdown": "@nandodmelo I don't think this is synthetic data, but it is some form of transformation on actual data. "
            }
          ]
        }
      ]
    },
    {
      "id": 3047067,
      "postDate": "2024-11-16T08:03:07.200Z",
      "content": "<p>Regarding 7), I hadn't checked that, but perhaps we're better served predicting 0 for those? (which yields 0 score)</p>",
      "rawMarkdown": "Regarding 7), I hadn't checked that, but perhaps we're better served predicting 0 for those? (which yields 0 score)"
    },
    {
      "id": 3045783,
      "postDate": "2024-11-14T19:20:16.400Z",
      "content": "<p>I believe we might see a private LB like the one in ISIC 2024, where the winners jumped hundreds (even more than 1,000 for some) of places compared to the public LB. Similarly, those who were at the top of the public LB ended up dropping hundreds of places in the private LB.</p>",
      "rawMarkdown": "I believe we might see a private LB like the one in ISIC 2024, where the winners jumped hundreds (even more than 1,000 for some) of places compared to the public LB. Similarly, those who were at the top of the public LB ended up dropping hundreds of places in the private LB.",
      "votes": 1,
      "replies": [
        {
          "id": 3046091,
          "postDate": "2024-11-15T05:28:41.597Z",
          "content": "<p><a href=\"https://www.kaggle.com/aymanallawi\" target=\"_blank\">@aymanallawi</a> this is most likely given the nature of the problem, possible bugs in one's code leading to issues in the forecast phase and no CV-LB relation for most participants</p>",
          "rawMarkdown": "@aymanallawi this is most likely given the nature of the problem, possible bugs in one's code leading to issues in the forecast phase and no CV-LB relation for most participants",
          "replies": [
            {
              "id": 3046130,
              "postDate": "2024-11-15T06:34:08.020Z",
              "content": "<p>ISIC2024 is because private ranking data is more difficult, and this game is due to market instability. Although they are both shake, the reasons are different</p>",
              "rawMarkdown": "ISIC2024 is because private ranking data is more difficult, and this game is due to market instability. Although they are both shake, the reasons are different",
              "votes": 2
            },
            {
              "id": 3047882,
              "postDate": "2024-11-17T09:53:17.713Z",
              "content": "<p><a href=\"https://www.kaggle.com/ravi20076\" target=\"_blank\">@ravi20076</a> <br>\nWhat do you mean by refering to  <code>CV-LB relation</code>? Is there an example?</p>",
              "rawMarkdown": "@ravi20076 \nWhat do you mean by refering to  `CV-LB relation`? Is there an example?",
              "votes": 1
            },
            {
              "id": 3049032,
              "postDate": "2024-11-18T16:07:25.473Z",
              "content": "<p>The CV-LB relationship refers to the connection between your model’s cross-validation (CV) score and its leaderboard (LB) score in competitive data science platforms like Kaggle.</p>",
              "rawMarkdown": "The CV-LB relationship refers to the connection between your model’s cross-validation (CV) score and its leaderboard (LB) score in competitive data science platforms like Kaggle.",
              "votes": 2
            }
          ]
        }
      ]
    },
    {
      "id": 3074892,
      "postDate": "2024-12-18T07:11:26.273Z",
      "content": "<p>I would like to know if the final leaderboard is the same as the ranking of the winners.</p>",
      "rawMarkdown": "I would like to know if the final leaderboard is the same as the ranking of the winners.",
      "votes": -1
    },
    {
      "id": 3048634,
      "postDate": "2024-11-18T07:54:31.257Z",
      "rawMarkdown": "",
      "votes": -7,
      "isDeleted": true
    }
  ],
  "comments": [
    {
      "id": 3045411,
      "author_name": "Timmy Juicehouse",
      "author_url": "",
      "post_date": "2024-11-14T13:13:16.463000",
      "content": "<p>Besides, labels fluctuates greatly over time, and it is quite possible that this is due to market factors rather than statistical factors in the next 6-month prediction phase after the end of submissions.</p>",
      "votes": 3,
      "replies": [
        {
          "id": 3045416,
          "author_name": "Timmy Juicehouse",
          "author_url": "",
          "post_date": "2024-11-14T13:17:49.897000",
          "content": "<p>Feature engineering and cv strategy are also critical due to the time limit of submission api, but as you mentioned, sometimes it's like a gamble when meeting a good cv correlating to a good public score.</p>",
          "votes": 2,
          "replies": [
            {
              "id": 3045509,
              "author_name": "Jack",
              "author_url": "",
              "post_date": "2024-11-14T14:25:28.447000",
              "content": "<p>Polars should be utilized if doing heavy feature engineering. I need to convert my pandas to polars in submission API.</p>",
              "votes": 1,
              "replies": []
            },
            {
              "id": 3045513,
              "author_name": "Timmy Juicehouse",
              "author_url": "",
              "post_date": "2024-11-14T14:28:35.177000",
              "content": "<p>I know that. I mean below</p>\n<blockquote>\n  <p>In my previous trials, EDA/feature engineering can mislead me into overfitting situation</p>\n</blockquote>",
              "votes": 1,
              "replies": []
            },
            {
              "id": 3045537,
              "author_name": "Jack",
              "author_url": "",
              "post_date": "2024-11-14T15:00:44.307000",
              "content": "<p>Agreed - just be careful that the patterns exhibited by a feature you engineer is shown throughout most time periods, especially your validation set. Or, suffer on submission 🫠 😭</p>",
              "votes": 1,
              "replies": []
            }
          ]
        },
        {
          "id": 3045417,
          "author_name": "Ravi Ramakrishnan",
          "author_url": "",
          "post_date": "2024-11-14T13:19:05.953000",
          "content": "<p><a href=\"https://www.kaggle.com/sweetyheehee\" target=\"_blank\">@sweetyheehee</a> We are not quite sure what these labels and features mean as well, making it difficult to structure secondary features as well. Some EDA indicates certain patterns and insights, but these are indicative rather than certain. </p>",
          "votes": 0,
          "replies": [
            {
              "id": 3045429,
              "author_name": "Timmy Juicehouse",
              "author_url": "",
              "post_date": "2024-11-14T13:30:28.820000",
              "content": "<p><a href=\"https://www.kaggle.com/ravi20076\" target=\"_blank\">@ravi20076</a> sure, even difficult to train secondly and more times. In my previous trials, EDA/feature engineering can mislead me into overfitting situation  </p>",
              "votes": 1,
              "replies": []
            }
          ]
        },
        {
          "id": 3045526,
          "author_name": "Fernando Melo",
          "author_url": "",
          "post_date": "2024-11-14T14:48:53.623000",
          "content": "<p>Out of curiosity, is your best model so far NN or Boosting?</p>",
          "votes": 1,
          "replies": [
            {
              "id": 3045540,
              "author_name": "Timmy Juicehouse",
              "author_url": "",
              "post_date": "2024-11-14T15:04:15.200000",
              "content": "<p>Not relating to my best cv but good one</p>",
              "votes": 1,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 3045413,
      "author_name": "yunsuxiaozi",
      "author_url": "",
      "post_date": "2024-11-14T13:15:51.260000",
      "content": "<p>Due to the huge amount of data, it is necessary to find a balance between the amount of data and the number of features. At present, I haven't done any feature engineering yet.</p>",
      "votes": 4,
      "replies": [
        {
          "id": 3046156,
          "author_name": "ironrro",
          "author_url": "",
          "post_date": "2024-11-15T07:03:22.003000",
          "content": "<p><a href=\"https://www.kaggle.com/yunsuxiaozi\" target=\"_blank\">@yunsuxiaozi</a> <br>\nHi. What kind of experiments have you done so far if feature engineering is not relavent at this stage?</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 3047029,
      "author_name": "snehal",
      "author_url": "",
      "post_date": "2024-11-16T07:00:08.257000",
      "content": "<p>I'm mostly curious about what the top 2 on LB discovered to get such a high score. yuanzhe and rib both made huge leaps in LB score very quickly. Seems like they must have discovered some magic the rest of us haven't</p>",
      "votes": 2,
      "replies": [
        {
          "id": 3047137,
          "author_name": "Gunes Evitan",
          "author_url": "",
          "post_date": "2024-11-16T09:36:19.290000",
          "content": "<p>I noticed that too. I'm suspecting of some kind of a post processing like people were normalizing each batch on last Optiver competition.</p>",
          "votes": 3,
          "replies": [
            {
              "id": 3047531,
              "author_name": "Ayman Allawi",
              "author_url": "",
              "post_date": "2024-11-16T19:35:34.533000",
              "content": "<p>Can you please explain what do you mean. Do you mean like target regularisation</p>",
              "votes": 1,
              "replies": []
            },
            {
              "id": 3047543,
              "author_name": "JM",
              "author_url": "",
              "post_date": "2024-11-16T20:08:22.993000",
              "content": "<p>In last the Optiver competition you got a small boost by zero-meaning your prediction batch on their metric, their score is quite higher than 4th onwards so any post processing trick must be very strong</p>\n<blockquote>\n  <p><code>preds = preds - preds.mean()</code></p>\n</blockquote>",
              "votes": 8,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 3045628,
      "author_name": "Ricardo Colomer",
      "author_url": "",
      "post_date": "2024-11-14T16:25:22.583000",
      "content": "<p>I'm still stuck with this from the beginning, it is not just CV-LB relationship, small changes in the local validation produce very different results.</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 3056060,
      "author_name": "Dear_Luna",
      "author_url": "",
      "post_date": "2024-11-26T14:11:53.447000",
      "content": "<ol>\n<li>Subject to GPU and memory size, is it better to use long-term training data (using part of the features) or use all features (intercept part of the training data), and how to balance the two</li>\n<li>The model’s inference time is very demanding.</li>\n<li>I think there are serious data distribution differences, which cause the huge difference in cv-lb</li>\n<li>How to reasonably construct feature engineering to capture more information (I currently see an operation similar to positional encoding in bert)</li>\n</ol>",
      "votes": 2,
      "replies": []
    },
    {
      "id": 3050605,
      "author_name": "Michael Timbs",
      "author_url": "",
      "post_date": "2024-11-20T11:24:11.847000",
      "content": "<p>For me</p>\n<ol>\n<li><p>This is actually a very hard problem when accounting for all the constraints and lack of clarity on the data. In some ways this is good and it has certainly been a lot of fun. The lack of any kind of potential for domain knowledge means everything has to be validated in the data.</p></li>\n<li><p>Spent way too long training/playing without getting a working e2e submission working. Ended up spending 2 weeks trying to use the Darts library which is a non starter for the submission requirements here.</p></li>\n<li><p>Use of lags is likely to be a distraction. I'm pretending they don't exist for now as the complexity of safely storing and infilling lags and handling new symbols seems very high for the unknown reward</p></li>\n<li><p>The size of the data has proven tricky. Without GPU access it often takes me 12 hours to do a single run and any bug is a major setback. There are limits to how many features I can even include or the depth of search. Often will have my training process killed after 3-4 hours.</p></li>\n<li><p>I think LB will shift dramatically. For example I haven't even managed to get a proper model submitted yet (other than a naive LGBM) to understand the kaggle competition due to complexityies with CV scores being unstable and missing subtle requirements. Definitely possibly many people are still struggling through to get to a foundational baseline that they can even build from (feature pruning, feature engineering, ensembling, bigger training runs et)</p></li>\n<li><p>One I am confident in model structure and assumptions it seems inevitable that I will have to rent a cloud A100 (at minimum) to train my final submission on</p></li>\n</ol>",
      "votes": 2,
      "replies": [
        {
          "id": 3050620,
          "author_name": "Natan Labarrère",
          "author_url": "",
          "post_date": "2024-11-20T11:36:43.233000",
          "content": "<p>Some suggestions based on what I learned so far:<br>\nFor 4. You can overcome that with random sampling. It will ultimately affect model accuracy, but non linearly (33% of samples will yield maybe 95% of results)<br>\nFor 5. What most people are doing and works fine is to separate the last 100-200 days of data for testing</p>",
          "votes": 2,
          "replies": [
            {
              "id": 3050630,
              "author_name": "Ayman Allawi",
              "author_url": "",
              "post_date": "2024-11-20T11:59:51.080000",
              "content": "<blockquote>\n  <p>For 4. You can overcome that with random sampling. It will ultimately affect model accuracy, but non linearly (33% of samples will yield maybe 95% of results)</p>\n</blockquote>\n<p>do you mean that if i used 33% of time_ids for every date_id i will finally get a model worse by only like let's say 5%<br>\ni started by training a model on the final 5 folds and the performance only decreased whenever i decrease the number of folds i.e. when i use 4 folds performance decreased  </p>",
              "votes": 1,
              "replies": []
            },
            {
              "id": 3050632,
              "author_name": "Michael Timbs",
              "author_url": "",
              "post_date": "2024-11-20T12:02:13.010000",
              "content": "<blockquote>\n  <p>What most people are doing and works fine is to separate the last 100-200 days of data for testing</p>\n</blockquote>\n<p>Yes I was doing that for almost 2 weeks before I realised that the models I were training could not actually be submitted for the competition and wasted a lot of time doing it.  Since I have not yet finished training a model that will work I don't know how well my holdout R2 will match submission R2 but many are saying they are quite different</p>\n<blockquote>\n  <p>For 4. You can overcome that with random sampling. It will ultimately affect model accuracy, but non linearly (33% of samples will yield maybe 95% of results)</p>\n</blockquote>\n<p>I think I will try sampling next once I have a baseline to compare against. At current pace it will take me 5-6 days per experiment which is not feasible</p>",
              "votes": 1,
              "replies": []
            },
            {
              "id": 3050665,
              "author_name": "Natan Labarrère",
              "author_url": "",
              "post_date": "2024-11-20T12:50:44.547000",
              "content": "<blockquote>\n  <p>Yes I was doing that for almost 2 weeks before I realised that the models I were training could not actually be submitted for the competition and wasted a lot of time doing it. Since I have not yet finished training a model that will work I don't know how well my holdout R2 will match submission R2 but many are saying they are quite different</p>\n</blockquote>\n<p>They do hold out quite well for me but the thing is that you have to retrain your model with that last 100-200 days for submission. The split is just for validation purposes. If you use the same model for submission you'll have a 100-200 lagged model which does not perform well in this scenario.</p>",
              "votes": 1,
              "replies": []
            },
            {
              "id": 3050669,
              "author_name": "Natan Labarrère",
              "author_url": "",
              "post_date": "2024-11-20T12:53:54.203000",
              "content": "<blockquote>\n  <p>do you mean that if i used 33% of time_ids for every date_id i will finally get a model worse by only like let's say 5%<br>\n  i started by training a model on the final 5 folds and the performance only decreased whenever i decrease the number of folds i.e. when i use 4 folds performance decreased</p>\n</blockquote>\n<p>No, I mean sampling rows, not sampling time_ids. And yes, if you remove partitions the performance will decrease because the partitions are from different periods with different behaviours. If you sample out of all partitions that shouldn't be an issue.</p>",
              "votes": 1,
              "replies": []
            },
            {
              "id": 3051288,
              "author_name": "Michael Timbs",
              "author_url": "",
              "post_date": "2024-11-21T06:00:52.017000",
              "content": "<p>Hopefully get an answer in the next 24 hours then as my first training run just finished with an R2 of 0.02 on the last 50 days holdout test set. Be interesting to see how much that craters over the evaluation set</p>",
              "votes": 1,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 3048113,
      "author_name": "I2nfinit3y",
      "author_url": "",
      "post_date": "2024-11-17T15:13:37.167000",
      "content": "<p>I found that just adding lag_1 feature simply will decrease my LB scores, we may need other method to handle the lags features. </p>",
      "votes": 2,
      "replies": []
    },
    {
      "id": 3046113,
      "author_name": "Gunes Evitan",
      "author_url": "",
      "post_date": "2024-11-15T06:06:18.500000",
      "content": "<p>I agree with the fact that you shouldn't bother too much with public notebooks. Take feature ideas and try them by yourself. I have somewhat decent validation/lb correlation using my own splits, but I still can't replicate the results of public notebooks.</p>",
      "votes": 2,
      "replies": [
        {
          "id": 3046164,
          "author_name": "Ravi Ramakrishnan",
          "author_url": "",
          "post_date": "2024-11-15T07:08:16.317000",
          "content": "<p>Agreed, most of the public notebooks are either absolute starters/ blind blends <a href=\"https://www.kaggle.com/gunesevitan\" target=\"_blank\">@gunesevitan</a> </p>",
          "votes": 0,
          "replies": [
            {
              "id": 3060227,
              "author_name": "ironrro",
              "author_url": "",
              "post_date": "2024-12-01T14:07:46.897000",
              "content": "<p>Hi <a href=\"https://www.kaggle.com/ravi20076\" target=\"_blank\">@ravi20076</a>  <br>\nI do I systematically improve the model? Can you give me some adivice?</p>",
              "votes": 0,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 3045418,
      "author_name": "Fernando Melo",
      "author_url": "",
      "post_date": "2024-11-14T13:21:43.750000",
      "content": "<p>My big problem so far is finding a good validation approach, the cv-lb relationship is erratic.</p>",
      "votes": 2,
      "replies": [
        {
          "id": 3045422,
          "author_name": "Ravi Ramakrishnan",
          "author_url": "",
          "post_date": "2024-11-14T13:24:33.097000",
          "content": "<p>Same here, cv-lb relation is erratic and very unpredictable <a href=\"https://www.kaggle.com/nandodmelo\" target=\"_blank\">@nandodmelo</a> </p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 3045428,
          "author_name": "yunsuxiaozi",
          "author_url": "",
          "post_date": "2024-11-14T13:29:18.713000",
          "content": "<p><a href=\"https://www.kaggle.com/code/yunsuxiaozi/js2024-synthetic-data-with-purgedkfold/notebook\">JS2024 synthetic data with purgedkfold</a></p>\n<p>I think this can be used. Due to Kaggle's memory limitations, the data split for each fold cross validation here is relatively close in time. If the time gap can be widened a bit for cross validation, the results obtained may be more accurate.</p>\n<p>For example, validate 4 with [0,1,2,3] data, validate 6 with [2,3,4,5] data, and validate 8 with [4,5,6,7] data.</p>\n<p>Considering the memory issue, it may be necessary to use multiple notebooks to save the training and validation sets for each fold, and then load them together for cross validation.</p>",
          "votes": 2,
          "replies": [
            {
              "id": 3045442,
              "author_name": "Ravi Ramakrishnan",
              "author_url": "",
              "post_date": "2024-11-14T13:40:20.663000",
              "content": "<p><a href=\"https://www.kaggle.com/yunsuxiaozi\" target=\"_blank\">@yunsuxiaozi</a> are you getting any correlation with the lb with this strategy?</p>",
              "votes": 1,
              "replies": []
            },
            {
              "id": 3045502,
              "author_name": "Fernando Melo",
              "author_url": "",
              "post_date": "2024-11-14T14:19:39.733000",
              "content": "<p>Due to my ignorance, I have some prejudice against synthetic data. Are your results good with this data? What is the correlation between cross-validation (CV) and leaderboard (LB) scores?</p>",
              "votes": 1,
              "replies": []
            },
            {
              "id": 3046092,
              "author_name": "Ravi Ramakrishnan",
              "author_url": "",
              "post_date": "2024-11-15T05:29:45.897000",
              "content": "<p><a href=\"https://www.kaggle.com/nandodmelo\" target=\"_blank\">@nandodmelo</a> I don't think this is synthetic data, but it is some form of transformation on actual data. </p>",
              "votes": 0,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 3047067,
      "author_name": "Natan Labarrère",
      "author_url": "",
      "post_date": "2024-11-16T08:03:07.200000",
      "content": "<p>Regarding 7), I hadn't checked that, but perhaps we're better served predicting 0 for those? (which yields 0 score)</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 3045783,
      "author_name": "Ayman Allawi",
      "author_url": "",
      "post_date": "2024-11-14T19:20:16.400000",
      "content": "<p>I believe we might see a private LB like the one in ISIC 2024, where the winners jumped hundreds (even more than 1,000 for some) of places compared to the public LB. Similarly, those who were at the top of the public LB ended up dropping hundreds of places in the private LB.</p>",
      "votes": 1,
      "replies": [
        {
          "id": 3046091,
          "author_name": "Ravi Ramakrishnan",
          "author_url": "",
          "post_date": "2024-11-15T05:28:41.597000",
          "content": "<p><a href=\"https://www.kaggle.com/aymanallawi\" target=\"_blank\">@aymanallawi</a> this is most likely given the nature of the problem, possible bugs in one's code leading to issues in the forecast phase and no CV-LB relation for most participants</p>",
          "votes": 0,
          "replies": [
            {
              "id": 3046130,
              "author_name": "yunsuxiaozi",
              "author_url": "",
              "post_date": "2024-11-15T06:34:08.020000",
              "content": "<p>ISIC2024 is because private ranking data is more difficult, and this game is due to market instability. Although they are both shake, the reasons are different</p>",
              "votes": 2,
              "replies": []
            },
            {
              "id": 3047882,
              "author_name": "ironrro",
              "author_url": "",
              "post_date": "2024-11-17T09:53:17.713000",
              "content": "<p><a href=\"https://www.kaggle.com/ravi20076\" target=\"_blank\">@ravi20076</a> <br>\nWhat do you mean by refering to  <code>CV-LB relation</code>? Is there an example?</p>",
              "votes": 1,
              "replies": []
            },
            {
              "id": 3049032,
              "author_name": "Nikhilesh Belulkar",
              "author_url": "",
              "post_date": "2024-11-18T16:07:25.473000",
              "content": "<p>The CV-LB relationship refers to the connection between your model’s cross-validation (CV) score and its leaderboard (LB) score in competitive data science platforms like Kaggle.</p>",
              "votes": 2,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 3074892,
      "author_name": "Qi",
      "author_url": "",
      "post_date": "2024-12-18T07:11:26.273000",
      "content": "<p>I would like to know if the final leaderboard is the same as the ranking of the winners.</p>",
      "votes": -1,
      "replies": []
    },
    {
      "id": 3048634,
      "author_name": "",
      "author_url": "",
      "post_date": "2024-11-18T07:54:31.257000",
      "content": "",
      "votes": -7,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "3045158": "Hello all,\n\nWe are exactly 1 month into the competition as on date (14-11-2024). I have learnt the below from my tryst with the competition insofar-\n\n1. This competition is very different from the rest of the time series competitions recently completed on Kaggle and in my opinion, far more difficult than the ones gone by\n2. Data size, timing constraints and inability to train models on Kaggle kernels with the entire training data is a big challenge\n3. It is very easy for one to get swayed with score clusters on the leaderboard and tread the wrong path. Such approaches are often met with disappointment later and don't work well with assignments like these\n4. CV-LB relation establishment is perhaps the biggest and herculean task here. Not all approaches have yielded good CV-Lb relations\n5. Conventional CV alignment with the LB is not likely to yield good results - one needs to be creative here\n6. The metric values (both CV score and LB score) are very low and perhaps not completely reliable, some creativity and out-of-box thinking is needed to make sense of these values \n7. Certain symbols are extremely hard to predict and are worse off with a model than conjecture\n\nThoughts? Comments?",
    "3045411": "Besides, labels fluctuates greatly over time, and it is quite possible that this is due to market factors rather than statistical factors in the next 6-month prediction phase after the end of submissions.",
    "3045413": "Due to the huge amount of data, it is necessary to find a balance between the amount of data and the number of features. At present, I haven't done any feature engineering yet.",
    "3047029": "I'm mostly curious about what the top 2 on LB discovered to get such a high score. yuanzhe and rib both made huge leaps in LB score very quickly. Seems like they must have discovered some magic the rest of us haven't",
    "3045628": "I'm still stuck with this from the beginning, it is not just CV-LB relationship, small changes in the local validation produce very different results.",
    "3056060": "1. Subject to GPU and memory size, is it better to use long-term training data (using part of the features) or use all features (intercept part of the training data), and how to balance the two\n2. The model’s inference time is very demanding.\n3. I think there are serious data distribution differences, which cause the huge difference in cv-lb\n4. How to reasonably construct feature engineering to capture more information (I currently see an operation similar to positional encoding in bert)",
    "3050605": "For me\n\n0. This is actually a very hard problem when accounting for all the constraints and lack of clarity on the data. In some ways this is good and it has certainly been a lot of fun. The lack of any kind of potential for domain knowledge means everything has to be validated in the data.\n\n1. Spent way too long training/playing without getting a working e2e submission working. Ended up spending 2 weeks trying to use the Darts library which is a non starter for the submission requirements here.\n\n2. Use of lags is likely to be a distraction. I'm pretending they don't exist for now as the complexity of safely storing and infilling lags and handling new symbols seems very high for the unknown reward\n\n3. The size of the data has proven tricky. Without GPU access it often takes me 12 hours to do a single run and any bug is a major setback. There are limits to how many features I can even include or the depth of search. Often will have my training process killed after 3-4 hours.\n\n4. I think LB will shift dramatically. For example I haven't even managed to get a proper model submitted yet (other than a naive LGBM) to understand the kaggle competition due to complexityies with CV scores being unstable and missing subtle requirements. Definitely possibly many people are still struggling through to get to a foundational baseline that they can even build from (feature pruning, feature engineering, ensembling, bigger training runs et)\n\n5. One I am confident in model structure and assumptions it seems inevitable that I will have to rent a cloud A100 (at minimum) to train my final submission on",
    "3048113": "I found that just adding lag_1 feature simply will decrease my LB scores, we may need other method to handle the lags features. ",
    "3046113": "I agree with the fact that you shouldn't bother too much with public notebooks. Take feature ideas and try them by yourself. I have somewhat decent validation/lb correlation using my own splits, but I still can't replicate the results of public notebooks.",
    "3045418": "My big problem so far is finding a good validation approach, the cv-lb relationship is erratic.",
    "3047067": "Regarding 7), I hadn't checked that, but perhaps we're better served predicting 0 for those? (which yields 0 score)",
    "3045783": "I believe we might see a private LB like the one in ISIC 2024, where the winners jumped hundreds (even more than 1,000 for some) of places compared to the public LB. Similarly, those who were at the top of the public LB ended up dropping hundreds of places in the private LB.",
    "3074892": "I would like to know if the final leaderboard is the same as the ranking of the winners.",
    "3048634": ""
  }
}