{
  "id": 543028,
  "title": "Validation Strategy & CV vs LB ",
  "url": "/competitions/jane-street-real-time-market-data-forecasting/discussion/543028",
  "author_name": "Ayman Allawi",
  "post_date": "2024-10-28T09:47:57.810000",
  "votes": 2,
  "comment_count": 11,
  "views": 0,
  "content": "<p>I used the last 6 million rows as my validation set. I picked this approach because top solutions in time series competitions often use it. My cross-validation score was R² = 0.8, but my leaderboard score dropped to 0.36. I’m wondering if this gap is normal or if I might have issues like data leakage or overfitting in my process that could cause this. I’d appreciate your thoughts on this gap and if you’ve faced a similar issue. Also, I’d like your opinion on my validation strategy and if it might need adjustment.</p>",
  "messages": [
    {
      "id": 3030217,
      "postDate": "2024-10-28T10:00:08.010Z",
      "content": "<p>This is quite normal in time series forecast (specially in financial time series forecast). If you use a different vintage as validation dataset and the data before that as train, you will see the validation score changes quite a bit. This is quite the nature of financial time series forecast. It's super volatile and not supposed to be easy to do. </p>",
      "rawMarkdown": "This is quite normal in time series forecast (specially in financial time series forecast). If you use a different vintage as validation dataset and the data before that as train, you will see the validation score changes quite a bit. This is quite the nature of financial time series forecast. It's super volatile and not supposed to be easy to do. ",
      "votes": 3,
      "replies": [
        {
          "id": 3030231,
          "postDate": "2024-10-28T10:20:16.583Z",
          "content": "<p>may I ask what is your cv and lb? you are right that such a gap is normal but i just want to make sure that there is no bug in my workflow that causing the gap and it is the result of time series data being time series data </p>",
          "rawMarkdown": "may I ask what is your cv and lb? you are right that such a gap is normal but i just want to make sure that there is no bug in my workflow that causing the gap and it is the result of time series data being time series data ",
          "replies": [
            {
              "id": 3030245,
              "postDate": "2024-10-28T10:37:52.313Z",
              "content": "<p>I used a very different train/val split compared to yours, so the score is not a really good reference for you to compare with. But my current score is achieved from a single model. </p>",
              "rawMarkdown": "I used a very different train/val split compared to yours, so the score is not a really good reference for you to compare with. But my current score is achieved from a single model. ",
              "votes": 1
            },
            {
              "id": 3030366,
              "postDate": "2024-10-28T13:07:52.760Z",
              "content": "<p>Hi <a href=\"https://www.kaggle.com/lihaorocky\" target=\"_blank\">@lihaorocky</a> </p>\n<p>It's quite amazing to see your ranking with a single model. May I ask if you have used only the native features, or there are also derivative features from feature engineering (assume you are using / have used GBDT models)? I am only curious if that make sense to make effort on FE with anonymous features.</p>\n<p>Thanks!</p>",
              "rawMarkdown": "Hi @lihaorocky \n\nIt's quite amazing to see your ranking with a single model. May I ask if you have used only the native features, or there are also derivative features from feature engineering (assume you are using / have used GBDT models)? I am only curious if that make sense to make effort on FE with anonymous features.\n\nThanks!"
            },
            {
              "id": 3030395,
              "postDate": "2024-10-28T13:37:23.043Z",
              "content": "<p>I think FE can still be one of the many possible keys in this competition. </p>",
              "rawMarkdown": "I think FE can still be one of the many possible keys in this competition. ",
              "votes": 2
            }
          ]
        }
      ]
    },
    {
      "id": 3030361,
      "postDate": "2024-10-28T13:01:00.150Z",
      "content": "<pre><code>df[] = np.nan\ndf.loc[df[] &lt; , ] = \ndf.loc[(df[] &gt;= ) &amp; (df[] &lt;= ), ] = \n</code></pre>\n<p>I think your gap is normal. This is one of my folds and it's very similar to yours. I have 5.6m samples in it and I have a very similar gap as well. That fold's validation score is 0.0089 and lb score is 0.0041.</p>",
      "rawMarkdown": "```python\ndf['fold1'] = np.nan\ndf.loc[df['date_id'] < 1548, 'fold1'] = 0\ndf.loc[(df['date_id'] >= 1548) & (df['date_id'] <= 1698), 'fold1'] = 1\n```\n\nI think your gap is normal. This is one of my folds and it's very similar to yours. I have 5.6m samples in it and I have a very similar gap as well. That fold's validation score is 0.0089 and lb score is 0.0041.",
      "votes": 1,
      "replies": [
        {
          "id": 3030378,
          "postDate": "2024-10-28T13:20:04.013Z",
          "content": "<p>May I ask how much training data and iterators do you usually use?</p>",
          "rawMarkdown": "May I ask how much training data and iterators do you usually use?"
        }
      ]
    },
    {
      "id": 3030239,
      "postDate": "2024-10-28T10:30:01.943Z",
      "content": "<p>from what I have seen, half of the training data is enough for others to produce a .007 cv .004 lb on a untuned baseline model<br>\ncan you try training your model with partition, say 4 to 8, then produce a cv score for partition 0 to 3 and 9?</p>\n<p>I have a bias against traditional CV split for temporal model too, and the 3rd place solution here <a href=\"https://www.kaggle.com/competitions/jane-street-market-prediction/discussion/224713\" target=\"_blank\">https://www.kaggle.com/competitions/jane-street-market-prediction/discussion/224713</a> uses a similar CV split strategy.(meaning it does work or it can work)</p>\n<p>I dont think what you are seeing is considered overfitting(CV vs lb score gap).</p>\n<p>generally I think this is because of a shift of pattern and data statistical properties, an existence of gap is normal but if your gap falls outside of other people's observation I think there is a problem with either the model or preprocessing code.</p>\n<p>the LARGEST gap that has been reported is around .003 gap on a baseline model against your .0043 gap <br>\n<a href=\"https://www.kaggle.com/competitions/jane-street-real-time-market-data-forecasting/discussion/540845\" target=\"_blank\">https://www.kaggle.com/competitions/jane-street-real-time-market-data-forecasting/discussion/540845</a></p>\n<p>Not addressing it can be problematic since the accuracy and direction of any optimization made to your model is more likely to be wrong than others </p>",
      "rawMarkdown": "from what I have seen, half of the training data is enough for others to produce a .007 cv .004 lb on a untuned baseline model\ncan you try training your model with partition, say 4 to 8, then produce a cv score for partition 0 to 3 and 9?\n\n\nI have a bias against traditional CV split for temporal model too, and the 3rd place solution here https://www.kaggle.com/competitions/jane-street-market-prediction/discussion/224713 uses a similar CV split strategy.(meaning it does work or it can work)\n\nI dont think what you are seeing is considered overfitting(CV vs lb score gap).\n\ngenerally I think this is because of a shift of pattern and data statistical properties, an existence of gap is normal but if your gap falls outside of other people's observation I think there is a problem with either the model or preprocessing code.\n\nthe LARGEST gap that has been reported is around .003 gap on a baseline model against your .0043 gap \nhttps://www.kaggle.com/competitions/jane-street-real-time-market-data-forecasting/discussion/540845\n\nNot addressing it can be problematic since the accuracy and direction of any optimization made to your model is more likely to be wrong than others ",
      "votes": 1
    },
    {
      "id": 3030202,
      "postDate": "2024-10-28T09:47:57.810Z",
      "content": "<p>I used the last 6 million rows as my validation set. I picked this approach because top solutions in time series competitions often use it. My cross-validation score was R² = 0.8, but my leaderboard score dropped to 0.36. I’m wondering if this gap is normal or if I might have issues like data leakage or overfitting in my process that could cause this. I’d appreciate your thoughts on this gap and if you’ve faced a similar issue. Also, I’d like your opinion on my validation strategy and if it might need adjustment.</p>",
      "rawMarkdown": "I used the last 6 million rows as my validation set. I picked this approach because top solutions in time series competitions often use it. My cross-validation score was R² = 0.8, but my leaderboard score dropped to 0.36. I’m wondering if this gap is normal or if I might have issues like data leakage or overfitting in my process that could cause this. I’d appreciate your thoughts on this gap and if you’ve faced a similar issue. Also, I’d like your opinion on my validation strategy and if it might need adjustment.",
      "votes": 2
    },
    {
      "id": 3030280,
      "postDate": "2024-10-28T11:16:54.043Z",
      "content": "<p>You can check this notebook from previous \"Jane Street Market Prediction\" competition<br>\n<a href=\"https://www.kaggle.com/code/marketneutral/purged-time-series-cv-xgboost-optuna\" target=\"_blank\">PurgedTimeSeriesSplit </a></p>",
      "rawMarkdown": "\nYou can check this notebook from previous \"Jane Street Market Prediction\" competition\n[PurgedTimeSeriesSplit ](https://www.kaggle.com/code/marketneutral/purged-time-series-cv-xgboost-optuna)"
    },
    {
      "id": 3030317,
      "postDate": "2024-10-28T12:18:51.237Z",
      "rawMarkdown": "",
      "votes": -2,
      "isDeleted": true,
      "replies": [
        {
          "id": 3030382,
          "postDate": "2024-10-28T13:23:57.210Z",
          "content": "<p>Ravi he took out 2 0s from his result : )</p>",
          "rawMarkdown": "Ravi he took out 2 0s from his result : )"
        }
      ]
    }
  ],
  "comments": [
    {
      "id": 3030217,
      "author_name": "HAO",
      "author_url": "",
      "post_date": "2024-10-28T10:00:08.010000",
      "content": "<p>This is quite normal in time series forecast (specially in financial time series forecast). If you use a different vintage as validation dataset and the data before that as train, you will see the validation score changes quite a bit. This is quite the nature of financial time series forecast. It's super volatile and not supposed to be easy to do. </p>",
      "votes": 3,
      "replies": [
        {
          "id": 3030231,
          "author_name": "Ayman Allawi",
          "author_url": "",
          "post_date": "2024-10-28T10:20:16.583000",
          "content": "<p>may I ask what is your cv and lb? you are right that such a gap is normal but i just want to make sure that there is no bug in my workflow that causing the gap and it is the result of time series data being time series data </p>",
          "votes": 0,
          "replies": [
            {
              "id": 3030245,
              "author_name": "HAO",
              "author_url": "",
              "post_date": "2024-10-28T10:37:52.313000",
              "content": "<p>I used a very different train/val split compared to yours, so the score is not a really good reference for you to compare with. But my current score is achieved from a single model. </p>",
              "votes": 1,
              "replies": []
            },
            {
              "id": 3030366,
              "author_name": "SLi",
              "author_url": "",
              "post_date": "2024-10-28T13:07:52.760000",
              "content": "<p>Hi <a href=\"https://www.kaggle.com/lihaorocky\" target=\"_blank\">@lihaorocky</a> </p>\n<p>It's quite amazing to see your ranking with a single model. May I ask if you have used only the native features, or there are also derivative features from feature engineering (assume you are using / have used GBDT models)? I am only curious if that make sense to make effort on FE with anonymous features.</p>\n<p>Thanks!</p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 3030395,
              "author_name": "HAO",
              "author_url": "",
              "post_date": "2024-10-28T13:37:23.043000",
              "content": "<p>I think FE can still be one of the many possible keys in this competition. </p>",
              "votes": 2,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 3030361,
      "author_name": "Gunes Evitan",
      "author_url": "",
      "post_date": "2024-10-28T13:01:00.150000",
      "content": "<pre><code>df[] = np.nan\ndf.loc[df[] &lt; , ] = \ndf.loc[(df[] &gt;= ) &amp; (df[] &lt;= ), ] = \n</code></pre>\n<p>I think your gap is normal. This is one of my folds and it's very similar to yours. I have 5.6m samples in it and I have a very similar gap as well. That fold's validation score is 0.0089 and lb score is 0.0041.</p>",
      "votes": 1,
      "replies": [
        {
          "id": 3030378,
          "author_name": "yunsuxiaozi",
          "author_url": "",
          "post_date": "2024-10-28T13:20:04.013000",
          "content": "<p>May I ask how much training data and iterators do you usually use?</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 3030239,
      "author_name": "dc260123",
      "author_url": "",
      "post_date": "2024-10-28T10:30:01.943000",
      "content": "<p>from what I have seen, half of the training data is enough for others to produce a .007 cv .004 lb on a untuned baseline model<br>\ncan you try training your model with partition, say 4 to 8, then produce a cv score for partition 0 to 3 and 9?</p>\n<p>I have a bias against traditional CV split for temporal model too, and the 3rd place solution here <a href=\"https://www.kaggle.com/competitions/jane-street-market-prediction/discussion/224713\" target=\"_blank\">https://www.kaggle.com/competitions/jane-street-market-prediction/discussion/224713</a> uses a similar CV split strategy.(meaning it does work or it can work)</p>\n<p>I dont think what you are seeing is considered overfitting(CV vs lb score gap).</p>\n<p>generally I think this is because of a shift of pattern and data statistical properties, an existence of gap is normal but if your gap falls outside of other people's observation I think there is a problem with either the model or preprocessing code.</p>\n<p>the LARGEST gap that has been reported is around .003 gap on a baseline model against your .0043 gap <br>\n<a href=\"https://www.kaggle.com/competitions/jane-street-real-time-market-data-forecasting/discussion/540845\" target=\"_blank\">https://www.kaggle.com/competitions/jane-street-real-time-market-data-forecasting/discussion/540845</a></p>\n<p>Not addressing it can be problematic since the accuracy and direction of any optimization made to your model is more likely to be wrong than others </p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 3030280,
      "author_name": "Ricardo Colomer",
      "author_url": "",
      "post_date": "2024-10-28T11:16:54.043000",
      "content": "<p>You can check this notebook from previous \"Jane Street Market Prediction\" competition<br>\n<a href=\"https://www.kaggle.com/code/marketneutral/purged-time-series-cv-xgboost-optuna\" target=\"_blank\">PurgedTimeSeriesSplit </a></p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 3030317,
      "author_name": "",
      "author_url": "",
      "post_date": "2024-10-28T12:18:51.237000",
      "content": "",
      "votes": -2,
      "replies": [
        {
          "id": 3030382,
          "author_name": "dc260123",
          "author_url": "",
          "post_date": "2024-10-28T13:23:57.210000",
          "content": "<p>Ravi he took out 2 0s from his result : )</p>",
          "votes": 0,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "3030217": "This is quite normal in time series forecast (specially in financial time series forecast). If you use a different vintage as validation dataset and the data before that as train, you will see the validation score changes quite a bit. This is quite the nature of financial time series forecast. It's super volatile and not supposed to be easy to do. ",
    "3030361": "```python\ndf['fold1'] = np.nan\ndf.loc[df['date_id'] < 1548, 'fold1'] = 0\ndf.loc[(df['date_id'] >= 1548) & (df['date_id'] <= 1698), 'fold1'] = 1\n```\n\nI think your gap is normal. This is one of my folds and it's very similar to yours. I have 5.6m samples in it and I have a very similar gap as well. That fold's validation score is 0.0089 and lb score is 0.0041.",
    "3030239": "from what I have seen, half of the training data is enough for others to produce a .007 cv .004 lb on a untuned baseline model\ncan you try training your model with partition, say 4 to 8, then produce a cv score for partition 0 to 3 and 9?\n\n\nI have a bias against traditional CV split for temporal model too, and the 3rd place solution here https://www.kaggle.com/competitions/jane-street-market-prediction/discussion/224713 uses a similar CV split strategy.(meaning it does work or it can work)\n\nI dont think what you are seeing is considered overfitting(CV vs lb score gap).\n\ngenerally I think this is because of a shift of pattern and data statistical properties, an existence of gap is normal but if your gap falls outside of other people's observation I think there is a problem with either the model or preprocessing code.\n\nthe LARGEST gap that has been reported is around .003 gap on a baseline model against your .0043 gap \nhttps://www.kaggle.com/competitions/jane-street-real-time-market-data-forecasting/discussion/540845\n\nNot addressing it can be problematic since the accuracy and direction of any optimization made to your model is more likely to be wrong than others ",
    "3030202": "I used the last 6 million rows as my validation set. I picked this approach because top solutions in time series competitions often use it. My cross-validation score was R² = 0.8, but my leaderboard score dropped to 0.36. I’m wondering if this gap is normal or if I might have issues like data leakage or overfitting in my process that could cause this. I’d appreciate your thoughts on this gap and if you’ve faced a similar issue. Also, I’d like your opinion on my validation strategy and if it might need adjustment.",
    "3030280": "\nYou can check this notebook from previous \"Jane Street Market Prediction\" competition\n[PurgedTimeSeriesSplit ](https://www.kaggle.com/code/marketneutral/purged-time-series-cv-xgboost-optuna)",
    "3030317": ""
  }
}