{
  "id": 548596,
  "title": "Why Does Training on the Full Dataset Lead to Worse Model Performance?",
  "url": "/competitions/jane-street-real-time-market-data-forecasting/discussion/548596",
  "author_name": "",
  "post_date": "2024-11-27T16:29:37.922308800Z",
  "votes": null,
  "comment_count": 5,
  "views": 0,
  "content": "<p>I am currently using the last 4.5M rows of the dataset as a validation set during training. This matches the number of rows we are required to predict during inference. Once I identify the best-performing model parameters using this validation strategy, I retrain my model on the entire dataset, including these last 4.5M rows.</p>\n<p>Surprisingly, the model trained on the full dataset (including the most recent 4.5M rows) underperforms compared to the model trained only on the earlier data (excluding the last 4.5M rows).</p>\n<ol>\n<li>Could this be due to overfitting, as the full-dataset model lacks a dedicated validation set to guide training?</li>\n<li>Would it make sense to use a sliding window approach instead? For example: Train on rows from dates 1–80, using 81–100 as validation. Then, train on rows from dates 1–90, using 91–100 as validation.<br>\nI would appreciate any insights into why this is happening and suggestions, thank you !!</li>\n</ol>",
  "messages": [
    {
      "id": "3057031",
      "postDate": "11/27/2024 16:29:37",
      "content": "<p>I am currently using the last 4.5M rows of the dataset as a validation set during training. This matches the number of rows we are required to predict during inference. Once I identify the best-performing model parameters using this validation strategy, I retrain my model on the entire dataset, including these last 4.5M rows.</p>\n<p>Surprisingly, the model trained on the full dataset (including the most recent 4.5M rows) underperforms compared to the model trained only on the earlier data (excluding the last 4.5M rows).</p>\n<ol>\n<li>Could this be due to overfitting, as the full-dataset model lacks a dedicated validation set to guide training?</li>\n<li>Would it make sense to use a sliding window approach instead? For example: Train on rows from dates 1–80, using 81–100 as validation. Then, train on rows from dates 1–90, using 91–100 as validation.<br>\nI would appreciate any insights into why this is happening and suggestions, thank you !!</li>\n</ol>",
      "rawMarkdown": "I am currently using the last 4.5M rows of the dataset as a validation set during training. This matches the number of rows we are required to predict during inference. Once I identify the best-performing model parameters using this validation strategy, I retrain my model on the entire dataset, including these last 4.5M rows.\n\nSurprisingly, the model trained on the full dataset (including the most recent 4.5M rows) underperforms compared to the model trained only on the earlier data (excluding the last 4.5M rows).\n1.  Could this be due to overfitting, as the full-dataset model lacks a dedicated validation set to guide training?\n2. Would it make sense to use a sliding window approach instead? For example: Train on rows from dates 1–80, using 81–100 as validation. Then, train on rows from dates 1–90, using 91–100 as validation.\nI would appreciate any insights into why this is happening and suggestions, thank you !!",
      "votes": null
    },
    {
      "id": "3057039",
      "postDate": "11/27/2024 16:33:42",
      "content": "<p>It's possible that the underlying distribution has shifted over time, so adding older data is pushing your model in the wrong direction</p>",
      "rawMarkdown": "It's possible that the underlying distribution has shifted over time, so adding older data is pushing your model in the wrong direction",
      "votes": null
    },
    {
      "id": "3057054",
      "postDate": "11/27/2024 16:52:04",
      "content": "<p>It's happening the exact opposite, that is: the model trained only on the older data (all dataset without the last 4.5M rows) Outperforms the model trained with all the data. So i'm not adding older data, I'm adding newer data, that's why I'm baffled.</p>",
      "rawMarkdown": "It's happening the exact opposite, that is: the model trained only on the older data (all dataset without the last 4.5M rows) Outperforms the model trained with all the data. So i'm not adding older data, I'm adding newer data, that's why I'm baffled.",
      "votes": null
    },
    {
      "id": "3057082",
      "postDate": "11/27/2024 17:25:32",
      "content": "<p>So a model trained without the last 4.5M rows has better LB score than a model trained with the last 4.5M rows appended to the training set? The reason could still be the same (training set has a closer distribution to the LB test set without the 4.5M rows) and/or the model was overtrained. </p>",
      "rawMarkdown": "So a model trained without the last 4.5M rows has better LB score than a model trained with the last 4.5M rows appended to the training set? The reason could still be the same (training set has a closer distribution to the LB test set without the 4.5M rows) and/or the model was overtrained.",
      "votes": null
    },
    {
      "id": "3057128",
      "postDate": "11/27/2024 18:29:55",
      "content": "<p>Isn't it peculiar that the distribution of past data aligns more closely with the LB test than when combining past data with recent data? Could this suggest that the last 4.5M rows contain some sort of market anomaly?</p>",
      "rawMarkdown": "Isn't it peculiar that the distribution of past data aligns more closely with the LB test than when combining past data with recent data? Could this suggest that the last 4.5M rows contain some sort of market anomaly?",
      "votes": null
    },
    {
      "id": "3057230",
      "postDate": "11/27/2024 22:10:42",
      "content": "<p>What you’re describing is the nature of non-stationary data; the distribution parameters change from time period to time period. Taking advantage of the fact that including the last 4.5M rows in the training set decreases model performance in LB sounds like curve fitting IMHO. </p>",
      "rawMarkdown": "What you’re describing is the nature of non-stationary data; the distribution parameters change from time period to time period. Taking advantage of the fact that including the last 4.5M rows in the training set decreases model performance in LB sounds like curve fitting IMHO.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 3057039,
      "author_name": "redfoongus",
      "author_url": "",
      "post_date": "11/27/2024 16:33:42",
      "content": "<p>It's possible that the underlying distribution has shifted over time, so adding older data is pushing your model in the wrong direction</p>",
      "votes": null,
      "replies": [
        {
          "id": 3057054,
          "author_name": "margaritelli",
          "author_url": "",
          "post_date": "11/27/2024 16:52:04",
          "content": "<p>It's happening the exact opposite, that is: the model trained only on the older data (all dataset without the last 4.5M rows) Outperforms the model trained with all the data. So i'm not adding older data, I'm adding newer data, that's why I'm baffled.</p>",
          "votes": null,
          "replies": [
            {
              "id": 3057082,
              "author_name": "redfoongus",
              "author_url": "",
              "post_date": "11/27/2024 17:25:32",
              "content": "<p>So a model trained without the last 4.5M rows has better LB score than a model trained with the last 4.5M rows appended to the training set? The reason could still be the same (training set has a closer distribution to the LB test set without the 4.5M rows) and/or the model was overtrained. </p>",
              "votes": null,
              "replies": [
                {
                  "id": 3057128,
                  "author_name": "margaritelli",
                  "author_url": "",
                  "post_date": "11/27/2024 18:29:55",
                  "content": "<p>Isn't it peculiar that the distribution of past data aligns more closely with the LB test than when combining past data with recent data? Could this suggest that the last 4.5M rows contain some sort of market anomaly?</p>",
                  "votes": null,
                  "replies": [
                    {
                      "id": 3057230,
                      "author_name": "maciejzawadzki",
                      "author_url": "",
                      "post_date": "11/27/2024 22:10:42",
                      "content": "<p>What you’re describing is the nature of non-stationary data; the distribution parameters change from time period to time period. Taking advantage of the fact that including the last 4.5M rows in the training set decreases model performance in LB sounds like curve fitting IMHO. </p>",
                      "votes": null,
                      "replies": []
                    }
                  ]
                }
              ]
            }
          ]
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "3057031": "I am currently using the last 4.5M rows of the dataset as a validation set during training. This matches the number of rows we are required to predict during inference. Once I identify the best-performing model parameters using this validation strategy, I retrain my model on the entire dataset, including these last 4.5M rows.\n\nSurprisingly, the model trained on the full dataset (including the most recent 4.5M rows) underperforms compared to the model trained only on the earlier data (excluding the last 4.5M rows).\n1.  Could this be due to overfitting, as the full-dataset model lacks a dedicated validation set to guide training?\n2. Would it make sense to use a sliding window approach instead? For example: Train on rows from dates 1–80, using 81–100 as validation. Then, train on rows from dates 1–90, using 91–100 as validation.\nI would appreciate any insights into why this is happening and suggestions, thank you !!",
    "3057039": "It's possible that the underlying distribution has shifted over time, so adding older data is pushing your model in the wrong direction",
    "3057054": "It's happening the exact opposite, that is: the model trained only on the older data (all dataset without the last 4.5M rows) Outperforms the model trained with all the data. So i'm not adding older data, I'm adding newer data, that's why I'm baffled.",
    "3057082": "So a model trained without the last 4.5M rows has better LB score than a model trained with the last 4.5M rows appended to the training set? The reason could still be the same (training set has a closer distribution to the LB test set without the 4.5M rows) and/or the model was overtrained.",
    "3057128": "Isn't it peculiar that the distribution of past data aligns more closely with the LB test than when combining past data with recent data? Could this suggest that the last 4.5M rows contain some sort of market anomaly?",
    "3057230": "What you’re describing is the nature of non-stationary data; the distribution parameters change from time period to time period. Taking advantage of the fact that including the last 4.5M rows in the training set decreases model performance in LB sounds like curve fitting IMHO."
  },
  "source": "meta"
}