{
  "id": 546392,
  "title": "Offline cross validation dataset",
  "url": "/competitions/jane-street-real-time-market-data-forecasting/discussion/546392",
  "author_name": "yunsuxiaozi",
  "post_date": "2024-11-15T13:46:08.816000",
  "votes": -7,
  "comment_count": 6,
  "views": 0,
  "content": "<p><a href=\"https://www.kaggle.com/code/yunsuxiaozi/js2024-synthetic-data-for-offline-cv/\">Here</a> is dataset.</p>\n<p><a href=\"https://www.kaggle.com/code/yunsuxiaozi/js2024-synthetic-data-with-purgedkfold/notebook\">Purgedkfold</a></p>",
  "messages": [
    {
      "id": 3046459,
      "postDate": "2024-11-15T13:46:08.817Z",
      "content": "<p><a href=\"https://www.kaggle.com/code/yunsuxiaozi/js2024-synthetic-data-for-offline-cv/\">Here</a> is dataset.</p>\n<p><a href=\"https://www.kaggle.com/code/yunsuxiaozi/js2024-synthetic-data-with-purgedkfold/notebook\">Purgedkfold</a></p>",
      "rawMarkdown": "<a href=\"https://www.kaggle.com/code/yunsuxiaozi/js2024-synthetic-data-for-offline-cv/\">Here</a> is dataset.\n\n<a href=\"https://www.kaggle.com/code/yunsuxiaozi/js2024-synthetic-data-with-purgedkfold/notebook\">Purgedkfold</a>",
      "votes": -7
    },
    {
      "id": 3046623,
      "postDate": "2024-11-15T16:55:20.107Z",
      "content": "<p>I don't know why this technique is so popular on Kaggle. I would never train on future and validate on past.</p>",
      "rawMarkdown": "I don't know why this technique is so popular on Kaggle. I would never train on future and validate on past.",
      "replies": [
        {
          "id": 3046847,
          "postDate": "2024-11-16T00:16:43.600Z",
          "content": "<p>May I ask why you said 'train on future and validate on past'? Is there a mistake in my code? I divided the training set and validation set in chronological order. Do you have any better cross validation methods?</p>",
          "rawMarkdown": "May I ask why you said 'train on future and validate on past'? Is there a mistake in my code? I divided the training set and validation set in chronological order. Do you have any better cross validation methods?",
          "replies": [
            {
              "id": 3046936,
              "postDate": "2024-11-16T04:16:34.627Z",
              "content": "<p>I saw your link to purged kfold and read this. <a href=\"https://stats.stackexchange.com/questions/443159/what-is-combinatorial-purged-cross-validation-for-time-series-data\" target=\"_blank\">https://stats.stackexchange.com/questions/443159/what-is-combinatorial-purged-cross-validation-for-time-series-data</a></p>\n<p>The explanation shows training data can be ahead of validation data.</p>",
              "rawMarkdown": "I saw your link to purged kfold and read this. https://stats.stackexchange.com/questions/443159/what-is-combinatorial-purged-cross-validation-for-time-series-data\n\nThe explanation shows training data can be ahead of validation data."
            }
          ]
        },
        {
          "id": 3047857,
          "postDate": "2024-11-17T09:21:18.260Z",
          "content": "<p>It was popularised in <a href=\"https://www.amazon.com/Advances-Financial-Machine-Learning-Marcos/dp/1119482089\" target=\"_blank\">https://www.amazon.com/Advances-Financial-Machine-Learning-Marcos/dp/1119482089</a> I think its usefullness might depends on the task. In a very low signal to noise environnement it might introduce more overfitting than it helps by building a more varied cv scheme. </p>",
          "rawMarkdown": "It was popularised in https://www.amazon.com/Advances-Financial-Machine-Learning-Marcos/dp/1119482089 I think its usefullness might depends on the task. In a very low signal to noise environnement it might introduce more overfitting than it helps by building a more varied cv scheme. ",
          "replies": [
            {
              "id": 3048228,
              "postDate": "2024-11-17T16:47:16.817Z",
              "content": "<p>By increasing the number of different folds used in training/validation it should reduce the variance of the prediction so I don't understand the remark about low signal to noise increasing the risk of overfitting with this strategy, can you expand your thoughts?</p>\n<p>The main argument against this strategy as far as I understand is not related to how much signal there is in the data, it argues there is long term dependency in the history that the purge is not efficiently removing or that somehow older history cannot repeat itself so you can't use if for validation of newer data (but then why train on it?).</p>",
              "rawMarkdown": "By increasing the number of different folds used in training/validation it should reduce the variance of the prediction so I don't understand the remark about low signal to noise increasing the risk of overfitting with this strategy, can you expand your thoughts?\n\nThe main argument against this strategy as far as I understand is not related to how much signal there is in the data, it argues there is long term dependency in the history that the purge is not efficiently removing or that somehow older history cannot repeat itself so you can't use if for validation of newer data (but then why train on it?)."
            }
          ]
        }
      ]
    },
    {
      "id": 3046574,
      "postDate": "2024-11-15T15:35:09.023Z",
      "rawMarkdown": "",
      "isDeleted": true
    }
  ],
  "comments": [
    {
      "id": 3046623,
      "author_name": "Gunes Evitan",
      "author_url": "",
      "post_date": "2024-11-15T16:55:20.107000",
      "content": "<p>I don't know why this technique is so popular on Kaggle. I would never train on future and validate on past.</p>",
      "votes": 0,
      "replies": [
        {
          "id": 3046847,
          "author_name": "yunsuxiaozi",
          "author_url": "",
          "post_date": "2024-11-16T00:16:43.600000",
          "content": "<p>May I ask why you said 'train on future and validate on past'? Is there a mistake in my code? I divided the training set and validation set in chronological order. Do you have any better cross validation methods?</p>",
          "votes": 0,
          "replies": [
            {
              "id": 3046936,
              "author_name": "Gunes Evitan",
              "author_url": "",
              "post_date": "2024-11-16T04:16:34.627000",
              "content": "<p>I saw your link to purged kfold and read this. <a href=\"https://stats.stackexchange.com/questions/443159/what-is-combinatorial-purged-cross-validation-for-time-series-data\" target=\"_blank\">https://stats.stackexchange.com/questions/443159/what-is-combinatorial-purged-cross-validation-for-time-series-data</a></p>\n<p>The explanation shows training data can be ahead of validation data.</p>",
              "votes": 0,
              "replies": []
            }
          ]
        },
        {
          "id": 3047857,
          "author_name": "Lucas Morin",
          "author_url": "",
          "post_date": "2024-11-17T09:21:18.260000",
          "content": "<p>It was popularised in <a href=\"https://www.amazon.com/Advances-Financial-Machine-Learning-Marcos/dp/1119482089\" target=\"_blank\">https://www.amazon.com/Advances-Financial-Machine-Learning-Marcos/dp/1119482089</a> I think its usefullness might depends on the task. In a very low signal to noise environnement it might introduce more overfitting than it helps by building a more varied cv scheme. </p>",
          "votes": 0,
          "replies": [
            {
              "id": 3048228,
              "author_name": "claudiu",
              "author_url": "",
              "post_date": "2024-11-17T16:47:16.817000",
              "content": "<p>By increasing the number of different folds used in training/validation it should reduce the variance of the prediction so I don't understand the remark about low signal to noise increasing the risk of overfitting with this strategy, can you expand your thoughts?</p>\n<p>The main argument against this strategy as far as I understand is not related to how much signal there is in the data, it argues there is long term dependency in the history that the purge is not efficiently removing or that somehow older history cannot repeat itself so you can't use if for validation of newer data (but then why train on it?).</p>",
              "votes": 0,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 3046574,
      "author_name": "",
      "author_url": "",
      "post_date": "2024-11-15T15:35:09.023000",
      "content": "",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "3046459": "<a href=\"https://www.kaggle.com/code/yunsuxiaozi/js2024-synthetic-data-for-offline-cv/\">Here</a> is dataset.\n\n<a href=\"https://www.kaggle.com/code/yunsuxiaozi/js2024-synthetic-data-with-purgedkfold/notebook\">Purgedkfold</a>",
    "3046623": "I don't know why this technique is so popular on Kaggle. I would never train on future and validate on past.",
    "3046574": ""
  }
}