{
  "id": 542492,
  "title": "Training Data Partitions",
  "url": "/competitions/jane-street-real-time-market-data-forecasting/discussion/542492",
  "author_name": "",
  "post_date": "2024-10-25T05:00:04.985529200Z",
  "votes": null,
  "comment_count": 1,
  "views": 0,
  "content": "<p>Should we use all 10 partitions of the training data? what I mean by that is should we combine all the 10 partitions into one training dataset for the full timeseries? I am quite new so any helpful answers would be greatly appreciated!</p>",
  "messages": [
    {
      "id": "3027654",
      "postDate": "10/25/2024 05:00:04",
      "content": "<p>Should we use all 10 partitions of the training data? what I mean by that is should we combine all the 10 partitions into one training dataset for the full timeseries? I am quite new so any helpful answers would be greatly appreciated!</p>",
      "rawMarkdown": "Should we use all 10 partitions of the training data? what I mean by that is should we combine all the 10 partitions into one training dataset for the full timeseries? I am quite new so any helpful answers would be greatly appreciated!",
      "votes": null
    },
    {
      "id": "3027664",
      "postDate": "10/25/2024 05:36:53",
      "content": "<p>There is a <a href=\"https://www.kaggle.com/competitions/jane-street-real-time-market-data-forecasting/discussion/541783#3026893\" target=\"_blank\">change in behaviour partway through the fourth file</a> (at date_id 698). I'm currently not using data before this point.</p>\n<p><a href=\"https://www.kaggle.com/competitions/jane-street-real-time-market-data-forecasting/discussion/542057\" target=\"_blank\">This discussion</a> has some approaches to CV.</p>\n<p>A good \"rough\" start would be to use partitions 7 &amp; 8 to predict 9, and to start looking for interesting signals, patterns, and oddities in the data.</p>",
      "rawMarkdown": "There is a [change in behaviour partway through the fourth file](https://www.kaggle.com/competitions/jane-street-real-time-market-data-forecasting/discussion/541783#3026893) (at date_id 698). I'm currently not using data before this point.\n\n[This discussion](https://www.kaggle.com/competitions/jane-street-real-time-market-data-forecasting/discussion/542057) has some approaches to CV.\n\nA good \"rough\" start would be to use partitions 7 & 8 to predict 9, and to start looking for interesting signals, patterns, and oddities in the data.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 3027664,
      "author_name": "paddykb",
      "author_url": "",
      "post_date": "10/25/2024 05:36:53",
      "content": "<p>There is a <a href=\"https://www.kaggle.com/competitions/jane-street-real-time-market-data-forecasting/discussion/541783#3026893\" target=\"_blank\">change in behaviour partway through the fourth file</a> (at date_id 698). I'm currently not using data before this point.</p>\n<p><a href=\"https://www.kaggle.com/competitions/jane-street-real-time-market-data-forecasting/discussion/542057\" target=\"_blank\">This discussion</a> has some approaches to CV.</p>\n<p>A good \"rough\" start would be to use partitions 7 &amp; 8 to predict 9, and to start looking for interesting signals, patterns, and oddities in the data.</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "3027654": "Should we use all 10 partitions of the training data? what I mean by that is should we combine all the 10 partitions into one training dataset for the full timeseries? I am quite new so any helpful answers would be greatly appreciated!",
    "3027664": "There is a [change in behaviour partway through the fourth file](https://www.kaggle.com/competitions/jane-street-real-time-market-data-forecasting/discussion/541783#3026893) (at date_id 698). I'm currently not using data before this point.\n\n[This discussion](https://www.kaggle.com/competitions/jane-street-real-time-market-data-forecasting/discussion/542057) has some approaches to CV.\n\nA good \"rough\" start would be to use partitions 7 & 8 to predict 9, and to start looking for interesting signals, patterns, and oddities in the data."
  },
  "source": "meta"
}