{
  "id": 581152,
  "title": "Are the training data enough?",
  "url": "/competitions/drw-crypto-market-prediction/discussion/581152",
  "author_name": "",
  "post_date": "2025-05-28T16:15:41.418770200Z",
  "votes": -3,
  "comment_count": 2,
  "views": 0,
  "content": "<p>In the recently concluded JS competition, we needed to use three years of data to predict six months of results. However, in this competition, we are required to use one year of data to predict one year of results. Should the host move some of the test set data to the training set to prevent shake?</p>",
  "messages": [
    {
      "id": "3211625",
      "postDate": "05/28/2025 16:15:41",
      "content": "<p>In the recently concluded JS competition, we needed to use three years of data to predict six months of results. However, in this competition, we are required to use one year of data to predict one year of results. Should the host move some of the test set data to the training set to prevent shake?</p>",
      "rawMarkdown": "In the recently concluded JS competition, we needed to use three years of data to predict six months of results. However, in this competition, we are required to use one year of data to predict one year of results. Should the host move some of the test set data to the training set to prevent shake?",
      "votes": null
    },
    {
      "id": "3213535",
      "postDate": "05/30/2025 06:31:01",
      "content": "<p>to be honest this competition has an unconvetional setup. the nature of the data is time-series but they remove that for test data and tell us to predict it like a tabular data. which cause a lot of  missing temporal pattern. for the amount of data yes it's really limited but you get what you get I dont think we can ask for more. even if we want to go collect online data it's almost impossible cause all 800 features are abstracted already. no way to know what it is.</p>",
      "rawMarkdown": "to be honest this competition has an unconvetional setup. the nature of the data is time-series but they remove that for test data and tell us to predict it like a tabular data. which cause a lot of  missing temporal pattern. for the amount of data yes it's really limited but you get what you get I dont think we can ask for more. even if we want to go collect online data it's almost impossible cause all 800 features are abstracted already. no way to know what it is.",
      "votes": null
    },
    {
      "id": "3213771",
      "postDate": "05/30/2025 12:20:48",
      "content": "<p>You are right. I just expect there is a big shake due to the lack of training data. I will decide whether to participate in this competition based on the further CV-LB results :)</p>",
      "rawMarkdown": "You are right. I just expect there is a big shake due to the lack of training data. I will decide whether to participate in this competition based on the further CV-LB results :)",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 3213535,
      "author_name": "theonetheonly",
      "author_url": "",
      "post_date": "05/30/2025 06:31:01",
      "content": "<p>to be honest this competition has an unconvetional setup. the nature of the data is time-series but they remove that for test data and tell us to predict it like a tabular data. which cause a lot of  missing temporal pattern. for the amount of data yes it's really limited but you get what you get I dont think we can ask for more. even if we want to go collect online data it's almost impossible cause all 800 features are abstracted already. no way to know what it is.</p>",
      "votes": null,
      "replies": [
        {
          "id": 3213771,
          "author_name": "zhusiqi2715",
          "author_url": "",
          "post_date": "05/30/2025 12:20:48",
          "content": "<p>You are right. I just expect there is a big shake due to the lack of training data. I will decide whether to participate in this competition based on the further CV-LB results :)</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "3211625": "In the recently concluded JS competition, we needed to use three years of data to predict six months of results. However, in this competition, we are required to use one year of data to predict one year of results. Should the host move some of the test set data to the training set to prevent shake?",
    "3213535": "to be honest this competition has an unconvetional setup. the nature of the data is time-series but they remove that for test data and tell us to predict it like a tabular data. which cause a lot of  missing temporal pattern. for the amount of data yes it's really limited but you get what you get I dont think we can ask for more. even if we want to go collect online data it's almost impossible cause all 800 features are abstracted already. no way to know what it is.",
    "3213771": "You are right. I just expect there is a big shake due to the lack of training data. I will decide whether to participate in this competition based on the further CV-LB results :)"
  },
  "source": "meta"
}