{
  "id": 540449,
  "title": "Question about competition data",
  "url": "/competitions/jane-street-real-time-market-data-forecasting/discussion/540449",
  "author_name": "",
  "post_date": "2024-10-14T16:08:36.488542300Z",
  "votes": 7,
  "comment_count": 3,
  "views": 0,
  "content": "<p>The train data are comprised of 9 parquet files, each in a unique, respective subfolder. The (sample) test data contain <strong>one</strong> parquet file. During submission with the evaluation API, will the test data still contain one file or have more than one files?</p>\n<p>Also, should participants incorporate the <code>lags.parquet</code> when processing the test data, or is the file only used by the API?</p>\n<p>Last but not least, what is the utility of the metadata (<code>features.csv</code> and <code>responders.csv</code>)</p>",
  "messages": [
    {
      "id": "3017185",
      "postDate": "10/14/2024 16:08:36",
      "content": "<p>The train data are comprised of 9 parquet files, each in a unique, respective subfolder. The (sample) test data contain <strong>one</strong> parquet file. During submission with the evaluation API, will the test data still contain one file or have more than one files?</p>\n<p>Also, should participants incorporate the <code>lags.parquet</code> when processing the test data, or is the file only used by the API?</p>\n<p>Last but not least, what is the utility of the metadata (<code>features.csv</code> and <code>responders.csv</code>)</p>",
      "rawMarkdown": "The train data are comprised of 9 parquet files, each in a unique, respective subfolder. The (sample) test data contain **one** parquet file. During submission with the evaluation API, will the test data still contain one file or have more than one files?\n\nAlso, should participants incorporate the `lags.parquet` when processing the test data, or is the file only used by the API?\n\nLast but not least, what is the utility of the metadata (`features.csv` and `responders.csv`)",
      "votes": null
    },
    {
      "id": "3017202",
      "postDate": "10/14/2024 16:20:41",
      "content": "<p>Hi <a href=\"https://www.kaggle.com/andreasbis\" target=\"_blank\">@andreasbis</a>,</p>\n<p>Your submission script will not be able to access the files for the test data directly. It is served through the evaluation API timestamp by timestamp. You should expect the actual test data to span many dates instead of just the one date in the sample file.</p>\n<p>The lags and metadata files are for your use. They may (or may not) be useful as inputs to your model.</p>",
      "rawMarkdown": "Hi @andreasbis,\n\nYour submission script will not be able to access the files for the test data directly. It is served through the evaluation API timestamp by timestamp. You should expect the actual test data to span many dates instead of just the one date in the sample file.\n\nThe lags and metadata files are for your use. They may (or may not) be useful as inputs to your model.",
      "votes": null
    },
    {
      "id": "3017274",
      "postDate": "10/14/2024 17:56:23",
      "content": "<p>About the metadata specifically, can we have an idea of what <em>kind</em> of information they're supposed to represent? They seem a lot of booleans, are these e.g. categories of the financial product (like what kind of product it is, or what industry sector the company is if it's a stock), or things like geographical location, or such?</p>",
      "rawMarkdown": "About the metadata specifically, can we have an idea of what *kind* of information they're supposed to represent? They seem a lot of booleans, are these e.g. categories of the financial product (like what kind of product it is, or what industry sector the company is if it's a stock), or things like geographical location, or such?",
      "votes": null
    },
    {
      "id": "3024821",
      "postDate": "10/22/2024 03:49:43",
      "content": "<p>I’ve used the metadata to guide some feature engineering and have some features that I think are going to be very helpful… once I can use them. How the heck do you join test with lags_? I have code that does it correctly with the whopping 38 or so rows of test data. Debugging is ridiculous</p>",
      "rawMarkdown": "I’ve used the metadata to guide some feature engineering and have some features that I think are going to be very helpful… once I can use them. How the heck do you join test with lags_? I have code that does it correctly with the whopping 38 or so rows of test data. Debugging is ridiculous",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 3017202,
      "author_name": "ryanholbrook",
      "author_url": "",
      "post_date": "10/14/2024 16:20:41",
      "content": "<p>Hi <a href=\"https://www.kaggle.com/andreasbis\" target=\"_blank\">@andreasbis</a>,</p>\n<p>Your submission script will not be able to access the files for the test data directly. It is served through the evaluation API timestamp by timestamp. You should expect the actual test data to span many dates instead of just the one date in the sample file.</p>\n<p>The lags and metadata files are for your use. They may (or may not) be useful as inputs to your model.</p>",
      "votes": null,
      "replies": [
        {
          "id": 3017274,
          "author_name": "sstur86",
          "author_url": "",
          "post_date": "10/14/2024 17:56:23",
          "content": "<p>About the metadata specifically, can we have an idea of what <em>kind</em> of information they're supposed to represent? They seem a lot of booleans, are these e.g. categories of the financial product (like what kind of product it is, or what industry sector the company is if it's a stock), or things like geographical location, or such?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 3024821,
          "author_name": "jackvd",
          "author_url": "",
          "post_date": "10/22/2024 03:49:43",
          "content": "<p>I’ve used the metadata to guide some feature engineering and have some features that I think are going to be very helpful… once I can use them. How the heck do you join test with lags_? I have code that does it correctly with the whopping 38 or so rows of test data. Debugging is ridiculous</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "3017185": "The train data are comprised of 9 parquet files, each in a unique, respective subfolder. The (sample) test data contain **one** parquet file. During submission with the evaluation API, will the test data still contain one file or have more than one files?\n\nAlso, should participants incorporate the `lags.parquet` when processing the test data, or is the file only used by the API?\n\nLast but not least, what is the utility of the metadata (`features.csv` and `responders.csv`)",
    "3017202": "Hi @andreasbis,\n\nYour submission script will not be able to access the files for the test data directly. It is served through the evaluation API timestamp by timestamp. You should expect the actual test data to span many dates instead of just the one date in the sample file.\n\nThe lags and metadata files are for your use. They may (or may not) be useful as inputs to your model.",
    "3017274": "About the metadata specifically, can we have an idea of what *kind* of information they're supposed to represent? They seem a lot of booleans, are these e.g. categories of the financial product (like what kind of product it is, or what industry sector the company is if it's a stock), or things like geographical location, or such?",
    "3024821": "I’ve used the metadata to guide some feature engineering and have some features that I think are going to be very helpful… once I can use them. How the heck do you join test with lags_? I have code that does it correctly with the whopping 38 or so rows of test data. Debugging is ridiculous"
  },
  "source": "meta"
}