{
  "id": 548658,
  "title": "Use \"time_id\" and \"symbol_id\" or not?",
  "url": "/competitions/jane-street-real-time-market-data-forecasting/discussion/548658",
  "author_name": "",
  "post_date": "2024-11-28T05:57:46.801958700Z",
  "votes": 1,
  "comment_count": 1,
  "views": 0,
  "content": "<p>I find removing \"time_id\" and \"symbol_id\" will decrease the CV significantly, but it seems not to hurt LB. So I am wondering if we should use two features or not? Any insights or opinions will be appreciated.</p>",
  "messages": [
    {
      "id": "3057407",
      "postDate": "11/28/2024 05:57:46",
      "content": "<p>I find removing \"time_id\" and \"symbol_id\" will decrease the CV significantly, but it seems not to hurt LB. So I am wondering if we should use two features or not? Any insights or opinions will be appreciated.</p>",
      "rawMarkdown": "I find removing \"time_id\" and \"symbol_id\" will decrease the CV significantly, but it seems not to hurt LB. So I am wondering if we should use two features or not? Any insights or opinions will be appreciated.",
      "votes": null
    },
    {
      "id": "3057814",
      "postDate": "11/28/2024 16:49:04",
      "content": "<p>My suspicion is that both features should not be used in their raw form, since they may behave differently in the hidden sets (e.g. due to additional time slots or additional symbols). However, one might consider tweaking them to improve a model's ability to extrapolate. For example, it has been suggested <a href=\"https://www.kaggle.com/competitions/jane-street-real-time-market-data-forecasting/discussion/548636\" target=\"_blank\">here</a> to use embdeddings for symbols.</p>",
      "rawMarkdown": "My suspicion is that both features should not be used in their raw form, since they may behave differently in the hidden sets (e.g. due to additional time slots or additional symbols). However, one might consider tweaking them to improve a model's ability to extrapolate. For example, it has been suggested [here](https://www.kaggle.com/competitions/jane-street-real-time-market-data-forecasting/discussion/548636) to use embdeddings for symbols.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 3057814,
      "author_name": "jonassend",
      "author_url": "",
      "post_date": "11/28/2024 16:49:04",
      "content": "<p>My suspicion is that both features should not be used in their raw form, since they may behave differently in the hidden sets (e.g. due to additional time slots or additional symbols). However, one might consider tweaking them to improve a model's ability to extrapolate. For example, it has been suggested <a href=\"https://www.kaggle.com/competitions/jane-street-real-time-market-data-forecasting/discussion/548636\" target=\"_blank\">here</a> to use embdeddings for symbols.</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "3057407": "I find removing \"time_id\" and \"symbol_id\" will decrease the CV significantly, but it seems not to hurt LB. So I am wondering if we should use two features or not? Any insights or opinions will be appreciated.",
    "3057814": "My suspicion is that both features should not be used in their raw form, since they may behave differently in the hidden sets (e.g. due to additional time slots or additional symbols). However, one might consider tweaking them to improve a model's ability to extrapolate. For example, it has been suggested [here](https://www.kaggle.com/competitions/jane-street-real-time-market-data-forecasting/discussion/548636) to use embdeddings for symbols."
  },
  "source": "meta"
}