{
  "id": 553893,
  "title": "Is groupby(\"symbol_id\") useful in preprocess?",
  "url": "/competitions/jane-street-real-time-market-data-forecasting/discussion/553893",
  "author_name": "yich",
  "post_date": "2024-12-29T05:51:29.992000",
  "votes": 1,
  "comment_count": 6,
  "views": 0,
  "content": "<p>The original train data seems to be sorted by date and time. Thus lines belong to each symbol are not consecutive. However, when I preprocess data(fillna) using groupby first, it performed worse than those methods not using groupby. What's the reason here?     !(^O^)y</p>",
  "messages": [
    {
      "id": 3083790,
      "postDate": "2024-12-30T03:11:45.153Z",
      "content": "<p>I am trying a method to use different models for each symbol_id.</p>",
      "rawMarkdown": "I am trying a method to use different models for each symbol_id.",
      "votes": 1,
      "replies": [
        {
          "id": 3083895,
          "postDate": "2024-12-30T06:21:32.700Z",
          "content": "<p>Be aware that new symbols may emerge  during the forecasting phase </p>",
          "rawMarkdown": "Be aware that new symbols may emerge  during the forecasting phase ",
          "votes": 4,
          "replies": [
            {
              "id": 3083912,
              "postDate": "2024-12-30T07:05:55.460Z",
              "content": "<p>Thanks a lot for the reminder! 😀</p>",
              "rawMarkdown": "Thanks a lot for the reminder! 😀"
            }
          ]
        }
      ]
    },
    {
      "id": 3083171,
      "postDate": "2024-12-29T05:51:29.993Z",
      "content": "<p>The original train data seems to be sorted by date and time. Thus lines belong to each symbol are not consecutive. However, when I preprocess data(fillna) using groupby first, it performed worse than those methods not using groupby. What's the reason here?     !(^O^)y</p>",
      "rawMarkdown": "The original train data seems to be sorted by date and time. Thus lines belong to each symbol are not consecutive. However, when I preprocess data(fillna) using groupby first, it performed worse than those methods not using groupby. What's the reason here?     !(^O^)y",
      "votes": 1
    },
    {
      "id": 3084743,
      "postDate": "2024-12-31T08:40:19.270Z",
      "content": "<p>In my local testing, this approach did not show a significant improvement in the local CV score.</p>",
      "rawMarkdown": "In my local testing, this approach did not show a significant improvement in the local CV score.",
      "replies": [
        {
          "id": 3084862,
          "postDate": "2024-12-31T12:15:23.467Z",
          "rawMarkdown": "",
          "isDeleted": true
        },
        {
          "id": 3084863,
          "postDate": "2024-12-31T12:16:09.487Z",
          "content": "<p>So do I, maybe that's not a good direction.</p>",
          "rawMarkdown": "So do I, maybe that's not a good direction."
        }
      ]
    }
  ],
  "comments": [
    {
      "id": 3083790,
      "author_name": "CodeHacker",
      "author_url": "",
      "post_date": "2024-12-30T03:11:45.153000",
      "content": "<p>I am trying a method to use different models for each symbol_id.</p>",
      "votes": 1,
      "replies": [
        {
          "id": 3083895,
          "author_name": "SLi",
          "author_url": "",
          "post_date": "2024-12-30T06:21:32.700000",
          "content": "<p>Be aware that new symbols may emerge  during the forecasting phase </p>",
          "votes": 4,
          "replies": [
            {
              "id": 3083912,
              "author_name": "yich",
              "author_url": "",
              "post_date": "2024-12-30T07:05:55.460000",
              "content": "<p>Thanks a lot for the reminder! 😀</p>",
              "votes": 0,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 3084743,
      "author_name": "Zhu Yuezhi",
      "author_url": "",
      "post_date": "2024-12-31T08:40:19.270000",
      "content": "<p>In my local testing, this approach did not show a significant improvement in the local CV score.</p>",
      "votes": 0,
      "replies": [
        {
          "id": 3084862,
          "author_name": "",
          "author_url": "",
          "post_date": "2024-12-31T12:15:23.467000",
          "content": "",
          "votes": 0,
          "replies": []
        },
        {
          "id": 3084863,
          "author_name": "yich",
          "author_url": "",
          "post_date": "2024-12-31T12:16:09.487000",
          "content": "<p>So do I, maybe that's not a good direction.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "3083790": "I am trying a method to use different models for each symbol_id.",
    "3083171": "The original train data seems to be sorted by date and time. Thus lines belong to each symbol are not consecutive. However, when I preprocess data(fillna) using groupby first, it performed worse than those methods not using groupby. What's the reason here?     !(^O^)y",
    "3084743": "In my local testing, this approach did not show a significant improvement in the local CV score."
  }
}