{
  "id": 542131,
  "title": "Question on making dataset for LSTM ",
  "url": "/competitions/jane-street-real-time-market-data-forecasting/discussion/542131",
  "author_name": "rivzzzzz",
  "post_date": "2024-10-23T05:48:18.358000",
  "votes": 2,
  "comment_count": 3,
  "views": 0,
  "content": "<p>If I want to do something similar to (<a href=\"https://www.kaggle.com/code/vmuzhichenko/g-research-parallel-lstm-training\" target=\"_blank\">https://www.kaggle.com/code/vmuzhichenko/g-research-parallel-lstm-training</a>) for creating dataset, like how it's shown in the image but Asset Id is replaced to Symbol id in this case, how do I implement it given that number of unique symbol_id per day is not always 14. </p>\n<p>More explicitly: df[df['date_id'] == 10]['symbol_id'].nunique() returns 8, while df[df['date_id'] == 1]['symbol_id'].nunique() returns 12, and the max is 20. How can I do similar thing when it's a variable instead of a set value 14</p>\n<p>edit: If anyone having same problem: just create a df with 0-38 all exist and merge it (dumb dumb way but works)</p>",
  "messages": [
    {
      "id": 3052078,
      "postDate": "2024-11-22T02:45:05.673Z",
      "content": "<p>Wassup chris</p>",
      "rawMarkdown": "Wassup chris",
      "votes": 1
    },
    {
      "id": 3025731,
      "postDate": "2024-10-23T05:48:18.357Z",
      "content": "<p>If I want to do something similar to (<a href=\"https://www.kaggle.com/code/vmuzhichenko/g-research-parallel-lstm-training\" target=\"_blank\">https://www.kaggle.com/code/vmuzhichenko/g-research-parallel-lstm-training</a>) for creating dataset, like how it's shown in the image but Asset Id is replaced to Symbol id in this case, how do I implement it given that number of unique symbol_id per day is not always 14. </p>\n<p>More explicitly: df[df['date_id'] == 10]['symbol_id'].nunique() returns 8, while df[df['date_id'] == 1]['symbol_id'].nunique() returns 12, and the max is 20. How can I do similar thing when it's a variable instead of a set value 14</p>\n<p>edit: If anyone having same problem: just create a df with 0-38 all exist and merge it (dumb dumb way but works)</p>",
      "rawMarkdown": "If I want to do something similar to (https://www.kaggle.com/code/vmuzhichenko/g-research-parallel-lstm-training) for creating dataset, like how it's shown in the image but Asset Id is replaced to Symbol id in this case, how do I implement it given that number of unique symbol_id per day is not always 14. \n\nMore explicitly: df[df['date_id'] == 10]['symbol_id'].nunique() returns 8, while df[df['date_id'] == 1]['symbol_id'].nunique() returns 12, and the max is 20. How can I do similar thing when it's a variable instead of a set value 14\n\n\n\nedit: If anyone having same problem: just create a df with 0-38 all exist and merge it (dumb dumb way but works)",
      "votes": 2
    },
    {
      "id": 3026538,
      "postDate": "2024-10-24T00:19:17.333Z",
      "rawMarkdown": "",
      "isDeleted": true
    },
    {
      "id": 3025891,
      "postDate": "2024-10-23T08:34:44.897Z",
      "rawMarkdown": "",
      "isDeleted": true
    }
  ],
  "comments": [
    {
      "id": 3052078,
      "author_name": "werus23",
      "author_url": "",
      "post_date": "2024-11-22T02:45:05.673000",
      "content": "<p>Wassup chris</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 3026538,
      "author_name": "",
      "author_url": "",
      "post_date": "2024-10-24T00:19:17.333000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 3025891,
      "author_name": "",
      "author_url": "",
      "post_date": "2024-10-23T08:34:44.897000",
      "content": "",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "3052078": "Wassup chris",
    "3025731": "If I want to do something similar to (https://www.kaggle.com/code/vmuzhichenko/g-research-parallel-lstm-training) for creating dataset, like how it's shown in the image but Asset Id is replaced to Symbol id in this case, how do I implement it given that number of unique symbol_id per day is not always 14. \n\nMore explicitly: df[df['date_id'] == 10]['symbol_id'].nunique() returns 8, while df[df['date_id'] == 1]['symbol_id'].nunique() returns 12, and the max is 20. How can I do similar thing when it's a variable instead of a set value 14\n\n\n\nedit: If anyone having same problem: just create a df with 0-38 all exist and merge it (dumb dumb way but works)",
    "3026538": "",
    "3025891": ""
  }
}