{
  "id": 544107,
  "title": "Test data format",
  "url": "/competitions/jane-street-real-time-market-data-forecasting/discussion/544107",
  "author_name": "",
  "post_date": "2024-11-03T08:55:06.135104300Z",
  "votes": 2,
  "comment_count": 2,
  "views": 0,
  "content": "<p>Hello competition host, thank you for hosting this competition, it's been exciting and tough at the same time :)<br>\nI have some questions regarding the test data so as to properly model the required edge cases.<br>\nFirst, i found that the total number of unique symbols in the train_dataset is 39 i.e 0-38, are this all the symbols that are expected in the test dataset? will there be new additions in the future that is not present in the train dataset (i.e new symbols that was not employed in training).<br>\nSecond, the date_ids for the train dataset stops at 1698, does the test data continue from this? i asked this because the test sample shows the date_id starts from 0 again.<br>\nThanks for your kind responses.</p>",
  "messages": [
    {
      "id": "3035302",
      "postDate": "11/03/2024 08:55:06",
      "content": "<p>Hello competition host, thank you for hosting this competition, it's been exciting and tough at the same time :)<br>\nI have some questions regarding the test data so as to properly model the required edge cases.<br>\nFirst, i found that the total number of unique symbols in the train_dataset is 39 i.e 0-38, are this all the symbols that are expected in the test dataset? will there be new additions in the future that is not present in the train dataset (i.e new symbols that was not employed in training).<br>\nSecond, the date_ids for the train dataset stops at 1698, does the test data continue from this? i asked this because the test sample shows the date_id starts from 0 again.<br>\nThanks for your kind responses.</p>",
      "rawMarkdown": "Hello competition host, thank you for hosting this competition, it's been exciting and tough at the same time :)\nI have some questions regarding the test data so as to properly model the required edge cases.\nFirst, i found that the total number of unique symbols in the train_dataset is 39 i.e 0-38, are this all the symbols that are expected in the test dataset? will there be new additions in the future that is not present in the train dataset (i.e new symbols that was not employed in training).\nSecond, the date_ids for the train dataset stops at 1698, does the test data continue from this? i asked this because the test sample shows the date_id starts from 0 again.\nThanks for your kind responses.",
      "votes": null
    },
    {
      "id": "3035997",
      "postDate": "11/04/2024 04:44:46",
      "content": "<p><a href=\"https://www.kaggle.com/oluwatobibetiku\" target=\"_blank\">@oluwatobibetiku</a> The description on the data page mentions that new symbols may be added.</p>\n<pre><code>The symbol_id  contains  identifiers.  symbol_id   guaranteed  appear   time_id  date_id combinations. Additionally,  symbol_id  may appear  future test sets.\n</code></pre>\n<p>The second question is maybe yes, although time_id starts from 0.</p>",
      "rawMarkdown": "oluwatobibetiku The description on the data page mentions that new symbols may be added.\n~~~\nThe symbol_id column contains encrypted identifiers. Each symbol_id is not guaranteed to appear in all time_id and date_id combinations. Additionally, new symbol_id values may appear in future test sets.\n~~~\nThe second question is maybe yes, although time_id starts from 0.",
      "votes": null
    },
    {
      "id": "3036097",
      "postDate": "11/04/2024 08:05:50",
      "content": "<p><a href=\"https://www.kaggle.com/chumajin\" target=\"_blank\">@chumajin</a> Thank you for the heads up :), i think i missed that in the description, but doesn't that make it impossible to apply sliding window since new symbols will be added and there won't be previous data memory to bank on or there is a totally new approach for addressing sequences as this, would really appreciate any thoughts you have on this. For the second question i think i would just offset the time_ids as this makes it cancels out all assumptions i guess. Thank you again for your input, i really appreciate it.</p>",
      "rawMarkdown": "chumajin Thank you for the heads up :), i think i missed that in the description, but doesn't that make it impossible to apply sliding window since new symbols will be added and there won't be previous data memory to bank on or there is a totally new approach for addressing sequences as this, would really appreciate any thoughts you have on this. For the second question i think i would just offset the time_ids as this makes it cancels out all assumptions i guess. Thank you again for your input, i really appreciate it.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 3035997,
      "author_name": "chumajin",
      "author_url": "",
      "post_date": "11/04/2024 04:44:46",
      "content": "<p><a href=\"https://www.kaggle.com/oluwatobibetiku\" target=\"_blank\">@oluwatobibetiku</a> The description on the data page mentions that new symbols may be added.</p>\n<pre><code>The symbol_id  contains  identifiers.  symbol_id   guaranteed  appear   time_id  date_id combinations. Additionally,  symbol_id  may appear  future test sets.\n</code></pre>\n<p>The second question is maybe yes, although time_id starts from 0.</p>",
      "votes": null,
      "replies": [
        {
          "id": 3036097,
          "author_name": "oluwatobibetiku",
          "author_url": "",
          "post_date": "11/04/2024 08:05:50",
          "content": "<p><a href=\"https://www.kaggle.com/chumajin\" target=\"_blank\">@chumajin</a> Thank you for the heads up :), i think i missed that in the description, but doesn't that make it impossible to apply sliding window since new symbols will be added and there won't be previous data memory to bank on or there is a totally new approach for addressing sequences as this, would really appreciate any thoughts you have on this. For the second question i think i would just offset the time_ids as this makes it cancels out all assumptions i guess. Thank you again for your input, i really appreciate it.</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "3035302": "Hello competition host, thank you for hosting this competition, it's been exciting and tough at the same time :)\nI have some questions regarding the test data so as to properly model the required edge cases.\nFirst, i found that the total number of unique symbols in the train_dataset is 39 i.e 0-38, are this all the symbols that are expected in the test dataset? will there be new additions in the future that is not present in the train dataset (i.e new symbols that was not employed in training).\nSecond, the date_ids for the train dataset stops at 1698, does the test data continue from this? i asked this because the test sample shows the date_id starts from 0 again.\nThanks for your kind responses.",
    "3035997": "oluwatobibetiku The description on the data page mentions that new symbols may be added.\n~~~\nThe symbol_id column contains encrypted identifiers. Each symbol_id is not guaranteed to appear in all time_id and date_id combinations. Additionally, new symbol_id values may appear in future test sets.\n~~~\nThe second question is maybe yes, although time_id starts from 0.",
    "3036097": "chumajin Thank you for the heads up :), i think i missed that in the description, but doesn't that make it impossible to apply sliding window since new symbols will be added and there won't be previous data memory to bank on or there is a totally new approach for addressing sequences as this, would really appreciate any thoughts you have on this. For the second question i think i would just offset the time_ids as this makes it cancels out all assumptions i guess. Thank you again for your input, i really appreciate it."
  },
  "source": "meta"
}