{
  "id": 544277,
  "title": "An assumption about the test data",
  "url": "/competitions/jane-street-real-time-market-data-forecasting/discussion/544277",
  "author_name": "",
  "post_date": "2024-11-04T08:29:13.717472600Z",
  "votes": 11,
  "comment_count": 5,
  "views": 0,
  "content": "<p>I have made a dummy submission to test whether we always have lag data for all symbol_ids.</p>\n<p>The answer is no, lags may be missing for symbol_ids present in the current test batch: since symbols may appear and disappear over the days, we have cases where no previous day lag data is present for a symbol_id. Same as in training data.</p>\n<p>Test condition:</p>\n<pre><code> = test.join(\n    self.lags, \n    =[, , ],\n    how=)\nassert test_with_lags.shape[] == test.shape[], \n</code></pre>",
  "messages": [
    {
      "id": "3036107",
      "postDate": "11/04/2024 08:29:13",
      "content": "<p>I have made a dummy submission to test whether we always have lag data for all symbol_ids.</p>\n<p>The answer is no, lags may be missing for symbol_ids present in the current test batch: since symbols may appear and disappear over the days, we have cases where no previous day lag data is present for a symbol_id. Same as in training data.</p>\n<p>Test condition:</p>\n<pre><code> = test.join(\n    self.lags, \n    =[, , ],\n    how=)\nassert test_with_lags.shape[] == test.shape[], \n</code></pre>",
      "rawMarkdown": "I have made a dummy submission to test whether we always have lag data for all symbol_ids.\n\nThe answer is no, lags may be missing for symbol_ids present in the current test batch: since symbols may appear and disappear over the days, we have cases where no previous day lag data is present for a symbol_id. Same as in training data.\n\nTest condition:\n```\ntest_with_lags = test.join(\n    self.lags, # Here my submission accumulates all previously collected lags.\n    on=['date_id', 'time_id', 'symbol_id'],\n    how='inner')\nassert test_with_lags.shape[0] == test.shape[0], \"Lags for some symbols aren't available!\"\n```",
      "votes": null
    },
    {
      "id": "3036109",
      "postDate": "11/04/2024 08:31:37",
      "content": "<p>This is a known condition - symbols may appear and leave at any time, making the lag data dynamic across symbols <a href=\"https://www.kaggle.com/shuthdar\" target=\"_blank\">@shuthdar</a> <br>\nLag data will always consider 1 day-old information</p>",
      "rawMarkdown": "This is a known condition - symbols may appear and leave at any time, making the lag data dynamic across symbols @shuthdar \nLag data will always consider 1 day-old information",
      "votes": null
    },
    {
      "id": "3036123",
      "postDate": "11/04/2024 08:49:31",
      "content": "<p>Thanks for sharing. They said it's unlikely to see the symbol_id change, but if it does, they'll inform us—if I understood them correctly. However, they didn't ay anything about it. I suspect this might be why the CV behavior is inconsistent compared to the LB.</p>",
      "rawMarkdown": "Thanks for sharing. They said it's unlikely to see the symbol_id change, but if it does, they'll inform us—if I understood them correctly. However, they didn't ay anything about it. I suspect this might be why the CV behavior is inconsistent compared to the LB.",
      "votes": null
    },
    {
      "id": "3036128",
      "postDate": "11/04/2024 08:51:33",
      "content": "<p>Yeah, but I considered it's worth checking and sharing, after seeing which questions people ask here. Also to avoid making wrong assumptions about the data when imitating test inputs from the training data: e.g. what if in case when symbol doesn't have the previous day's data we still get respective records in the lags df but they're nulls? But now I know that the values will be just missing.</p>",
      "rawMarkdown": "Yeah, but I considered it's worth checking and sharing, after seeing which questions people ask here. Also to avoid making wrong assumptions about the data when imitating test inputs from the training data: e.g. what if in case when symbol doesn't have the previous day's data we still get respective records in the lags df but they're nulls? But now I know that the values will be just missing.",
      "votes": null
    },
    {
      "id": "3036136",
      "postDate": "11/04/2024 08:55:50",
      "content": "<p>Yes, this is also observed in the training dataset. Symbols are not always present in every day. A robust nan handling approach is necessary here if one would like to use the lagged responders.</p>",
      "rawMarkdown": "Yes, this is also observed in the training dataset. Symbols are not always present in every day. A robust nan handling approach is necessary here if one would like to use the lagged responders.",
      "votes": null
    },
    {
      "id": "3037449",
      "postDate": "11/05/2024 18:09:32",
      "content": "<p>If the proportion of missing lag data is not high, perhaps we can put this problem aside by filling nan by 0.</p>",
      "rawMarkdown": "If the proportion of missing lag data is not high, perhaps we can put this problem aside by filling nan by 0.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 3036109,
      "author_name": "ravi20076",
      "author_url": "",
      "post_date": "11/04/2024 08:31:37",
      "content": "<p>This is a known condition - symbols may appear and leave at any time, making the lag data dynamic across symbols <a href=\"https://www.kaggle.com/shuthdar\" target=\"_blank\">@shuthdar</a> <br>\nLag data will always consider 1 day-old information</p>",
      "votes": null,
      "replies": [
        {
          "id": 3036128,
          "author_name": "shuthdar",
          "author_url": "",
          "post_date": "11/04/2024 08:51:33",
          "content": "<p>Yeah, but I considered it's worth checking and sharing, after seeing which questions people ask here. Also to avoid making wrong assumptions about the data when imitating test inputs from the training data: e.g. what if in case when symbol doesn't have the previous day's data we still get respective records in the lags df but they're nulls? But now I know that the values will be just missing.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 3036123,
      "author_name": "aymanallawi",
      "author_url": "",
      "post_date": "11/04/2024 08:49:31",
      "content": "<p>Thanks for sharing. They said it's unlikely to see the symbol_id change, but if it does, they'll inform us—if I understood them correctly. However, they didn't ay anything about it. I suspect this might be why the CV behavior is inconsistent compared to the LB.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 3036136,
      "author_name": "shiyili",
      "author_url": "",
      "post_date": "11/04/2024 08:55:50",
      "content": "<p>Yes, this is also observed in the training dataset. Symbols are not always present in every day. A robust nan handling approach is necessary here if one would like to use the lagged responders.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 3037449,
      "author_name": "voix97",
      "author_url": "",
      "post_date": "11/05/2024 18:09:32",
      "content": "<p>If the proportion of missing lag data is not high, perhaps we can put this problem aside by filling nan by 0.</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "3036107": "I have made a dummy submission to test whether we always have lag data for all symbol_ids.\n\nThe answer is no, lags may be missing for symbol_ids present in the current test batch: since symbols may appear and disappear over the days, we have cases where no previous day lag data is present for a symbol_id. Same as in training data.\n\nTest condition:\n```\ntest_with_lags = test.join(\n    self.lags, # Here my submission accumulates all previously collected lags.\n    on=['date_id', 'time_id', 'symbol_id'],\n    how='inner')\nassert test_with_lags.shape[0] == test.shape[0], \"Lags for some symbols aren't available!\"\n```",
    "3036109": "This is a known condition - symbols may appear and leave at any time, making the lag data dynamic across symbols @shuthdar \nLag data will always consider 1 day-old information",
    "3036123": "Thanks for sharing. They said it's unlikely to see the symbol_id change, but if it does, they'll inform us—if I understood them correctly. However, they didn't ay anything about it. I suspect this might be why the CV behavior is inconsistent compared to the LB.",
    "3036128": "Yeah, but I considered it's worth checking and sharing, after seeing which questions people ask here. Also to avoid making wrong assumptions about the data when imitating test inputs from the training data: e.g. what if in case when symbol doesn't have the previous day's data we still get respective records in the lags df but they're nulls? But now I know that the values will be just missing.",
    "3036136": "Yes, this is also observed in the training dataset. Symbols are not always present in every day. A robust nan handling approach is necessary here if one would like to use the lagged responders.",
    "3037449": "If the proportion of missing lag data is not high, perhaps we can put this problem aside by filling nan by 0."
  },
  "source": "meta"
}