{
  "id": 547866,
  "title": "Missing symbols while joining lags and test df in test run",
  "url": "/competitions/jane-street-real-time-market-data-forecasting/discussion/547866",
  "author_name": "",
  "post_date": "2024-11-23T19:09:05.035750300Z",
  "votes": 3,
  "comment_count": 4,
  "views": 0,
  "content": "<p>I am working with a model which joins the lags and test data to extract the lagged responder values for the given symbol_id for the same time_id of the previous date. </p>\n<p>I have used <a href=\"this\" target=\"_blank\">https://www.kaggle.com/code/shiyili/js24-rmf-submission-api-debug-with-synthetic-test?scriptVersionId=201579302</a> notebook to create synthetic data for testing locally before submitting.</p>\n<p>Everything works well when you only use the features in your model without using the lagged responder values. However, when you use join operation for test and lags df as follows</p>\n<p><code>combined_data = test.join(lags_, on=[\"date_id\", \"time_id\", \"symbol_id\"], how=\"inner\")</code></p>\n<p>There are cases where the number of unique symbols in the lags df is 38 while the symbols we're looking to predict given in the test df are 39 (one symbol_id is missing from the lags df). This causes the join to reduce the dimensions from the expected <code>(39, 95)</code> to <code>(38,95)</code> due to a single missing symbol_id.</p>\n<p>I have verified this by printing the dimensions of all the dfs and also manually inspecting the data for the given test entry.</p>\n<pre><code>prediction dimension mismatch 38 != 39\nCombined df shape: (38, 95)\ndf shape: (39, 85)\nLag df shape: (36784, 12)\nSymbols in lag df: shape: (38, 1)\n</code></pre>\n<p>This is the output from which shows symbol_id for 31 is missing.<br>\n<code>lags_.filter((pl.col(\"date_id\") == 2) &amp; (pl.col(\"time_id\") == 0) &amp; (pl.col(\"symbol_id\") &gt; 25) &amp; (pl.col(\"symbol_id\") &lt; 34))</code></p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F13387107%2F451518e58aa69cb6e4da04189abd0446%2FScreenshot%202024-11-24%20at%2012.38.01AM.png?generation=1732388897175690&amp;alt=media\" alt=\"image_output\"></p>\n<p>I wanted to ask if this is expected or is this some sort of error in the data? As the data specifications mention, the lags_ should provide us with all the lagged responder values for the previous date_id. </p>\n<p>If not, is it possible to handle missing lagged values for a single symbol_id randomly in the test run?</p>\n<p>I have also tried submitting the code for evaluation and it threw an exception while processing (as expected due to prediction and test dimensions mismatch)</p>\n<p>My entire prediction code is as below for reference.</p>\n<pre><code>lags_ : pl.DataFrame |  = \ntest_: pl.DataFrame |  = \n\n\n\n\n xgboost  xgb\n\nloaded_model = \n\n () -&gt; pl.DataFrame | pd.DataFrame:\n    \n\n    \n    \n     lags_, loaded_model, test_\n     lags   :\n        lags_ = lags\n\n    test_ = test\n\n     loaded_model  :\n        \n        loaded_model = xgb.Booster()\n        loaded_model.load_model()\n\n\n    combined_data = test.join(lags_, on=[, , ], how=)\n\n    \n    feature_columns = [col  col  combined_data.columns  col.startswith()  col.endswith()]\n\n    \n    X_pred = combined_data.select(feature_columns).to_pandas()  \n\n    \n    dmatrix_pred = xgb.DMatrix(X_pred)\n\n    \n    predictions = loaded_model.predict(dmatrix_pred)\n\n    combined_data = combined_data.with_columns(pl.Series(name=, values=predictions))\n\n    \n    predictions = combined_data.select(\n        , \n    )\n\n     (predictions) != (test):\n        ()\n        ()\n        ()\n        ()\n        ()\n\n     (predictions, pl.DataFrame):\n         predictions.columns == [, ]\n     (predictions, pd.DataFrame):\n         (predictions.columns == [, ]).()\n    :\n         TypeError()\n    \n     (predictions) == (test)\n\n     predictions\n</code></pre>",
  "messages": [
    {
      "id": "3053701",
      "postDate": "11/23/2024 19:09:05",
      "content": "<p>I am working with a model which joins the lags and test data to extract the lagged responder values for the given symbol_id for the same time_id of the previous date. </p>\n<p>I have used <a href=\"this\" target=\"_blank\">https://www.kaggle.com/code/shiyili/js24-rmf-submission-api-debug-with-synthetic-test?scriptVersionId=201579302</a> notebook to create synthetic data for testing locally before submitting.</p>\n<p>Everything works well when you only use the features in your model without using the lagged responder values. However, when you use join operation for test and lags df as follows</p>\n<p><code>combined_data = test.join(lags_, on=[\"date_id\", \"time_id\", \"symbol_id\"], how=\"inner\")</code></p>\n<p>There are cases where the number of unique symbols in the lags df is 38 while the symbols we're looking to predict given in the test df are 39 (one symbol_id is missing from the lags df). This causes the join to reduce the dimensions from the expected <code>(39, 95)</code> to <code>(38,95)</code> due to a single missing symbol_id.</p>\n<p>I have verified this by printing the dimensions of all the dfs and also manually inspecting the data for the given test entry.</p>\n<pre><code>prediction dimension mismatch 38 != 39\nCombined df shape: (38, 95)\ndf shape: (39, 85)\nLag df shape: (36784, 12)\nSymbols in lag df: shape: (38, 1)\n</code></pre>\n<p>This is the output from which shows symbol_id for 31 is missing.<br>\n<code>lags_.filter((pl.col(\"date_id\") == 2) &amp; (pl.col(\"time_id\") == 0) &amp; (pl.col(\"symbol_id\") &gt; 25) &amp; (pl.col(\"symbol_id\") &lt; 34))</code></p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F13387107%2F451518e58aa69cb6e4da04189abd0446%2FScreenshot%202024-11-24%20at%2012.38.01AM.png?generation=1732388897175690&amp;alt=media\" alt=\"image_output\"></p>\n<p>I wanted to ask if this is expected or is this some sort of error in the data? As the data specifications mention, the lags_ should provide us with all the lagged responder values for the previous date_id. </p>\n<p>If not, is it possible to handle missing lagged values for a single symbol_id randomly in the test run?</p>\n<p>I have also tried submitting the code for evaluation and it threw an exception while processing (as expected due to prediction and test dimensions mismatch)</p>\n<p>My entire prediction code is as below for reference.</p>\n<pre><code>lags_ : pl.DataFrame |  = \ntest_: pl.DataFrame |  = \n\n\n\n\n xgboost  xgb\n\nloaded_model = \n\n () -&gt; pl.DataFrame | pd.DataFrame:\n    \n\n    \n    \n     lags_, loaded_model, test_\n     lags   :\n        lags_ = lags\n\n    test_ = test\n\n     loaded_model  :\n        \n        loaded_model = xgb.Booster()\n        loaded_model.load_model()\n\n\n    combined_data = test.join(lags_, on=[, , ], how=)\n\n    \n    feature_columns = [col  col  combined_data.columns  col.startswith()  col.endswith()]\n\n    \n    X_pred = combined_data.select(feature_columns).to_pandas()  \n\n    \n    dmatrix_pred = xgb.DMatrix(X_pred)\n\n    \n    predictions = loaded_model.predict(dmatrix_pred)\n\n    combined_data = combined_data.with_columns(pl.Series(name=, values=predictions))\n\n    \n    predictions = combined_data.select(\n        , \n    )\n\n     (predictions) != (test):\n        ()\n        ()\n        ()\n        ()\n        ()\n\n     (predictions, pl.DataFrame):\n         predictions.columns == [, ]\n     (predictions, pd.DataFrame):\n         (predictions.columns == [, ]).()\n    :\n         TypeError()\n    \n     (predictions) == (test)\n\n     predictions\n</code></pre>",
      "rawMarkdown": "I am working with a model which joins the lags and test data to extract the lagged responder values for the given symbol_id for the same time_id of the previous date. \n\nI have used [https://www.kaggle.com/code/shiyili/js24-rmf-submission-api-debug-with-synthetic-test?scriptVersionId=201579302](this) notebook to create synthetic data for testing locally before submitting.\n\nEverything works well when you only use the features in your model without using the lagged responder values. However, when you use join operation for test and lags df as follows\n\n`combined_data = test.join(lags_, on=[\"date_id\", \"time_id\", \"symbol_id\"], how=\"inner\")`\n\nThere are cases where the number of unique symbols in the lags df is 38 while the symbols we're looking to predict given in the test df are 39 (one symbol_id is missing from the lags df). This causes the join to reduce the dimensions from the expected `(39, 95)` to `(38,95)` due to a single missing symbol_id.\n\nI have verified this by printing the dimensions of all the dfs and also manually inspecting the data for the given test entry.\n\n```\nERROR: prediction dimension mismatch 38 != 39\nCombined df shape: (38, 95)\nTest df shape: (39, 85)\nLag df shape: (36784, 12)\nSymbols in lag df: shape: (38, 1)\n```\n\nThis is the output from which shows symbol_id for 31 is missing.\n`lags_.filter((pl.col(\"date_id\") == 2) & (pl.col(\"time_id\") == 0) & (pl.col(\"symbol_id\") > 25) & (pl.col(\"symbol_id\") < 34))`\n\n![image_output](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F13387107%2F451518e58aa69cb6e4da04189abd0446%2FScreenshot%202024-11-24%20at%2012.38.01AM.png?generation=1732388897175690&alt=media)\n\nI wanted to ask if this is expected or is this some sort of error in the data? As the data specifications mention, the lags_ should provide us with all the lagged responder values for the previous date_id. \n\nIf not, is it possible to handle missing lagged values for a single symbol_id randomly in the test run?\n\nI have also tried submitting the code for evaluation and it threw an exception while processing (as expected due to prediction and test dimensions mismatch)\n\nMy entire prediction code is as below for reference.\n\n```\nlags_ : pl.DataFrame | None = None\ntest_: pl.DataFrame | None = None\n\n# Replace this function with your inference code.\n# You can return either a Pandas or Polars dataframe, though Polars is recommended.\n# Each batch of predictions (except the very first) must be returned within 1 minute of the batch features being provided.\nimport xgboost as xgb\n\nloaded_model = None\n\ndef predict(test: pl.DataFrame, lags: pl.DataFrame | None) -> pl.DataFrame | pd.DataFrame:\n    \"\"\"Make a prediction.\"\"\"\n\n    # All the responders from the previous day are passed in at time_id == 0. We save them in a global variable for access at every time_id.\n    # Use them as extra features, if you like.\n    global lags_, loaded_model, test_\n    if lags is not None:\n        lags_ = lags\n\n    test_ = test\n\n    if loaded_model is None:\n        # Load model\n        loaded_model = xgb.Booster()\n        loaded_model.load_model('/kaggle/input/xgboost-lags-model-basic/other/basic/1/xgbbost_lags_1.json')\n\n\n    combined_data = test.join(lags_, on=[\"date_id\", \"time_id\", \"symbol_id\"], how=\"inner\")\n\n    # Assuming `features_*` are all columns starting with \"feature_\" or ending with \"_lag\"\n    feature_columns = [col for col in combined_data.columns if col.startswith('feature_') or col.endswith(\"lag_1\")]\n\n    # Prepare the input features\n    X_pred = combined_data.select(feature_columns).to_pandas()  # Convert Polars DataFrame to Pandas for XGBoost compatibility\n    \n    # Convert to DMatrix (required format for XGBoost predictions)\n    dmatrix_pred = xgb.DMatrix(X_pred)\n    \n    # Get predictions from the loaded model\n    predictions = loaded_model.predict(dmatrix_pred)\n\n    combined_data = combined_data.with_columns(pl.Series(name=\"responder_6\", values=predictions))\n\n    # Replace this section with your own predictions\n    predictions = combined_data.select(\n        'row_id', 'responder_6'\n    )\n\n    if len(predictions) != len(test):\n        print(f\"ERROR: prediction dimension mismatch {len(predictions)} != {len(test)}\")\n        print(f\"Combined df shape: {combined_data.shape}\")\n        print(f\"Test df shape: {test.shape}\")\n        print(f\"Lag df shape: {lags_.shape}\")\n        print(f\"Symbols in lag df: {lags_.select(pl.col('symbol_id').unique())}\")\n\n    if isinstance(predictions, pl.DataFrame):\n        assert predictions.columns == ['row_id', 'responder_6']\n    elif isinstance(predictions, pd.DataFrame):\n        assert (predictions.columns == ['row_id', 'responder_6']).all()\n    else:\n        raise TypeError('The predict function must return a DataFrame')\n    # Confirm has as many rows as the test data.\n    assert len(predictions) == len(test)\n\n    return predictions\n```",
      "votes": null
    },
    {
      "id": "3053724",
      "postDate": "11/23/2024 19:45:12",
      "content": "<p>I am facing the same issue! </p>",
      "rawMarkdown": "I am facing the same issue!",
      "votes": null
    },
    {
      "id": "3053778",
      "postDate": "11/23/2024 22:22:02",
      "content": "<p>Glad to know that my notebook is useful for you.</p>\n<p>In short, yes, it is expected that the symbol id changes from date to date. So when you join lags, it is better to use left join to ensure that symbols traded today but not yesterday do not get removed. Then you need to have an extra nan-filling to deal with the absence of lags for these missing symbols.</p>",
      "rawMarkdown": "Glad to know that my notebook is useful for you.\n\nIn short, yes, it is expected that the symbol id changes from date to date. So when you join lags, it is better to use left join to ensure that symbols traded today but not yesterday do not get removed. Then you need to have an extra nan-filling to deal with the absence of lags for these missing symbols.",
      "votes": null
    },
    {
      "id": "3053781",
      "postDate": "11/23/2024 22:32:26",
      "content": "<p>Understood. Thanks a lot for your notebook, has cleared out a lot for me!</p>",
      "rawMarkdown": "Understood. Thanks a lot for your notebook, has cleared out a lot for me!",
      "votes": null
    },
    {
      "id": "3056114",
      "postDate": "11/26/2024 15:13:24",
      "content": "<p>I wasted a lot of time on this before declaring it a dead end. There is so much nan/zero handling and padding issues. Then the windowing issue as well gets harder with missing values. Then the performance hit of having to concatenate and pad all the time made this just not worth pursuing for me but hopefully you figure it out!</p>",
      "rawMarkdown": "I wasted a lot of time on this before declaring it a dead end. There is so much nan/zero handling and padding issues. Then the windowing issue as well gets harder with missing values. Then the performance hit of having to concatenate and pad all the time made this just not worth pursuing for me but hopefully you figure it out!",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 3053724,
      "author_name": "dipitgolechha7",
      "author_url": "",
      "post_date": "11/23/2024 19:45:12",
      "content": "<p>I am facing the same issue! </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 3053778,
      "author_name": "shiyili",
      "author_url": "",
      "post_date": "11/23/2024 22:22:02",
      "content": "<p>Glad to know that my notebook is useful for you.</p>\n<p>In short, yes, it is expected that the symbol id changes from date to date. So when you join lags, it is better to use left join to ensure that symbols traded today but not yesterday do not get removed. Then you need to have an extra nan-filling to deal with the absence of lags for these missing symbols.</p>",
      "votes": null,
      "replies": [
        {
          "id": 3053781,
          "author_name": "deveshshah2003",
          "author_url": "",
          "post_date": "11/23/2024 22:32:26",
          "content": "<p>Understood. Thanks a lot for your notebook, has cleared out a lot for me!</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 3056114,
      "author_name": "michaeltimbs",
      "author_url": "",
      "post_date": "11/26/2024 15:13:24",
      "content": "<p>I wasted a lot of time on this before declaring it a dead end. There is so much nan/zero handling and padding issues. Then the windowing issue as well gets harder with missing values. Then the performance hit of having to concatenate and pad all the time made this just not worth pursuing for me but hopefully you figure it out!</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "3053701": "I am working with a model which joins the lags and test data to extract the lagged responder values for the given symbol_id for the same time_id of the previous date. \n\nI have used [https://www.kaggle.com/code/shiyili/js24-rmf-submission-api-debug-with-synthetic-test?scriptVersionId=201579302](this) notebook to create synthetic data for testing locally before submitting.\n\nEverything works well when you only use the features in your model without using the lagged responder values. However, when you use join operation for test and lags df as follows\n\n`combined_data = test.join(lags_, on=[\"date_id\", \"time_id\", \"symbol_id\"], how=\"inner\")`\n\nThere are cases where the number of unique symbols in the lags df is 38 while the symbols we're looking to predict given in the test df are 39 (one symbol_id is missing from the lags df). This causes the join to reduce the dimensions from the expected `(39, 95)` to `(38,95)` due to a single missing symbol_id.\n\nI have verified this by printing the dimensions of all the dfs and also manually inspecting the data for the given test entry.\n\n```\nERROR: prediction dimension mismatch 38 != 39\nCombined df shape: (38, 95)\nTest df shape: (39, 85)\nLag df shape: (36784, 12)\nSymbols in lag df: shape: (38, 1)\n```\n\nThis is the output from which shows symbol_id for 31 is missing.\n`lags_.filter((pl.col(\"date_id\") == 2) & (pl.col(\"time_id\") == 0) & (pl.col(\"symbol_id\") > 25) & (pl.col(\"symbol_id\") < 34))`\n\n![image_output](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F13387107%2F451518e58aa69cb6e4da04189abd0446%2FScreenshot%202024-11-24%20at%2012.38.01AM.png?generation=1732388897175690&alt=media)\n\nI wanted to ask if this is expected or is this some sort of error in the data? As the data specifications mention, the lags_ should provide us with all the lagged responder values for the previous date_id. \n\nIf not, is it possible to handle missing lagged values for a single symbol_id randomly in the test run?\n\nI have also tried submitting the code for evaluation and it threw an exception while processing (as expected due to prediction and test dimensions mismatch)\n\nMy entire prediction code is as below for reference.\n\n```\nlags_ : pl.DataFrame | None = None\ntest_: pl.DataFrame | None = None\n\n# Replace this function with your inference code.\n# You can return either a Pandas or Polars dataframe, though Polars is recommended.\n# Each batch of predictions (except the very first) must be returned within 1 minute of the batch features being provided.\nimport xgboost as xgb\n\nloaded_model = None\n\ndef predict(test: pl.DataFrame, lags: pl.DataFrame | None) -> pl.DataFrame | pd.DataFrame:\n    \"\"\"Make a prediction.\"\"\"\n\n    # All the responders from the previous day are passed in at time_id == 0. We save them in a global variable for access at every time_id.\n    # Use them as extra features, if you like.\n    global lags_, loaded_model, test_\n    if lags is not None:\n        lags_ = lags\n\n    test_ = test\n\n    if loaded_model is None:\n        # Load model\n        loaded_model = xgb.Booster()\n        loaded_model.load_model('/kaggle/input/xgboost-lags-model-basic/other/basic/1/xgbbost_lags_1.json')\n\n\n    combined_data = test.join(lags_, on=[\"date_id\", \"time_id\", \"symbol_id\"], how=\"inner\")\n\n    # Assuming `features_*` are all columns starting with \"feature_\" or ending with \"_lag\"\n    feature_columns = [col for col in combined_data.columns if col.startswith('feature_') or col.endswith(\"lag_1\")]\n\n    # Prepare the input features\n    X_pred = combined_data.select(feature_columns).to_pandas()  # Convert Polars DataFrame to Pandas for XGBoost compatibility\n    \n    # Convert to DMatrix (required format for XGBoost predictions)\n    dmatrix_pred = xgb.DMatrix(X_pred)\n    \n    # Get predictions from the loaded model\n    predictions = loaded_model.predict(dmatrix_pred)\n\n    combined_data = combined_data.with_columns(pl.Series(name=\"responder_6\", values=predictions))\n\n    # Replace this section with your own predictions\n    predictions = combined_data.select(\n        'row_id', 'responder_6'\n    )\n\n    if len(predictions) != len(test):\n        print(f\"ERROR: prediction dimension mismatch {len(predictions)} != {len(test)}\")\n        print(f\"Combined df shape: {combined_data.shape}\")\n        print(f\"Test df shape: {test.shape}\")\n        print(f\"Lag df shape: {lags_.shape}\")\n        print(f\"Symbols in lag df: {lags_.select(pl.col('symbol_id').unique())}\")\n\n    if isinstance(predictions, pl.DataFrame):\n        assert predictions.columns == ['row_id', 'responder_6']\n    elif isinstance(predictions, pd.DataFrame):\n        assert (predictions.columns == ['row_id', 'responder_6']).all()\n    else:\n        raise TypeError('The predict function must return a DataFrame')\n    # Confirm has as many rows as the test data.\n    assert len(predictions) == len(test)\n\n    return predictions\n```",
    "3053724": "I am facing the same issue!",
    "3053778": "Glad to know that my notebook is useful for you.\n\nIn short, yes, it is expected that the symbol id changes from date to date. So when you join lags, it is better to use left join to ensure that symbols traded today but not yesterday do not get removed. Then you need to have an extra nan-filling to deal with the absence of lags for these missing symbols.",
    "3053781": "Understood. Thanks a lot for your notebook, has cleared out a lot for me!",
    "3056114": "I wasted a lot of time on this before declaring it a dead end. There is so much nan/zero handling and padding issues. Then the windowing issue as well gets harder with missing values. Then the performance hit of having to concatenate and pad all the time made this just not worth pursuing for me but hopefully you figure it out!"
  },
  "source": "meta"
}