{
  "id": 545024,
  "title": "Submission Scoring Error with Lags",
  "url": "/competitions/jane-street-real-time-market-data-forecasting/discussion/545024",
  "author_name": "",
  "post_date": "2024-11-08T05:57:07.845033600Z",
  "votes": 1,
  "comment_count": 5,
  "views": 0,
  "content": "<p>when ever I included lags_ feature during submission, it report sumission scoring error. <br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4858569%2Fcf554f35254b147e62947d77ac00a3ed%2Fscore%20error.png?generation=1731045114615013&amp;alt=media\" alt=\"\"></p>\n<p>I did debug, it all passed</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4858569%2Fc2fbb705772dc60e38787e4822f35b63%2Fdebug.png?generation=1731045149268263&amp;alt=media\" alt=\"\"></p>\n<p>the log shows nothing wrong</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4858569%2F0f36f9cfb63d59f9cb895debe3387979%2Flog.png?generation=1731045175150478&amp;alt=media\" alt=\"\"></p>\n<pre><code> () -&gt; pl.DataFrame:\n     lags_, mlp_model  \n\n\n    \n     lags   :\n        lags_ = lags\n\n\n    \n    test_df = test.to_pandas()\n    lags_df = lags_.to_pandas()\n\n\n    X_test = pd.merge(test_df, lags_df, on=[, ], how=, suffixes = (,))\n\n    \n    X_test = X_test[feature_names].values\n\n    prediction  = model.predict(X_test)\n\n    prediction_df = pd.DataFrame(prediction)\n    responder_6 = prediction_df.iloc[:, ].values\n\n    output_df = pd.DataFrame({: test_df[], : responder_6})\n     pl.from_pandas(output_df)\n</code></pre>\n<p>any ideas?</p>",
  "messages": [
    {
      "id": "3039508",
      "postDate": "11/08/2024 05:57:07",
      "content": "<p>when ever I included lags_ feature during submission, it report sumission scoring error. <br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4858569%2Fcf554f35254b147e62947d77ac00a3ed%2Fscore%20error.png?generation=1731045114615013&amp;alt=media\" alt=\"\"></p>\n<p>I did debug, it all passed</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4858569%2Fc2fbb705772dc60e38787e4822f35b63%2Fdebug.png?generation=1731045149268263&amp;alt=media\" alt=\"\"></p>\n<p>the log shows nothing wrong</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4858569%2F0f36f9cfb63d59f9cb895debe3387979%2Flog.png?generation=1731045175150478&amp;alt=media\" alt=\"\"></p>\n<pre><code> () -&gt; pl.DataFrame:\n     lags_, mlp_model  \n\n\n    \n     lags   :\n        lags_ = lags\n\n\n    \n    test_df = test.to_pandas()\n    lags_df = lags_.to_pandas()\n\n\n    X_test = pd.merge(test_df, lags_df, on=[, ], how=, suffixes = (,))\n\n    \n    X_test = X_test[feature_names].values\n\n    prediction  = model.predict(X_test)\n\n    prediction_df = pd.DataFrame(prediction)\n    responder_6 = prediction_df.iloc[:, ].values\n\n    output_df = pd.DataFrame({: test_df[], : responder_6})\n     pl.from_pandas(output_df)\n</code></pre>\n<p>any ideas?</p>",
      "rawMarkdown": "when ever I included lags_ feature during submission, it report sumission scoring error. \n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4858569%2Fcf554f35254b147e62947d77ac00a3ed%2Fscore%20error.png?generation=1731045114615013&alt=media)\n\n\n\nI did debug, it all passed\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4858569%2Fc2fbb705772dc60e38787e4822f35b63%2Fdebug.png?generation=1731045149268263&alt=media)\n\nthe log shows nothing wrong\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4858569%2F0f36f9cfb63d59f9cb895debe3387979%2Flog.png?generation=1731045175150478&alt=media)\n\n\n\n```python\ndef predict(test: pl.DataFrame, lags: pl.DataFrame | None) -> pl.DataFrame:\n    global lags_, mlp_model  # Declare models as global\n    \n    \n    # Logic for saving or loading lags\n    if lags is not None:\n        lags_ = lags\n\n\n    # Convert to Pandas DataFrame if needed (since model expects numpy arrays)\n    test_df = test.to_pandas()\n    lags_df = lags_.to_pandas()\n    \n\n    X_test = pd.merge(test_df, lags_df, on=['time_id', 'symbol_id'], how='left', suffixes = (\"\",\"_drop\"))\n\n    # Extract features for prediction\n    X_test = X_test[feature_names].values\n\n    prediction  = model.predict(X_test)\n\n    prediction_df = pd.DataFrame(prediction)\n    responder_6 = prediction_df.iloc[:, 6].values\n\n    output_df = pd.DataFrame({\"row_id\": test_df['row_id'], \"responder_6\": responder_6})\n    return pl.from_pandas(output_df)\n```\n\nany ideas?",
      "votes": null
    },
    {
      "id": "3039538",
      "postDate": "11/08/2024 06:51:38",
      "content": "<p>In this line:<br>\n<code>X_test = pd.merge(test_df, lags_df, on=['time_id', 'symbol_id'], how='left', suffixes = (\"\",\"_drop\"))</code><br>\nsome (time_id, symbol_id) pairs could only appear in test_df (not showing in lags_df), so for the features in lags_df used in feature_names, they will be missing values, which is apparently not acceptable for neural network models. Action needed: do something to deal with this, for example imputing missing values. </p>",
      "rawMarkdown": "In this line:\n` X_test = pd.merge(test_df, lags_df, on=['time_id', 'symbol_id'], how='left', suffixes = (\"\",\"_drop\"))`\nsome (time_id, symbol_id) pairs could only appear in test_df (not showing in lags_df), so for the features in lags_df used in feature_names, they will be missing values, which is apparently not acceptable for neural network models. Action needed: do something to deal with this, for example imputing missing values.",
      "votes": null
    },
    {
      "id": "3039546",
      "postDate": "11/08/2024 07:07:45",
      "content": "<p>Will the \"lag\" not having the (time_id, symbol_id) at all or only some rows may be NaN?<br>\nso far I know is that, the lag was given at the beginning of each day, like</p>\n<p>date_id==0, time_id==0, we will get \"lag\"  and \"test\"<br>\ndate_id==0, time_id==1, we will get only \"test\" <br>\n….<br>\nuntil<br>\ndate_id==1, time_id==0, we will get \"lag\"  and \"test\" again.</p>\n<p>if the \"lag\" (time_id, symbol_id) feature is entirly missing. it would be hard to make one.<br>\nif just some NaNs, I suppose we could do forward fill.</p>\n<p>but I think you are right. <br>\nit could be lots NaNs so that even I ran the model, I got NaN predictions, which led to the model unable to give me a score</p>",
      "rawMarkdown": "Will the \"lag\" not having the (time_id, symbol_id) at all or only some rows may be NaN?\nso far I know is that, the lag was given at the beginning of each day, like\n\ndate_id==0, time_id==0, we will get \"lag\"  and \"test\"\ndate_id==0, time_id==1, we will get only \"test\" \n....\nuntil\ndate_id==1, time_id==0, we will get \"lag\"  and \"test\" again.\n\nif the \"lag\" (time_id, symbol_id) feature is entirly missing. it would be hard to make one.\nif just some NaNs, I suppose we could do forward fill.\n\nbut I think you are right. \nit could be lots NaNs so that even I ran the model, I got NaN predictions, which led to the model unable to give me a score",
      "votes": null
    },
    {
      "id": "3039551",
      "postDate": "11/08/2024 07:14:46",
      "content": "<p>For example in date_id == 2, if there are in total 38 symbol_ids, then when date_id == 3 and time_id == 0, what you could have is test having 39 symbol_ids and lags having only 38 coresponding symbol_ids to previous date_id == 2. So clearly there is a mismatch. </p>",
      "rawMarkdown": "For example in date_id == 2, if there are in total 38 symbol_ids, then when date_id == 3 and time_id == 0, what you could have is test having 39 symbol_ids and lags having only 38 coresponding symbol_ids to previous date_id == 2. So clearly there is a mismatch.",
      "votes": null
    },
    {
      "id": "3039570",
      "postDate": "11/08/2024 07:24:31",
      "content": "<p>I see! thats what I missed. Thanks a lot! 😁</p>",
      "rawMarkdown": "I see! thats what I missed. Thanks a lot! 😁",
      "votes": null
    },
    {
      "id": "3040111",
      "postDate": "11/08/2024 17:42:12",
      "content": "<p>I suggest using the synthetic test data to debug. The test.parquet contains only one time_id and can not effectively cover all coner cases. </p>",
      "rawMarkdown": "I suggest using the synthetic test data to debug. The test.parquet contains only one time_id and can not effectively cover all coner cases.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 3039538,
      "author_name": "lihaorocky",
      "author_url": "",
      "post_date": "11/08/2024 06:51:38",
      "content": "<p>In this line:<br>\n<code>X_test = pd.merge(test_df, lags_df, on=['time_id', 'symbol_id'], how='left', suffixes = (\"\",\"_drop\"))</code><br>\nsome (time_id, symbol_id) pairs could only appear in test_df (not showing in lags_df), so for the features in lags_df used in feature_names, they will be missing values, which is apparently not acceptable for neural network models. Action needed: do something to deal with this, for example imputing missing values. </p>",
      "votes": null,
      "replies": [
        {
          "id": 3039546,
          "author_name": "zoutain",
          "author_url": "",
          "post_date": "11/08/2024 07:07:45",
          "content": "<p>Will the \"lag\" not having the (time_id, symbol_id) at all or only some rows may be NaN?<br>\nso far I know is that, the lag was given at the beginning of each day, like</p>\n<p>date_id==0, time_id==0, we will get \"lag\"  and \"test\"<br>\ndate_id==0, time_id==1, we will get only \"test\" <br>\n….<br>\nuntil<br>\ndate_id==1, time_id==0, we will get \"lag\"  and \"test\" again.</p>\n<p>if the \"lag\" (time_id, symbol_id) feature is entirly missing. it would be hard to make one.<br>\nif just some NaNs, I suppose we could do forward fill.</p>\n<p>but I think you are right. <br>\nit could be lots NaNs so that even I ran the model, I got NaN predictions, which led to the model unable to give me a score</p>",
          "votes": null,
          "replies": [
            {
              "id": 3039551,
              "author_name": "lihaorocky",
              "author_url": "",
              "post_date": "11/08/2024 07:14:46",
              "content": "<p>For example in date_id == 2, if there are in total 38 symbol_ids, then when date_id == 3 and time_id == 0, what you could have is test having 39 symbol_ids and lags having only 38 coresponding symbol_ids to previous date_id == 2. So clearly there is a mismatch. </p>",
              "votes": null,
              "replies": [
                {
                  "id": 3039570,
                  "author_name": "zoutain",
                  "author_url": "",
                  "post_date": "11/08/2024 07:24:31",
                  "content": "<p>I see! thats what I missed. Thanks a lot! 😁</p>",
                  "votes": null,
                  "replies": []
                }
              ]
            }
          ]
        }
      ]
    },
    {
      "id": 3040111,
      "author_name": "shiyili",
      "author_url": "",
      "post_date": "11/08/2024 17:42:12",
      "content": "<p>I suggest using the synthetic test data to debug. The test.parquet contains only one time_id and can not effectively cover all coner cases. </p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "3039508": "when ever I included lags_ feature during submission, it report sumission scoring error. \n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4858569%2Fcf554f35254b147e62947d77ac00a3ed%2Fscore%20error.png?generation=1731045114615013&alt=media)\n\n\n\nI did debug, it all passed\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4858569%2Fc2fbb705772dc60e38787e4822f35b63%2Fdebug.png?generation=1731045149268263&alt=media)\n\nthe log shows nothing wrong\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4858569%2F0f36f9cfb63d59f9cb895debe3387979%2Flog.png?generation=1731045175150478&alt=media)\n\n\n\n```python\ndef predict(test: pl.DataFrame, lags: pl.DataFrame | None) -> pl.DataFrame:\n    global lags_, mlp_model  # Declare models as global\n    \n    \n    # Logic for saving or loading lags\n    if lags is not None:\n        lags_ = lags\n\n\n    # Convert to Pandas DataFrame if needed (since model expects numpy arrays)\n    test_df = test.to_pandas()\n    lags_df = lags_.to_pandas()\n    \n\n    X_test = pd.merge(test_df, lags_df, on=['time_id', 'symbol_id'], how='left', suffixes = (\"\",\"_drop\"))\n\n    # Extract features for prediction\n    X_test = X_test[feature_names].values\n\n    prediction  = model.predict(X_test)\n\n    prediction_df = pd.DataFrame(prediction)\n    responder_6 = prediction_df.iloc[:, 6].values\n\n    output_df = pd.DataFrame({\"row_id\": test_df['row_id'], \"responder_6\": responder_6})\n    return pl.from_pandas(output_df)\n```\n\nany ideas?",
    "3039538": "In this line:\n` X_test = pd.merge(test_df, lags_df, on=['time_id', 'symbol_id'], how='left', suffixes = (\"\",\"_drop\"))`\nsome (time_id, symbol_id) pairs could only appear in test_df (not showing in lags_df), so for the features in lags_df used in feature_names, they will be missing values, which is apparently not acceptable for neural network models. Action needed: do something to deal with this, for example imputing missing values.",
    "3039546": "Will the \"lag\" not having the (time_id, symbol_id) at all or only some rows may be NaN?\nso far I know is that, the lag was given at the beginning of each day, like\n\ndate_id==0, time_id==0, we will get \"lag\"  and \"test\"\ndate_id==0, time_id==1, we will get only \"test\" \n....\nuntil\ndate_id==1, time_id==0, we will get \"lag\"  and \"test\" again.\n\nif the \"lag\" (time_id, symbol_id) feature is entirly missing. it would be hard to make one.\nif just some NaNs, I suppose we could do forward fill.\n\nbut I think you are right. \nit could be lots NaNs so that even I ran the model, I got NaN predictions, which led to the model unable to give me a score",
    "3039551": "For example in date_id == 2, if there are in total 38 symbol_ids, then when date_id == 3 and time_id == 0, what you could have is test having 39 symbol_ids and lags having only 38 coresponding symbol_ids to previous date_id == 2. So clearly there is a mismatch.",
    "3039570": "I see! thats what I missed. Thanks a lot! 😁",
    "3040111": "I suggest using the synthetic test data to debug. The test.parquet contains only one time_id and can not effectively cover all coner cases."
  },
  "source": "meta"
}