{
  "id": 548100,
  "title": "I haven't yet figure out how the inputs are provided in inference stage..",
  "url": "/competitions/jane-street-real-time-market-data-forecasting/discussion/548100",
  "author_name": "",
  "post_date": "2024-11-25T06:55:45.175167600Z",
  "votes": null,
  "comment_count": 7,
  "views": 0,
  "content": "<p>I am confusing yet. Which case is right?<br>\nI think that case 1 is more reasonable considering leakage. (ex. imputation nan with next day in lag)<br>\n[Case1]<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F6610354%2Ffd930e9643fd4d9508fbecaf94ba5a88%2Fcase1.png?generation=1732517447727717&amp;alt=media\" alt=\"\"><br>\n[Case2]<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F6610354%2Fb358c8ec93f321fb756eeadf33e11308%2Fcase2.png?generation=1732517454750318&amp;alt=media\" alt=\"\"></p>",
  "messages": [
    {
      "id": "3054786",
      "postDate": "11/25/2024 06:55:45",
      "content": "<p>I am confusing yet. Which case is right?<br>\nI think that case 1 is more reasonable considering leakage. (ex. imputation nan with next day in lag)<br>\n[Case1]<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F6610354%2Ffd930e9643fd4d9508fbecaf94ba5a88%2Fcase1.png?generation=1732517447727717&amp;alt=media\" alt=\"\"><br>\n[Case2]<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F6610354%2Fb358c8ec93f321fb756eeadf33e11308%2Fcase2.png?generation=1732517454750318&amp;alt=media\" alt=\"\"></p>",
      "rawMarkdown": "I am confusing yet. Which case is right?\nI think that case 1 is more reasonable considering leakage. (ex. imputation nan with next day in lag)\n[Case1]\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F6610354%2Ffd930e9643fd4d9508fbecaf94ba5a88%2Fcase1.png?generation=1732517447727717&alt=media)\n[Case2]\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F6610354%2Fb358c8ec93f321fb756eeadf33e11308%2Fcase2.png?generation=1732517454750318&alt=media)",
      "votes": null
    },
    {
      "id": "3055139",
      "postDate": "11/25/2024 15:15:15",
      "content": "<p>It is the first one.</p>\n<p>For example:</p>\n<p>Say your predict function is being called for <code>date_id = D</code>. Then, inside your predict function, at the first <code>time_id = T</code> of <code>date_id = D</code>, you are given:</p>\n<ol>\n<li><code>test</code> dataframe - data (feature_01, 02, weight… etc) corresponding to <code>time_id = T</code> of <code>date_id D</code> for all <code>symbol_ids</code> available at that <code>date_id_time_id</code>.</li>\n<li><code>lags</code> dataframe - 1 day lagged responder values, as well as <code>time_id, symbol_id</code> for all <code>time_ids</code>, and all available <code>symbol_ids</code> for <code>date_id D-1</code>. The unique values of <code>time_ids</code> and <code>symbol_ids</code> are not guaranteed to be exactly the same as what is given in the <code>test</code> dataframe. Note that in this <code>lags</code> df, the date_id = D i.e. same as test df, even though the columns represent the lagged responders of previous day.</li>\n</ol>\n<p>On all other time_ids for date_id = D, lags will be None. i.e. lags df will be non-null only on the first time_id of a given date_id.</p>",
      "rawMarkdown": "It is the first one.\n\nFor example:\n\nSay your predict function is being called for ```date_id = D```. Then, inside your predict function, at the first ```time_id = T``` of ```date_id = D```, you are given:\n1. ```test``` dataframe - data (feature_01, 02, weight... etc) corresponding to ```time_id = T``` of ```date_id D``` for all ```symbol_ids``` available at that ```date_id_time_id```.\n2. ```lags``` dataframe - 1 day lagged responder values, as well as ```time_id, symbol_id``` for all ```time_ids```, and all available ```symbol_ids``` for ```date_id D-1```. The unique values of ```time_ids``` and ```symbol_ids``` are not guaranteed to be exactly the same as what is given in the ```test``` dataframe. Note that in this ```lags``` df, the date_id = D i.e. same as test df, even though the columns represent the lagged responders of previous day.\n\n\nOn all other time_ids for date_id = D, lags will be None. i.e. lags df will be non-null only on the first time_id of a given date_id.",
      "votes": null
    },
    {
      "id": "3055177",
      "postDate": "11/25/2024 15:59:59",
      "content": "<p>however, the lags will be saved in a global variable lags_ which we can access when the time_id is not 0, correct?<br>\ni`m getting an error as i try to merge the lags_ dataframe with the test when time_id is not zero.<br>\nAny tips? </p>\n<p>`lags_ : pl.DataFrame | None = None</p>\n<p>def predict(test: pl.DataFrame, lags: pl.DataFrame | None) -&gt; pl.DataFrame | pd.DataFrame:</p>\n<pre><code> lags is not None:\n    lags_ = lags \n\ntest1= test(lags_, on=, how=)\n)\nfeat = test1()\npred = model(feat)\npredictions = test()(pl(, pred()))\nassert (predictions, (pl, pd.DataFrame))\nassert (predictions.) == [, ]\nassert (predictions) == (test)\nreturn predictions\n</code></pre>\n<p>`</p>",
      "rawMarkdown": "however, the lags will be saved in a global variable lags_ which we can access when the time_id is not 0, correct?\ni`m getting an error as i try to merge the lags_ dataframe with the test when time_id is not zero.\nAny tips? \n\n`lags_ : pl.DataFrame | None = None\n\ndef predict(test: pl.DataFrame, lags: pl.DataFrame | None) -> pl.DataFrame | pd.DataFrame:\n\n    if lags is not None:\n        lags_ = lags \n\n    test1= test.join(lags_, on=['date_id', 'time_id', 'symbol_id'], how='left')\n    print(test1.head(10))\n    feat = test1[feature_names].to_numpy()\n    pred = model.predict(feat)\n    predictions = test.select('row_id').with_columns(pl.Series('responder_6', pred.ravel()))\n    assert isinstance(predictions, (pl.DataFrame, pd.DataFrame))\n    assert list(predictions.columns) == ['row_id', 'responder_6']\n    assert len(predictions) == len(test)\n    return predictions\n`",
      "votes": null
    },
    {
      "id": "3055205",
      "postDate": "11/25/2024 16:27:45",
      "content": "<p>what was your error?</p>\n<p>your code to join the lags with test seems ok. However, your issue might be related to handling of NaNs when calling model.predict.</p>",
      "rawMarkdown": "what was your error?\n\nyour code to join the lags with test seems ok. However, your issue might be related to handling of NaNs when calling model.predict.",
      "votes": null
    },
    {
      "id": "3055245",
      "postDate": "11/25/2024 16:44:36",
      "content": "<p>lgbm, no issues with predict. Funny thing is that I look at the logs and there is no errors there. Only at the submissions page:<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F23236619%2Fee9b3f4289f6de1cf298fa9a2219a04b%2Fsub.png?generation=1732552997114172&amp;alt=media\" alt=\"\"></p>\n<p>And the log (I was trying to print some shapes to see what I was missing) No errors:<br>\n32.3s    1   CPU times: user 13.8 s, sys: 16 s, total: 29.8 s<br>\n32.3s    2   Wall time: 18 s<br>\n43.9s    3   /opt/conda/lib/python3.10/site-packages/lightgbm/engine.py:172: UserWarning: Found <code>n_estimators</code> in params. Will use it instead of argument<br>\n43.9s    4     <em>log_warning(f\"Found <code>{alias}</code> in params. Will use it instead of argument\")\n843.2s    5   CPU times: user 48min 19s, sys: 13.7 s, total: 48min 32s\n843.2s    6   Wall time: 13min 19s\n843.9s    7   (39, 12)\n843.9s    8   shape: (1, 1)\n843.9s    9   ┌─────────┐\n843.9s    10  │ date_id │\n843.9s    11  │ ---     │\n843.9s    12  │ i16     │\n843.9s    13  ╞═════════╡\n843.9s    14  │ 0       │\n843.9s    15  └─────────┘\n843.9s    16  shape: (1, 1)\n843.9s    17  ┌─────────┐\n843.9s    18  │ time_id │\n843.9s    19  │ ---     │\n843.9s    20  │ i16     │\n843.9s    21  ╞═════════╡\n843.9s    22  │ 0       │\n843.9s    23  └─────────┘\n843.9s    24  shape: (39, 1)\n843.9s    25  ┌───────────┐\n843.9s    26  │ symbol_id │\n843.9s    27  │ ---       │\n843.9s    28  │ i8        │\n843.9s    29  ╞═══════════╡\n843.9s    30  │ 23        │\n843.9s    31  │ 4         │\n843.9s    32  │ 26        │\n843.9s    33  │ 32        │\n843.9s    34  │ 21        │\n843.9s    35  │ …         │\n843.9s    36  │ 31        │\n843.9s    37  │ 33        │\n843.9s    38  │ 28        │\n843.9s    39  │ 17        │\n843.9s    40  │ 11        │\n843.9s    41  └───────────┘\n843.9s    42  test1\n843.9s    43  shape: (10, 94)\n843.9s    44  ┌────────┬─────────┬─────────┬───────────┬───┬─────────────┬─────────────┬────────────┬────────────┐\n843.9s    45  │ row_id ┆ date_id ┆ time_id ┆ symbol_id ┆ … ┆ responder_5 ┆ responder_6 ┆ responder</em> ┆ responder_ │<br>\n843.9s    46  │ ---    ┆ ---     ┆ ---     ┆ ---       ┆   ┆ _lag_1      ┆ _lag_1      ┆ 7_lag_1    ┆ 8_lag_1    │<br>\n843.9s    47  │ i64    ┆ i16     ┆ i16     ┆ i8        ┆   ┆ ---         ┆ ---         ┆ ---        ┆ ---        │<br>\n843.9s    48  │        ┆         ┆         ┆           ┆   ┆ f32         ┆ f32         ┆ f32        ┆ f32        │<br>\n843.9s    49  ╞════════╪═════════╪═════════╪═══════════╪═══╪═════════════╪═════════════╪════════════╪════════════╡<br>\n843.9s    50  │ 0      ┆ 0       ┆ 0       ┆ 0         ┆ … ┆ -0.036595   ┆ -1.305746   ┆ -0.795677  ┆ -0.143724  │<br>\n843.9s    51  │ 1      ┆ 0       ┆ 0       ┆ 1         ┆ … ┆ -0.615652   ┆ -1.162801   ┆ -1.205924  ┆ -1.245934  │<br>\n843.9s    52  │ 2      ┆ 0       ┆ 0       ┆ 2         ┆ … ┆ -0.378265   ┆ -1.57429    ┆ -1.863071  ┆ -0.027343  │<br>\n843.9s    53  │ 3      ┆ 0       ┆ 0       ┆ 3         ┆ … ┆ -0.054984   ┆ 0.329152    ┆ -0.965471  ┆ 0.576635   │<br>\n843.9s    54  │ 4      ┆ 0       ┆ 0       ┆ 4         ┆ … ┆ -0.597093   ┆ 0.219856    ┆ -0.276356  ┆ -0.90479   │<br>\n843.9s    55  │ 5      ┆ 0       ┆ 0       ┆ 5         ┆ … ┆ -0.24035    ┆ -0.913801   ┆ -0.548867  ┆ -1.283726  │<br>\n843.9s    56  │ 6      ┆ 0       ┆ 0       ┆ 6         ┆ … ┆ 0.517728    ┆ 0.896325    ┆ 1.068884   ┆ 1.57929    │<br>\n843.9s    57  │ 7      ┆ 0       ┆ 0       ┆ 7         ┆ … ┆ -0.182887   ┆ -0.492168   ┆ -0.142915  ┆ -0.202081  │<br>\n843.9s    58  │ 8      ┆ 0       ┆ 0       ┆ 8         ┆ … ┆ -0.225042   ┆ 0.95646     ┆ 2.185598   ┆ -0.435856  │<br>\n843.9s    59  │ 9      ┆ 0       ┆ 0       ┆ 9         ┆ … ┆ -0.774299   ┆ -0.716492   ┆ -1.471419  ┆ -1.107083  │<br>\n843.9s    60  └────────┴─────────┴─────────┴───────────┴───┴─────────────┴─────────────┴────────────┴────────────┘<br>\n849.4s    61  [NbConvertApp] Converting notebook <strong>notebook</strong>.ipynb to notebook<br>\n850.0s    62  [NbConvertApp] Writing 40602 bytes to <strong>notebook</strong>.ipynb<br>\n852.4s    63  [NbConvertApp] Converting notebook <strong>notebook</strong>.ipynb to html<br>\n853.6s    64  [NbConvertApp] Writing 332474 bytes to <strong>results</strong>.html</p>",
      "rawMarkdown": "lgbm, no issues with predict. Funny thing is that I look at the logs and there is no errors there. Only at the submissions page:\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F23236619%2Fee9b3f4289f6de1cf298fa9a2219a04b%2Fsub.png?generation=1732552997114172&alt=media)\n\nAnd the log (I was trying to print some shapes to see what I was missing) No errors:\n32.3s\t1\tCPU times: user 13.8 s, sys: 16 s, total: 29.8 s\n32.3s\t2\tWall time: 18 s\n43.9s\t3\t/opt/conda/lib/python3.10/site-packages/lightgbm/engine.py:172: UserWarning: Found `n_estimators` in params. Will use it instead of argument\n43.9s\t4\t  _log_warning(f\"Found `{alias}` in params. Will use it instead of argument\")\n843.2s\t5\tCPU times: user 48min 19s, sys: 13.7 s, total: 48min 32s\n843.2s\t6\tWall time: 13min 19s\n843.9s\t7\t(39, 12)\n843.9s\t8\tshape: (1, 1)\n843.9s\t9\t┌─────────┐\n843.9s\t10\t│ date_id │\n843.9s\t11\t│ ---     │\n843.9s\t12\t│ i16     │\n843.9s\t13\t╞═════════╡\n843.9s\t14\t│ 0       │\n843.9s\t15\t└─────────┘\n843.9s\t16\tshape: (1, 1)\n843.9s\t17\t┌─────────┐\n843.9s\t18\t│ time_id │\n843.9s\t19\t│ ---     │\n843.9s\t20\t│ i16     │\n843.9s\t21\t╞═════════╡\n843.9s\t22\t│ 0       │\n843.9s\t23\t└─────────┘\n843.9s\t24\tshape: (39, 1)\n843.9s\t25\t┌───────────┐\n843.9s\t26\t│ symbol_id │\n843.9s\t27\t│ ---       │\n843.9s\t28\t│ i8        │\n843.9s\t29\t╞═══════════╡\n843.9s\t30\t│ 23        │\n843.9s\t31\t│ 4         │\n843.9s\t32\t│ 26        │\n843.9s\t33\t│ 32        │\n843.9s\t34\t│ 21        │\n843.9s\t35\t│ …         │\n843.9s\t36\t│ 31        │\n843.9s\t37\t│ 33        │\n843.9s\t38\t│ 28        │\n843.9s\t39\t│ 17        │\n843.9s\t40\t│ 11        │\n843.9s\t41\t└───────────┘\n843.9s\t42\ttest1\n843.9s\t43\tshape: (10, 94)\n843.9s\t44\t┌────────┬─────────┬─────────┬───────────┬───┬─────────────┬─────────────┬────────────┬────────────┐\n843.9s\t45\t│ row_id ┆ date_id ┆ time_id ┆ symbol_id ┆ … ┆ responder_5 ┆ responder_6 ┆ responder_ ┆ responder_ │\n843.9s\t46\t│ ---    ┆ ---     ┆ ---     ┆ ---       ┆   ┆ _lag_1      ┆ _lag_1      ┆ 7_lag_1    ┆ 8_lag_1    │\n843.9s\t47\t│ i64    ┆ i16     ┆ i16     ┆ i8        ┆   ┆ ---         ┆ ---         ┆ ---        ┆ ---        │\n843.9s\t48\t│        ┆         ┆         ┆           ┆   ┆ f32         ┆ f32         ┆ f32        ┆ f32        │\n843.9s\t49\t╞════════╪═════════╪═════════╪═══════════╪═══╪═════════════╪═════════════╪════════════╪════════════╡\n843.9s\t50\t│ 0      ┆ 0       ┆ 0       ┆ 0         ┆ … ┆ -0.036595   ┆ -1.305746   ┆ -0.795677  ┆ -0.143724  │\n843.9s\t51\t│ 1      ┆ 0       ┆ 0       ┆ 1         ┆ … ┆ -0.615652   ┆ -1.162801   ┆ -1.205924  ┆ -1.245934  │\n843.9s\t52\t│ 2      ┆ 0       ┆ 0       ┆ 2         ┆ … ┆ -0.378265   ┆ -1.57429    ┆ -1.863071  ┆ -0.027343  │\n843.9s\t53\t│ 3      ┆ 0       ┆ 0       ┆ 3         ┆ … ┆ -0.054984   ┆ 0.329152    ┆ -0.965471  ┆ 0.576635   │\n843.9s\t54\t│ 4      ┆ 0       ┆ 0       ┆ 4         ┆ … ┆ -0.597093   ┆ 0.219856    ┆ -0.276356  ┆ -0.90479   │\n843.9s\t55\t│ 5      ┆ 0       ┆ 0       ┆ 5         ┆ … ┆ -0.24035    ┆ -0.913801   ┆ -0.548867  ┆ -1.283726  │\n843.9s\t56\t│ 6      ┆ 0       ┆ 0       ┆ 6         ┆ … ┆ 0.517728    ┆ 0.896325    ┆ 1.068884   ┆ 1.57929    │\n843.9s\t57\t│ 7      ┆ 0       ┆ 0       ┆ 7         ┆ … ┆ -0.182887   ┆ -0.492168   ┆ -0.142915  ┆ -0.202081  │\n843.9s\t58\t│ 8      ┆ 0       ┆ 0       ┆ 8         ┆ … ┆ -0.225042   ┆ 0.95646     ┆ 2.185598   ┆ -0.435856  │\n843.9s\t59\t│ 9      ┆ 0       ┆ 0       ┆ 9         ┆ … ┆ -0.774299   ┆ -0.716492   ┆ -1.471419  ┆ -1.107083  │\n843.9s\t60\t└────────┴─────────┴─────────┴───────────┴───┴─────────────┴─────────────┴────────────┴────────────┘\n849.4s\t61\t[NbConvertApp] Converting notebook __notebook__.ipynb to notebook\n850.0s\t62\t[NbConvertApp] Writing 40602 bytes to __notebook__.ipynb\n852.4s\t63\t[NbConvertApp] Converting notebook __notebook__.ipynb to html\n853.6s\t64\t[NbConvertApp] Writing 332474 bytes to __results__.html",
      "votes": null
    },
    {
      "id": "3055889",
      "postDate": "11/26/2024 08:50:30",
      "content": "<p>It is the code below I did test.<br>\nI found that <strong>I explicitly select the samples on time_id=0 in lag data</strong>.<br>\nI don't know why I do this… (since I don't know how the lag is provided actually)</p>\n<pre><code> () -&gt; pl.DataFrame | pd.DataFrame:\n     lags_\n     lags   :\n        \n        lags_ = lags.group_by([, ], maintain_order=).first().drop()\n        \n        \n     lags_  :\n        lags_ = test.select(\n            pl.col(), pl.col(),\n            *[pl.lit().alias()  col  CFG.responder_cols]\n        )\n\n    \n    df_test = preprocessor.run(test.drop([, ]), lags_, mode=)\n    df_test = df_test.drop(CFG.metadata_cols).to_pandas()\n\n    fold_pred = []\n     fold  (CFG.n_folds):\n        ()\n        \n        test_x = df_test.copy()\n        \n        scaler = fold_output[fold][]\n        test_x[var_info[]] = scaler.transform(test_x[var_info[]])\n        encoder = fold_output[fold][]\n        ohe_cols = fold_output[fold][]\n        test_x = pd.concat([\n            test_x.drop(var_info[], axis=),\n            pd.DataFrame(encoder.transform(test_x[var_info[]]), index=test_x.index.copy(), columns=ohe_cols)\n        ], axis=)\n        (, test_x.shape)\n        \n        model = fold_output[fold][]\n        \n        test_pred = model.predict(test_x)\n        \n        test_pred = np.clip(test_pred, a_min=-, a_max=)\n        fold_pred.append(test_pred)\n\n    fold_pred = np.stack(fold_pred, axis=).mean(axis=)\n    predictions = test.select(, pl.Series(CFG.target_col, fold_pred, dtype=pl.datatypes.Float32))\n     predictions\n</code></pre>",
      "rawMarkdown": "It is the code below I did test.\nI found that **I explicitly select the samples on time_id=0 in lag data**.\nI don't know why I do this... (since I don't know how the lag is provided actually)\n\n```python\ndef predict(test: pl.DataFrame, lags: pl.DataFrame | None) -> pl.DataFrame | pd.DataFrame:\n    global lags_\n    if lags is not None:\n        # IMPORTANT -> select the samples time_id=0 with group_by before merge with test data\n        lags_ = lags.group_by([\"date_id\", \"symbol_id\"], maintain_order=True).first().drop(\"time_id\")\n        # Below lags_ cause error when merging with test\n        # lags_ = lags.drop(\"time_id\")\n    if lags_ is None:\n        lags_ = test.select(\n            pl.col(\"date_id\"), pl.col(\"symbol_id\"),\n            *[pl.lit(None).alias(f\"{col}_lag_1\") for col in CFG.responder_cols]\n        )\n    \n    \"\"\"Make a prediction.\"\"\"\n    df_test = preprocessor.run(test.drop([\"row_id\", \"is_scored\"]), lags_, mode=\"test\")\n    df_test = df_test.drop(CFG.metadata_cols).to_pandas()\n    \n    fold_pred = []\n    for fold in range(CFG.n_folds):\n        print(f\"\\n=== FOLD {fold} ===\")\n        # split train & test\n        test_x = df_test.copy()\n        # scaling & encoding\n        scaler = fold_output[fold][\"scaler\"]\n        test_x[var_info[\"num\"]] = scaler.transform(test_x[var_info[\"num\"]])\n        encoder = fold_output[fold][\"encoder\"]\n        ohe_cols = fold_output[fold][\"ohe_cols\"]\n        test_x = pd.concat([\n            test_x.drop(var_info[\"cat\"], axis=1),\n            pd.DataFrame(encoder.transform(test_x[var_info[\"cat\"]]), index=test_x.index.copy(), columns=ohe_cols)\n        ], axis=1)\n        print(\"shape info ->\", test_x.shape)\n        # create model\n        model = fold_output[fold][\"model\"]\n        # inference\n        test_pred = model.predict(test_x)\n        # target clipping\n        test_pred = np.clip(test_pred, a_min=-5.0, a_max=5.0)\n        fold_pred.append(test_pred)\n\n    fold_pred = np.stack(fold_pred, axis=0).mean(axis=0)\n    predictions = test.select(\"row_id\", pl.Series(CFG.target_col, fold_pred, dtype=pl.datatypes.Float32))\n    return predictions\n```",
      "votes": null
    },
    {
      "id": "3055890",
      "postDate": "11/26/2024 08:50:36",
      "content": "<p>Thanks for your explanation :)</p>",
      "rawMarkdown": "Thanks for your explanation :)",
      "votes": null
    },
    {
      "id": "3055969",
      "postDate": "11/26/2024 11:14:45",
      "content": "<p>Thanks! Really helpful! </p>",
      "rawMarkdown": "Thanks! Really helpful!",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 3055139,
      "author_name": "alamjs",
      "author_url": "",
      "post_date": "11/25/2024 15:15:15",
      "content": "<p>It is the first one.</p>\n<p>For example:</p>\n<p>Say your predict function is being called for <code>date_id = D</code>. Then, inside your predict function, at the first <code>time_id = T</code> of <code>date_id = D</code>, you are given:</p>\n<ol>\n<li><code>test</code> dataframe - data (feature_01, 02, weight… etc) corresponding to <code>time_id = T</code> of <code>date_id D</code> for all <code>symbol_ids</code> available at that <code>date_id_time_id</code>.</li>\n<li><code>lags</code> dataframe - 1 day lagged responder values, as well as <code>time_id, symbol_id</code> for all <code>time_ids</code>, and all available <code>symbol_ids</code> for <code>date_id D-1</code>. The unique values of <code>time_ids</code> and <code>symbol_ids</code> are not guaranteed to be exactly the same as what is given in the <code>test</code> dataframe. Note that in this <code>lags</code> df, the date_id = D i.e. same as test df, even though the columns represent the lagged responders of previous day.</li>\n</ol>\n<p>On all other time_ids for date_id = D, lags will be None. i.e. lags df will be non-null only on the first time_id of a given date_id.</p>",
      "votes": null,
      "replies": [
        {
          "id": 3055177,
          "author_name": "tomazrochadamota",
          "author_url": "",
          "post_date": "11/25/2024 15:59:59",
          "content": "<p>however, the lags will be saved in a global variable lags_ which we can access when the time_id is not 0, correct?<br>\ni`m getting an error as i try to merge the lags_ dataframe with the test when time_id is not zero.<br>\nAny tips? </p>\n<p>`lags_ : pl.DataFrame | None = None</p>\n<p>def predict(test: pl.DataFrame, lags: pl.DataFrame | None) -&gt; pl.DataFrame | pd.DataFrame:</p>\n<pre><code> lags is not None:\n    lags_ = lags \n\ntest1= test(lags_, on=, how=)\n)\nfeat = test1()\npred = model(feat)\npredictions = test()(pl(, pred()))\nassert (predictions, (pl, pd.DataFrame))\nassert (predictions.) == [, ]\nassert (predictions) == (test)\nreturn predictions\n</code></pre>\n<p>`</p>",
          "votes": null,
          "replies": [
            {
              "id": 3055205,
              "author_name": "alamjs",
              "author_url": "",
              "post_date": "11/25/2024 16:27:45",
              "content": "<p>what was your error?</p>\n<p>your code to join the lags with test seems ok. However, your issue might be related to handling of NaNs when calling model.predict.</p>",
              "votes": null,
              "replies": [
                {
                  "id": 3055245,
                  "author_name": "tomazrochadamota",
                  "author_url": "",
                  "post_date": "11/25/2024 16:44:36",
                  "content": "<p>lgbm, no issues with predict. Funny thing is that I look at the logs and there is no errors there. Only at the submissions page:<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F23236619%2Fee9b3f4289f6de1cf298fa9a2219a04b%2Fsub.png?generation=1732552997114172&amp;alt=media\" alt=\"\"></p>\n<p>And the log (I was trying to print some shapes to see what I was missing) No errors:<br>\n32.3s    1   CPU times: user 13.8 s, sys: 16 s, total: 29.8 s<br>\n32.3s    2   Wall time: 18 s<br>\n43.9s    3   /opt/conda/lib/python3.10/site-packages/lightgbm/engine.py:172: UserWarning: Found <code>n_estimators</code> in params. Will use it instead of argument<br>\n43.9s    4     <em>log_warning(f\"Found <code>{alias}</code> in params. Will use it instead of argument\")\n843.2s    5   CPU times: user 48min 19s, sys: 13.7 s, total: 48min 32s\n843.2s    6   Wall time: 13min 19s\n843.9s    7   (39, 12)\n843.9s    8   shape: (1, 1)\n843.9s    9   ┌─────────┐\n843.9s    10  │ date_id │\n843.9s    11  │ ---     │\n843.9s    12  │ i16     │\n843.9s    13  ╞═════════╡\n843.9s    14  │ 0       │\n843.9s    15  └─────────┘\n843.9s    16  shape: (1, 1)\n843.9s    17  ┌─────────┐\n843.9s    18  │ time_id │\n843.9s    19  │ ---     │\n843.9s    20  │ i16     │\n843.9s    21  ╞═════════╡\n843.9s    22  │ 0       │\n843.9s    23  └─────────┘\n843.9s    24  shape: (39, 1)\n843.9s    25  ┌───────────┐\n843.9s    26  │ symbol_id │\n843.9s    27  │ ---       │\n843.9s    28  │ i8        │\n843.9s    29  ╞═══════════╡\n843.9s    30  │ 23        │\n843.9s    31  │ 4         │\n843.9s    32  │ 26        │\n843.9s    33  │ 32        │\n843.9s    34  │ 21        │\n843.9s    35  │ …         │\n843.9s    36  │ 31        │\n843.9s    37  │ 33        │\n843.9s    38  │ 28        │\n843.9s    39  │ 17        │\n843.9s    40  │ 11        │\n843.9s    41  └───────────┘\n843.9s    42  test1\n843.9s    43  shape: (10, 94)\n843.9s    44  ┌────────┬─────────┬─────────┬───────────┬───┬─────────────┬─────────────┬────────────┬────────────┐\n843.9s    45  │ row_id ┆ date_id ┆ time_id ┆ symbol_id ┆ … ┆ responder_5 ┆ responder_6 ┆ responder</em> ┆ responder_ │<br>\n843.9s    46  │ ---    ┆ ---     ┆ ---     ┆ ---       ┆   ┆ _lag_1      ┆ _lag_1      ┆ 7_lag_1    ┆ 8_lag_1    │<br>\n843.9s    47  │ i64    ┆ i16     ┆ i16     ┆ i8        ┆   ┆ ---         ┆ ---         ┆ ---        ┆ ---        │<br>\n843.9s    48  │        ┆         ┆         ┆           ┆   ┆ f32         ┆ f32         ┆ f32        ┆ f32        │<br>\n843.9s    49  ╞════════╪═════════╪═════════╪═══════════╪═══╪═════════════╪═════════════╪════════════╪════════════╡<br>\n843.9s    50  │ 0      ┆ 0       ┆ 0       ┆ 0         ┆ … ┆ -0.036595   ┆ -1.305746   ┆ -0.795677  ┆ -0.143724  │<br>\n843.9s    51  │ 1      ┆ 0       ┆ 0       ┆ 1         ┆ … ┆ -0.615652   ┆ -1.162801   ┆ -1.205924  ┆ -1.245934  │<br>\n843.9s    52  │ 2      ┆ 0       ┆ 0       ┆ 2         ┆ … ┆ -0.378265   ┆ -1.57429    ┆ -1.863071  ┆ -0.027343  │<br>\n843.9s    53  │ 3      ┆ 0       ┆ 0       ┆ 3         ┆ … ┆ -0.054984   ┆ 0.329152    ┆ -0.965471  ┆ 0.576635   │<br>\n843.9s    54  │ 4      ┆ 0       ┆ 0       ┆ 4         ┆ … ┆ -0.597093   ┆ 0.219856    ┆ -0.276356  ┆ -0.90479   │<br>\n843.9s    55  │ 5      ┆ 0       ┆ 0       ┆ 5         ┆ … ┆ -0.24035    ┆ -0.913801   ┆ -0.548867  ┆ -1.283726  │<br>\n843.9s    56  │ 6      ┆ 0       ┆ 0       ┆ 6         ┆ … ┆ 0.517728    ┆ 0.896325    ┆ 1.068884   ┆ 1.57929    │<br>\n843.9s    57  │ 7      ┆ 0       ┆ 0       ┆ 7         ┆ … ┆ -0.182887   ┆ -0.492168   ┆ -0.142915  ┆ -0.202081  │<br>\n843.9s    58  │ 8      ┆ 0       ┆ 0       ┆ 8         ┆ … ┆ -0.225042   ┆ 0.95646     ┆ 2.185598   ┆ -0.435856  │<br>\n843.9s    59  │ 9      ┆ 0       ┆ 0       ┆ 9         ┆ … ┆ -0.774299   ┆ -0.716492   ┆ -1.471419  ┆ -1.107083  │<br>\n843.9s    60  └────────┴─────────┴─────────┴───────────┴───┴─────────────┴─────────────┴────────────┴────────────┘<br>\n849.4s    61  [NbConvertApp] Converting notebook <strong>notebook</strong>.ipynb to notebook<br>\n850.0s    62  [NbConvertApp] Writing 40602 bytes to <strong>notebook</strong>.ipynb<br>\n852.4s    63  [NbConvertApp] Converting notebook <strong>notebook</strong>.ipynb to html<br>\n853.6s    64  [NbConvertApp] Writing 332474 bytes to <strong>results</strong>.html</p>",
                  "votes": null,
                  "replies": []
                }
              ]
            },
            {
              "id": 3055889,
              "author_name": "cafelatte1",
              "author_url": "",
              "post_date": "11/26/2024 08:50:30",
              "content": "<p>It is the code below I did test.<br>\nI found that <strong>I explicitly select the samples on time_id=0 in lag data</strong>.<br>\nI don't know why I do this… (since I don't know how the lag is provided actually)</p>\n<pre><code> () -&gt; pl.DataFrame | pd.DataFrame:\n     lags_\n     lags   :\n        \n        lags_ = lags.group_by([, ], maintain_order=).first().drop()\n        \n        \n     lags_  :\n        lags_ = test.select(\n            pl.col(), pl.col(),\n            *[pl.lit().alias()  col  CFG.responder_cols]\n        )\n\n    \n    df_test = preprocessor.run(test.drop([, ]), lags_, mode=)\n    df_test = df_test.drop(CFG.metadata_cols).to_pandas()\n\n    fold_pred = []\n     fold  (CFG.n_folds):\n        ()\n        \n        test_x = df_test.copy()\n        \n        scaler = fold_output[fold][]\n        test_x[var_info[]] = scaler.transform(test_x[var_info[]])\n        encoder = fold_output[fold][]\n        ohe_cols = fold_output[fold][]\n        test_x = pd.concat([\n            test_x.drop(var_info[], axis=),\n            pd.DataFrame(encoder.transform(test_x[var_info[]]), index=test_x.index.copy(), columns=ohe_cols)\n        ], axis=)\n        (, test_x.shape)\n        \n        model = fold_output[fold][]\n        \n        test_pred = model.predict(test_x)\n        \n        test_pred = np.clip(test_pred, a_min=-, a_max=)\n        fold_pred.append(test_pred)\n\n    fold_pred = np.stack(fold_pred, axis=).mean(axis=)\n    predictions = test.select(, pl.Series(CFG.target_col, fold_pred, dtype=pl.datatypes.Float32))\n     predictions\n</code></pre>",
              "votes": null,
              "replies": [
                {
                  "id": 3055969,
                  "author_name": "tomazrochadamota",
                  "author_url": "",
                  "post_date": "11/26/2024 11:14:45",
                  "content": "<p>Thanks! Really helpful! </p>",
                  "votes": null,
                  "replies": []
                }
              ]
            }
          ]
        },
        {
          "id": 3055890,
          "author_name": "cafelatte1",
          "author_url": "",
          "post_date": "11/26/2024 08:50:36",
          "content": "<p>Thanks for your explanation :)</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "3054786": "I am confusing yet. Which case is right?\nI think that case 1 is more reasonable considering leakage. (ex. imputation nan with next day in lag)\n[Case1]\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F6610354%2Ffd930e9643fd4d9508fbecaf94ba5a88%2Fcase1.png?generation=1732517447727717&alt=media)\n[Case2]\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F6610354%2Fb358c8ec93f321fb756eeadf33e11308%2Fcase2.png?generation=1732517454750318&alt=media)",
    "3055139": "It is the first one.\n\nFor example:\n\nSay your predict function is being called for ```date_id = D```. Then, inside your predict function, at the first ```time_id = T``` of ```date_id = D```, you are given:\n1. ```test``` dataframe - data (feature_01, 02, weight... etc) corresponding to ```time_id = T``` of ```date_id D``` for all ```symbol_ids``` available at that ```date_id_time_id```.\n2. ```lags``` dataframe - 1 day lagged responder values, as well as ```time_id, symbol_id``` for all ```time_ids```, and all available ```symbol_ids``` for ```date_id D-1```. The unique values of ```time_ids``` and ```symbol_ids``` are not guaranteed to be exactly the same as what is given in the ```test``` dataframe. Note that in this ```lags``` df, the date_id = D i.e. same as test df, even though the columns represent the lagged responders of previous day.\n\n\nOn all other time_ids for date_id = D, lags will be None. i.e. lags df will be non-null only on the first time_id of a given date_id.",
    "3055177": "however, the lags will be saved in a global variable lags_ which we can access when the time_id is not 0, correct?\ni`m getting an error as i try to merge the lags_ dataframe with the test when time_id is not zero.\nAny tips? \n\n`lags_ : pl.DataFrame | None = None\n\ndef predict(test: pl.DataFrame, lags: pl.DataFrame | None) -> pl.DataFrame | pd.DataFrame:\n\n    if lags is not None:\n        lags_ = lags \n\n    test1= test.join(lags_, on=['date_id', 'time_id', 'symbol_id'], how='left')\n    print(test1.head(10))\n    feat = test1[feature_names].to_numpy()\n    pred = model.predict(feat)\n    predictions = test.select('row_id').with_columns(pl.Series('responder_6', pred.ravel()))\n    assert isinstance(predictions, (pl.DataFrame, pd.DataFrame))\n    assert list(predictions.columns) == ['row_id', 'responder_6']\n    assert len(predictions) == len(test)\n    return predictions\n`",
    "3055205": "what was your error?\n\nyour code to join the lags with test seems ok. However, your issue might be related to handling of NaNs when calling model.predict.",
    "3055245": "lgbm, no issues with predict. Funny thing is that I look at the logs and there is no errors there. Only at the submissions page:\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F23236619%2Fee9b3f4289f6de1cf298fa9a2219a04b%2Fsub.png?generation=1732552997114172&alt=media)\n\nAnd the log (I was trying to print some shapes to see what I was missing) No errors:\n32.3s\t1\tCPU times: user 13.8 s, sys: 16 s, total: 29.8 s\n32.3s\t2\tWall time: 18 s\n43.9s\t3\t/opt/conda/lib/python3.10/site-packages/lightgbm/engine.py:172: UserWarning: Found `n_estimators` in params. Will use it instead of argument\n43.9s\t4\t  _log_warning(f\"Found `{alias}` in params. Will use it instead of argument\")\n843.2s\t5\tCPU times: user 48min 19s, sys: 13.7 s, total: 48min 32s\n843.2s\t6\tWall time: 13min 19s\n843.9s\t7\t(39, 12)\n843.9s\t8\tshape: (1, 1)\n843.9s\t9\t┌─────────┐\n843.9s\t10\t│ date_id │\n843.9s\t11\t│ ---     │\n843.9s\t12\t│ i16     │\n843.9s\t13\t╞═════════╡\n843.9s\t14\t│ 0       │\n843.9s\t15\t└─────────┘\n843.9s\t16\tshape: (1, 1)\n843.9s\t17\t┌─────────┐\n843.9s\t18\t│ time_id │\n843.9s\t19\t│ ---     │\n843.9s\t20\t│ i16     │\n843.9s\t21\t╞═════════╡\n843.9s\t22\t│ 0       │\n843.9s\t23\t└─────────┘\n843.9s\t24\tshape: (39, 1)\n843.9s\t25\t┌───────────┐\n843.9s\t26\t│ symbol_id │\n843.9s\t27\t│ ---       │\n843.9s\t28\t│ i8        │\n843.9s\t29\t╞═══════════╡\n843.9s\t30\t│ 23        │\n843.9s\t31\t│ 4         │\n843.9s\t32\t│ 26        │\n843.9s\t33\t│ 32        │\n843.9s\t34\t│ 21        │\n843.9s\t35\t│ …         │\n843.9s\t36\t│ 31        │\n843.9s\t37\t│ 33        │\n843.9s\t38\t│ 28        │\n843.9s\t39\t│ 17        │\n843.9s\t40\t│ 11        │\n843.9s\t41\t└───────────┘\n843.9s\t42\ttest1\n843.9s\t43\tshape: (10, 94)\n843.9s\t44\t┌────────┬─────────┬─────────┬───────────┬───┬─────────────┬─────────────┬────────────┬────────────┐\n843.9s\t45\t│ row_id ┆ date_id ┆ time_id ┆ symbol_id ┆ … ┆ responder_5 ┆ responder_6 ┆ responder_ ┆ responder_ │\n843.9s\t46\t│ ---    ┆ ---     ┆ ---     ┆ ---       ┆   ┆ _lag_1      ┆ _lag_1      ┆ 7_lag_1    ┆ 8_lag_1    │\n843.9s\t47\t│ i64    ┆ i16     ┆ i16     ┆ i8        ┆   ┆ ---         ┆ ---         ┆ ---        ┆ ---        │\n843.9s\t48\t│        ┆         ┆         ┆           ┆   ┆ f32         ┆ f32         ┆ f32        ┆ f32        │\n843.9s\t49\t╞════════╪═════════╪═════════╪═══════════╪═══╪═════════════╪═════════════╪════════════╪════════════╡\n843.9s\t50\t│ 0      ┆ 0       ┆ 0       ┆ 0         ┆ … ┆ -0.036595   ┆ -1.305746   ┆ -0.795677  ┆ -0.143724  │\n843.9s\t51\t│ 1      ┆ 0       ┆ 0       ┆ 1         ┆ … ┆ -0.615652   ┆ -1.162801   ┆ -1.205924  ┆ -1.245934  │\n843.9s\t52\t│ 2      ┆ 0       ┆ 0       ┆ 2         ┆ … ┆ -0.378265   ┆ -1.57429    ┆ -1.863071  ┆ -0.027343  │\n843.9s\t53\t│ 3      ┆ 0       ┆ 0       ┆ 3         ┆ … ┆ -0.054984   ┆ 0.329152    ┆ -0.965471  ┆ 0.576635   │\n843.9s\t54\t│ 4      ┆ 0       ┆ 0       ┆ 4         ┆ … ┆ -0.597093   ┆ 0.219856    ┆ -0.276356  ┆ -0.90479   │\n843.9s\t55\t│ 5      ┆ 0       ┆ 0       ┆ 5         ┆ … ┆ -0.24035    ┆ -0.913801   ┆ -0.548867  ┆ -1.283726  │\n843.9s\t56\t│ 6      ┆ 0       ┆ 0       ┆ 6         ┆ … ┆ 0.517728    ┆ 0.896325    ┆ 1.068884   ┆ 1.57929    │\n843.9s\t57\t│ 7      ┆ 0       ┆ 0       ┆ 7         ┆ … ┆ -0.182887   ┆ -0.492168   ┆ -0.142915  ┆ -0.202081  │\n843.9s\t58\t│ 8      ┆ 0       ┆ 0       ┆ 8         ┆ … ┆ -0.225042   ┆ 0.95646     ┆ 2.185598   ┆ -0.435856  │\n843.9s\t59\t│ 9      ┆ 0       ┆ 0       ┆ 9         ┆ … ┆ -0.774299   ┆ -0.716492   ┆ -1.471419  ┆ -1.107083  │\n843.9s\t60\t└────────┴─────────┴─────────┴───────────┴───┴─────────────┴─────────────┴────────────┴────────────┘\n849.4s\t61\t[NbConvertApp] Converting notebook __notebook__.ipynb to notebook\n850.0s\t62\t[NbConvertApp] Writing 40602 bytes to __notebook__.ipynb\n852.4s\t63\t[NbConvertApp] Converting notebook __notebook__.ipynb to html\n853.6s\t64\t[NbConvertApp] Writing 332474 bytes to __results__.html",
    "3055889": "It is the code below I did test.\nI found that **I explicitly select the samples on time_id=0 in lag data**.\nI don't know why I do this... (since I don't know how the lag is provided actually)\n\n```python\ndef predict(test: pl.DataFrame, lags: pl.DataFrame | None) -> pl.DataFrame | pd.DataFrame:\n    global lags_\n    if lags is not None:\n        # IMPORTANT -> select the samples time_id=0 with group_by before merge with test data\n        lags_ = lags.group_by([\"date_id\", \"symbol_id\"], maintain_order=True).first().drop(\"time_id\")\n        # Below lags_ cause error when merging with test\n        # lags_ = lags.drop(\"time_id\")\n    if lags_ is None:\n        lags_ = test.select(\n            pl.col(\"date_id\"), pl.col(\"symbol_id\"),\n            *[pl.lit(None).alias(f\"{col}_lag_1\") for col in CFG.responder_cols]\n        )\n    \n    \"\"\"Make a prediction.\"\"\"\n    df_test = preprocessor.run(test.drop([\"row_id\", \"is_scored\"]), lags_, mode=\"test\")\n    df_test = df_test.drop(CFG.metadata_cols).to_pandas()\n    \n    fold_pred = []\n    for fold in range(CFG.n_folds):\n        print(f\"\\n=== FOLD {fold} ===\")\n        # split train & test\n        test_x = df_test.copy()\n        # scaling & encoding\n        scaler = fold_output[fold][\"scaler\"]\n        test_x[var_info[\"num\"]] = scaler.transform(test_x[var_info[\"num\"]])\n        encoder = fold_output[fold][\"encoder\"]\n        ohe_cols = fold_output[fold][\"ohe_cols\"]\n        test_x = pd.concat([\n            test_x.drop(var_info[\"cat\"], axis=1),\n            pd.DataFrame(encoder.transform(test_x[var_info[\"cat\"]]), index=test_x.index.copy(), columns=ohe_cols)\n        ], axis=1)\n        print(\"shape info ->\", test_x.shape)\n        # create model\n        model = fold_output[fold][\"model\"]\n        # inference\n        test_pred = model.predict(test_x)\n        # target clipping\n        test_pred = np.clip(test_pred, a_min=-5.0, a_max=5.0)\n        fold_pred.append(test_pred)\n\n    fold_pred = np.stack(fold_pred, axis=0).mean(axis=0)\n    predictions = test.select(\"row_id\", pl.Series(CFG.target_col, fold_pred, dtype=pl.datatypes.Float32))\n    return predictions\n```",
    "3055890": "Thanks for your explanation :)",
    "3055969": "Thanks! Really helpful!"
  },
  "source": "meta"
}