{
  "id": 545073,
  "title": "The way to use Lag from API",
  "url": "/competitions/jane-street-real-time-market-data-forecasting/discussion/545073",
  "author_name": "",
  "post_date": "2024-11-08T08:33:34.924624900Z",
  "votes": 15,
  "comment_count": 9,
  "views": 0,
  "content": "<p>Thanks to <a href=\"https://www.kaggle.com/lihaorocky\" target=\"_blank\">@lihaorocky</a> and <a href=\"https://www.kaggle.com/younesbenalia\" target=\"_blank\">@younesbenalia</a> , they had deep insights and taught me how this works.<br>\nSo that I was able to come up with the following content.</p>\n<p>the following content is for the lag dataset in API.</p>\n<hr>\n<h2>1. Column Names</h2>\n<p>[\"date_id\", \"time_id\", \"symbol_id\"] + [f'responder_{i}_lag_1' for i in range(1, 9)]</p>\n<hr>\n<h2><strong>2. You have been served</strong></h2>\n<p>its not served every batch with test data</p>\n<p>date_id==0, time_id==0, we will get \"lag\" and \"test\"<br>\ndate_id==0, time_id==1, we will get only \"test\"<br>\n….<br>\nuntil<br>\ndate_id==1, time_id==0, we will get \"lag\" and \"test\" again.</p>\n<p>thats why we have this function:</p>\n<pre><code>    \n     lags   :\n        lags_ = lags\n</code></pre>\n<p>meaning we save the lags to lags_, if its been served</p>\n<hr>\n<h2><strong>3. Mis-match</strong></h2>\n<p>date_id and time_id and symbol_id may not have all that test data has.</p>\n<p>for example, test set may have symbol_id from 0 to 38, but lag may miss the 27th symbol</p>\n<p>or </p>\n<p>test set may miss one date_id or random time_id</p>\n<p>this is artificially generated data for you to visualize<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4858569%2F7acdf38cc6113bb4f2da16067cafc32d%2F_20241108164054.png?generation=1731055292641276&amp;alt=media\" alt=\"\"></p>\n<hr>\n<h2><strong>A not really Solution</strong></h2>\n<p>(this did not work, check my comments)<br>\nI fill the NaN after merge with median values.</p>\n<pre><code>    \n    test_df = test.to_pandas()\n    lags_df = lags_.to_pandas()\n    X_test = pd.merge(test_df, lags_df, on=[, ], how=, suffixes = (,))\n\n\n    \n    responder_columns = [  i  (, )]\n    \n     col  responder_columns:\n        median_value = X_test [col].median(skipna=)  \n        X_test [col] = X_test [col].fillna(median_value)  \n</code></pre>\n<p>there could be other ways, like use yesterdays value.. </p>\n<hr>\n<h2><strong>4. May be a Solution</strong></h2>\n<p>I know the public notebook use last row of last date_id to fill missing values. That could work and I believe its a better solution.<br>\nHowever, to provide you something to reflect on, at least on my failed tries, this is filling base on time_id. which isnt work well. I'd like to see a better solution if you have any. let me know.</p>\n<p>I have reconstructed the lags that include 0-40 symbol id.. there were 38 but just incase. and time_id from 0 to 1000. 978 something actual but just incase im using 1000. <br>\nFillin the data with lag_ from API and merge it with X_test. <br>\nThis way we dont leave blanks or NaN after merge with the test. All data will be accounted for. </p>\n<pre><code> () -&gt; pl.DataFrame:\n     lags_, lags_df, mlp_model  \n\n\n\n    full_symbols = np.arange(, )  \n    full_time_ids = np.arange(, )  \n    full_index = pd.MultiIndex.from_product([full_symbols, full_time_ids], names=[, ])\n    full_df = pd.DataFrame(index=full_index).reset_index()\n\n    \n     lags   :\n        lags_ = lags\n        \n        lags_df = lags_.to_pandas()\n        lags_df = pd.merge(full_df, lags_df, on=[, ], how=)\n        lags_df = lags_df.groupby().apply( group: group.bfill().ffill())\n        lags_df = lags_df.bfill().ffill()\n        lags_df.reset_index(drop=, inplace=)\n\n\n    test_df = test.to_pandas()\n    test_df = test_df.bfill().ffill()\n    test_df = test_df.fillna()\n    X_test = pd.merge(test_df, lags_df, on=[, ], how=, suffixes = (,))\n\n    \n    responder_columns = [  i  (, )]\n     col  responder_columns:\n        median_value = X_test[col].median(skipna=)  \n        X_test[col] = X_test[col].fillna(median_value)  \n\n\n    \n    X_test = X_test[feature_names].values\n\n    predictions1, predictions2, predictions3 = mlp_model.predict(X_test)\n\n    prediction = (predictions2 + predictions3) /   \n\n    prediction_df = pd.DataFrame(prediction)\n    responder_6 = prediction_df.iloc[:, ].values\n\n    fin_predict = responder_6\n    output_df = pd.DataFrame({: test_df[], : fin_predict})\n     pl.from_pandas(output_df)\n</code></pre>\n<h2><strong>Simulation of mismatch</strong></h2>\n<pre><code> pandas  pd\n numpy  np\n\n\nsymbol_ids = ()  \ntime_ids = ()   \ntest_data = [(symbol, time)  symbol  symbol_ids  time  time_ids]\ntest_df = pd.DataFrame(test_data, columns=[, ])\n\n\nnum_features = \n i  (num_features):\n    feature_name = \n    test_df[feature_name] = np.random.rand((test_df))  \n\n\nlags_data = []\nrng = np.random.default_rng(seed=)\n\n symbol  symbol_ids:\n     symbol == :\n          \n     time  time_ids:\n        \n        responder_values = [rng.random()  _  ()]\n        lags_data.append((symbol, time, *responder_values))\n\ncolumns = [, ] + [  i  (, )]\nlags_df = pd.DataFrame(lags_data, columns=columns)\n\n\nmissing_indices = rng.choice(lags_df.index, size=((lags_df) * ), replace=)\n col  columns[:]:  \n    lags_df.loc[missing_indices, col] = np.nan\n\n\n\nresponder_columns = [  i  (, )]\n\n col  responder_columns:\n    median_value = lags_df[col].median(skipna=)  \n    lags_df[col].fillna(median_value, inplace=)\n</code></pre>",
  "messages": [
    {
      "id": "3039654",
      "postDate": "11/08/2024 08:33:34",
      "content": "<p>Thanks to <a href=\"https://www.kaggle.com/lihaorocky\" target=\"_blank\">@lihaorocky</a> and <a href=\"https://www.kaggle.com/younesbenalia\" target=\"_blank\">@younesbenalia</a> , they had deep insights and taught me how this works.<br>\nSo that I was able to come up with the following content.</p>\n<p>the following content is for the lag dataset in API.</p>\n<hr>\n<h2>1. Column Names</h2>\n<p>[\"date_id\", \"time_id\", \"symbol_id\"] + [f'responder_{i}_lag_1' for i in range(1, 9)]</p>\n<hr>\n<h2><strong>2. You have been served</strong></h2>\n<p>its not served every batch with test data</p>\n<p>date_id==0, time_id==0, we will get \"lag\" and \"test\"<br>\ndate_id==0, time_id==1, we will get only \"test\"<br>\n….<br>\nuntil<br>\ndate_id==1, time_id==0, we will get \"lag\" and \"test\" again.</p>\n<p>thats why we have this function:</p>\n<pre><code>    \n     lags   :\n        lags_ = lags\n</code></pre>\n<p>meaning we save the lags to lags_, if its been served</p>\n<hr>\n<h2><strong>3. Mis-match</strong></h2>\n<p>date_id and time_id and symbol_id may not have all that test data has.</p>\n<p>for example, test set may have symbol_id from 0 to 38, but lag may miss the 27th symbol</p>\n<p>or </p>\n<p>test set may miss one date_id or random time_id</p>\n<p>this is artificially generated data for you to visualize<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4858569%2F7acdf38cc6113bb4f2da16067cafc32d%2F_20241108164054.png?generation=1731055292641276&amp;alt=media\" alt=\"\"></p>\n<hr>\n<h2><strong>A not really Solution</strong></h2>\n<p>(this did not work, check my comments)<br>\nI fill the NaN after merge with median values.</p>\n<pre><code>    \n    test_df = test.to_pandas()\n    lags_df = lags_.to_pandas()\n    X_test = pd.merge(test_df, lags_df, on=[, ], how=, suffixes = (,))\n\n\n    \n    responder_columns = [  i  (, )]\n    \n     col  responder_columns:\n        median_value = X_test [col].median(skipna=)  \n        X_test [col] = X_test [col].fillna(median_value)  \n</code></pre>\n<p>there could be other ways, like use yesterdays value.. </p>\n<hr>\n<h2><strong>4. May be a Solution</strong></h2>\n<p>I know the public notebook use last row of last date_id to fill missing values. That could work and I believe its a better solution.<br>\nHowever, to provide you something to reflect on, at least on my failed tries, this is filling base on time_id. which isnt work well. I'd like to see a better solution if you have any. let me know.</p>\n<p>I have reconstructed the lags that include 0-40 symbol id.. there were 38 but just incase. and time_id from 0 to 1000. 978 something actual but just incase im using 1000. <br>\nFillin the data with lag_ from API and merge it with X_test. <br>\nThis way we dont leave blanks or NaN after merge with the test. All data will be accounted for. </p>\n<pre><code> () -&gt; pl.DataFrame:\n     lags_, lags_df, mlp_model  \n\n\n\n    full_symbols = np.arange(, )  \n    full_time_ids = np.arange(, )  \n    full_index = pd.MultiIndex.from_product([full_symbols, full_time_ids], names=[, ])\n    full_df = pd.DataFrame(index=full_index).reset_index()\n\n    \n     lags   :\n        lags_ = lags\n        \n        lags_df = lags_.to_pandas()\n        lags_df = pd.merge(full_df, lags_df, on=[, ], how=)\n        lags_df = lags_df.groupby().apply( group: group.bfill().ffill())\n        lags_df = lags_df.bfill().ffill()\n        lags_df.reset_index(drop=, inplace=)\n\n\n    test_df = test.to_pandas()\n    test_df = test_df.bfill().ffill()\n    test_df = test_df.fillna()\n    X_test = pd.merge(test_df, lags_df, on=[, ], how=, suffixes = (,))\n\n    \n    responder_columns = [  i  (, )]\n     col  responder_columns:\n        median_value = X_test[col].median(skipna=)  \n        X_test[col] = X_test[col].fillna(median_value)  \n\n\n    \n    X_test = X_test[feature_names].values\n\n    predictions1, predictions2, predictions3 = mlp_model.predict(X_test)\n\n    prediction = (predictions2 + predictions3) /   \n\n    prediction_df = pd.DataFrame(prediction)\n    responder_6 = prediction_df.iloc[:, ].values\n\n    fin_predict = responder_6\n    output_df = pd.DataFrame({: test_df[], : fin_predict})\n     pl.from_pandas(output_df)\n</code></pre>\n<h2><strong>Simulation of mismatch</strong></h2>\n<pre><code> pandas  pd\n numpy  np\n\n\nsymbol_ids = ()  \ntime_ids = ()   \ntest_data = [(symbol, time)  symbol  symbol_ids  time  time_ids]\ntest_df = pd.DataFrame(test_data, columns=[, ])\n\n\nnum_features = \n i  (num_features):\n    feature_name = \n    test_df[feature_name] = np.random.rand((test_df))  \n\n\nlags_data = []\nrng = np.random.default_rng(seed=)\n\n symbol  symbol_ids:\n     symbol == :\n          \n     time  time_ids:\n        \n        responder_values = [rng.random()  _  ()]\n        lags_data.append((symbol, time, *responder_values))\n\ncolumns = [, ] + [  i  (, )]\nlags_df = pd.DataFrame(lags_data, columns=columns)\n\n\nmissing_indices = rng.choice(lags_df.index, size=((lags_df) * ), replace=)\n col  columns[:]:  \n    lags_df.loc[missing_indices, col] = np.nan\n\n\n\nresponder_columns = [  i  (, )]\n\n col  responder_columns:\n    median_value = lags_df[col].median(skipna=)  \n    lags_df[col].fillna(median_value, inplace=)\n</code></pre>",
      "rawMarkdown": "Thanks to @lihaorocky and @younesbenalia , they had deep insights and taught me how this works.\nSo that I was able to come up with the following content.\n\n\nthe following content is for the lag dataset in API.\n\n---------------------------------------\n\n## 1. Column Names \n\n[\"date_id\", \"time_id\", \"symbol_id\"] + [f'responder_{i}_lag_1' for i in range(1, 9)]\n\n---------------------------------------\n\n\n## **2. You have been served** \n\nits not served every batch with test data\n\ndate_id==0, time_id==0, we will get \"lag\" and \"test\"\ndate_id==0, time_id==1, we will get only \"test\"\n….\nuntil\ndate_id==1, time_id==0, we will get \"lag\" and \"test\" again.\n\n\nthats why we have this function:\n\n```python\n    # Logic for saving or loading lags\n    if lags is not None:\n        lags_ = lags\n```\n\nmeaning we save the lags to lags_, if its been served\n\n---------------------------------------\n\n\n\n## **3. Mis-match**\n\ndate_id and time_id and symbol_id may not have all that test data has.\n\nfor example, test set may have symbol_id from 0 to 38, but lag may miss the 27th symbol\n\nor \n\ntest set may miss one date_id or random time_id\n\nthis is artificially generated data for you to visualize\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4858569%2F7acdf38cc6113bb4f2da16067cafc32d%2F_20241108164054.png?generation=1731055292641276&alt=media)\n\n---------------------------------------\n\n\n##**A not really Solution**\n\n(this did not work, check my comments)\nI fill the NaN after merge with median values.\n\n```python\n    # Convert to Pandas DataFrame if needed (since model expects numpy arrays)\n    test_df = test.to_pandas()\n    lags_df = lags_.to_pandas()\n    X_test = pd.merge(test_df, lags_df, on=['time_id', 'symbol_id'], how='left', suffixes = (\"\",\"_drop\"))\n    \n    \n    # List of specific columns to fill with median values\n    responder_columns = [f'responder_{i}_lag_1' for i in range(1, 9)]\n    # Now we fill the missing values with the median of each column\n    for col in responder_columns:\n        median_value = X_test [col].median(skipna=True)  # Calculate the median ignoring NaNs\n        X_test [col] = X_test [col].fillna(median_value)  # Use assignment to fill NaNs\n```\n\nthere could be other ways, like use yesterdays value.. \n\n---------------------------------------\n\n\n\n\n## **4. May be a Solution**\n\nI know the public notebook use last row of last date_id to fill missing values. That could work and I believe its a better solution.\nHowever, to provide you something to reflect on, at least on my failed tries, this is filling base on time_id. which isnt work well. I'd like to see a better solution if you have any. let me know.\n\nI have reconstructed the lags that include 0-40 symbol id.. there were 38 but just incase. and time_id from 0 to 1000. 978 something actual but just incase im using 1000. \nFillin the data with lag_ from API and merge it with X_test. \nThis way we dont leave blanks or NaN after merge with the test. All data will be accounted for. \n\n\n```python\ndef predict(test: pl.DataFrame, lags: pl.DataFrame | None) -> pl.DataFrame:\n    global lags_, lags_df, mlp_model  # Declare models as global\n    \n    \n    \n    full_symbols = np.arange(0, 41)  # 0 to 40\n    full_time_ids = np.arange(0, 1001)  # 0 to 1000\n    full_index = pd.MultiIndex.from_product([full_symbols, full_time_ids], names=['symbol_id', 'time_id'])\n    full_df = pd.DataFrame(index=full_index).reset_index()\n\n    # Logic for saving or loading lags\n    if lags is not None:\n        lags_ = lags\n        # Convert to Pandas DataFrame if needed (since model expects numpy arrays)\n        lags_df = lags_.to_pandas()\n        lags_df = pd.merge(full_df, lags_df, on=['symbol_id', 'time_id'], how='outer')\n        lags_df = lags_df.groupby('symbol_id').apply(lambda group: group.bfill().ffill())\n        lags_df = lags_df.bfill().ffill()\n        lags_df.reset_index(drop=True, inplace=True)\n\n\n    test_df = test.to_pandas()\n    test_df = test_df.bfill().ffill()\n    test_df = test_df.fillna(0)\n    X_test = pd.merge(test_df, lags_df, on=['time_id', 'symbol_id'], how='left', suffixes = (\"\",\"_drop\"))\n\n    # List of specific columns to fill with median values\n    responder_columns = [f'responder_{i}_lag_1' for i in range(1, 9)]\n    for col in responder_columns:\n        median_value = X_test[col].median(skipna=True)  # Calculate the median ignoring NaNs\n        X_test[col] = X_test[col].fillna(median_value)  # Use assignment to fill NaNs\n\n\n    # Extract features for prediction\n    X_test = X_test[feature_names].values\n\n    predictions1, predictions2, predictions3 = mlp_model.predict(X_test)\n\n    prediction = (predictions2 + predictions3) / 2  # Ensemble predictions\n\n    prediction_df = pd.DataFrame(prediction)\n    responder_6 = prediction_df.iloc[:, 6].values\n\n    fin_predict = responder_6\n    output_df = pd.DataFrame({\"row_id\": test_df['row_id'], \"responder_6\": fin_predict})\n    return pl.from_pandas(output_df)\n\n```\n\n\n\n##**Simulation of mismatch**\n\n\n```python\nimport pandas as pd\nimport numpy as np\n\n# Generate test_df with symbol_id from 0 to 5, time_id from 0 to 24\nsymbol_ids = range(6)  # 0 to 5\ntime_ids = range(25)   # 0 to 24\ntest_data = [(symbol, time) for symbol in symbol_ids for time in time_ids]\ntest_df = pd.DataFrame(test_data, columns=['symbol_id', 'time_id'])\n\n# Add 5 features to test_df\nnum_features = 5\nfor i in range(num_features):\n    feature_name = f\"feature_{i:02d}\"\n    test_df[feature_name] = np.random.rand(len(test_df))  # Random data for each feature\n\n# Create example data for lags_df\nlags_data = []\nrng = np.random.default_rng(seed=42)\n\nfor symbol in symbol_ids:\n    if symbol == 3:\n        continue  # Exclude symbol_id 3\n    for time in time_ids:\n        # Generate random lag values for each responder\n        responder_values = [rng.random() for _ in range(8)]\n        lags_data.append((symbol, time, *responder_values))\n\ncolumns = ['symbol_id', 'time_id'] + [f'responder_{i}_lag_1' for i in range(1, 9)]\nlags_df = pd.DataFrame(lags_data, columns=columns)\n\n# Introduce missing values: randomly set some values to NaN\nmissing_indices = rng.choice(lags_df.index, size=int(len(lags_df) * 0.2), replace=False)\nfor col in columns[2:]:  # Exclude 'symbol_id' and 'time_id'\n    lags_df.loc[missing_indices, col] = np.nan\n\n\n# List of specific columns to fill with median values\nresponder_columns = [f'responder_{i}_lag_1' for i in range(1, 9)]\n# Now we fill the missing values with the median of each column\nfor col in responder_columns:\n    median_value = lags_df[col].median(skipna=True)  # Calculate the median ignoring NaNs\n    lags_df[col].fillna(median_value, inplace=True)\n```",
      "votes": null
    },
    {
      "id": "3039678",
      "postDate": "11/08/2024 08:58:23",
      "content": "<p>The idea is clear but you should consider the efficiency of your code. Another point is, does it make sense to use the data points from the same time of yesterday as features? Any correlation or predictive power between yesterday’s time T and today’s time T?</p>",
      "rawMarkdown": "The idea is clear but you should consider the efficiency of your code. Another point is, does it make sense to use the data points from the same time of yesterday as features? Any correlation or predictive power between yesterday’s time T and today’s time T?",
      "votes": null
    },
    {
      "id": "3039682",
      "postDate": "11/08/2024 09:04:29",
      "content": "<p>for efficiency, any suggestions? we could compare test's id and lag id see whats missing each time and fill it using numpy… but its a lot steps.<br>\nso i assume you would not use yesterday's data, but fill with today's value?</p>",
      "rawMarkdown": "for efficiency, any suggestions? we could compare test's id and lag id see whats missing each time and fill it using numpy... but its a lot steps.\nso i assume you would not use yesterday's data, but fill with today's value?",
      "votes": null
    },
    {
      "id": "3040687",
      "postDate": "11/09/2024 13:09:42",
      "content": "<p>After merging, I did fill by median number… but still no score.<br>\nI suspect that after the merge, the responders are all NaNs….<br>\nforward, backward fill will not pass either!<br>\nSo I did a fill NaN by 0, which I passed. </p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4858569%2F88f2306917a12779bfad8737d97014d2%2F_20241109211055.png?generation=1731158394809445&amp;alt=media\" alt=\"\"><br>\ncheck my recent results link:<br>\n<a href=\"https://www.kaggle.com/competitions/jane-street-real-time-market-data-forecasting/discussion/545315\" target=\"_blank\">https://www.kaggle.com/competitions/jane-street-real-time-market-data-forecasting/discussion/545315</a></p>",
      "rawMarkdown": "After merging, I did fill by median number... but still no score.\nI suspect that after the merge, the responders are all NaNs....\nforward, backward fill will not pass either!\nSo I did a fill NaN by 0, which I passed. \n\n\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4858569%2F88f2306917a12779bfad8737d97014d2%2F_20241109211055.png?generation=1731158394809445&alt=media)\ncheck my recent results link:\nhttps://www.kaggle.com/competitions/jane-street-real-time-market-data-forecasting/discussion/545315",
      "votes": null
    },
    {
      "id": "3040696",
      "postDate": "11/09/2024 13:16:57",
      "content": "<p>Thank you for the information..</p>",
      "rawMarkdown": "Thank you for the information..",
      "votes": null
    },
    {
      "id": "3043020",
      "postDate": "11/12/2024 02:37:46",
      "content": "<p>I think it makes sense that the model can know the trend of the data change. There are many time-based alpha in quantization, although in this case all are anonymous features.</p>",
      "rawMarkdown": "I think it makes sense that the model can know the trend of the data change. There are many time-based alpha in quantization, although in this case all are anonymous features.",
      "votes": null
    },
    {
      "id": "3045120",
      "postDate": "11/14/2024 07:26:54",
      "content": "<p>In the end, I found a way to make it work. altough the results was terrible.<br>\nwe need to reconstruct the dataframe of lag_ in the prediction function.<br>\nget a zero like dataframe with time_id from 0 to 1000, and symbol_id from 0 to 40.<br>\nthen fill in the data from lag_ and do a merge afterwards. so that we wont get NaNs and we will have a score…</p>",
      "rawMarkdown": "In the end, I found a way to make it work. altough the results was terrible.\nwe need to reconstruct the dataframe of lag_ in the prediction function.\nget a zero like dataframe with time_id from 0 to 1000, and symbol_id from 0 to 40.\nthen fill in the data from lag_ and do a merge afterwards. so that we wont get NaNs and we will have a score...",
      "votes": null
    },
    {
      "id": "3046265",
      "postDate": "11/15/2024 09:34:57",
      "content": "<p><a href=\"https://www.kaggle.com/shiyili\" target=\"_blank\">@shiyili</a>, I had the same question and to be clear, while at training time we have access to the values at previous time_id's, at test we have only access to the last time_id of the previous day. Correct? </p>",
      "rawMarkdown": "shiyili, I had the same question and to be clear, while at training time we have access to the values at previous time_id's, at test we have only access to the last time_id of the previous day. Correct?",
      "votes": null
    },
    {
      "id": "3053197",
      "postDate": "11/23/2024 08:45:04",
      "content": "<p>Thanks for your information !</p>\n<p>1.There may be some missing data in the test set.<br>\n2.There may be some missing data in the lag.</p>\n<p>Do you mean that both of these problems could occur?<br>\nIf so, for the missing data in the lag (which served at next day),  the prediction of corresponding data will not be scored  that day?<br>\n(Only the data that exists in the lag will be scored, even if the missing data in the lag (served the next day) is predicted.)</p>",
      "rawMarkdown": "Thanks for your information !\n\n1.There may be some missing data in the test set.\n2.There may be some missing data in the lag.\n\nDo you mean that both of these problems could occur?\nIf so, for the missing data in the lag (which served at next day),  the prediction of corresponding data will not be scored  that day?\n(Only the data that exists in the lag will be scored, even if the missing data in the lag (served the next day) is predicted.)",
      "votes": null
    },
    {
      "id": "3053262",
      "postDate": "11/23/2024 10:56:06",
      "content": "<p>I cant say for sure what will be missing. but yeah, I believe there could be stuff missing on both test and lag during submission. date_id i cant say for sure, but time_id data for certain symbols can be missing entirely.</p>",
      "rawMarkdown": "I cant say for sure what will be missing. but yeah, I believe there could be stuff missing on both test and lag during submission. date_id i cant say for sure, but time_id data for certain symbols can be missing entirely.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 3039678,
      "author_name": "shiyili",
      "author_url": "",
      "post_date": "11/08/2024 08:58:23",
      "content": "<p>The idea is clear but you should consider the efficiency of your code. Another point is, does it make sense to use the data points from the same time of yesterday as features? Any correlation or predictive power between yesterday’s time T and today’s time T?</p>",
      "votes": null,
      "replies": [
        {
          "id": 3039682,
          "author_name": "zoutain",
          "author_url": "",
          "post_date": "11/08/2024 09:04:29",
          "content": "<p>for efficiency, any suggestions? we could compare test's id and lag id see whats missing each time and fill it using numpy… but its a lot steps.<br>\nso i assume you would not use yesterday's data, but fill with today's value?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 3043020,
          "author_name": "chenboluo",
          "author_url": "",
          "post_date": "11/12/2024 02:37:46",
          "content": "<p>I think it makes sense that the model can know the trend of the data change. There are many time-based alpha in quantization, although in this case all are anonymous features.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 3046265,
          "author_name": "rghiglia",
          "author_url": "",
          "post_date": "11/15/2024 09:34:57",
          "content": "<p><a href=\"https://www.kaggle.com/shiyili\" target=\"_blank\">@shiyili</a>, I had the same question and to be clear, while at training time we have access to the values at previous time_id's, at test we have only access to the last time_id of the previous day. Correct? </p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 3040687,
      "author_name": "zoutain",
      "author_url": "",
      "post_date": "11/09/2024 13:09:42",
      "content": "<p>After merging, I did fill by median number… but still no score.<br>\nI suspect that after the merge, the responders are all NaNs….<br>\nforward, backward fill will not pass either!<br>\nSo I did a fill NaN by 0, which I passed. </p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4858569%2F88f2306917a12779bfad8737d97014d2%2F_20241109211055.png?generation=1731158394809445&amp;alt=media\" alt=\"\"><br>\ncheck my recent results link:<br>\n<a href=\"https://www.kaggle.com/competitions/jane-street-real-time-market-data-forecasting/discussion/545315\" target=\"_blank\">https://www.kaggle.com/competitions/jane-street-real-time-market-data-forecasting/discussion/545315</a></p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 3040696,
      "author_name": "sumit08",
      "author_url": "",
      "post_date": "11/09/2024 13:16:57",
      "content": "<p>Thank you for the information..</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 3045120,
      "author_name": "zoutain",
      "author_url": "",
      "post_date": "11/14/2024 07:26:54",
      "content": "<p>In the end, I found a way to make it work. altough the results was terrible.<br>\nwe need to reconstruct the dataframe of lag_ in the prediction function.<br>\nget a zero like dataframe with time_id from 0 to 1000, and symbol_id from 0 to 40.<br>\nthen fill in the data from lag_ and do a merge afterwards. so that we wont get NaNs and we will have a score…</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 3053197,
      "author_name": "minamiiiiiiiiiii",
      "author_url": "",
      "post_date": "11/23/2024 08:45:04",
      "content": "<p>Thanks for your information !</p>\n<p>1.There may be some missing data in the test set.<br>\n2.There may be some missing data in the lag.</p>\n<p>Do you mean that both of these problems could occur?<br>\nIf so, for the missing data in the lag (which served at next day),  the prediction of corresponding data will not be scored  that day?<br>\n(Only the data that exists in the lag will be scored, even if the missing data in the lag (served the next day) is predicted.)</p>",
      "votes": null,
      "replies": [
        {
          "id": 3053262,
          "author_name": "zoutain",
          "author_url": "",
          "post_date": "11/23/2024 10:56:06",
          "content": "<p>I cant say for sure what will be missing. but yeah, I believe there could be stuff missing on both test and lag during submission. date_id i cant say for sure, but time_id data for certain symbols can be missing entirely.</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "3039654": "Thanks to @lihaorocky and @younesbenalia , they had deep insights and taught me how this works.\nSo that I was able to come up with the following content.\n\n\nthe following content is for the lag dataset in API.\n\n---------------------------------------\n\n## 1. Column Names \n\n[\"date_id\", \"time_id\", \"symbol_id\"] + [f'responder_{i}_lag_1' for i in range(1, 9)]\n\n---------------------------------------\n\n\n## **2. You have been served** \n\nits not served every batch with test data\n\ndate_id==0, time_id==0, we will get \"lag\" and \"test\"\ndate_id==0, time_id==1, we will get only \"test\"\n….\nuntil\ndate_id==1, time_id==0, we will get \"lag\" and \"test\" again.\n\n\nthats why we have this function:\n\n```python\n    # Logic for saving or loading lags\n    if lags is not None:\n        lags_ = lags\n```\n\nmeaning we save the lags to lags_, if its been served\n\n---------------------------------------\n\n\n\n## **3. Mis-match**\n\ndate_id and time_id and symbol_id may not have all that test data has.\n\nfor example, test set may have symbol_id from 0 to 38, but lag may miss the 27th symbol\n\nor \n\ntest set may miss one date_id or random time_id\n\nthis is artificially generated data for you to visualize\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4858569%2F7acdf38cc6113bb4f2da16067cafc32d%2F_20241108164054.png?generation=1731055292641276&alt=media)\n\n---------------------------------------\n\n\n##**A not really Solution**\n\n(this did not work, check my comments)\nI fill the NaN after merge with median values.\n\n```python\n    # Convert to Pandas DataFrame if needed (since model expects numpy arrays)\n    test_df = test.to_pandas()\n    lags_df = lags_.to_pandas()\n    X_test = pd.merge(test_df, lags_df, on=['time_id', 'symbol_id'], how='left', suffixes = (\"\",\"_drop\"))\n    \n    \n    # List of specific columns to fill with median values\n    responder_columns = [f'responder_{i}_lag_1' for i in range(1, 9)]\n    # Now we fill the missing values with the median of each column\n    for col in responder_columns:\n        median_value = X_test [col].median(skipna=True)  # Calculate the median ignoring NaNs\n        X_test [col] = X_test [col].fillna(median_value)  # Use assignment to fill NaNs\n```\n\nthere could be other ways, like use yesterdays value.. \n\n---------------------------------------\n\n\n\n\n## **4. May be a Solution**\n\nI know the public notebook use last row of last date_id to fill missing values. That could work and I believe its a better solution.\nHowever, to provide you something to reflect on, at least on my failed tries, this is filling base on time_id. which isnt work well. I'd like to see a better solution if you have any. let me know.\n\nI have reconstructed the lags that include 0-40 symbol id.. there were 38 but just incase. and time_id from 0 to 1000. 978 something actual but just incase im using 1000. \nFillin the data with lag_ from API and merge it with X_test. \nThis way we dont leave blanks or NaN after merge with the test. All data will be accounted for. \n\n\n```python\ndef predict(test: pl.DataFrame, lags: pl.DataFrame | None) -> pl.DataFrame:\n    global lags_, lags_df, mlp_model  # Declare models as global\n    \n    \n    \n    full_symbols = np.arange(0, 41)  # 0 to 40\n    full_time_ids = np.arange(0, 1001)  # 0 to 1000\n    full_index = pd.MultiIndex.from_product([full_symbols, full_time_ids], names=['symbol_id', 'time_id'])\n    full_df = pd.DataFrame(index=full_index).reset_index()\n\n    # Logic for saving or loading lags\n    if lags is not None:\n        lags_ = lags\n        # Convert to Pandas DataFrame if needed (since model expects numpy arrays)\n        lags_df = lags_.to_pandas()\n        lags_df = pd.merge(full_df, lags_df, on=['symbol_id', 'time_id'], how='outer')\n        lags_df = lags_df.groupby('symbol_id').apply(lambda group: group.bfill().ffill())\n        lags_df = lags_df.bfill().ffill()\n        lags_df.reset_index(drop=True, inplace=True)\n\n\n    test_df = test.to_pandas()\n    test_df = test_df.bfill().ffill()\n    test_df = test_df.fillna(0)\n    X_test = pd.merge(test_df, lags_df, on=['time_id', 'symbol_id'], how='left', suffixes = (\"\",\"_drop\"))\n\n    # List of specific columns to fill with median values\n    responder_columns = [f'responder_{i}_lag_1' for i in range(1, 9)]\n    for col in responder_columns:\n        median_value = X_test[col].median(skipna=True)  # Calculate the median ignoring NaNs\n        X_test[col] = X_test[col].fillna(median_value)  # Use assignment to fill NaNs\n\n\n    # Extract features for prediction\n    X_test = X_test[feature_names].values\n\n    predictions1, predictions2, predictions3 = mlp_model.predict(X_test)\n\n    prediction = (predictions2 + predictions3) / 2  # Ensemble predictions\n\n    prediction_df = pd.DataFrame(prediction)\n    responder_6 = prediction_df.iloc[:, 6].values\n\n    fin_predict = responder_6\n    output_df = pd.DataFrame({\"row_id\": test_df['row_id'], \"responder_6\": fin_predict})\n    return pl.from_pandas(output_df)\n\n```\n\n\n\n##**Simulation of mismatch**\n\n\n```python\nimport pandas as pd\nimport numpy as np\n\n# Generate test_df with symbol_id from 0 to 5, time_id from 0 to 24\nsymbol_ids = range(6)  # 0 to 5\ntime_ids = range(25)   # 0 to 24\ntest_data = [(symbol, time) for symbol in symbol_ids for time in time_ids]\ntest_df = pd.DataFrame(test_data, columns=['symbol_id', 'time_id'])\n\n# Add 5 features to test_df\nnum_features = 5\nfor i in range(num_features):\n    feature_name = f\"feature_{i:02d}\"\n    test_df[feature_name] = np.random.rand(len(test_df))  # Random data for each feature\n\n# Create example data for lags_df\nlags_data = []\nrng = np.random.default_rng(seed=42)\n\nfor symbol in symbol_ids:\n    if symbol == 3:\n        continue  # Exclude symbol_id 3\n    for time in time_ids:\n        # Generate random lag values for each responder\n        responder_values = [rng.random() for _ in range(8)]\n        lags_data.append((symbol, time, *responder_values))\n\ncolumns = ['symbol_id', 'time_id'] + [f'responder_{i}_lag_1' for i in range(1, 9)]\nlags_df = pd.DataFrame(lags_data, columns=columns)\n\n# Introduce missing values: randomly set some values to NaN\nmissing_indices = rng.choice(lags_df.index, size=int(len(lags_df) * 0.2), replace=False)\nfor col in columns[2:]:  # Exclude 'symbol_id' and 'time_id'\n    lags_df.loc[missing_indices, col] = np.nan\n\n\n# List of specific columns to fill with median values\nresponder_columns = [f'responder_{i}_lag_1' for i in range(1, 9)]\n# Now we fill the missing values with the median of each column\nfor col in responder_columns:\n    median_value = lags_df[col].median(skipna=True)  # Calculate the median ignoring NaNs\n    lags_df[col].fillna(median_value, inplace=True)\n```",
    "3039678": "The idea is clear but you should consider the efficiency of your code. Another point is, does it make sense to use the data points from the same time of yesterday as features? Any correlation or predictive power between yesterday’s time T and today’s time T?",
    "3039682": "for efficiency, any suggestions? we could compare test's id and lag id see whats missing each time and fill it using numpy... but its a lot steps.\nso i assume you would not use yesterday's data, but fill with today's value?",
    "3040687": "After merging, I did fill by median number... but still no score.\nI suspect that after the merge, the responders are all NaNs....\nforward, backward fill will not pass either!\nSo I did a fill NaN by 0, which I passed. \n\n\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4858569%2F88f2306917a12779bfad8737d97014d2%2F_20241109211055.png?generation=1731158394809445&alt=media)\ncheck my recent results link:\nhttps://www.kaggle.com/competitions/jane-street-real-time-market-data-forecasting/discussion/545315",
    "3040696": "Thank you for the information..",
    "3043020": "I think it makes sense that the model can know the trend of the data change. There are many time-based alpha in quantization, although in this case all are anonymous features.",
    "3045120": "In the end, I found a way to make it work. altough the results was terrible.\nwe need to reconstruct the dataframe of lag_ in the prediction function.\nget a zero like dataframe with time_id from 0 to 1000, and symbol_id from 0 to 40.\nthen fill in the data from lag_ and do a merge afterwards. so that we wont get NaNs and we will have a score...",
    "3046265": "shiyili, I had the same question and to be clear, while at training time we have access to the values at previous time_id's, at test we have only access to the last time_id of the previous day. Correct?",
    "3053197": "Thanks for your information !\n\n1.There may be some missing data in the test set.\n2.There may be some missing data in the lag.\n\nDo you mean that both of these problems could occur?\nIf so, for the missing data in the lag (which served at next day),  the prediction of corresponding data will not be scored  that day?\n(Only the data that exists in the lag will be scored, even if the missing data in the lag (served the next day) is predicted.)",
    "3053262": "I cant say for sure what will be missing. but yeah, I believe there could be stuff missing on both test and lag during submission. date_id i cant say for sure, but time_id data for certain symbols can be missing entirely."
  },
  "source": "meta"
}