{
  "id": 551345,
  "title": "Notebook timeout, when trying to use lagged features.",
  "url": "/competitions/jane-street-real-time-market-data-forecasting/discussion/551345",
  "author_name": "",
  "post_date": "2024-12-12T17:15:57.184274Z",
  "votes": null,
  "comment_count": 2,
  "views": 0,
  "content": "<p>So I have trained my model to keep storing history cache, so that I can use lagged values of some features as featuers, I am using a very basic linear regression model. But everytime I run the following code it leads to Notebook timeout and I am not able to understand why, I am attaching my testing code below if you have any improvements or suggestions to solve this issue, please let me know. </p>\n<p>`import pickle<br>\nimport time<br>\nimport polars as pl<br>\nimport pandas as pd<br>\nfrom sklearn.linear_model import LinearRegression</p>\n<h1>Global variables</h1>\n<p>lags_: pl.DataFrame | None = None<br>\nhistory_cache = pl.DataFrame()<br>\nlags_cache = pl.DataFrame()<br>\nloaded_model = None<br>\n'<br>\n'</p>\n<h1>Feature and responder configurations</h1>\n<p>features_to_lag = [<br>\n    'feature_06', 'feature_60', 'feature_49', 'feature_04', 'feature_07',<br>\n    'feature_58', 'feature_59', 'feature_47', 'feature_51', 'feature_36',<br>\n    'feature_52', 'feature_68', 'feature_13', 'feature_02', 'feature_05']<br>\nlag_range = range(1, 11)  # Lags from t-1 to t-10<br>\nresponder_columns = [f\"responder_{i}\" for i in range(9)]<br>\n''</p>\n<h1>Feature columns for prediction</h1>\n<p>lagged_feature_cols = [f\"{feature}<em>lag</em>{lag}\" for feature in features_to_lag for lag in lag_range]<br>\nlagged_responder_cols = [f\"{responder}_lag_1\" for responder in responder_columns]<br>\nfeature_columns = features_to_lag + lagged_feature_cols + lagged_responder_cols + [\"time_id\"]<br>\nmax_lag = 11</p>\n<p>with open(\"/kaggle/input/linregwithlagged/other/default/1/linreg_model_cv_nonrandom (1) (1).pkl\", \"rb\") as f:<br>\n    loaded_model = pickle.load(f)<br>\n''<br>\ndef predict(test: pl.DataFrame, lags: pl.DataFrame | None) -&gt; pl.DataFrame | pd.DataFrame:<br>\n    \"\"\"Make a prediction for the test set.\"\"\"<br>\n    global lags_, loaded_model, history_cache, lags_cache, max_lag</p>\n<pre><code>\n test  :\n    test = pl.DataFrame()\n\n\n lags   :\n    lags_ = lags\n    lags_cache = pl.concat([lags_cache, lags])\n\n\n test.is_empty():\n    \n    test = pl.DataFrame({col: []  col  [, , , ] + features_to_lag})\n\ndate_id = test[][]   test.is_empty()  \n\n\n\n\n\nhistory_cache = pl.concat([history_cache, test])\n\n\nmin_date_id = (date_id - , )\nhistory_cache = history_cache.(pl.col() &gt;= min_date_id)\nlags_cache = lags_cache.(pl.col() &gt;= min_date_id)\n\n\nfiltered_df = history_cache\n\n\ncurrent_df = filtered_df.(pl.col() == date_id)\n\n\n lag  (, max_lag + ):\n    lagged_cols = [\n        pl.col(feature).shift(lag).alias()\n         feature  features_to_lag\n    ]\n    current_df = current_df.with_columns(lagged_cols)\n\n\nfiltered_df = filtered_df.join(\n    current_df.select([, , ] + [ \n                                                              feature  features_to_lag \n                                                              lag  (, max_lag + )]),\n    on=[, , ],\n    how=\n)\n\n\n\n\n\ncombined_data = test.join(\n    filtered_df,\n    on=[, , ],\n    how=\n)\n\n\n lags_   :\n    combined_data = combined_data.join(\n        lags_,\n        on=[, ],\n        how=\n    )\n\n\ncolumns_to_keep = [col  col  combined_data.columns   col.endswith()]\ncombined_data = combined_data.select(columns_to_keep)\n\n\ncombined_data = combined_data.fill_null()\n\n\nmissing_columns = (feature_columns) - (combined_data.columns)\n missing_columns:\n     missing_col  missing_columns:\n        combined_data = combined_data.with_columns(\n            pl.lit().alias(missing_col)\n        )\n\n\nX_pred = combined_data.select(feature_columns).to_pandas()\nX_pred.fillna(, inplace=)\n\n\n (loaded_model, ):\n    expected_features = loaded_model.feature_names_in_\n    X_pred = X_pred.reindex(columns=expected_features, fill_value=)\n:\n    X_pred = X_pred[feature_columns]\n\n\npredictions = loaded_model.predict(X_pred)\n\n\ncombined_data = combined_data.with_columns(pl.Series(name=, values=predictions))\n\n\npredictions_df = combined_data.select([, ])\n\n\n (predictions_df, pl.DataFrame):\n     predictions_df.columns == [, ]\n (predictions_df, pd.DataFrame):\n     (predictions_df.columns) == [, ]\n:\n    \n    predictions_df = pd.DataFrame(columns=[, ])\n\n\n (predictions_df) != (test):\n    \n    predictions_df = test.select([]).with_columns(\n        pl.lit().alias()\n    )\n\n predictions_df\n</code></pre>\n<p>`</p>",
  "messages": [
    {
      "id": "3070455",
      "postDate": "12/12/2024 17:15:57",
      "content": "<p>So I have trained my model to keep storing history cache, so that I can use lagged values of some features as featuers, I am using a very basic linear regression model. But everytime I run the following code it leads to Notebook timeout and I am not able to understand why, I am attaching my testing code below if you have any improvements or suggestions to solve this issue, please let me know. </p>\n<p>`import pickle<br>\nimport time<br>\nimport polars as pl<br>\nimport pandas as pd<br>\nfrom sklearn.linear_model import LinearRegression</p>\n<h1>Global variables</h1>\n<p>lags_: pl.DataFrame | None = None<br>\nhistory_cache = pl.DataFrame()<br>\nlags_cache = pl.DataFrame()<br>\nloaded_model = None<br>\n'<br>\n'</p>\n<h1>Feature and responder configurations</h1>\n<p>features_to_lag = [<br>\n    'feature_06', 'feature_60', 'feature_49', 'feature_04', 'feature_07',<br>\n    'feature_58', 'feature_59', 'feature_47', 'feature_51', 'feature_36',<br>\n    'feature_52', 'feature_68', 'feature_13', 'feature_02', 'feature_05']<br>\nlag_range = range(1, 11)  # Lags from t-1 to t-10<br>\nresponder_columns = [f\"responder_{i}\" for i in range(9)]<br>\n''</p>\n<h1>Feature columns for prediction</h1>\n<p>lagged_feature_cols = [f\"{feature}<em>lag</em>{lag}\" for feature in features_to_lag for lag in lag_range]<br>\nlagged_responder_cols = [f\"{responder}_lag_1\" for responder in responder_columns]<br>\nfeature_columns = features_to_lag + lagged_feature_cols + lagged_responder_cols + [\"time_id\"]<br>\nmax_lag = 11</p>\n<p>with open(\"/kaggle/input/linregwithlagged/other/default/1/linreg_model_cv_nonrandom (1) (1).pkl\", \"rb\") as f:<br>\n    loaded_model = pickle.load(f)<br>\n''<br>\ndef predict(test: pl.DataFrame, lags: pl.DataFrame | None) -&gt; pl.DataFrame | pd.DataFrame:<br>\n    \"\"\"Make a prediction for the test set.\"\"\"<br>\n    global lags_, loaded_model, history_cache, lags_cache, max_lag</p>\n<pre><code>\n test  :\n    test = pl.DataFrame()\n\n\n lags   :\n    lags_ = lags\n    lags_cache = pl.concat([lags_cache, lags])\n\n\n test.is_empty():\n    \n    test = pl.DataFrame({col: []  col  [, , , ] + features_to_lag})\n\ndate_id = test[][]   test.is_empty()  \n\n\n\n\n\nhistory_cache = pl.concat([history_cache, test])\n\n\nmin_date_id = (date_id - , )\nhistory_cache = history_cache.(pl.col() &gt;= min_date_id)\nlags_cache = lags_cache.(pl.col() &gt;= min_date_id)\n\n\nfiltered_df = history_cache\n\n\ncurrent_df = filtered_df.(pl.col() == date_id)\n\n\n lag  (, max_lag + ):\n    lagged_cols = [\n        pl.col(feature).shift(lag).alias()\n         feature  features_to_lag\n    ]\n    current_df = current_df.with_columns(lagged_cols)\n\n\nfiltered_df = filtered_df.join(\n    current_df.select([, , ] + [ \n                                                              feature  features_to_lag \n                                                              lag  (, max_lag + )]),\n    on=[, , ],\n    how=\n)\n\n\n\n\n\ncombined_data = test.join(\n    filtered_df,\n    on=[, , ],\n    how=\n)\n\n\n lags_   :\n    combined_data = combined_data.join(\n        lags_,\n        on=[, ],\n        how=\n    )\n\n\ncolumns_to_keep = [col  col  combined_data.columns   col.endswith()]\ncombined_data = combined_data.select(columns_to_keep)\n\n\ncombined_data = combined_data.fill_null()\n\n\nmissing_columns = (feature_columns) - (combined_data.columns)\n missing_columns:\n     missing_col  missing_columns:\n        combined_data = combined_data.with_columns(\n            pl.lit().alias(missing_col)\n        )\n\n\nX_pred = combined_data.select(feature_columns).to_pandas()\nX_pred.fillna(, inplace=)\n\n\n (loaded_model, ):\n    expected_features = loaded_model.feature_names_in_\n    X_pred = X_pred.reindex(columns=expected_features, fill_value=)\n:\n    X_pred = X_pred[feature_columns]\n\n\npredictions = loaded_model.predict(X_pred)\n\n\ncombined_data = combined_data.with_columns(pl.Series(name=, values=predictions))\n\n\npredictions_df = combined_data.select([, ])\n\n\n (predictions_df, pl.DataFrame):\n     predictions_df.columns == [, ]\n (predictions_df, pd.DataFrame):\n     (predictions_df.columns) == [, ]\n:\n    \n    predictions_df = pd.DataFrame(columns=[, ])\n\n\n (predictions_df) != (test):\n    \n    predictions_df = test.select([]).with_columns(\n        pl.lit().alias()\n    )\n\n predictions_df\n</code></pre>\n<p>`</p>",
      "rawMarkdown": "So I have trained my model to keep storing history cache, so that I can use lagged values of some features as featuers, I am using a very basic linear regression model. But everytime I run the following code it leads to Notebook timeout and I am not able to understand why, I am attaching my testing code below if you have any improvements or suggestions to solve this issue, please let me know. \n\n`import pickle\nimport time\nimport polars as pl\nimport pandas as pd\nfrom sklearn.linear_model import LinearRegression\n\n# Global variables\nlags_: pl.DataFrame | None = None\nhistory_cache = pl.DataFrame()\nlags_cache = pl.DataFrame()\nloaded_model = None\n'\n'\n# Feature and responder configurations\nfeatures_to_lag = [\n    'feature_06', 'feature_60', 'feature_49', 'feature_04', 'feature_07',\n    'feature_58', 'feature_59', 'feature_47', 'feature_51', 'feature_36',\n    'feature_52', 'feature_68', 'feature_13', 'feature_02', 'feature_05']\nlag_range = range(1, 11)  # Lags from t-1 to t-10\nresponder_columns = [f\"responder_{i}\" for i in range(9)]\n''\n# Feature columns for prediction\nlagged_feature_cols = [f\"{feature}_lag_{lag}\" for feature in features_to_lag for lag in lag_range]\nlagged_responder_cols = [f\"{responder}_lag_1\" for responder in responder_columns]\nfeature_columns = features_to_lag + lagged_feature_cols + lagged_responder_cols + [\"time_id\"]\nmax_lag = 11\n\nwith open(\"/kaggle/input/linregwithlagged/other/default/1/linreg_model_cv_nonrandom (1) (1).pkl\", \"rb\") as f:\n    loaded_model = pickle.load(f)\n''\ndef predict(test: pl.DataFrame, lags: pl.DataFrame | None) -> pl.DataFrame | pd.DataFrame:\n    \"\"\"Make a prediction for the test set.\"\"\"\n    global lags_, loaded_model, history_cache, lags_cache, max_lag\n\n    # Ensure test DataFrame is not None\n    if test is None:\n        test = pl.DataFrame()\n\n    # If lags DataFrame is provided, store the lagged responder data\n    if lags is not None:\n        lags_ = lags\n        lags_cache = pl.concat([lags_cache, lags])\n\n    # If test DataFrame is empty, create necessary columns with zeros\n    if test.is_empty():\n        # Create empty DataFrame with required columns\n        test = pl.DataFrame({col: [] for col in ['date_id', 'time_id', 'symbol_id', 'row_id'] + features_to_lag})\n\n    date_id = test[\"date_id\"][0] if not test.is_empty() else 0\n\n    # Load the saved Linear Regression model\n    \n\n    # Update history_cache with test data\n    history_cache = pl.concat([history_cache, test])\n\n    # Filter history_cache and lags_cache to keep only recent data\n    min_date_id = max(date_id - 2, 0)\n    history_cache = history_cache.filter(pl.col(\"date_id\") >= min_date_id)\n    lags_cache = lags_cache.filter(pl.col(\"date_id\") >= min_date_id)\n\n    # Start with filtered_df as a copy of history_cache\n    filtered_df = history_cache\n\n    # Filter to only the rows for the current date_id\n    current_df = filtered_df.filter(pl.col(\"date_id\") == date_id)\n    \n    # Create lagged features only for the current date_id\n    for lag in range(1, max_lag + 1):\n        lagged_cols = [\n            pl.col(feature).shift(lag).alias(f\"{feature}_lag_{lag}\")\n            for feature in features_to_lag\n        ]\n        current_df = current_df.with_columns(lagged_cols)\n    \n    # Join the lagged columns back into the main filtered_df\n    filtered_df = filtered_df.join(\n        current_df.select([\"date_id\", \"time_id\", \"symbol_id\"] + [f\"{feature}_lag_{lag}\" \n                                                                 for feature in features_to_lag \n                                                                 for lag in range(1, max_lag + 1)]),\n        on=[\"date_id\", \"time_id\", \"symbol_id\"],\n        how=\"left\"\n    )\n\n\n    # Proceed with joining filtered_df with test data\n\n    # Step 1: Join `test` with `filtered_df` on `date_id`, `time_id`, and `symbol_id`\n    combined_data = test.join(\n        filtered_df,\n        on=[\"date_id\", \"time_id\", \"symbol_id\"],\n        how=\"left\"\n    )\n\n    # Step 2: Join the result with `lags_` on `date_id` and `symbol_id`, if lags_ is not None\n    if lags_ is not None:\n        combined_data = combined_data.join(\n            lags_,\n            on=[\"date_id\", \"symbol_id\"],\n            how=\"left\"\n        )\n\n    # Remove columns ending with '_right' if any\n    columns_to_keep = [col for col in combined_data.columns if not col.endswith('_right')]\n    combined_data = combined_data.select(columns_to_keep)\n\n    # Fill nulls for missing features dynamically\n    combined_data = combined_data.fill_null(0)\n\n    # Check if all feature columns exist in the combined data\n    missing_columns = set(feature_columns) - set(combined_data.columns)\n    if missing_columns:\n        for missing_col in missing_columns:\n            combined_data = combined_data.with_columns(\n                pl.lit(0).alias(missing_col)\n            )\n\n    # Prepare input features for prediction\n    X_pred = combined_data.select(feature_columns).to_pandas()\n    X_pred.fillna(0, inplace=True)\n\n    # Ensure the feature columns are in the same order as during training\n    if hasattr(loaded_model, 'feature_names_in_'):\n        expected_features = loaded_model.feature_names_in_\n        X_pred = X_pred.reindex(columns=expected_features, fill_value=0)\n    else:\n        X_pred = X_pred[feature_columns]\n\n    # Predict using the Linear Regression model\n    predictions = loaded_model.predict(X_pred)\n\n    # Add predictions to the DataFrame\n    combined_data = combined_data.with_columns(pl.Series(name=\"responder_6\", values=predictions))\n\n    # Prepare the final predictions DataFrame\n    predictions_df = combined_data.select(['row_id', 'responder_6'])\n\n    # Ensure predictions_df is a DataFrame and has the correct columns\n    if isinstance(predictions_df, pl.DataFrame):\n        assert predictions_df.columns == ['row_id', 'responder_6']\n    elif isinstance(predictions_df, pd.DataFrame):\n        assert list(predictions_df.columns) == ['row_id', 'responder_6']\n    else:\n        # If predictions_df is empty, create an empty DataFrame with required columns\n        predictions_df = pd.DataFrame(columns=['row_id', 'responder_6'])\n\n    # Confirm that predictions match the test data length\n    if len(predictions_df) != len(test):\n        # If lengths do not match, adjust predictions_df to match test data\n        predictions_df = test.select(['row_id']).with_columns(\n            pl.lit(0).alias('responder_6')\n        )\n\n    return predictions_df\n`",
      "votes": null
    },
    {
      "id": "3070592",
      "postDate": "12/12/2024 21:36:34",
      "content": "<p>If I correctly understand your code:</p>\n<pre><code>date_id = test[][]   test.is_empty()  \nmin_date_id = (date_id - , )\n</code></pre>\n<p>means that min_date_id will be -2 and ascending, because test date_id starts at 0 and not from 1699 if that you were thinking here.</p>\n<p>And if that's the problem - it means you are filtering out just the begining of the whole df here and not keeping it's end like you wanted:</p>\n<pre><code>history_cache = history_cache.(pl.col() &gt;= min_date_id)\nlags_cache = lags_cache.(pl.col() &gt;= min_date_id)\n</code></pre>\n<p>With all that in mind - you are dealing with a huge df and trying to manipulate with it at every step, so you are out of time</p>",
      "rawMarkdown": "If I correctly understand your code:\n```python\ndate_id = test[\"date_id\"][0] if not test.is_empty() else 0\nmin_date_id = max(date_id - 2, 0)\n```\nmeans that min_date_id will be -2 and ascending, because test date_id starts at 0 and not from 1699 if that you were thinking here.\n\nAnd if that's the problem - it means you are filtering out just the begining of the whole df here and not keeping it's end like you wanted:\n```python\nhistory_cache = history_cache.filter(pl.col(\"date_id\") >= min_date_id)\nlags_cache = lags_cache.filter(pl.col(\"date_id\") >= min_date_id)\n```\n\nWith all that in mind - you are dealing with a huge df and trying to manipulate with it at every step, so you are out of time",
      "votes": null
    },
    {
      "id": "3070596",
      "postDate": "12/12/2024 21:44:39",
      "content": "<p>Thank you for pointing out this issue, I will try to update my code according to the changes you suggested and hopefully have a working submission. </p>",
      "rawMarkdown": "Thank you for pointing out this issue, I will try to update my code according to the changes you suggested and hopefully have a working submission.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 3070592,
      "author_name": "eu1234",
      "author_url": "",
      "post_date": "12/12/2024 21:36:34",
      "content": "<p>If I correctly understand your code:</p>\n<pre><code>date_id = test[][]   test.is_empty()  \nmin_date_id = (date_id - , )\n</code></pre>\n<p>means that min_date_id will be -2 and ascending, because test date_id starts at 0 and not from 1699 if that you were thinking here.</p>\n<p>And if that's the problem - it means you are filtering out just the begining of the whole df here and not keeping it's end like you wanted:</p>\n<pre><code>history_cache = history_cache.(pl.col() &gt;= min_date_id)\nlags_cache = lags_cache.(pl.col() &gt;= min_date_id)\n</code></pre>\n<p>With all that in mind - you are dealing with a huge df and trying to manipulate with it at every step, so you are out of time</p>",
      "votes": null,
      "replies": [
        {
          "id": 3070596,
          "author_name": "dipitgolechha7",
          "author_url": "",
          "post_date": "12/12/2024 21:44:39",
          "content": "<p>Thank you for pointing out this issue, I will try to update my code according to the changes you suggested and hopefully have a working submission. </p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "3070455": "So I have trained my model to keep storing history cache, so that I can use lagged values of some features as featuers, I am using a very basic linear regression model. But everytime I run the following code it leads to Notebook timeout and I am not able to understand why, I am attaching my testing code below if you have any improvements or suggestions to solve this issue, please let me know. \n\n`import pickle\nimport time\nimport polars as pl\nimport pandas as pd\nfrom sklearn.linear_model import LinearRegression\n\n# Global variables\nlags_: pl.DataFrame | None = None\nhistory_cache = pl.DataFrame()\nlags_cache = pl.DataFrame()\nloaded_model = None\n'\n'\n# Feature and responder configurations\nfeatures_to_lag = [\n    'feature_06', 'feature_60', 'feature_49', 'feature_04', 'feature_07',\n    'feature_58', 'feature_59', 'feature_47', 'feature_51', 'feature_36',\n    'feature_52', 'feature_68', 'feature_13', 'feature_02', 'feature_05']\nlag_range = range(1, 11)  # Lags from t-1 to t-10\nresponder_columns = [f\"responder_{i}\" for i in range(9)]\n''\n# Feature columns for prediction\nlagged_feature_cols = [f\"{feature}_lag_{lag}\" for feature in features_to_lag for lag in lag_range]\nlagged_responder_cols = [f\"{responder}_lag_1\" for responder in responder_columns]\nfeature_columns = features_to_lag + lagged_feature_cols + lagged_responder_cols + [\"time_id\"]\nmax_lag = 11\n\nwith open(\"/kaggle/input/linregwithlagged/other/default/1/linreg_model_cv_nonrandom (1) (1).pkl\", \"rb\") as f:\n    loaded_model = pickle.load(f)\n''\ndef predict(test: pl.DataFrame, lags: pl.DataFrame | None) -> pl.DataFrame | pd.DataFrame:\n    \"\"\"Make a prediction for the test set.\"\"\"\n    global lags_, loaded_model, history_cache, lags_cache, max_lag\n\n    # Ensure test DataFrame is not None\n    if test is None:\n        test = pl.DataFrame()\n\n    # If lags DataFrame is provided, store the lagged responder data\n    if lags is not None:\n        lags_ = lags\n        lags_cache = pl.concat([lags_cache, lags])\n\n    # If test DataFrame is empty, create necessary columns with zeros\n    if test.is_empty():\n        # Create empty DataFrame with required columns\n        test = pl.DataFrame({col: [] for col in ['date_id', 'time_id', 'symbol_id', 'row_id'] + features_to_lag})\n\n    date_id = test[\"date_id\"][0] if not test.is_empty() else 0\n\n    # Load the saved Linear Regression model\n    \n\n    # Update history_cache with test data\n    history_cache = pl.concat([history_cache, test])\n\n    # Filter history_cache and lags_cache to keep only recent data\n    min_date_id = max(date_id - 2, 0)\n    history_cache = history_cache.filter(pl.col(\"date_id\") >= min_date_id)\n    lags_cache = lags_cache.filter(pl.col(\"date_id\") >= min_date_id)\n\n    # Start with filtered_df as a copy of history_cache\n    filtered_df = history_cache\n\n    # Filter to only the rows for the current date_id\n    current_df = filtered_df.filter(pl.col(\"date_id\") == date_id)\n    \n    # Create lagged features only for the current date_id\n    for lag in range(1, max_lag + 1):\n        lagged_cols = [\n            pl.col(feature).shift(lag).alias(f\"{feature}_lag_{lag}\")\n            for feature in features_to_lag\n        ]\n        current_df = current_df.with_columns(lagged_cols)\n    \n    # Join the lagged columns back into the main filtered_df\n    filtered_df = filtered_df.join(\n        current_df.select([\"date_id\", \"time_id\", \"symbol_id\"] + [f\"{feature}_lag_{lag}\" \n                                                                 for feature in features_to_lag \n                                                                 for lag in range(1, max_lag + 1)]),\n        on=[\"date_id\", \"time_id\", \"symbol_id\"],\n        how=\"left\"\n    )\n\n\n    # Proceed with joining filtered_df with test data\n\n    # Step 1: Join `test` with `filtered_df` on `date_id`, `time_id`, and `symbol_id`\n    combined_data = test.join(\n        filtered_df,\n        on=[\"date_id\", \"time_id\", \"symbol_id\"],\n        how=\"left\"\n    )\n\n    # Step 2: Join the result with `lags_` on `date_id` and `symbol_id`, if lags_ is not None\n    if lags_ is not None:\n        combined_data = combined_data.join(\n            lags_,\n            on=[\"date_id\", \"symbol_id\"],\n            how=\"left\"\n        )\n\n    # Remove columns ending with '_right' if any\n    columns_to_keep = [col for col in combined_data.columns if not col.endswith('_right')]\n    combined_data = combined_data.select(columns_to_keep)\n\n    # Fill nulls for missing features dynamically\n    combined_data = combined_data.fill_null(0)\n\n    # Check if all feature columns exist in the combined data\n    missing_columns = set(feature_columns) - set(combined_data.columns)\n    if missing_columns:\n        for missing_col in missing_columns:\n            combined_data = combined_data.with_columns(\n                pl.lit(0).alias(missing_col)\n            )\n\n    # Prepare input features for prediction\n    X_pred = combined_data.select(feature_columns).to_pandas()\n    X_pred.fillna(0, inplace=True)\n\n    # Ensure the feature columns are in the same order as during training\n    if hasattr(loaded_model, 'feature_names_in_'):\n        expected_features = loaded_model.feature_names_in_\n        X_pred = X_pred.reindex(columns=expected_features, fill_value=0)\n    else:\n        X_pred = X_pred[feature_columns]\n\n    # Predict using the Linear Regression model\n    predictions = loaded_model.predict(X_pred)\n\n    # Add predictions to the DataFrame\n    combined_data = combined_data.with_columns(pl.Series(name=\"responder_6\", values=predictions))\n\n    # Prepare the final predictions DataFrame\n    predictions_df = combined_data.select(['row_id', 'responder_6'])\n\n    # Ensure predictions_df is a DataFrame and has the correct columns\n    if isinstance(predictions_df, pl.DataFrame):\n        assert predictions_df.columns == ['row_id', 'responder_6']\n    elif isinstance(predictions_df, pd.DataFrame):\n        assert list(predictions_df.columns) == ['row_id', 'responder_6']\n    else:\n        # If predictions_df is empty, create an empty DataFrame with required columns\n        predictions_df = pd.DataFrame(columns=['row_id', 'responder_6'])\n\n    # Confirm that predictions match the test data length\n    if len(predictions_df) != len(test):\n        # If lengths do not match, adjust predictions_df to match test data\n        predictions_df = test.select(['row_id']).with_columns(\n            pl.lit(0).alias('responder_6')\n        )\n\n    return predictions_df\n`",
    "3070592": "If I correctly understand your code:\n```python\ndate_id = test[\"date_id\"][0] if not test.is_empty() else 0\nmin_date_id = max(date_id - 2, 0)\n```\nmeans that min_date_id will be -2 and ascending, because test date_id starts at 0 and not from 1699 if that you were thinking here.\n\nAnd if that's the problem - it means you are filtering out just the begining of the whole df here and not keeping it's end like you wanted:\n```python\nhistory_cache = history_cache.filter(pl.col(\"date_id\") >= min_date_id)\nlags_cache = lags_cache.filter(pl.col(\"date_id\") >= min_date_id)\n```\n\nWith all that in mind - you are dealing with a huge df and trying to manipulate with it at every step, so you are out of time",
    "3070596": "Thank you for pointing out this issue, I will try to update my code according to the changes you suggested and hopefully have a working submission."
  },
  "source": "meta"
}