{
  "id": 549915,
  "title": "Notebook inference error",
  "url": "/competitions/jane-street-real-time-market-data-forecasting/discussion/549915",
  "author_name": "",
  "post_date": "2024-12-04T14:34:17.691864700Z",
  "votes": null,
  "comment_count": 2,
  "views": 0,
  "content": "<p>When I submit my notebook, I get this error</p>\n<blockquote>\n  <p>Notebook Inference Server Error<br>\n  Your submission notebook may not have started the inference server that is called to obtain predictions. This could mean you forgot to start it, or the notebook crashed.</p>\n</blockquote>\n<p>The code runs successfully when I run it locally. And when I submit it, the logs show that most of the code runs sucesfully, in well under 15 minutes. How could this be the case? It uses much less than 30gb ram as well.</p>\n<p>I've submitted effectively the exact same code before, and it has worked.</p>\n<p>I've added the notebook along.</p>",
  "messages": [
    {
      "id": "3063472",
      "postDate": "12/04/2024 14:34:17",
      "content": "<p>When I submit my notebook, I get this error</p>\n<blockquote>\n  <p>Notebook Inference Server Error<br>\n  Your submission notebook may not have started the inference server that is called to obtain predictions. This could mean you forgot to start it, or the notebook crashed.</p>\n</blockquote>\n<p>The code runs successfully when I run it locally. And when I submit it, the logs show that most of the code runs sucesfully, in well under 15 minutes. How could this be the case? It uses much less than 30gb ram as well.</p>\n<p>I've submitted effectively the exact same code before, and it has worked.</p>\n<p>I've added the notebook along.</p>",
      "rawMarkdown": "When I submit my notebook, I get this error\n\n>Notebook Inference Server Error\nYour submission notebook may not have started the inference server that is called to obtain predictions. This could mean you forgot to start it, or the notebook crashed.\n\n\nThe code runs successfully when I run it locally. And when I submit it, the logs show that most of the code runs sucesfully, in well under 15 minutes. How could this be the case? It uses much less than 30gb ram as well.\n\nI've submitted effectively the exact same code before, and it has worked.\n\nI've added the notebook along.",
      "votes": null
    },
    {
      "id": "3063628",
      "postDate": "12/04/2024 17:21:04",
      "content": "<p>Try to debug your inference code with the following code snippet:</p>\n<pre><code>test_data = pl.scan_parquet()\ntest_data = test_data.(\n    pl.col() &gt; \n).collect()\n\nlag_data_columns = {\n    :,\n    :,\n    :,\n    :,\n    :,\n    :,\n    :,\n    :,\n    :,\n    :,\n    :,\n    :\n}\n\n\njane_predictor = JanePredictor(\n    initial_data = data,\n    y = ,\n    train_window = train_size,\n    forecast_window = ,\n    lgbm_params = lgbm_params\n)\n\ntime_list = []\n row  test_data.select([,]).unique().sort([,]).iter_rows(named=):\n\n    (row[], row[])\n\n    test = test_data.(\n        ( pl.col() == row[] ) &amp;\n        ( pl.col() == row[] )\n    ).with_row_count()\n\n     row[] == :\n        lag = test.select(\n            ( lag_data_columns.keys() )\n        ).rename(lag_data_columns)\n\n        start_time = time.perf_counter()\n        jane_predictor.predict(test, lag) \n        end_time = time.perf_counter()\n        delta_time = delta_time = end_time - start_time\n        time_list.append(delta_time)\n        ()\n\n    :\n        start_time = time.perf_counter()\n        jane_predictor.predict(test, ) \n        end_time = time.perf_counter()\n        delta_time = delta_time = end_time - start_time\n        time_list.append(delta_time)\n        ()\n</code></pre>\n<p>Pay attention to the comments where I'm pointing out what you should replace.</p>",
      "rawMarkdown": "Try to debug your inference code with the following code snippet:\n\n```python\ntest_data = pl.scan_parquet('/kaggle/input/jane-street-real-time-market-data-forecasting/train.parquet')\ntest_data = test_data.filter(\n    pl.col('date_id') > 1572\n).collect()\n\nlag_data_columns = {\n    'date_id':'date_id',\n    'time_id':'time_id',\n    'symbol_id':'symbol_id',\n    'responder_0':'responder_0_lag_1',\n    'responder_1':'responder_1_lag_1',\n    'responder_2':'responder_2_lag_1',\n    'responder_3':'responder_3_lag_1',\n    'responder_4':'responder_4_lag_1',\n    'responder_5':'responder_5_lag_1',\n    'responder_6':'responder_6_lag_1',\n    'responder_7':'responder_7_lag_1',\n    'responder_8':'responder_8_lag_1'\n}\n\n# If you are not using class delete this\njane_predictor = JanePredictor(\n    initial_data = data,\n    y = 'responder_6',\n    train_window = train_size,\n    forecast_window = 1,\n    lgbm_params = lgbm_params\n)\n\ntime_list = []\nfor row in test_data.select(['date_id','time_id']).unique().sort(['date_id','time_id']).iter_rows(named=True):\n    \n    print(row['date_id'], row['time_id'])\n    \n    test = test_data.filter(\n        ( pl.col('date_id') == row['date_id'] ) &\n        ( pl.col('time_id') == row['time_id'] )\n    ).with_row_count(\"row_id\")\n    \n    if row['time_id'] == 0:\n        lag = test.select(\n            list( lag_data_columns.keys() )\n        ).rename(lag_data_columns)\n        \n        start_time = time.perf_counter()\n        jane_predictor.predict(test, lag) # replace this with your prediction method\n        end_time = time.perf_counter()\n        delta_time = delta_time = end_time - start_time\n        time_list.append(delta_time)\n        print(f\"Elapsed time: {delta_time:.6f} seconds\")\n\n    else:\n        start_time = time.perf_counter()\n        jane_predictor.predict(test, None) # replace this with your prediction method\n        end_time = time.perf_counter()\n        delta_time = delta_time = end_time - start_time\n        time_list.append(delta_time)\n        print(f\"Elapsed time: {delta_time:.6f} seconds\")\n```\n\nPay attention to the comments where I'm pointing out what you should replace.",
      "votes": null
    },
    {
      "id": "3064261",
      "postDate": "12/05/2024 12:45:39",
      "content": "<p>I figured it out. It was just taking too long to do inference. But the error's when submitting are not so informative I have to admit.</p>",
      "rawMarkdown": "I figured it out. It was just taking too long to do inference. But the error's when submitting are not so informative I have to admit.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 3063628,
      "author_name": "serjhenrique",
      "author_url": "",
      "post_date": "12/04/2024 17:21:04",
      "content": "<p>Try to debug your inference code with the following code snippet:</p>\n<pre><code>test_data = pl.scan_parquet()\ntest_data = test_data.(\n    pl.col() &gt; \n).collect()\n\nlag_data_columns = {\n    :,\n    :,\n    :,\n    :,\n    :,\n    :,\n    :,\n    :,\n    :,\n    :,\n    :,\n    :\n}\n\n\njane_predictor = JanePredictor(\n    initial_data = data,\n    y = ,\n    train_window = train_size,\n    forecast_window = ,\n    lgbm_params = lgbm_params\n)\n\ntime_list = []\n row  test_data.select([,]).unique().sort([,]).iter_rows(named=):\n\n    (row[], row[])\n\n    test = test_data.(\n        ( pl.col() == row[] ) &amp;\n        ( pl.col() == row[] )\n    ).with_row_count()\n\n     row[] == :\n        lag = test.select(\n            ( lag_data_columns.keys() )\n        ).rename(lag_data_columns)\n\n        start_time = time.perf_counter()\n        jane_predictor.predict(test, lag) \n        end_time = time.perf_counter()\n        delta_time = delta_time = end_time - start_time\n        time_list.append(delta_time)\n        ()\n\n    :\n        start_time = time.perf_counter()\n        jane_predictor.predict(test, ) \n        end_time = time.perf_counter()\n        delta_time = delta_time = end_time - start_time\n        time_list.append(delta_time)\n        ()\n</code></pre>\n<p>Pay attention to the comments where I'm pointing out what you should replace.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 3064261,
      "author_name": "reeeeeeeeeeeeeee",
      "author_url": "",
      "post_date": "12/05/2024 12:45:39",
      "content": "<p>I figured it out. It was just taking too long to do inference. But the error's when submitting are not so informative I have to admit.</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "3063472": "When I submit my notebook, I get this error\n\n>Notebook Inference Server Error\nYour submission notebook may not have started the inference server that is called to obtain predictions. This could mean you forgot to start it, or the notebook crashed.\n\n\nThe code runs successfully when I run it locally. And when I submit it, the logs show that most of the code runs sucesfully, in well under 15 minutes. How could this be the case? It uses much less than 30gb ram as well.\n\nI've submitted effectively the exact same code before, and it has worked.\n\nI've added the notebook along.",
    "3063628": "Try to debug your inference code with the following code snippet:\n\n```python\ntest_data = pl.scan_parquet('/kaggle/input/jane-street-real-time-market-data-forecasting/train.parquet')\ntest_data = test_data.filter(\n    pl.col('date_id') > 1572\n).collect()\n\nlag_data_columns = {\n    'date_id':'date_id',\n    'time_id':'time_id',\n    'symbol_id':'symbol_id',\n    'responder_0':'responder_0_lag_1',\n    'responder_1':'responder_1_lag_1',\n    'responder_2':'responder_2_lag_1',\n    'responder_3':'responder_3_lag_1',\n    'responder_4':'responder_4_lag_1',\n    'responder_5':'responder_5_lag_1',\n    'responder_6':'responder_6_lag_1',\n    'responder_7':'responder_7_lag_1',\n    'responder_8':'responder_8_lag_1'\n}\n\n# If you are not using class delete this\njane_predictor = JanePredictor(\n    initial_data = data,\n    y = 'responder_6',\n    train_window = train_size,\n    forecast_window = 1,\n    lgbm_params = lgbm_params\n)\n\ntime_list = []\nfor row in test_data.select(['date_id','time_id']).unique().sort(['date_id','time_id']).iter_rows(named=True):\n    \n    print(row['date_id'], row['time_id'])\n    \n    test = test_data.filter(\n        ( pl.col('date_id') == row['date_id'] ) &\n        ( pl.col('time_id') == row['time_id'] )\n    ).with_row_count(\"row_id\")\n    \n    if row['time_id'] == 0:\n        lag = test.select(\n            list( lag_data_columns.keys() )\n        ).rename(lag_data_columns)\n        \n        start_time = time.perf_counter()\n        jane_predictor.predict(test, lag) # replace this with your prediction method\n        end_time = time.perf_counter()\n        delta_time = delta_time = end_time - start_time\n        time_list.append(delta_time)\n        print(f\"Elapsed time: {delta_time:.6f} seconds\")\n\n    else:\n        start_time = time.perf_counter()\n        jane_predictor.predict(test, None) # replace this with your prediction method\n        end_time = time.perf_counter()\n        delta_time = delta_time = end_time - start_time\n        time_list.append(delta_time)\n        print(f\"Elapsed time: {delta_time:.6f} seconds\")\n```\n\nPay attention to the comments where I'm pointing out what you should replace.",
    "3064261": "I figured it out. It was just taking too long to do inference. But the error's when submitting are not so informative I have to admit."
  },
  "source": "meta"
}