{
  "id": 545561,
  "title": "Notebook Inference Server Error",
  "url": "/competitions/jane-street-real-time-market-data-forecasting/discussion/545561",
  "author_name": "",
  "post_date": "2024-11-11T00:56:52.570768300Z",
  "votes": 1,
  "comment_count": 4,
  "views": 0,
  "content": "<p>What could cause the scoring process to fail <strong>within 30 seconds</strong>⁉️</p>\n<blockquote>\n  <p>Notebook Inference Server Error<br>\n  Your submission notebook may not have started the inference server that is called to obtain predictions. This could mean you forgot to start it, or the notebook crashed. </p>\n</blockquote>\n<p>My code is as follows, almost <strong>exactly the same</strong> as the demo: </p>\n<pre><code> autogluon.core.dataset  TabularDataset\n\nlags_ : pl.DataFrame |  = \n\n\n\n\n () -&gt; pl.DataFrame | pd.DataFrame:\n    \n    \n    \n\n     lags_\n     lags   :\n        lags_ = lags\n\n    \n    predictions = test.select(\n        ,\n        pl.lit().alias(),\n    )\n\n    feature_names = [  i  ()]\n    feat = test[feature_names].to_pandas()\n    pred = pretrained_model.predict(feat, model=)\n    predictions = predictions.with_columns(pl.Series(, pred.to_numpy()))\n    (predictions)\n\n     (predictions, pl.DataFrame):\n         predictions.columns == [, ]\n     (predictions, pd.DataFrame):\n         (predictions.columns == [, ]).()\n    :\n         TypeError()\n    \n     (predictions) == (test)\n\n     predictions\n</code></pre>\n<pre><code>inference_server = kaggle_evaluation(predict)\n\n os():\n    inference_server()\n:\n    inference_server(\n        (\n            ,\n            ,\n        )\n    )\n</code></pre>\n<p>The notebook runs <strong>successfully</strong> every time, and the last cell runs only about <strong>600 milliseconds</strong> for each time_id. What have I overlooked?</p>\n<p>Thank you very much for your help.</p>",
  "messages": [
    {
      "id": "3041880",
      "postDate": "11/11/2024 00:56:52",
      "content": "<p>What could cause the scoring process to fail <strong>within 30 seconds</strong>⁉️</p>\n<blockquote>\n  <p>Notebook Inference Server Error<br>\n  Your submission notebook may not have started the inference server that is called to obtain predictions. This could mean you forgot to start it, or the notebook crashed. </p>\n</blockquote>\n<p>My code is as follows, almost <strong>exactly the same</strong> as the demo: </p>\n<pre><code> autogluon.core.dataset  TabularDataset\n\nlags_ : pl.DataFrame |  = \n\n\n\n\n () -&gt; pl.DataFrame | pd.DataFrame:\n    \n    \n    \n\n     lags_\n     lags   :\n        lags_ = lags\n\n    \n    predictions = test.select(\n        ,\n        pl.lit().alias(),\n    )\n\n    feature_names = [  i  ()]\n    feat = test[feature_names].to_pandas()\n    pred = pretrained_model.predict(feat, model=)\n    predictions = predictions.with_columns(pl.Series(, pred.to_numpy()))\n    (predictions)\n\n     (predictions, pl.DataFrame):\n         predictions.columns == [, ]\n     (predictions, pd.DataFrame):\n         (predictions.columns == [, ]).()\n    :\n         TypeError()\n    \n     (predictions) == (test)\n\n     predictions\n</code></pre>\n<pre><code>inference_server = kaggle_evaluation(predict)\n\n os():\n    inference_server()\n:\n    inference_server(\n        (\n            ,\n            ,\n        )\n    )\n</code></pre>\n<p>The notebook runs <strong>successfully</strong> every time, and the last cell runs only about <strong>600 milliseconds</strong> for each time_id. What have I overlooked?</p>\n<p>Thank you very much for your help.</p>",
      "rawMarkdown": "What could cause the scoring process to fail **within 30 seconds**⁉️\n\n>Notebook Inference Server Error\nYour submission notebook may not have started the inference server that is called to obtain predictions. This could mean you forgot to start it, or the notebook crashed. \n\n\nMy code is as follows, almost **exactly the same** as the demo: \n```\nfrom autogluon.core.dataset import TabularDataset\n\nlags_ : pl.DataFrame | None = None\n\n# Replace this function with your inference code.\n# You can return either a Pandas or Polars dataframe, though Polars is recommended.\n# Each batch of predictions (except the very first) must be returned within 1 minute of the batch features being provided.\ndef predict(test: pl.DataFrame, lags: pl.DataFrame | None) -> pl.DataFrame | pd.DataFrame:\n    \"\"\"Make a prediction.\"\"\"\n    # All the responders from the previous day are passed in at time_id == 0. We save them in a global variable for access at every time_id.\n    # Use them as extra features, if you like.\n\n    global lags_\n    if lags is not None:\n        lags_ = lags\n\n    # Replace this section with your own predictions\n    predictions = test.select(\n        'row_id',\n        pl.lit(0.0).alias('responder_6'),\n    )\n\n    feature_names = [f\"feature_{i:02d}\" for i in range(79)]\n    feat = test[feature_names].to_pandas()\n    pred = pretrained_model.predict(feat, model=\"WeightedEnsemble_2_L2\")\n    predictions = predictions.with_columns(pl.Series('responder_6', pred.to_numpy()))\n    print(predictions)\n\n    if isinstance(predictions, pl.DataFrame):\n        assert predictions.columns == ['row_id', 'responder_6']\n    elif isinstance(predictions, pd.DataFrame):\n        assert (predictions.columns == ['row_id', 'responder_6']).all()\n    else:\n        raise TypeError('The predict function must return a DataFrame')\n    # Confirm has as many rows as the test data.\n    assert len(predictions) == len(test)\n\n    return predictions\n```\n```\ninference_server = kaggle_evaluation.jane_street_inference_server.JSInferenceServer(predict)\n\nif os.getenv('KAGGLE_IS_COMPETITION_RERUN'):\n    inference_server.serve()\nelse:\n    inference_server.run_local_gateway(\n        (\n            '/kaggle/input/jane-street-real-time-market-data-forecasting/test.parquet',\n            '/kaggle/input/jane-street-real-time-market-data-forecasting/lags.parquet',\n        )\n    )\n```\nThe notebook runs **successfully** every time, and the last cell runs only about **600 milliseconds** for each time_id. What have I overlooked?\n\nThank you very much for your help.",
      "votes": null
    },
    {
      "id": "3041908",
      "postDate": "11/11/2024 01:51:34",
      "content": "<blockquote>\n  <p>When your notebook is run on the hidden test set, inference_server.serve must be called within 15 minutes of the notebook starting or the gateway will throw an error. If you need more than 15 minutes to load your model you can do so during the very first predict call, which does not have the usual 1 minute response deadline.</p>\n</blockquote>\n<p>Maybe this part?</p>\n<p><code>pred = pretrained_model.predict(feat, model=\"WeightedEnsemble_2_L2\")</code></p>",
      "rawMarkdown": ">When your notebook is run on the hidden test set, inference_server.serve must be called within 15 minutes of the notebook starting or the gateway will throw an error. If you need more than 15 minutes to load your model you can do so during the very first predict call, which does not have the usual 1 minute response deadline.\n\nMaybe this part?\n\n`pred = pretrained_model.predict(feat, model=\"WeightedEnsemble_2_L2\")`",
      "votes": null
    },
    {
      "id": "3041939",
      "postDate": "11/11/2024 03:04:15",
      "content": "<p>Thanks for the reply. I directly load the pre-trained model from the uploaded input. After testing, it only takes approximately 350 milliseconds to complete this operation. So it really confuses me because I am even far from the 1-minute limit.😂</p>",
      "rawMarkdown": "Thanks for the reply. I directly load the pre-trained model from the uploaded input. After testing, it only takes approximately 350 milliseconds to complete this operation. So it really confuses me because I am even far from the 1-minute limit.😂",
      "votes": null
    },
    {
      "id": "3042053",
      "postDate": "11/11/2024 06:02:28",
      "content": "<p>I think I have figured out where the problem lies. 😅</p>\n<p>I cannot use \"pip install autogluon.tabular\" to install autogluon.tabular even though I place the code in the correct add-on area specifically designed for \"Install Packages\". This is due to the offline requirement for submission. </p>\n<p>Instead, I should use the input dataset to install the packages: <a href=\"https://www.kaggle.com/datasets/minhaozhang1/autogluon-1-1-1-tabular\" target=\"_blank\">https://www.kaggle.com/datasets/minhaozhang1/autogluon-1-1-1-tabular</a>. Many thanks to Minhao Zhang who contributed the dataset! </p>\n<p>However, strangely, when I submit and turn off the internet option, the notebook can run successfully, and the submission can also sometimes be successful (perhaps 1 out of 20). I think this must be a bug that the officials should pay attention to.</p>",
      "rawMarkdown": "I think I have figured out where the problem lies. 😅\n\nI cannot use \"pip install autogluon.tabular\" to install autogluon.tabular even though I place the code in the correct add-on area specifically designed for \"Install Packages\". This is due to the offline requirement for submission. \n\nInstead, I should use the input dataset to install the packages: https://www.kaggle.com/datasets/minhaozhang1/autogluon-1-1-1-tabular. Many thanks to Minhao Zhang who contributed the dataset! \n\nHowever, strangely, when I submit and turn off the internet option, the notebook can run successfully, and the submission can also sometimes be successful (perhaps 1 out of 20). I think this must be a bug that the officials should pay attention to.",
      "votes": null
    },
    {
      "id": "3042094",
      "postDate": "11/11/2024 07:34:32",
      "content": "<p>could be NaN values in test make it fail.</p>",
      "rawMarkdown": "could be NaN values in test make it fail.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 3041908,
      "author_name": "sweetyheehee",
      "author_url": "",
      "post_date": "11/11/2024 01:51:34",
      "content": "<blockquote>\n  <p>When your notebook is run on the hidden test set, inference_server.serve must be called within 15 minutes of the notebook starting or the gateway will throw an error. If you need more than 15 minutes to load your model you can do so during the very first predict call, which does not have the usual 1 minute response deadline.</p>\n</blockquote>\n<p>Maybe this part?</p>\n<p><code>pred = pretrained_model.predict(feat, model=\"WeightedEnsemble_2_L2\")</code></p>",
      "votes": null,
      "replies": [
        {
          "id": 3041939,
          "author_name": "carrotwait",
          "author_url": "",
          "post_date": "11/11/2024 03:04:15",
          "content": "<p>Thanks for the reply. I directly load the pre-trained model from the uploaded input. After testing, it only takes approximately 350 milliseconds to complete this operation. So it really confuses me because I am even far from the 1-minute limit.😂</p>",
          "votes": null,
          "replies": [
            {
              "id": 3042094,
              "author_name": "shiyili",
              "author_url": "",
              "post_date": "11/11/2024 07:34:32",
              "content": "<p>could be NaN values in test make it fail.</p>",
              "votes": null,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 3042053,
      "author_name": "carrotwait",
      "author_url": "",
      "post_date": "11/11/2024 06:02:28",
      "content": "<p>I think I have figured out where the problem lies. 😅</p>\n<p>I cannot use \"pip install autogluon.tabular\" to install autogluon.tabular even though I place the code in the correct add-on area specifically designed for \"Install Packages\". This is due to the offline requirement for submission. </p>\n<p>Instead, I should use the input dataset to install the packages: <a href=\"https://www.kaggle.com/datasets/minhaozhang1/autogluon-1-1-1-tabular\" target=\"_blank\">https://www.kaggle.com/datasets/minhaozhang1/autogluon-1-1-1-tabular</a>. Many thanks to Minhao Zhang who contributed the dataset! </p>\n<p>However, strangely, when I submit and turn off the internet option, the notebook can run successfully, and the submission can also sometimes be successful (perhaps 1 out of 20). I think this must be a bug that the officials should pay attention to.</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "3041880": "What could cause the scoring process to fail **within 30 seconds**⁉️\n\n>Notebook Inference Server Error\nYour submission notebook may not have started the inference server that is called to obtain predictions. This could mean you forgot to start it, or the notebook crashed. \n\n\nMy code is as follows, almost **exactly the same** as the demo: \n```\nfrom autogluon.core.dataset import TabularDataset\n\nlags_ : pl.DataFrame | None = None\n\n# Replace this function with your inference code.\n# You can return either a Pandas or Polars dataframe, though Polars is recommended.\n# Each batch of predictions (except the very first) must be returned within 1 minute of the batch features being provided.\ndef predict(test: pl.DataFrame, lags: pl.DataFrame | None) -> pl.DataFrame | pd.DataFrame:\n    \"\"\"Make a prediction.\"\"\"\n    # All the responders from the previous day are passed in at time_id == 0. We save them in a global variable for access at every time_id.\n    # Use them as extra features, if you like.\n\n    global lags_\n    if lags is not None:\n        lags_ = lags\n\n    # Replace this section with your own predictions\n    predictions = test.select(\n        'row_id',\n        pl.lit(0.0).alias('responder_6'),\n    )\n\n    feature_names = [f\"feature_{i:02d}\" for i in range(79)]\n    feat = test[feature_names].to_pandas()\n    pred = pretrained_model.predict(feat, model=\"WeightedEnsemble_2_L2\")\n    predictions = predictions.with_columns(pl.Series('responder_6', pred.to_numpy()))\n    print(predictions)\n\n    if isinstance(predictions, pl.DataFrame):\n        assert predictions.columns == ['row_id', 'responder_6']\n    elif isinstance(predictions, pd.DataFrame):\n        assert (predictions.columns == ['row_id', 'responder_6']).all()\n    else:\n        raise TypeError('The predict function must return a DataFrame')\n    # Confirm has as many rows as the test data.\n    assert len(predictions) == len(test)\n\n    return predictions\n```\n```\ninference_server = kaggle_evaluation.jane_street_inference_server.JSInferenceServer(predict)\n\nif os.getenv('KAGGLE_IS_COMPETITION_RERUN'):\n    inference_server.serve()\nelse:\n    inference_server.run_local_gateway(\n        (\n            '/kaggle/input/jane-street-real-time-market-data-forecasting/test.parquet',\n            '/kaggle/input/jane-street-real-time-market-data-forecasting/lags.parquet',\n        )\n    )\n```\nThe notebook runs **successfully** every time, and the last cell runs only about **600 milliseconds** for each time_id. What have I overlooked?\n\nThank you very much for your help.",
    "3041908": ">When your notebook is run on the hidden test set, inference_server.serve must be called within 15 minutes of the notebook starting or the gateway will throw an error. If you need more than 15 minutes to load your model you can do so during the very first predict call, which does not have the usual 1 minute response deadline.\n\nMaybe this part?\n\n`pred = pretrained_model.predict(feat, model=\"WeightedEnsemble_2_L2\")`",
    "3041939": "Thanks for the reply. I directly load the pre-trained model from the uploaded input. After testing, it only takes approximately 350 milliseconds to complete this operation. So it really confuses me because I am even far from the 1-minute limit.😂",
    "3042053": "I think I have figured out where the problem lies. 😅\n\nI cannot use \"pip install autogluon.tabular\" to install autogluon.tabular even though I place the code in the correct add-on area specifically designed for \"Install Packages\". This is due to the offline requirement for submission. \n\nInstead, I should use the input dataset to install the packages: https://www.kaggle.com/datasets/minhaozhang1/autogluon-1-1-1-tabular. Many thanks to Minhao Zhang who contributed the dataset! \n\nHowever, strangely, when I submit and turn off the internet option, the notebook can run successfully, and the submission can also sometimes be successful (perhaps 1 out of 20). I think this must be a bug that the officials should pay attention to.",
    "3042094": "could be NaN values in test make it fail."
  },
  "source": "meta"
}