{
  "id": 543571,
  "title": "How to use H2O models? Notebook Inference Server Error.",
  "url": "/competitions/jane-street-real-time-market-data-forecasting/discussion/543571",
  "author_name": "Abhi",
  "post_date": "2024-10-31T10:43:00.993000",
  "votes": 1,
  "comment_count": 0,
  "views": 0,
  "content": "<p>I have trained a model from h2o but transforming test data to h2o frame is a problem, anybody has idea about how to efficiently do this, i got a Notebook Inference Server Error when doing submission might be due to that it takes time to convert pandas to H2OFrame or the sys memory got full and the kernel crashed.<br>\nDirectly converting pandas to h2oframe does not work !<br>\n this is my current code:    </p>\n<p>lags_ : pl.DataFrame | None = None</p>\n<p>def predict(test: pl.DataFrame, lags: pl.DataFrame | None) -&gt; pl.DataFrame | pd.DataFrame:</p>\n<pre><code>global lags_\n\nif lags is not :\n\n    lags_ = lags\n\ntest1=test[feature_names]\n\ntest1.().()\n\ntest1=h2o.()\n\npred=model.(test1)\n\npred = pred.().values.()  # Extract predictions as D array\n\npredictions = test.(\n    ,\n    pl.(\n        name   = , \n        values = np.(pred, a_min = -, a_max = ),\n        dtype  = pl.Float64,\n    )\n)\n\nassert (predictions, pl.DataFrame | pd.DataFrame)\nassert predictions.columns == [, ]\nassert (predictions) == (test)\n\nreturn predictions\n</code></pre>",
  "messages": [
    {
      "id": 3032757,
      "postDate": "2024-10-31T10:43:00.993Z",
      "content": "<p>I have trained a model from h2o but transforming test data to h2o frame is a problem, anybody has idea about how to efficiently do this, i got a Notebook Inference Server Error when doing submission might be due to that it takes time to convert pandas to H2OFrame or the sys memory got full and the kernel crashed.<br>\nDirectly converting pandas to h2oframe does not work !<br>\n this is my current code:    </p>\n<p>lags_ : pl.DataFrame | None = None</p>\n<p>def predict(test: pl.DataFrame, lags: pl.DataFrame | None) -&gt; pl.DataFrame | pd.DataFrame:</p>\n<pre><code>global lags_\n\nif lags is not :\n\n    lags_ = lags\n\ntest1=test[feature_names]\n\ntest1.().()\n\ntest1=h2o.()\n\npred=model.(test1)\n\npred = pred.().values.()  # Extract predictions as D array\n\npredictions = test.(\n    ,\n    pl.(\n        name   = , \n        values = np.(pred, a_min = -, a_max = ),\n        dtype  = pl.Float64,\n    )\n)\n\nassert (predictions, pl.DataFrame | pd.DataFrame)\nassert predictions.columns == [, ]\nassert (predictions) == (test)\n\nreturn predictions\n</code></pre>",
      "rawMarkdown": "I have trained a model from h2o but transforming test data to h2o frame is a problem, anybody has idea about how to efficiently do this, i got a Notebook Inference Server Error when doing submission might be due to that it takes time to convert pandas to H2OFrame or the sys memory got full and the kernel crashed.\nDirectly converting pandas to h2oframe does not work !\n this is my current code:    \n\n\n                                                   \nlags_ : pl.DataFrame | None = None\n\n\ndef predict(test: pl.DataFrame, lags: pl.DataFrame | None) -> pl.DataFrame | pd.DataFrame:\n\n    global lags_\n\n    if lags is not None:\n\n        lags_ = lags\n\n    test1=test[feature_names]\n\n    test1.to_pandas().to_csv('./test.csv')\n\n    test1=h2o.import_file('./test.csv')\n\n    pred=model.predict(test1)\n\n    pred = pred.as_data_frame().values.flatten()  # Extract predictions as 1D array\n\n    predictions = test.select(\n        'row_id',\n        pl.Series(\n            name   = 'responder_6', \n            values = np.clip(pred, a_min = -5, a_max = 5),\n            dtype  = pl.Float64,\n        )\n    )\n\n    assert isinstance(predictions, pl.DataFrame | pd.DataFrame)\n    assert predictions.columns == ['row_id', 'responder_6']\n    assert len(predictions) == len(test)\n\n    return predictions",
      "votes": 1
    }
  ],
  "comments": [],
  "raw_markdown_by_id": {
    "3032757": "I have trained a model from h2o but transforming test data to h2o frame is a problem, anybody has idea about how to efficiently do this, i got a Notebook Inference Server Error when doing submission might be due to that it takes time to convert pandas to H2OFrame or the sys memory got full and the kernel crashed.\nDirectly converting pandas to h2oframe does not work !\n this is my current code:    \n\n\n                                                   \nlags_ : pl.DataFrame | None = None\n\n\ndef predict(test: pl.DataFrame, lags: pl.DataFrame | None) -> pl.DataFrame | pd.DataFrame:\n\n    global lags_\n\n    if lags is not None:\n\n        lags_ = lags\n\n    test1=test[feature_names]\n\n    test1.to_pandas().to_csv('./test.csv')\n\n    test1=h2o.import_file('./test.csv')\n\n    pred=model.predict(test1)\n\n    pred = pred.as_data_frame().values.flatten()  # Extract predictions as 1D array\n\n    predictions = test.select(\n        'row_id',\n        pl.Series(\n            name   = 'responder_6', \n            values = np.clip(pred, a_min = -5, a_max = 5),\n            dtype  = pl.Float64,\n        )\n    )\n\n    assert isinstance(predictions, pl.DataFrame | pd.DataFrame)\n    assert predictions.columns == ['row_id', 'responder_6']\n    assert len(predictions) == len(test)\n\n    return predictions"
  }
}