{
  "id": 551301,
  "title": "Can the ragged tensor approach lead to OOM (RAM)?",
  "url": "/competitions/jane-street-real-time-market-data-forecasting/discussion/551301",
  "author_name": "",
  "post_date": "2024-12-12T13:00:14.380196600Z",
  "votes": 1,
  "comment_count": 5,
  "views": 0,
  "content": "<p>I have known about ragged tensors for some time, but this competition is the first time I have tried to apply this approach.</p>\n<p>Every time I tried to submit, it returned a Notebook Inference Server Error. So, I downloaded the offline API notebook to test, and it gave me a RAM OOM error (not VRAM). It works fine on my personal computer, where I have 64GB of RAM.<br>\nSince it's a loop, RAM shouldn't increase if it doesn't cache/append features while looping.</p>\n<p>I tried using the garbage collector and deleting tests in each loop, but it keeps giving me OOM errors. My question is, could it be some kind of caching of ragged tensors or something like that?</p>\n<p>Tf and Keras version:</p>\n<pre><code>(tf.__version__)\n(keras.__version__)\n\n\n</code></pre>\n<p>Network Architecture: Some MultiHead layers, layer normalization, nothing new.</p>\n<p>My prediction code:</p>\n<pre><code>lags_ : pl.DataFrame |  = \n gc\n\n\n () -&gt; pl.DataFrame | pd.DataFrame:\n\n\n    X_test = test[features].fill_null(-).to_numpy()\n\n\n\n\n\n\n    \n    y_pred = np.squeeze(model.predict(np.expand_dims(X_test, )))\n    \n    predictions = test.select().with_columns(\n        pl.Series(, y_pred.flatten())\n    )\n    \n     (predictions, pl.DataFrame | pd.DataFrame)\n    \n     predictions.columns == [, ]\n    \n     (predictions) == (test)\n    gc.collect()\n     X_test, test\n     predictions\n</code></pre>",
  "messages": [
    {
      "id": "3070230",
      "postDate": "12/12/2024 13:00:14",
      "content": "<p>I have known about ragged tensors for some time, but this competition is the first time I have tried to apply this approach.</p>\n<p>Every time I tried to submit, it returned a Notebook Inference Server Error. So, I downloaded the offline API notebook to test, and it gave me a RAM OOM error (not VRAM). It works fine on my personal computer, where I have 64GB of RAM.<br>\nSince it's a loop, RAM shouldn't increase if it doesn't cache/append features while looping.</p>\n<p>I tried using the garbage collector and deleting tests in each loop, but it keeps giving me OOM errors. My question is, could it be some kind of caching of ragged tensors or something like that?</p>\n<p>Tf and Keras version:</p>\n<pre><code>(tf.__version__)\n(keras.__version__)\n\n\n</code></pre>\n<p>Network Architecture: Some MultiHead layers, layer normalization, nothing new.</p>\n<p>My prediction code:</p>\n<pre><code>lags_ : pl.DataFrame |  = \n gc\n\n\n () -&gt; pl.DataFrame | pd.DataFrame:\n\n\n    X_test = test[features].fill_null(-).to_numpy()\n\n\n\n\n\n\n    \n    y_pred = np.squeeze(model.predict(np.expand_dims(X_test, )))\n    \n    predictions = test.select().with_columns(\n        pl.Series(, y_pred.flatten())\n    )\n    \n     (predictions, pl.DataFrame | pd.DataFrame)\n    \n     predictions.columns == [, ]\n    \n     (predictions) == (test)\n    gc.collect()\n     X_test, test\n     predictions\n</code></pre>",
      "rawMarkdown": "I have known about ragged tensors for some time, but this competition is the first time I have tried to apply this approach.\n\nEvery time I tried to submit, it returned a Notebook Inference Server Error. So, I downloaded the offline API notebook to test, and it gave me a RAM OOM error (not VRAM). It works fine on my personal computer, where I have 64GB of RAM.\nSince it's a loop, RAM shouldn't increase if it doesn't cache/append features while looping.\n\nI tried using the garbage collector and deleting tests in each loop, but it keeps giving me OOM errors. My question is, could it be some kind of caching of ragged tensors or something like that?\n\nTf and Keras version:\n```python\nprint(tf.__version__)\nprint(keras.__version__)\n2.13.0\n2.13.1\n```\n\nNetwork Architecture: Some MultiHead layers, layer normalization, nothing new.\n\n\nMy prediction code:\n```\nlags_ : pl.DataFrame | None = None\nimport gc\n\n\ndef predict(test: pl.DataFrame, lags: pl.DataFrame | None) -> pl.DataFrame | pd.DataFrame:\n\n\n    X_test = test[features].fill_null(-1).to_numpy()\n\n\n    \n\n    \n    \n    # 2. Make predictions using the Keras model\n    y_pred = np.squeeze(model.predict(np.expand_dims(X_test, 0)))\n    # 3. Prepare the DataFrame for output\n    predictions = test.select('row_id').with_columns(\n        pl.Series(\"responder_6\", y_pred.flatten())\n    )\n    # The predict function must return a DataFrame\n    assert isinstance(predictions, pl.DataFrame | pd.DataFrame)\n    # with columns 'row_id', 'responer_6'\n    assert predictions.columns == ['row_id', 'responder_6']\n    # and as many rows as the test data.\n    assert len(predictions) == len(test)\n    gc.collect()\n    del X_test, test\n    return predictions\n```",
      "votes": null
    },
    {
      "id": "3070534",
      "postDate": "12/12/2024 19:59:53",
      "content": "<p>Have you tried setting a batch size for your model's prediction method?</p>\n<p>I think each time_id in the prediction loop has more than 100k rows.</p>",
      "rawMarkdown": "Have you tried setting a batch size for your model's prediction method?\n\nI think each time_id in the prediction loop has more than 100k rows.",
      "votes": null
    },
    {
      "id": "3071206",
      "postDate": "12/13/2024 13:36:54",
      "content": "<p>I think each time_id has symbol_ids number of rows, normally from 8 to almost 40.</p>",
      "rawMarkdown": "I think each time_id has symbol_ids number of rows, normally from 8 to almost 40.",
      "votes": null
    },
    {
      "id": "3071569",
      "postDate": "12/14/2024 00:08:32",
      "content": "<p>My experience with the API: the Notebook Inference Server Error usually comes from timeout. There is a 60 sec request timeout limit. Plus on average it's better to reach under 0.15s per request (at least from the previous estimate), otherwise hitting the 8 hour total limit. Is gc.collect() really needed here for each time step? It can be quite slow. </p>",
      "rawMarkdown": "My experience with the API: the Notebook Inference Server Error usually comes from timeout. There is a 60 sec request timeout limit. Plus on average it's better to reach under 0.15s per request (at least from the previous estimate), otherwise hitting the 8 hour total limit. Is gc.collect() really needed here for each time step? It can be quite slow.",
      "votes": null
    },
    {
      "id": "3072706",
      "postDate": "12/15/2024 15:23:00",
      "content": "<p>I tried using the offline API, and it gave an OOM error.</p>\n<p>Yesterday, I ran a test where I removed everything from the predict function and filled y_pred with a dummy variable (lit(0)), without performing any inference from the neural network.<br>\nHowever, I still encountered an error.<br>\nI think there might be a bug because I am changing the TensorFlow and Keras versions.</p>",
      "rawMarkdown": "I tried using the offline API, and it gave an OOM error.\n\nYesterday, I ran a test where I removed everything from the predict function and filled y_pred with a dummy variable (lit(0)), without performing any inference from the neural network.\nHowever, I still encountered an error.\nI think there might be a bug because I am changing the TensorFlow and Keras versions.",
      "votes": null
    },
    {
      "id": "3073387",
      "postDate": "12/16/2024 13:10:55",
      "content": "<p>I think my notebook build is broken.</p>\n<p>Yesterday, in the same notebook that gives me OOM (Out of Memory) or Timeout errors, I removed the \"install dependencies\" section, didn’t load any packages, and simply ran this code:<br>\n<code>predictions = test.select(\n        'row_id',\n        pl.lit(0.0).alias('responder_6'),\n    )</code></p>\n<p>And it still gives me a Timeout error now.<br>\nI'll try using a new notebook.</p>",
      "rawMarkdown": "I think my notebook build is broken.\n\nYesterday, in the same notebook that gives me OOM (Out of Memory) or Timeout errors, I removed the \"install dependencies\" section, didn’t load any packages, and simply ran this code:\n`predictions = test.select(\n        'row_id',\n        pl.lit(0.0).alias('responder_6'),\n    )`\n\nAnd it still gives me a Timeout error now.\nI'll try using a new notebook.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 3070534,
      "author_name": "serjhenrique",
      "author_url": "",
      "post_date": "12/12/2024 19:59:53",
      "content": "<p>Have you tried setting a batch size for your model's prediction method?</p>\n<p>I think each time_id in the prediction loop has more than 100k rows.</p>",
      "votes": null,
      "replies": [
        {
          "id": 3071206,
          "author_name": "nandodmelo",
          "author_url": "",
          "post_date": "12/13/2024 13:36:54",
          "content": "<p>I think each time_id has symbol_ids number of rows, normally from 8 to almost 40.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 3071569,
      "author_name": "benlai",
      "author_url": "",
      "post_date": "12/14/2024 00:08:32",
      "content": "<p>My experience with the API: the Notebook Inference Server Error usually comes from timeout. There is a 60 sec request timeout limit. Plus on average it's better to reach under 0.15s per request (at least from the previous estimate), otherwise hitting the 8 hour total limit. Is gc.collect() really needed here for each time step? It can be quite slow. </p>",
      "votes": null,
      "replies": [
        {
          "id": 3072706,
          "author_name": "nandodmelo",
          "author_url": "",
          "post_date": "12/15/2024 15:23:00",
          "content": "<p>I tried using the offline API, and it gave an OOM error.</p>\n<p>Yesterday, I ran a test where I removed everything from the predict function and filled y_pred with a dummy variable (lit(0)), without performing any inference from the neural network.<br>\nHowever, I still encountered an error.<br>\nI think there might be a bug because I am changing the TensorFlow and Keras versions.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 3073387,
      "author_name": "nandodmelo",
      "author_url": "",
      "post_date": "12/16/2024 13:10:55",
      "content": "<p>I think my notebook build is broken.</p>\n<p>Yesterday, in the same notebook that gives me OOM (Out of Memory) or Timeout errors, I removed the \"install dependencies\" section, didn’t load any packages, and simply ran this code:<br>\n<code>predictions = test.select(\n        'row_id',\n        pl.lit(0.0).alias('responder_6'),\n    )</code></p>\n<p>And it still gives me a Timeout error now.<br>\nI'll try using a new notebook.</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "3070230": "I have known about ragged tensors for some time, but this competition is the first time I have tried to apply this approach.\n\nEvery time I tried to submit, it returned a Notebook Inference Server Error. So, I downloaded the offline API notebook to test, and it gave me a RAM OOM error (not VRAM). It works fine on my personal computer, where I have 64GB of RAM.\nSince it's a loop, RAM shouldn't increase if it doesn't cache/append features while looping.\n\nI tried using the garbage collector and deleting tests in each loop, but it keeps giving me OOM errors. My question is, could it be some kind of caching of ragged tensors or something like that?\n\nTf and Keras version:\n```python\nprint(tf.__version__)\nprint(keras.__version__)\n2.13.0\n2.13.1\n```\n\nNetwork Architecture: Some MultiHead layers, layer normalization, nothing new.\n\n\nMy prediction code:\n```\nlags_ : pl.DataFrame | None = None\nimport gc\n\n\ndef predict(test: pl.DataFrame, lags: pl.DataFrame | None) -> pl.DataFrame | pd.DataFrame:\n\n\n    X_test = test[features].fill_null(-1).to_numpy()\n\n\n    \n\n    \n    \n    # 2. Make predictions using the Keras model\n    y_pred = np.squeeze(model.predict(np.expand_dims(X_test, 0)))\n    # 3. Prepare the DataFrame for output\n    predictions = test.select('row_id').with_columns(\n        pl.Series(\"responder_6\", y_pred.flatten())\n    )\n    # The predict function must return a DataFrame\n    assert isinstance(predictions, pl.DataFrame | pd.DataFrame)\n    # with columns 'row_id', 'responer_6'\n    assert predictions.columns == ['row_id', 'responder_6']\n    # and as many rows as the test data.\n    assert len(predictions) == len(test)\n    gc.collect()\n    del X_test, test\n    return predictions\n```",
    "3070534": "Have you tried setting a batch size for your model's prediction method?\n\nI think each time_id in the prediction loop has more than 100k rows.",
    "3071206": "I think each time_id has symbol_ids number of rows, normally from 8 to almost 40.",
    "3071569": "My experience with the API: the Notebook Inference Server Error usually comes from timeout. There is a 60 sec request timeout limit. Plus on average it's better to reach under 0.15s per request (at least from the previous estimate), otherwise hitting the 8 hour total limit. Is gc.collect() really needed here for each time step? It can be quite slow.",
    "3072706": "I tried using the offline API, and it gave an OOM error.\n\nYesterday, I ran a test where I removed everything from the predict function and filled y_pred with a dummy variable (lit(0)), without performing any inference from the neural network.\nHowever, I still encountered an error.\nI think there might be a bug because I am changing the TensorFlow and Keras versions.",
    "3073387": "I think my notebook build is broken.\n\nYesterday, in the same notebook that gives me OOM (Out of Memory) or Timeout errors, I removed the \"install dependencies\" section, didn’t load any packages, and simply ran this code:\n`predictions = test.select(\n        'row_id',\n        pl.lit(0.0).alias('responder_6'),\n    )`\n\nAnd it still gives me a Timeout error now.\nI'll try using a new notebook."
  },
  "source": "meta"
}