{
  "id": 553291,
  "title": "Is Tensorflow.keras inference speed very slow?",
  "url": "/competitions/jane-street-real-time-market-data-forecasting/discussion/553291",
  "author_name": "",
  "post_date": "2024-12-25T04:20:08.163754Z",
  "votes": 6,
  "comment_count": 11,
  "views": 0,
  "content": "<p>Hi,I am just starting this competition,I found Tensorflow.keras inference is very slow, even I run a mlp model without feature engineering it takes 3.5 hours, if ensemble two models, it takes about 6 hours.</p>\n<p>The kaggle notebook's tensorflow version is 2.17.0, I tried two versions for my local(2.13.1, 2.17.0) training, no too much difference,local inference seems very fast, 5 seconds with 4.5 million rows,but I don't know why online inference is so slow.  </p>",
  "messages": [
    {
      "id": "3080339",
      "postDate": "12/25/2024 04:20:08",
      "content": "<p>Hi,I am just starting this competition,I found Tensorflow.keras inference is very slow, even I run a mlp model without feature engineering it takes 3.5 hours, if ensemble two models, it takes about 6 hours.</p>\n<p>The kaggle notebook's tensorflow version is 2.17.0, I tried two versions for my local(2.13.1, 2.17.0) training, no too much difference,local inference seems very fast, 5 seconds with 4.5 million rows,but I don't know why online inference is so slow.  </p>",
      "rawMarkdown": "Hi,I am just starting this competition,I found Tensorflow.keras inference is very slow, even I run a mlp model without feature engineering it takes 3.5 hours, if ensemble two models, it takes about 6 hours.\n\nThe kaggle notebook's tensorflow version is 2.17.0, I tried two versions for my local(2.13.1, 2.17.0) training, no too much difference,local inference seems very fast, 5 seconds with 4.5 million rows,but I don't know why online inference is so slow.",
      "votes": null
    },
    {
      "id": "3080369",
      "postDate": "12/25/2024 05:53:10",
      "content": "<p>I have 5 LightGBM models with some feature engineering and it takes 1.5 hours for me. 3.5 hours for a single MLP sounds too much. Is your MLP too wide? How many parameters does it have?</p>",
      "rawMarkdown": "I have 5 LightGBM models with some feature engineering and it takes 1.5 hours for me. 3.5 hours for a single MLP sounds too much. Is your MLP too wide? How many parameters does it have?",
      "votes": null
    },
    {
      "id": "3080376",
      "postDate": "12/25/2024 06:18:44",
      "content": "<p>I suspect that there might be an issue with the server. Here's the situation: I submitted two notebooks that are nearly identical. The first one took approximately 2 hours to complete the inference process yesterday. However, it has been 6 hours since I submitted the second notebook, and it still hasn't finished. This significant time difference leads me to believe that something is amiss with the server.</p>",
      "rawMarkdown": "I suspect that there might be an issue with the server. Here's the situation: I submitted two notebooks that are nearly identical. The first one took approximately 2 hours to complete the inference process yesterday. However, it has been 6 hours since I submitted the second notebook, and it still hasn't finished. This significant time difference leads me to believe that something is amiss with the server.",
      "votes": null
    },
    {
      "id": "3080632",
      "postDate": "12/25/2024 13:47:31",
      "content": "<p>yes, I think it could happen if kaggle cloud resource allocation was busy,so many competitions right now. </p>",
      "rawMarkdown": "yes, I think it could happen if kaggle cloud resource allocation was busy,so many competitions right now.",
      "votes": null
    },
    {
      "id": "3080640",
      "postDate": "12/25/2024 13:55:04",
      "content": "<p>good to know LightGBM only takes 1.5h, mlp is not wide, only 50k params</p>",
      "rawMarkdown": "good to know LightGBM only takes 1.5h, mlp is not wide, only 50k params",
      "votes": null
    },
    {
      "id": "3080652",
      "postDate": "12/25/2024 14:15:17",
      "content": "<p>I am nearly certain that the issue lies with the GPU. In my notebook, merely a day ago, the GPU was able to complete 968 rounds of inferencing in just 17 seconds. However, presently, it is taking an astounding 6 minutes and 30 seconds. But it is rather strange that, aside from this topic, there has been no other discussion regarding this particular problem.</p>",
      "rawMarkdown": "I am nearly certain that the issue lies with the GPU. In my notebook, merely a day ago, the GPU was able to complete 968 rounds of inferencing in just 17 seconds. However, presently, it is taking an astounding 6 minutes and 30 seconds. But it is rather strange that, aside from this topic, there has been no other discussion regarding this particular problem.",
      "votes": null
    },
    {
      "id": "3080765",
      "postDate": "12/25/2024 17:13:21",
      "content": "<p>I've also faced the same issue. Are there any updates on how to address it?</p>",
      "rawMarkdown": "I've also faced the same issue. Are there any updates on how to address it?",
      "votes": null
    },
    {
      "id": "3080834",
      "postDate": "12/25/2024 19:54:47",
      "content": "<p>Not able to get submission results at all. Wasted 8 subs during today and yesterday. Successfully inference-run notebook with no change takes Notebook Timeout error. Christmas workload.</p>",
      "rawMarkdown": "Not able to get submission results at all. Wasted 8 subs during today and yesterday. Successfully inference-run notebook with no change takes Notebook Timeout error. Christmas workload.",
      "votes": null
    },
    {
      "id": "3080854",
      "postDate": "12/25/2024 20:51:08",
      "content": "<p>Same here! </p>",
      "rawMarkdown": "Same here!",
      "votes": null
    },
    {
      "id": "3080965",
      "postDate": "12/26/2024 02:47:52",
      "content": "<p>You are right! There is no discussion around this. I wonder why! A submission that took around 2 hours two weeks ago now takes 8 hours for me, and then I get a timeout error. </p>",
      "rawMarkdown": "You are right! There is no discussion around this. I wonder why! A submission that took around 2 hours two weeks ago now takes 8 hours for me, and then I get a timeout error.",
      "votes": null
    },
    {
      "id": "3081341",
      "postDate": "12/26/2024 16:15:02",
      "content": "<p>You can try to:<br>\n1)  Call model without built-in .predict function: model(data, training=False)<br>\n2) Cast input data explicitly to tf tensors right before calling model: tf.convert_to_tensor(numpy_array, dtype=tf.float32)<br>\n3) Use tf.dataset api to load data: dataset = tf.data.Dataset.from_tensor_slices(numpy_array).batch(X)<br>\n                                                         model.predict(dataset) or model(dataset, training=False)</p>",
      "rawMarkdown": "You can try to:\n1)  Call model without built-in .predict function: model(data, training=False)\n2) Cast input data explicitly to tf tensors right before calling model: tf.convert_to_tensor(numpy_array, dtype=tf.float32)\n3) Use tf.dataset api to load data: dataset = tf.data.Dataset.from_tensor_slices(numpy_array).batch(X)\n                                                         model.predict(dataset) or model(dataset, training=False)",
      "votes": null
    },
    {
      "id": "3081587",
      "postDate": "12/27/2024 00:51:16",
      "content": "<p>The problem has been resolved! And I think it's all due to the workload during Christmas. LOL.</p>",
      "rawMarkdown": "The problem has been resolved! And I think it's all due to the workload during Christmas. LOL.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 3080369,
      "author_name": "gunesevitan",
      "author_url": "",
      "post_date": "12/25/2024 05:53:10",
      "content": "<p>I have 5 LightGBM models with some feature engineering and it takes 1.5 hours for me. 3.5 hours for a single MLP sounds too much. Is your MLP too wide? How many parameters does it have?</p>",
      "votes": null,
      "replies": [
        {
          "id": 3080640,
          "author_name": "senkin13",
          "author_url": "",
          "post_date": "12/25/2024 13:55:04",
          "content": "<p>good to know LightGBM only takes 1.5h, mlp is not wide, only 50k params</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 3080376,
      "author_name": "carrotwait",
      "author_url": "",
      "post_date": "12/25/2024 06:18:44",
      "content": "<p>I suspect that there might be an issue with the server. Here's the situation: I submitted two notebooks that are nearly identical. The first one took approximately 2 hours to complete the inference process yesterday. However, it has been 6 hours since I submitted the second notebook, and it still hasn't finished. This significant time difference leads me to believe that something is amiss with the server.</p>",
      "votes": null,
      "replies": [
        {
          "id": 3080632,
          "author_name": "senkin13",
          "author_url": "",
          "post_date": "12/25/2024 13:47:31",
          "content": "<p>yes, I think it could happen if kaggle cloud resource allocation was busy,so many competitions right now. </p>",
          "votes": null,
          "replies": [
            {
              "id": 3080652,
              "author_name": "carrotwait",
              "author_url": "",
              "post_date": "12/25/2024 14:15:17",
              "content": "<p>I am nearly certain that the issue lies with the GPU. In my notebook, merely a day ago, the GPU was able to complete 968 rounds of inferencing in just 17 seconds. However, presently, it is taking an astounding 6 minutes and 30 seconds. But it is rather strange that, aside from this topic, there has been no other discussion regarding this particular problem.</p>",
              "votes": null,
              "replies": [
                {
                  "id": 3080965,
                  "author_name": "arashab",
                  "author_url": "",
                  "post_date": "12/26/2024 02:47:52",
                  "content": "<p>You are right! There is no discussion around this. I wonder why! A submission that took around 2 hours two weeks ago now takes 8 hours for me, and then I get a timeout error. </p>",
                  "votes": null,
                  "replies": []
                }
              ]
            }
          ]
        },
        {
          "id": 3080765,
          "author_name": "arashab",
          "author_url": "",
          "post_date": "12/25/2024 17:13:21",
          "content": "<p>I've also faced the same issue. Are there any updates on how to address it?</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 3080834,
      "author_name": "yekenot",
      "author_url": "",
      "post_date": "12/25/2024 19:54:47",
      "content": "<p>Not able to get submission results at all. Wasted 8 subs during today and yesterday. Successfully inference-run notebook with no change takes Notebook Timeout error. Christmas workload.</p>",
      "votes": null,
      "replies": [
        {
          "id": 3080854,
          "author_name": "arashab",
          "author_url": "",
          "post_date": "12/25/2024 20:51:08",
          "content": "<p>Same here! </p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 3081341,
      "author_name": "adrianwiniewski",
      "author_url": "",
      "post_date": "12/26/2024 16:15:02",
      "content": "<p>You can try to:<br>\n1)  Call model without built-in .predict function: model(data, training=False)<br>\n2) Cast input data explicitly to tf tensors right before calling model: tf.convert_to_tensor(numpy_array, dtype=tf.float32)<br>\n3) Use tf.dataset api to load data: dataset = tf.data.Dataset.from_tensor_slices(numpy_array).batch(X)<br>\n                                                         model.predict(dataset) or model(dataset, training=False)</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 3081587,
      "author_name": "carrotwait",
      "author_url": "",
      "post_date": "12/27/2024 00:51:16",
      "content": "<p>The problem has been resolved! And I think it's all due to the workload during Christmas. LOL.</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "3080339": "Hi,I am just starting this competition,I found Tensorflow.keras inference is very slow, even I run a mlp model without feature engineering it takes 3.5 hours, if ensemble two models, it takes about 6 hours.\n\nThe kaggle notebook's tensorflow version is 2.17.0, I tried two versions for my local(2.13.1, 2.17.0) training, no too much difference,local inference seems very fast, 5 seconds with 4.5 million rows,but I don't know why online inference is so slow.",
    "3080369": "I have 5 LightGBM models with some feature engineering and it takes 1.5 hours for me. 3.5 hours for a single MLP sounds too much. Is your MLP too wide? How many parameters does it have?",
    "3080376": "I suspect that there might be an issue with the server. Here's the situation: I submitted two notebooks that are nearly identical. The first one took approximately 2 hours to complete the inference process yesterday. However, it has been 6 hours since I submitted the second notebook, and it still hasn't finished. This significant time difference leads me to believe that something is amiss with the server.",
    "3080632": "yes, I think it could happen if kaggle cloud resource allocation was busy,so many competitions right now.",
    "3080640": "good to know LightGBM only takes 1.5h, mlp is not wide, only 50k params",
    "3080652": "I am nearly certain that the issue lies with the GPU. In my notebook, merely a day ago, the GPU was able to complete 968 rounds of inferencing in just 17 seconds. However, presently, it is taking an astounding 6 minutes and 30 seconds. But it is rather strange that, aside from this topic, there has been no other discussion regarding this particular problem.",
    "3080765": "I've also faced the same issue. Are there any updates on how to address it?",
    "3080834": "Not able to get submission results at all. Wasted 8 subs during today and yesterday. Successfully inference-run notebook with no change takes Notebook Timeout error. Christmas workload.",
    "3080854": "Same here!",
    "3080965": "You are right! There is no discussion around this. I wonder why! A submission that took around 2 hours two weeks ago now takes 8 hours for me, and then I get a timeout error.",
    "3081341": "You can try to:\n1)  Call model without built-in .predict function: model(data, training=False)\n2) Cast input data explicitly to tf tensors right before calling model: tf.convert_to_tensor(numpy_array, dtype=tf.float32)\n3) Use tf.dataset api to load data: dataset = tf.data.Dataset.from_tensor_slices(numpy_array).batch(X)\n                                                         model.predict(dataset) or model(dataset, training=False)",
    "3081587": "The problem has been resolved! And I think it's all due to the workload during Christmas. LOL."
  },
  "source": "meta"
}