{
  "id": 432511,
  "title": "Tflite float16 quantization not compatiable with LSTM?",
  "url": "/competitions/asl-fingerspelling/discussion/432511",
  "author_name": "",
  "post_date": "2023-08-17T18:42:01.818030600Z",
  "votes": 3,
  "comment_count": 4,
  "views": 0,
  "content": "<p>I'm using the following code to quantize my model:</p>\n<pre><code>keras_model_converter = tf(tflite_keras_model)\nkeras_model_converter = \nkeras_model_converter = \ntflite_model = keras_model_converter()\n</code></pre>\n<p>However, it always fails by exhausting all memory after a few minutes even with a toy model + a LSTM layer:</p>\n<pre><code>def get:\n    inp = tf.keras.layers.)\n\n    x = inp\n    x = tf.keras.layers.(x)\n    x = tf.keras.layers.(x)\n\n    x = tf.keras.layers.(x)\n    x = tf.keras.layers.(x)\n    model = tf.keras.\n    return model\n</code></pre>\n<p>Removing <code>keras_model_converter.target_spec.supported_types = [tf.float16]</code> or removing LSTM resolves this problem.</p>\n<p>Does anybody using LSTM and quantization have this problem and do you solve this?</p>\n<p>Thank you</p>",
  "messages": [
    {
      "id": "2395712",
      "postDate": "08/17/2023 18:42:01",
      "content": "<p>I'm using the following code to quantize my model:</p>\n<pre><code>keras_model_converter = tf(tflite_keras_model)\nkeras_model_converter = \nkeras_model_converter = \ntflite_model = keras_model_converter()\n</code></pre>\n<p>However, it always fails by exhausting all memory after a few minutes even with a toy model + a LSTM layer:</p>\n<pre><code>def get:\n    inp = tf.keras.layers.)\n\n    x = inp\n    x = tf.keras.layers.(x)\n    x = tf.keras.layers.(x)\n\n    x = tf.keras.layers.(x)\n    x = tf.keras.layers.(x)\n    model = tf.keras.\n    return model\n</code></pre>\n<p>Removing <code>keras_model_converter.target_spec.supported_types = [tf.float16]</code> or removing LSTM resolves this problem.</p>\n<p>Does anybody using LSTM and quantization have this problem and do you solve this?</p>\n<p>Thank you</p>",
      "rawMarkdown": "I'm using the following code to quantize my model:\n```\nkeras_model_converter = tf.lite.TFLiteConverter.from_keras_model(tflite_keras_model)\nkeras_model_converter.optimizations = [tf.lite.Optimize.DEFAULT]\nkeras_model_converter.target_spec.supported_types = [tf.float16]\ntflite_model = keras_model_converter.convert()\n```\n\n\n\nHowever, it always fails by exhausting all memory after a few minutes even with a toy model + a LSTM layer:\n\n```\ndef get_model():\n    inp = tf.keras.layers.Input(shape=(TEMPORAL_DIM,INPUT_DIM))\n    \n    x = inp\n    x = tf.keras.layers.Dense(512)(x)\n    x = tf.keras.layers.BatchNormalization()(x)\n\n    x = tf.keras.layers.LSTM(512,return_sequences=True)(x)\n    x = tf.keras.layers.Dense(model_cfg['num_classes'])(x)\n    model = tf.keras.Model(inputs=inp,outputs=x)\n    return model\n```\nRemoving `keras_model_converter.target_spec.supported_types = [tf.float16]` or removing LSTM resolves this problem.\n\nDoes anybody using LSTM and quantization have this problem and do you solve this?\n\nThank you",
      "votes": null
    },
    {
      "id": "2395822",
      "postDate": "08/17/2023 21:06:49",
      "content": "<p>I had exactly same memory issue as you when doing the quantization of LSTM. I tried different running units of LSTM(64, 128, etc…) but none of them work. There seems that Tflite just no support the LSTM layer quantization. Check out this page, maybe there are some ways to resolve the issue but I haven't try it. <a href=\"https://github.com/tensorflow/tensorflow/issues/25563\" target=\"_blank\">https://github.com/tensorflow/tensorflow/issues/25563</a></p>\n<p>When I switched to GRU instead of LSTM, it can be perfectly quantized. However, my LB score drop after adding the GRU layer. If you are going to try it, please let me know. Thanks</p>",
      "rawMarkdown": "I had exactly same memory issue as you when doing the quantization of LSTM. I tried different running units of LSTM(64, 128, etc...) but none of them work. There seems that Tflite just no support the LSTM layer quantization. Check out this page, maybe there are some ways to resolve the issue but I haven't try it. https://github.com/tensorflow/tensorflow/issues/25563\n\nWhen I switched to GRU instead of LSTM, it can be perfectly quantized. However, my LB score drop after adding the GRU layer. If you are going to try it, please let me know. Thanks",
      "votes": null
    },
    {
      "id": "2395862",
      "postDate": "08/17/2023 22:04:05",
      "content": "<p>I also used GRU before but CV score dropped. But if LSTM can't be quantized, I think I would have to try GRU anyway 😥</p>",
      "rawMarkdown": "I also used GRU before but CV score dropped. But if LSTM can't be quantized, I think I would have to try GRU anyway 😥",
      "votes": null
    },
    {
      "id": "2399197",
      "postDate": "08/20/2023 07:45:16",
      "content": "<p>how about writing LSTM from scratch?<br>\nsince we only have one input for tfite in server evaluation, speed may not be an issue</p>",
      "rawMarkdown": "how about writing LSTM from scratch?\nsince we only have one input for tfite in server evaluation, speed may not be an issue",
      "votes": null
    },
    {
      "id": "2399253",
      "postDate": "08/20/2023 08:21:15",
      "content": "<p>I guess it should work.  But as GRU's performance is acceptable for me and I'm still working on masking…<br>\nso I probably won't have time to do it </p>",
      "rawMarkdown": "I guess it should work.  But as GRU's performance is acceptable for me and I'm still working on masking...\nso I probably won't have time to do it",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 2395822,
      "author_name": "wuhongrui",
      "author_url": "",
      "post_date": "08/17/2023 21:06:49",
      "content": "<p>I had exactly same memory issue as you when doing the quantization of LSTM. I tried different running units of LSTM(64, 128, etc…) but none of them work. There seems that Tflite just no support the LSTM layer quantization. Check out this page, maybe there are some ways to resolve the issue but I haven't try it. <a href=\"https://github.com/tensorflow/tensorflow/issues/25563\" target=\"_blank\">https://github.com/tensorflow/tensorflow/issues/25563</a></p>\n<p>When I switched to GRU instead of LSTM, it can be perfectly quantized. However, my LB score drop after adding the GRU layer. If you are going to try it, please let me know. Thanks</p>",
      "votes": null,
      "replies": [
        {
          "id": 2395862,
          "author_name": "nightsh4de",
          "author_url": "",
          "post_date": "08/17/2023 22:04:05",
          "content": "<p>I also used GRU before but CV score dropped. But if LSTM can't be quantized, I think I would have to try GRU anyway 😥</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 2399197,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "08/20/2023 07:45:16",
      "content": "<p>how about writing LSTM from scratch?<br>\nsince we only have one input for tfite in server evaluation, speed may not be an issue</p>",
      "votes": null,
      "replies": [
        {
          "id": 2399253,
          "author_name": "nightsh4de",
          "author_url": "",
          "post_date": "08/20/2023 08:21:15",
          "content": "<p>I guess it should work.  But as GRU's performance is acceptable for me and I'm still working on masking…<br>\nso I probably won't have time to do it </p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "2395712": "I'm using the following code to quantize my model:\n```\nkeras_model_converter = tf.lite.TFLiteConverter.from_keras_model(tflite_keras_model)\nkeras_model_converter.optimizations = [tf.lite.Optimize.DEFAULT]\nkeras_model_converter.target_spec.supported_types = [tf.float16]\ntflite_model = keras_model_converter.convert()\n```\n\n\n\nHowever, it always fails by exhausting all memory after a few minutes even with a toy model + a LSTM layer:\n\n```\ndef get_model():\n    inp = tf.keras.layers.Input(shape=(TEMPORAL_DIM,INPUT_DIM))\n    \n    x = inp\n    x = tf.keras.layers.Dense(512)(x)\n    x = tf.keras.layers.BatchNormalization()(x)\n\n    x = tf.keras.layers.LSTM(512,return_sequences=True)(x)\n    x = tf.keras.layers.Dense(model_cfg['num_classes'])(x)\n    model = tf.keras.Model(inputs=inp,outputs=x)\n    return model\n```\nRemoving `keras_model_converter.target_spec.supported_types = [tf.float16]` or removing LSTM resolves this problem.\n\nDoes anybody using LSTM and quantization have this problem and do you solve this?\n\nThank you",
    "2395822": "I had exactly same memory issue as you when doing the quantization of LSTM. I tried different running units of LSTM(64, 128, etc...) but none of them work. There seems that Tflite just no support the LSTM layer quantization. Check out this page, maybe there are some ways to resolve the issue but I haven't try it. https://github.com/tensorflow/tensorflow/issues/25563\n\nWhen I switched to GRU instead of LSTM, it can be perfectly quantized. However, my LB score drop after adding the GRU layer. If you are going to try it, please let me know. Thanks",
    "2395862": "I also used GRU before but CV score dropped. But if LSTM can't be quantized, I think I would have to try GRU anyway 😥",
    "2399197": "how about writing LSTM from scratch?\nsince we only have one input for tfite in server evaluation, speed may not be an issue",
    "2399253": "I guess it should work.  But as GRU's performance is acceptable for me and I'm still working on masking...\nso I probably won't have time to do it"
  },
  "source": "meta"
}