{
  "id": 395425,
  "title": "Quantization leads to time out? ",
  "url": "/competitions/asl-signs/discussion/395425",
  "author_name": "hyd",
  "post_date": "2023-03-17T08:31:14.155000",
  "votes": 4,
  "comment_count": 14,
  "views": 0,
  "content": "<p>The former submision is alright, but got time out after added this line:<br>\n<code>converter.optimizations = [tf.lite.Optimize.DEFAULT]</code><br>\n😧</p>",
  "messages": [
    {
      "id": 2186271,
      "postDate": "2023-03-17T17:15:58.177Z",
      "content": "<p>tflite quantized models are optimized for ARM devices. For x86 CPU's, FP32 tflite model is faster than INT8 quantized tflite.</p>\n<p><a href=\"https://github.com/tensorflow/tensorflow/issues/21698#issuecomment-414764709\" target=\"_blank\">https://github.com/tensorflow/tensorflow/issues/21698#issuecomment-414764709</a></p>",
      "rawMarkdown": "tflite quantized models are optimized for ARM devices. For x86 CPU's, FP32 tflite model is faster than INT8 quantized tflite.\n\nhttps://github.com/tensorflow/tensorflow/issues/21698#issuecomment-414764709",
      "votes": 9
    },
    {
      "id": 2199612,
      "postDate": "2023-03-27T21:11:58.913Z",
      "content": "<p>float16 quantization tends to works well with no speed or accuracy loss.</p>\n<blockquote>\n  <p>keras_model_converter.optimizations = [tf.lite.Optimize.DEFAULT]<br>\n  keras_model_converter.target_spec.supported_types = [tf.float16]<br>\n  tflite_model = keras_model_converter.convert() </p>\n</blockquote>",
      "rawMarkdown": "float16 quantization tends to works well with no speed or accuracy loss.\n>keras_model_converter.optimizations = [tf.lite.Optimize.DEFAULT]\nkeras_model_converter.target_spec.supported_types = [tf.float16]\ntflite_model = keras_model_converter.convert() ",
      "votes": 6,
      "replies": [
        {
          "id": 2199648,
          "postDate": "2023-03-27T22:41:41.800Z",
          "content": "<p>for me it drops performance by ~10% and accuracy by 1-3%</p>\n<p>Even with converting dataset. But maybe it will help someone</p>\n<pre><code>num_calibration_steps = \n ():\n   i  (num_calibration_steps):\n     [load_relevant_data_subset(BASE_DIR + train_df.loc[i, ])]\n\nkeras_model_converter.optimizations = [tf.lite.Optimize.DEFAULT]\nkeras_model_converter.representative_dataset = representative_dataset_gen\nkeras_model_converter.target_spec.supported_types = [tf.float16]\ntflite_model = keras_model_converter.convert()\n</code></pre>",
          "rawMarkdown": "for me it drops performance by ~10% and accuracy by 1-3%\n\nEven with converting dataset. But maybe it will help someone\n\n```python\n\nnum_calibration_steps = 1000\ndef representative_dataset_gen():\n  for i in range(num_calibration_steps):\n    yield [load_relevant_data_subset(BASE_DIR + train_df.loc[i, 'path'])]\n\nkeras_model_converter.optimizations = [tf.lite.Optimize.DEFAULT]\nkeras_model_converter.representative_dataset = representative_dataset_gen\nkeras_model_converter.target_spec.supported_types = [tf.float16]\ntflite_model = keras_model_converter.convert()\n``` ",
          "votes": 2,
          "replies": [
            {
              "id": 2199699,
              "postDate": "2023-03-28T00:31:50.793Z",
              "content": "<p>I don't think you need any calibration for fp16, for typical networks, model weight range should be well within the fp16 range.<br>\nfp16 conversion has worked well for past competitions in kaggle and during my work too (using pytorch), and you shouldn't need any calibrations.<br>\nWhat will happen if you exclude <code>keras_model_converter.representative_dataset = representative_dataset_gen</code>?</p>",
              "rawMarkdown": "I don't think you need any calibration for fp16, for typical networks, model weight range should be well within the fp16 range.\nfp16 conversion has worked well for past competitions in kaggle and during my work too (using pytorch), and you shouldn't need any calibrations.\nWhat will happen if you exclude `keras_model_converter.representative_dataset = representative_dataset_gen`?",
              "votes": 1
            },
            {
              "id": 2199705,
              "postDate": "2023-03-28T00:43:30.673Z",
              "content": "<p>With my model and fp16 - absolutely nothing. It helps me a little with int8, but integer quantization performs much worse anyway (as <a href=\"https://www.kaggle.com/samfc10\" target=\"_blank\">@samfc10</a> have mentioned below).</p>\n<p>I just wanted to leave it here because maybe it will improve someone's model that is okay with quantization.</p>",
              "rawMarkdown": "With my model and fp16 - absolutely nothing. It helps me a little with int8, but integer quantization performs much worse anyway (as @samfc10 have mentioned below).\n\nI just wanted to leave it here because maybe it will improve someone's model that is okay with quantization."
            }
          ]
        },
        {
          "id": 2200439,
          "postDate": "2023-03-28T14:58:28.117Z",
          "rawMarkdown": "",
          "isDeleted": true
        },
        {
          "id": 2218419,
          "postDate": "2023-04-11T16:38:52.007Z",
          "content": "<p>In my case, fp16 quantization causes severe accuracy drop (CV: 0.77955 -&gt; 0.00421).<br>\nOn the other hand, int8 quantization works fine (CV: 0.77955 -&gt; 0.77948).</p>\n<pre><code>    converter = tf.lite.TFLiteConverter.from_keras_model(model)\n     QUANTIZE  [, ]:\n        converter.optimizations = [tf.lite.Optimize.DEFAULT]\n         QUANTIZE == :\n            converter.target_spec.supported_types = [tf.float16]\n    tflite_model = converter.convert()\n</code></pre>",
          "rawMarkdown": "In my case, fp16 quantization causes severe accuracy drop (CV: 0.77955 -> 0.00421).\nOn the other hand, int8 quantization works fine (CV: 0.77955 -> 0.77948).\n```python\n    converter = tf.lite.TFLiteConverter.from_keras_model(model)\n    if QUANTIZE in [\"int8\", \"fp16\"]:\n        converter.optimizations = [tf.lite.Optimize.DEFAULT]\n        if QUANTIZE == \"fp16\":\n            converter.target_spec.supported_types = [tf.float16]\n    tflite_model = converter.convert()\n``` ",
          "replies": [
            {
              "id": 2218744,
              "postDate": "2023-04-12T01:47:54.147Z",
              "content": "<p>Thank you for sharing!<br>\nI wonder how did you covert fp32 into int8?</p>",
              "rawMarkdown": "Thank you for sharing!\nI wonder how did you covert fp32 into int8?"
            },
            {
              "id": 2218752,
              "postDate": "2023-04-12T02:01:05.123Z",
              "content": "<p><a href=\"https://www.kaggle.com/heoynwoo\" target=\"_blank\">@heoynwoo</a> I used the same code shared above (you should set QUANTIZE = \"int8\").</p>\n<p>To cralify, int8 quantization result in about 20% slower inference time.</p>\n<p>local evaluation time is:</p>\n<ul>\n<li>float32: 167 s</li>\n<li>int8: 200 s</li>\n</ul>",
              "rawMarkdown": "@heoynwoo I used the same code shared above (you should set QUANTIZE = \"int8\").\n\nTo cralify, int8 quantization result in about 20% slower inference time.\n\nlocal evaluation time is:\n- float32: 167 s\n- int8: 200 s"
            },
            {
              "id": 2218813,
              "postDate": "2023-04-12T03:59:26.910Z",
              "content": "<p>int8 quantization turned out to be seriously slower in the Kaggle's evaluation environment.<br>\nit took far more longer time compared to local evaluation.</p>\n<p>below is a comparison of the actual submission time:</p>\n<table>\n<thead>\n<tr>\n<th>quantization</th>\n<th>local eval time (sec)</th>\n<th>estimated submission time (min)</th>\n<th>actual submission time (min)</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>tf.float32</td>\n<td>167</td>\n<td>-</td>\n<td>44</td>\n</tr>\n<tr>\n<td>tf.int8</td>\n<td>257</td>\n<td>67.7</td>\n<td>93 (timeout errror)</td>\n</tr>\n</tbody>\n</table>",
              "rawMarkdown": "int8 quantization turned out to be seriously slower in the Kaggle's evaluation environment.\nit took far more longer time compared to local evaluation.\n\nbelow is a comparison of the actual submission time:\n\n| quantization | local eval time (sec) | estimated submission time (min) | actual submission time (min) |\n|--------------|-----------------------|---------------------------------|------------------------------|\n| tf.float32   | 167                   | -                               | 44                           |\n| tf.int8      | 257                   | 67.7                            | 93 (timeout errror)          |",
              "votes": 2
            },
            {
              "id": 2218822,
              "postDate": "2023-04-12T04:28:21.790Z",
              "content": "<p>Thank you! how did you measure actual submission time??</p>",
              "rawMarkdown": "Thank you! how did you measure actual submission time??"
            },
            {
              "id": 2218894,
              "postDate": "2023-04-12T05:47:48.570Z",
              "content": "<p><a href=\"https://www.kaggle.com/heoynwoo\" target=\"_blank\">@heoynwoo</a> You can poll the status of submission using Kaggle API.</p>\n<p>Various public notebooks exists:<br>\ne.g. <a href=\"https://www.kaggle.com/code/yasufuminakama/fb3-submission-time?scriptVersionId=104649325\" target=\"_blank\">https://www.kaggle.com/code/yasufuminakama/fb3-submission-time?scriptVersionId=104649325</a></p>",
              "rawMarkdown": "@heoynwoo You can poll the status of submission using Kaggle API.\n\nVarious public notebooks exists:\ne.g. https://www.kaggle.com/code/yasufuminakama/fb3-submission-time?scriptVersionId=104649325",
              "votes": 2
            }
          ]
        }
      ]
    },
    {
      "id": 2185696,
      "postDate": "2023-03-17T08:31:14.157Z",
      "content": "<p>The former submision is alright, but got time out after added this line:<br>\n<code>converter.optimizations = [tf.lite.Optimize.DEFAULT]</code><br>\n😧</p>",
      "rawMarkdown": "The former submision is alright, but got time out after added this line:\n`converter.optimizations = [tf.lite.Optimize.DEFAULT]`\n😧\n",
      "votes": 4
    },
    {
      "id": 2199444,
      "postDate": "2023-03-27T17:47:02.220Z",
      "content": "<p>In my experience, it all depends on the model. There are models that compress well and work well. Quantization can help in some cases, plus this is only the basic version, there are more complex options.</p>",
      "rawMarkdown": "In my experience, it all depends on the model. There are models that compress well and work well. Quantization can help in some cases, plus this is only the basic version, there are more complex options.",
      "votes": 2
    },
    {
      "id": 2235463,
      "postDate": "2023-04-26T04:23:33.440Z",
      "content": "<p>In my specific case, submitting without any quantization results in normal result in about 45-60 mins while default (int8 quantization) optimization gave an output error in my model. Screenshot below:<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1303569%2F13b47c4858b34bae505ac82662d28df4%2FSubmission%20Error%20Screenshot.jpg?generation=1682478808560391&amp;alt=media\" alt=\"\"></p>",
      "rawMarkdown": "In my specific case, submitting without any quantization results in normal result in about 45-60 mins while default (int8 quantization) optimization gave an output error in my model. Screenshot below:\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1303569%2F13b47c4858b34bae505ac82662d28df4%2FSubmission%20Error%20Screenshot.jpg?generation=1682478808560391&alt=media)"
    }
  ],
  "comments": [
    {
      "id": 2186271,
      "author_name": "Jebastin Nadar",
      "author_url": "",
      "post_date": "2023-03-17T17:15:58.177000",
      "content": "<p>tflite quantized models are optimized for ARM devices. For x86 CPU's, FP32 tflite model is faster than INT8 quantized tflite.</p>\n<p><a href=\"https://github.com/tensorflow/tensorflow/issues/21698#issuecomment-414764709\" target=\"_blank\">https://github.com/tensorflow/tensorflow/issues/21698#issuecomment-414764709</a></p>",
      "votes": 9,
      "replies": []
    },
    {
      "id": 2199612,
      "author_name": "arutema47",
      "author_url": "",
      "post_date": "2023-03-27T21:11:58.913000",
      "content": "<p>float16 quantization tends to works well with no speed or accuracy loss.</p>\n<blockquote>\n  <p>keras_model_converter.optimizations = [tf.lite.Optimize.DEFAULT]<br>\n  keras_model_converter.target_spec.supported_types = [tf.float16]<br>\n  tflite_model = keras_model_converter.convert() </p>\n</blockquote>",
      "votes": 6,
      "replies": [
        {
          "id": 2199648,
          "author_name": "Kolya Forrat",
          "author_url": "",
          "post_date": "2023-03-27T22:41:41.800000",
          "content": "<p>for me it drops performance by ~10% and accuracy by 1-3%</p>\n<p>Even with converting dataset. But maybe it will help someone</p>\n<pre><code>num_calibration_steps = \n ():\n   i  (num_calibration_steps):\n     [load_relevant_data_subset(BASE_DIR + train_df.loc[i, ])]\n\nkeras_model_converter.optimizations = [tf.lite.Optimize.DEFAULT]\nkeras_model_converter.representative_dataset = representative_dataset_gen\nkeras_model_converter.target_spec.supported_types = [tf.float16]\ntflite_model = keras_model_converter.convert()\n</code></pre>",
          "votes": 2,
          "replies": [
            {
              "id": 2199699,
              "author_name": "arutema47",
              "author_url": "",
              "post_date": "2023-03-28T00:31:50.793000",
              "content": "<p>I don't think you need any calibration for fp16, for typical networks, model weight range should be well within the fp16 range.<br>\nfp16 conversion has worked well for past competitions in kaggle and during my work too (using pytorch), and you shouldn't need any calibrations.<br>\nWhat will happen if you exclude <code>keras_model_converter.representative_dataset = representative_dataset_gen</code>?</p>",
              "votes": 1,
              "replies": []
            },
            {
              "id": 2199705,
              "author_name": "Kolya Forrat",
              "author_url": "",
              "post_date": "2023-03-28T00:43:30.673000",
              "content": "<p>With my model and fp16 - absolutely nothing. It helps me a little with int8, but integer quantization performs much worse anyway (as <a href=\"https://www.kaggle.com/samfc10\" target=\"_blank\">@samfc10</a> have mentioned below).</p>\n<p>I just wanted to leave it here because maybe it will improve someone's model that is okay with quantization.</p>",
              "votes": 0,
              "replies": []
            }
          ]
        },
        {
          "id": 2200439,
          "author_name": "",
          "author_url": "",
          "post_date": "2023-03-28T14:58:28.117000",
          "content": "",
          "votes": 0,
          "replies": []
        },
        {
          "id": 2218419,
          "author_name": "Bilzard",
          "author_url": "",
          "post_date": "2023-04-11T16:38:52.007000",
          "content": "<p>In my case, fp16 quantization causes severe accuracy drop (CV: 0.77955 -&gt; 0.00421).<br>\nOn the other hand, int8 quantization works fine (CV: 0.77955 -&gt; 0.77948).</p>\n<pre><code>    converter = tf.lite.TFLiteConverter.from_keras_model(model)\n     QUANTIZE  [, ]:\n        converter.optimizations = [tf.lite.Optimize.DEFAULT]\n         QUANTIZE == :\n            converter.target_spec.supported_types = [tf.float16]\n    tflite_model = converter.convert()\n</code></pre>",
          "votes": 0,
          "replies": [
            {
              "id": 2218744,
              "author_name": "Hyeonwoo Cho",
              "author_url": "",
              "post_date": "2023-04-12T01:47:54.147000",
              "content": "<p>Thank you for sharing!<br>\nI wonder how did you covert fp32 into int8?</p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 2218752,
              "author_name": "Bilzard",
              "author_url": "",
              "post_date": "2023-04-12T02:01:05.123000",
              "content": "<p><a href=\"https://www.kaggle.com/heoynwoo\" target=\"_blank\">@heoynwoo</a> I used the same code shared above (you should set QUANTIZE = \"int8\").</p>\n<p>To cralify, int8 quantization result in about 20% slower inference time.</p>\n<p>local evaluation time is:</p>\n<ul>\n<li>float32: 167 s</li>\n<li>int8: 200 s</li>\n</ul>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 2218813,
              "author_name": "Bilzard",
              "author_url": "",
              "post_date": "2023-04-12T03:59:26.910000",
              "content": "<p>int8 quantization turned out to be seriously slower in the Kaggle's evaluation environment.<br>\nit took far more longer time compared to local evaluation.</p>\n<p>below is a comparison of the actual submission time:</p>\n<table>\n<thead>\n<tr>\n<th>quantization</th>\n<th>local eval time (sec)</th>\n<th>estimated submission time (min)</th>\n<th>actual submission time (min)</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>tf.float32</td>\n<td>167</td>\n<td>-</td>\n<td>44</td>\n</tr>\n<tr>\n<td>tf.int8</td>\n<td>257</td>\n<td>67.7</td>\n<td>93 (timeout errror)</td>\n</tr>\n</tbody>\n</table>",
              "votes": 2,
              "replies": []
            },
            {
              "id": 2218822,
              "author_name": "Hyeonwoo Cho",
              "author_url": "",
              "post_date": "2023-04-12T04:28:21.790000",
              "content": "<p>Thank you! how did you measure actual submission time??</p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 2218894,
              "author_name": "Bilzard",
              "author_url": "",
              "post_date": "2023-04-12T05:47:48.570000",
              "content": "<p><a href=\"https://www.kaggle.com/heoynwoo\" target=\"_blank\">@heoynwoo</a> You can poll the status of submission using Kaggle API.</p>\n<p>Various public notebooks exists:<br>\ne.g. <a href=\"https://www.kaggle.com/code/yasufuminakama/fb3-submission-time?scriptVersionId=104649325\" target=\"_blank\">https://www.kaggle.com/code/yasufuminakama/fb3-submission-time?scriptVersionId=104649325</a></p>",
              "votes": 2,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 2199444,
      "author_name": "Andrij",
      "author_url": "",
      "post_date": "2023-03-27T17:47:02.220000",
      "content": "<p>In my experience, it all depends on the model. There are models that compress well and work well. Quantization can help in some cases, plus this is only the basic version, there are more complex options.</p>",
      "votes": 2,
      "replies": []
    },
    {
      "id": 2235463,
      "author_name": "coderRKJ",
      "author_url": "",
      "post_date": "2023-04-26T04:23:33.440000",
      "content": "<p>In my specific case, submitting without any quantization results in normal result in about 45-60 mins while default (int8 quantization) optimization gave an output error in my model. Screenshot below:<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1303569%2F13b47c4858b34bae505ac82662d28df4%2FSubmission%20Error%20Screenshot.jpg?generation=1682478808560391&amp;alt=media\" alt=\"\"></p>",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "2186271": "tflite quantized models are optimized for ARM devices. For x86 CPU's, FP32 tflite model is faster than INT8 quantized tflite.\n\nhttps://github.com/tensorflow/tensorflow/issues/21698#issuecomment-414764709",
    "2199612": "float16 quantization tends to works well with no speed or accuracy loss.\n>keras_model_converter.optimizations = [tf.lite.Optimize.DEFAULT]\nkeras_model_converter.target_spec.supported_types = [tf.float16]\ntflite_model = keras_model_converter.convert() ",
    "2185696": "The former submision is alright, but got time out after added this line:\n`converter.optimizations = [tf.lite.Optimize.DEFAULT]`\n😧\n",
    "2199444": "In my experience, it all depends on the model. There are models that compress well and work well. Quantization can help in some cases, plus this is only the basic version, there are more complex options.",
    "2235463": "In my specific case, submitting without any quantization results in normal result in about 45-60 mins while default (int8 quantization) optimization gave an output error in my model. Screenshot below:\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1303569%2F13b47c4858b34bae505ac82662d28df4%2FSubmission%20Error%20Screenshot.jpg?generation=1682478808560391&alt=media)"
  }
}