{
  "id": 393431,
  "title": "Dynamic range quantization Submission Scoring Error",
  "url": "/competitions/asl-signs/discussion/393431",
  "author_name": "",
  "post_date": "2023-03-09T11:56:18.492363Z",
  "votes": 1,
  "comment_count": 9,
  "views": 0,
  "content": "<p>Hi to all!<br>\nI keep trying to reduce the size of the model.<br>\nThis time I used Dynamic range quantization (<a href=\"https://www.tensorflow.org/lite/performance/post_training_quantization)\" target=\"_blank\">https://www.tensorflow.org/lite/performance/post_training_quantization)</a>.<br>\nHere is the code and results:</p>\n<hr>\n<p>keras_model_converter = tf.lite.TFLiteConverter.from_keras_model(tflite_keras_model)<br>\nkeras_model_converter.optimizations = [tf.lite.Optimize.DEFAULT]<br>\ntflite_model = keras_model_converter.convert()<br>\nwith open('/kaggle/working/models/model.tflite', 'wb') as f:<br>\n    f.write(tflite_model)<br>\n!zip submission.zip /kaggle/working/models/model.tflite</p>\n<p>!pip install tflite-runtime<br>\nimport tflite_runtime.interpreter as tflite</p>\n<p>interpreter = tflite.Interpreter(\"/kaggle/working/models/model.tflite\")<br>\nfound_signatures = list(interpreter.get_signature_list().keys())<br>\nprediction_fn = interpreter.get_signature_runner(\"serving_default\")</p>\n<p>output = prediction_fn(inputs=load_relevant_data_subset(train_df.path[0]))<br>\nsign = np.argmax(output[\"outputs\"])</p>\n<p>print(\"PRED : \", decoder(sign))<br>\nprint(\"GT   : \", train_df.sign[0])<br>\n  adding: kaggle/working/models/model.tflite (deflated 37%)<br>\nCollecting tflite-runtime<br>\n  Downloading tflite_runtime-2.11.0-cp37-cp37m-manylinux2014_x86_64.whl (2.5 MB)<br>\n     ━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━ 2.5/2.5 MB 34.5 MB/s eta 0:00:00<br>\nRequirement already satisfied: numpy&gt;=1.19.2 in /opt/conda/lib/python3.7/site-packages (from tflite-runtime) (1.21.6)<br>\nInstalling collected packages: tflite-runtime<br>\nSuccessfully installed tflite-runtime-2.11.0<br>\nWARNING: Running pip as the 'root' user can result in broken permissions and conflicting behaviour with the system package manager. It is recommended to use a virtual environment instead: <a href=\"https://pip.pypa.io/warnings/venv\" target=\"_blank\">https://pip.pypa.io/warnings/venv</a><br>\nPRED :  blow<br>\nGT   :  blow</p>\n<hr>\n<p>I added keras_model_converter.optimizations = [tf.lite.Optimize.DEFAULT]. <br>\nThis optimization reduced the size of my model by about 5 times, submission.zip is now about 30 mb in size. <br>\nAs with prunig, the model on the test works fine, but I get Submission Scoring Error.</p>\n<p>Any ideas how to fix this problem?</p>",
  "messages": [
    {
      "id": "2174802",
      "postDate": "03/09/2023 11:56:18",
      "content": "<p>Hi to all!<br>\nI keep trying to reduce the size of the model.<br>\nThis time I used Dynamic range quantization (<a href=\"https://www.tensorflow.org/lite/performance/post_training_quantization)\" target=\"_blank\">https://www.tensorflow.org/lite/performance/post_training_quantization)</a>.<br>\nHere is the code and results:</p>\n<hr>\n<p>keras_model_converter = tf.lite.TFLiteConverter.from_keras_model(tflite_keras_model)<br>\nkeras_model_converter.optimizations = [tf.lite.Optimize.DEFAULT]<br>\ntflite_model = keras_model_converter.convert()<br>\nwith open('/kaggle/working/models/model.tflite', 'wb') as f:<br>\n    f.write(tflite_model)<br>\n!zip submission.zip /kaggle/working/models/model.tflite</p>\n<p>!pip install tflite-runtime<br>\nimport tflite_runtime.interpreter as tflite</p>\n<p>interpreter = tflite.Interpreter(\"/kaggle/working/models/model.tflite\")<br>\nfound_signatures = list(interpreter.get_signature_list().keys())<br>\nprediction_fn = interpreter.get_signature_runner(\"serving_default\")</p>\n<p>output = prediction_fn(inputs=load_relevant_data_subset(train_df.path[0]))<br>\nsign = np.argmax(output[\"outputs\"])</p>\n<p>print(\"PRED : \", decoder(sign))<br>\nprint(\"GT   : \", train_df.sign[0])<br>\n  adding: kaggle/working/models/model.tflite (deflated 37%)<br>\nCollecting tflite-runtime<br>\n  Downloading tflite_runtime-2.11.0-cp37-cp37m-manylinux2014_x86_64.whl (2.5 MB)<br>\n     ━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━ 2.5/2.5 MB 34.5 MB/s eta 0:00:00<br>\nRequirement already satisfied: numpy&gt;=1.19.2 in /opt/conda/lib/python3.7/site-packages (from tflite-runtime) (1.21.6)<br>\nInstalling collected packages: tflite-runtime<br>\nSuccessfully installed tflite-runtime-2.11.0<br>\nWARNING: Running pip as the 'root' user can result in broken permissions and conflicting behaviour with the system package manager. It is recommended to use a virtual environment instead: <a href=\"https://pip.pypa.io/warnings/venv\" target=\"_blank\">https://pip.pypa.io/warnings/venv</a><br>\nPRED :  blow<br>\nGT   :  blow</p>\n<hr>\n<p>I added keras_model_converter.optimizations = [tf.lite.Optimize.DEFAULT]. <br>\nThis optimization reduced the size of my model by about 5 times, submission.zip is now about 30 mb in size. <br>\nAs with prunig, the model on the test works fine, but I get Submission Scoring Error.</p>\n<p>Any ideas how to fix this problem?</p>",
      "rawMarkdown": "Hi to all!\nI keep trying to reduce the size of the model.\nThis time I used Dynamic range quantization (https://www.tensorflow.org/lite/performance/post_training_quantization).\nHere is the code and results:\n_______________________________________________________________________\nkeras_model_converter = tf.lite.TFLiteConverter.from_keras_model(tflite_keras_model)\nkeras_model_converter.optimizations = [tf.lite.Optimize.DEFAULT]\ntflite_model = keras_model_converter.convert()\nwith open('/kaggle/working/models/model.tflite', 'wb') as f:\n    f.write(tflite_model)\n!zip submission.zip /kaggle/working/models/model.tflite\n\n!pip install tflite-runtime\nimport tflite_runtime.interpreter as tflite\n\ninterpreter = tflite.Interpreter(\"/kaggle/working/models/model.tflite\")\nfound_signatures = list(interpreter.get_signature_list().keys())\nprediction_fn = interpreter.get_signature_runner(\"serving_default\")\n\noutput = prediction_fn(inputs=load_relevant_data_subset(train_df.path[0]))\nsign = np.argmax(output[\"outputs\"])\n\nprint(\"PRED : \", decoder(sign))\nprint(\"GT   : \", train_df.sign[0])\n  adding: kaggle/working/models/model.tflite (deflated 37%)\nCollecting tflite-runtime\n  Downloading tflite_runtime-2.11.0-cp37-cp37m-manylinux2014_x86_64.whl (2.5 MB)\n     ━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━ 2.5/2.5 MB 34.5 MB/s eta 0:00:00\nRequirement already satisfied: numpy>=1.19.2 in /opt/conda/lib/python3.7/site-packages (from tflite-runtime) (1.21.6)\nInstalling collected packages: tflite-runtime\nSuccessfully installed tflite-runtime-2.11.0\nWARNING: Running pip as the 'root' user can result in broken permissions and conflicting behaviour with the system package manager. It is recommended to use a virtual environment instead: https://pip.pypa.io/warnings/venv\nPRED :  blow\nGT   :  blow\n___________________________________________________________________________________________________________\nI added keras_model_converter.optimizations = [tf.lite.Optimize.DEFAULT]. \nThis optimization reduced the size of my model by about 5 times, submission.zip is now about 30 mb in size. \nAs with prunig, the model on the test works fine, but I get Submission Scoring Error.\n\nAny ideas how to fix this problem?",
      "votes": null
    },
    {
      "id": "2174902",
      "postDate": "03/09/2023 13:52:45",
      "content": "<p>Check whether your previous submission still works. I suspect that they changed the evaluation script and your error might not be because of pruning but other errors like not meeting the time requirement.</p>",
      "rawMarkdown": "Check whether your previous submission still works. I suspect that they changed the evaluation script and your error might not be because of pruning but other errors like not meeting the time requirement.",
      "votes": null
    },
    {
      "id": "2174907",
      "postDate": "03/09/2023 14:00:49",
      "content": "<p>I really complicated the model, I'll try to check tomorrow. Note that I got an error with pruning yesterday and today with keras_model_converter.optimizations = [tf.lite.Optimize.DEFAULT]. Between these two experiments there were successful subs.</p>",
      "rawMarkdown": "I really complicated the model, I'll try to check tomorrow. Note that I got an error with pruning yesterday and today with keras_model_converter.optimizations = [tf.lite.Optimize.DEFAULT]. Between these two experiments there were successful subs.",
      "votes": null
    },
    {
      "id": "2174915",
      "postDate": "03/09/2023 14:07:59",
      "content": "<p>Have you verified that the model works with TensorFlow Lite Runtime v2.9.1 ?</p>\n<p>Perhaps the optimizers use features that were added in 2.11.0</p>",
      "rawMarkdown": "Have you verified that the model works with TensorFlow Lite Runtime v2.9.1 ?\n\nPerhaps the optimizers use features that were added in 2.11.0",
      "votes": null
    },
    {
      "id": "2174924",
      "postDate": "03/09/2023 14:14:17",
      "content": "<p>Ah okay, then please disregard my comment.</p>",
      "rawMarkdown": "Ah okay, then please disregard my comment.",
      "votes": null
    },
    {
      "id": "2174934",
      "postDate": "03/09/2023 14:21:04",
      "content": "<p>I have TENSORFLOW VERSION: 2.11.0 installed. This version works.<br>\nI only added keras_model_converter.optimizations = [tf.lite.Optimize.DEFAULT].<br>\nAre you suggesting to install tf version 2.9.1? I'll give it a try, but it's strange that it works without optimization</p>",
      "rawMarkdown": "I have TENSORFLOW VERSION: 2.11.0 installed. This version works.\nI only added keras_model_converter.optimizations = [tf.lite.Optimize.DEFAULT].\nAre you suggesting to install tf version 2.9.1? I'll give it a try, but it's strange that it works without optimization",
      "votes": null
    },
    {
      "id": "2174952",
      "postDate": "03/09/2023 14:32:37",
      "content": "<p>I only mention tflite-runtime-2.9.1 because it's what the competition uses.  I don't have any specific reason to think it's the issue, but it should be easy to rule out</p>",
      "rawMarkdown": "I only mention tflite-runtime-2.9.1 because it's what the competition uses.  I don't have any specific reason to think it's the issue, but it should be easy to rule out",
      "votes": null
    },
    {
      "id": "2174957",
      "postDate": "03/09/2023 14:34:47",
      "content": "<p>I think I understood what the problem is, today I have no free subs. So I will check it tomorrow, I will write about the result</p>",
      "rawMarkdown": "I think I understood what the problem is, today I have no free subs. So I will check it tomorrow, I will write about the result",
      "votes": null
    },
    {
      "id": "2175979",
      "postDate": "03/10/2023 09:51:48",
      "content": "<p>Everything works, the problem was the size of the model. Model size should be less than 40mb when unzipped, not when zipped</p>",
      "rawMarkdown": "Everything works, the problem was the size of the model. Model size should be less than 40mb when unzipped, not when zipped",
      "votes": null
    },
    {
      "id": "2176961",
      "postDate": "03/11/2023 04:49:35",
      "content": "<p>Have you calculated the mean time in milliseconds it takes for the quantized model to process a sample? I've also tried using this quantization - it shrunk the # of megabytes the model takes, but increased inference time from ~10 ms to ~300 ms, which is beyond the competition's limit. Perhaps the same issue is occurring for you? The TF tutorial also states that you are supposed to use a representative dataset to calibrate the optimization, but when I tried this the accuracy dropped to zero (even with ~5000 samples) and the inference time stayed at 300 ms. If anyone here figures out how to do this quantization properly, please let me know!</p>",
      "rawMarkdown": "Have you calculated the mean time in milliseconds it takes for the quantized model to process a sample? I've also tried using this quantization - it shrunk the # of megabytes the model takes, but increased inference time from ~10 ms to ~300 ms, which is beyond the competition's limit. Perhaps the same issue is occurring for you? The TF tutorial also states that you are supposed to use a representative dataset to calibrate the optimization, but when I tried this the accuracy dropped to zero (even with ~5000 samples) and the inference time stayed at 300 ms. If anyone here figures out how to do this quantization properly, please let me know!",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 2174902,
      "author_name": "nyakasko",
      "author_url": "",
      "post_date": "03/09/2023 13:52:45",
      "content": "<p>Check whether your previous submission still works. I suspect that they changed the evaluation script and your error might not be because of pruning but other errors like not meeting the time requirement.</p>",
      "votes": null,
      "replies": [
        {
          "id": 2174907,
          "author_name": "aikhmelnytskyy",
          "author_url": "",
          "post_date": "03/09/2023 14:00:49",
          "content": "<p>I really complicated the model, I'll try to check tomorrow. Note that I got an error with pruning yesterday and today with keras_model_converter.optimizations = [tf.lite.Optimize.DEFAULT]. Between these two experiments there were successful subs.</p>",
          "votes": null,
          "replies": [
            {
              "id": 2174924,
              "author_name": "nyakasko",
              "author_url": "",
              "post_date": "03/09/2023 14:14:17",
              "content": "<p>Ah okay, then please disregard my comment.</p>",
              "votes": null,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 2174915,
      "author_name": "quillen",
      "author_url": "",
      "post_date": "03/09/2023 14:07:59",
      "content": "<p>Have you verified that the model works with TensorFlow Lite Runtime v2.9.1 ?</p>\n<p>Perhaps the optimizers use features that were added in 2.11.0</p>",
      "votes": null,
      "replies": [
        {
          "id": 2174934,
          "author_name": "aikhmelnytskyy",
          "author_url": "",
          "post_date": "03/09/2023 14:21:04",
          "content": "<p>I have TENSORFLOW VERSION: 2.11.0 installed. This version works.<br>\nI only added keras_model_converter.optimizations = [tf.lite.Optimize.DEFAULT].<br>\nAre you suggesting to install tf version 2.9.1? I'll give it a try, but it's strange that it works without optimization</p>",
          "votes": null,
          "replies": [
            {
              "id": 2174952,
              "author_name": "quillen",
              "author_url": "",
              "post_date": "03/09/2023 14:32:37",
              "content": "<p>I only mention tflite-runtime-2.9.1 because it's what the competition uses.  I don't have any specific reason to think it's the issue, but it should be easy to rule out</p>",
              "votes": null,
              "replies": [
                {
                  "id": 2174957,
                  "author_name": "aikhmelnytskyy",
                  "author_url": "",
                  "post_date": "03/09/2023 14:34:47",
                  "content": "<p>I think I understood what the problem is, today I have no free subs. So I will check it tomorrow, I will write about the result</p>",
                  "votes": null,
                  "replies": [
                    {
                      "id": 2175979,
                      "author_name": "aikhmelnytskyy",
                      "author_url": "",
                      "post_date": "03/10/2023 09:51:48",
                      "content": "<p>Everything works, the problem was the size of the model. Model size should be less than 40mb when unzipped, not when zipped</p>",
                      "votes": null,
                      "replies": []
                    }
                  ]
                }
              ]
            }
          ]
        }
      ]
    },
    {
      "id": 2176961,
      "author_name": "megaray",
      "author_url": "",
      "post_date": "03/11/2023 04:49:35",
      "content": "<p>Have you calculated the mean time in milliseconds it takes for the quantized model to process a sample? I've also tried using this quantization - it shrunk the # of megabytes the model takes, but increased inference time from ~10 ms to ~300 ms, which is beyond the competition's limit. Perhaps the same issue is occurring for you? The TF tutorial also states that you are supposed to use a representative dataset to calibrate the optimization, but when I tried this the accuracy dropped to zero (even with ~5000 samples) and the inference time stayed at 300 ms. If anyone here figures out how to do this quantization properly, please let me know!</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "2174802": "Hi to all!\nI keep trying to reduce the size of the model.\nThis time I used Dynamic range quantization (https://www.tensorflow.org/lite/performance/post_training_quantization).\nHere is the code and results:\n_______________________________________________________________________\nkeras_model_converter = tf.lite.TFLiteConverter.from_keras_model(tflite_keras_model)\nkeras_model_converter.optimizations = [tf.lite.Optimize.DEFAULT]\ntflite_model = keras_model_converter.convert()\nwith open('/kaggle/working/models/model.tflite', 'wb') as f:\n    f.write(tflite_model)\n!zip submission.zip /kaggle/working/models/model.tflite\n\n!pip install tflite-runtime\nimport tflite_runtime.interpreter as tflite\n\ninterpreter = tflite.Interpreter(\"/kaggle/working/models/model.tflite\")\nfound_signatures = list(interpreter.get_signature_list().keys())\nprediction_fn = interpreter.get_signature_runner(\"serving_default\")\n\noutput = prediction_fn(inputs=load_relevant_data_subset(train_df.path[0]))\nsign = np.argmax(output[\"outputs\"])\n\nprint(\"PRED : \", decoder(sign))\nprint(\"GT   : \", train_df.sign[0])\n  adding: kaggle/working/models/model.tflite (deflated 37%)\nCollecting tflite-runtime\n  Downloading tflite_runtime-2.11.0-cp37-cp37m-manylinux2014_x86_64.whl (2.5 MB)\n     ━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━ 2.5/2.5 MB 34.5 MB/s eta 0:00:00\nRequirement already satisfied: numpy>=1.19.2 in /opt/conda/lib/python3.7/site-packages (from tflite-runtime) (1.21.6)\nInstalling collected packages: tflite-runtime\nSuccessfully installed tflite-runtime-2.11.0\nWARNING: Running pip as the 'root' user can result in broken permissions and conflicting behaviour with the system package manager. It is recommended to use a virtual environment instead: https://pip.pypa.io/warnings/venv\nPRED :  blow\nGT   :  blow\n___________________________________________________________________________________________________________\nI added keras_model_converter.optimizations = [tf.lite.Optimize.DEFAULT]. \nThis optimization reduced the size of my model by about 5 times, submission.zip is now about 30 mb in size. \nAs with prunig, the model on the test works fine, but I get Submission Scoring Error.\n\nAny ideas how to fix this problem?",
    "2174902": "Check whether your previous submission still works. I suspect that they changed the evaluation script and your error might not be because of pruning but other errors like not meeting the time requirement.",
    "2174907": "I really complicated the model, I'll try to check tomorrow. Note that I got an error with pruning yesterday and today with keras_model_converter.optimizations = [tf.lite.Optimize.DEFAULT]. Between these two experiments there were successful subs.",
    "2174915": "Have you verified that the model works with TensorFlow Lite Runtime v2.9.1 ?\n\nPerhaps the optimizers use features that were added in 2.11.0",
    "2174924": "Ah okay, then please disregard my comment.",
    "2174934": "I have TENSORFLOW VERSION: 2.11.0 installed. This version works.\nI only added keras_model_converter.optimizations = [tf.lite.Optimize.DEFAULT].\nAre you suggesting to install tf version 2.9.1? I'll give it a try, but it's strange that it works without optimization",
    "2174952": "I only mention tflite-runtime-2.9.1 because it's what the competition uses.  I don't have any specific reason to think it's the issue, but it should be easy to rule out",
    "2174957": "I think I understood what the problem is, today I have no free subs. So I will check it tomorrow, I will write about the result",
    "2175979": "Everything works, the problem was the size of the model. Model size should be less than 40mb when unzipped, not when zipped",
    "2176961": "Have you calculated the mean time in milliseconds it takes for the quantized model to process a sample? I've also tried using this quantization - it shrunk the # of megabytes the model takes, but increased inference time from ~10 ms to ~300 ms, which is beyond the competition's limit. Perhaps the same issue is occurring for you? The TF tutorial also states that you are supposed to use a representative dataset to calibrate the optimization, but when I tried this the accuracy dropped to zero (even with ~5000 samples) and the inference time stayed at 300 ms. If anyone here figures out how to do this quantization properly, please let me know!"
  },
  "source": "meta"
}