{
  "id": 394371,
  "title": "How to make sub models larger than 100 mb",
  "url": "/competitions/asl-signs/discussion/394371",
  "author_name": "",
  "post_date": "2023-03-13T09:14:39.759903500Z",
  "votes": 15,
  "comment_count": 8,
  "views": 0,
  "content": "<p>Hi to all! <br>\nI've spent the last week looking for a solution to downsize the model, here are the results:</p>\n<ol>\n<li>Reducing the size of the model will not necessarily help us, because there are 60 min. execution time limit.</li>\n<li>Remember the limit of 40 mb per model means the limit of 40 mb per unpacked model and not submission. zip archive, even if the archive is smaller, but the unpacked model will be more than mb, you will get a Submission Scoring Error.</li>\n</ol>\n<p>Now for the results</p>\n<ol>\n<li>Pruning (<a href=\"https://www.tensorflow.org/model_optimization/guide/pruning/pruning_with_keras\" target=\"_blank\">https://www.tensorflow.org/model_optimization/guide/pruning/pruning_with_keras</a>) at first glance reduces the size of the model, but in fact it does not help because the unpacked model is still large (see point two);</li>\n<li>Knowledge Distillation (<a href=\"https://keras.io/examples/vision/knowledge_distillation/\" target=\"_blank\">https://keras.io/examples/vision/knowledge_distillation/</a>) significantly degrades CV, plus requires additional moves, which is a waste of time for me.</li>\n<li>Dynamic range quantization (<a href=\"https://www.tensorflow.org/lite/performance/post_training_quantization\" target=\"_blank\">https://www.tensorflow.org/lite/performance/post_training_quantization</a>) is the best solution I found. You add one line of code:<br>\nconverter.optimizations = [tf.lite.Optimize.DEFAULT]<br>\nand all your model will be reduced several times (up to 5). But remember that the execution time remains, this method does not help here, I saw a comment that it even increases the execution time, but it all depends on the model. Regarding model accuracy, I see that this method slightly degrades LB(&lt;0.01).<br>\nYou can see an example in my public notebook. Just remove converter.optimizations = [tf.lite.Optimize.DEFAULT]<br>\nand see the difference in the size of the submission.zip.</li>\n</ol>\n<p>General summary:<br>\nUsing Dynamic range quantization can help you successfully use +100mb size model, LB price. But from what I see, you should still try to build many different models with different preprocessing, their simple ensemble will give better results. Good luck to everyone!</p>",
  "messages": [
    {
      "id": "2179577",
      "postDate": "03/13/2023 09:14:39",
      "content": "<p>Hi to all! <br>\nI've spent the last week looking for a solution to downsize the model, here are the results:</p>\n<ol>\n<li>Reducing the size of the model will not necessarily help us, because there are 60 min. execution time limit.</li>\n<li>Remember the limit of 40 mb per model means the limit of 40 mb per unpacked model and not submission. zip archive, even if the archive is smaller, but the unpacked model will be more than mb, you will get a Submission Scoring Error.</li>\n</ol>\n<p>Now for the results</p>\n<ol>\n<li>Pruning (<a href=\"https://www.tensorflow.org/model_optimization/guide/pruning/pruning_with_keras\" target=\"_blank\">https://www.tensorflow.org/model_optimization/guide/pruning/pruning_with_keras</a>) at first glance reduces the size of the model, but in fact it does not help because the unpacked model is still large (see point two);</li>\n<li>Knowledge Distillation (<a href=\"https://keras.io/examples/vision/knowledge_distillation/\" target=\"_blank\">https://keras.io/examples/vision/knowledge_distillation/</a>) significantly degrades CV, plus requires additional moves, which is a waste of time for me.</li>\n<li>Dynamic range quantization (<a href=\"https://www.tensorflow.org/lite/performance/post_training_quantization\" target=\"_blank\">https://www.tensorflow.org/lite/performance/post_training_quantization</a>) is the best solution I found. You add one line of code:<br>\nconverter.optimizations = [tf.lite.Optimize.DEFAULT]<br>\nand all your model will be reduced several times (up to 5). But remember that the execution time remains, this method does not help here, I saw a comment that it even increases the execution time, but it all depends on the model. Regarding model accuracy, I see that this method slightly degrades LB(&lt;0.01).<br>\nYou can see an example in my public notebook. Just remove converter.optimizations = [tf.lite.Optimize.DEFAULT]<br>\nand see the difference in the size of the submission.zip.</li>\n</ol>\n<p>General summary:<br>\nUsing Dynamic range quantization can help you successfully use +100mb size model, LB price. But from what I see, you should still try to build many different models with different preprocessing, their simple ensemble will give better results. Good luck to everyone!</p>",
      "rawMarkdown": "Hi to all! \nI've spent the last week looking for a solution to downsize the model, here are the results:\n1. Reducing the size of the model will not necessarily help us, because there are 60 min. execution time limit.\n2. Remember the limit of 40 mb per model means the limit of 40 mb per unpacked model and not submission. zip archive, even if the archive is smaller, but the unpacked model will be more than mb, you will get a Submission Scoring Error.\n\nNow for the results\n\n1. Pruning (https://www.tensorflow.org/model_optimization/guide/pruning/pruning_with_keras) at first glance reduces the size of the model, but in fact it does not help because the unpacked model is still large (see point two);\n2. Knowledge Distillation (https://keras.io/examples/vision/knowledge_distillation/) significantly degrades CV, plus requires additional moves, which is a waste of time for me.\n3. Dynamic range quantization (https://www.tensorflow.org/lite/performance/post_training_quantization) is the best solution I found. You add one line of code:\nconverter.optimizations = [tf.lite.Optimize.DEFAULT]\nand all your model will be reduced several times (up to 5). But remember that the execution time remains, this method does not help here, I saw a comment that it even increases the execution time, but it all depends on the model. Regarding model accuracy, I see that this method slightly degrades LB(<0.01).\n  You can see an example in my public notebook. Just remove converter.optimizations = [tf.lite.Optimize.DEFAULT]\n  and see the difference in the size of the submission.zip.\n\nGeneral summary:\nUsing Dynamic range quantization can help you successfully use +100mb size model, LB price. But from what I see, you should still try to build many different models with different preprocessing, their simple ensemble will give better results. Good luck to everyone!",
      "votes": null
    },
    {
      "id": "2179697",
      "postDate": "03/13/2023 10:34:01",
      "content": "<p>For dynamic range quantization, did you have to use a representative dataset to get the execution time down?</p>",
      "rawMarkdown": "For dynamic range quantization, did you have to use a representative dataset to get the execution time down?",
      "votes": null
    },
    {
      "id": "2179702",
      "postDate": "03/13/2023 10:40:17",
      "content": "<p>I haven't tried it, it needs to be tested. There are several approaches to how to do this in the documentation. I used the simplest method. I think if you dig into the settings there, you can pick up a small drop in lb</p>",
      "rawMarkdown": "I haven't tried it, it needs to be tested. There are several approaches to how to do this in the documentation. I used the simplest method. I think if you dig into the settings there, you can pick up a small drop in lb",
      "votes": null
    },
    {
      "id": "2179827",
      "postDate": "03/13/2023 12:48:38",
      "content": "<p>Interesting. I’ll add this into my ensemble notebook. Thanks!</p>",
      "rawMarkdown": "Interesting. I’ll add this into my ensemble notebook. Thanks!",
      "votes": null
    },
    {
      "id": "2179847",
      "postDate": "03/13/2023 12:59:57",
      "content": "<p>Note that this slightly affects lb and the overall accuracy of the model. I think the influence there is minimal, but it is due to the fact that: This type of quantization, statically quantizes only the weights from floating point to integer at conversion time, which provides 8-bits of precision. </p>",
      "rawMarkdown": "Note that this slightly affects lb and the overall accuracy of the model. I think the influence there is minimal, but it is due to the fact that: This type of quantization, statically quantizes only the weights from floating point to integer at conversion time, which provides 8-bits of precision.",
      "votes": null
    },
    {
      "id": "2179862",
      "postDate": "03/13/2023 13:05:56",
      "content": "<p>I’ll add it in as an option to the NB. I’ll credit you and Hengck (I saw him use it earlier but didn’t know enough to add it). </p>\n<p>I’ll quote you and link this as additional context. I think it’s a valid option, but it should only be used where appropriate due to the performance degradation you described.</p>",
      "rawMarkdown": "I’ll add it in as an option to the NB. I’ll credit you and Hengck (I saw him use it earlier but didn’t know enough to add it). \n\nI’ll quote you and link this as additional context. I think it’s a valid option, but it should only be used where appropriate due to the performance degradation you described.",
      "votes": null
    },
    {
      "id": "2180565",
      "postDate": "03/14/2023 00:04:27",
      "content": "<p>\"I see that this method slightly degrades LB(&lt;0.01).\" </p>\n<p>this can be avoided:</p>\n<ul>\n<li>understand the quantization and measure the numerical error (i.e difference or orginal and quantised values)</li>\n<li>add quantization noise  during training (e.g. gussian noise)</li>\n</ul>\n<p>this is an easy fix. the quantization is just to save space. values are converted to fp32 during computation.<br>\nyou can also google some work on better quantization error modeling.<br>\n(use search keyword: Quantization Methods for Neural Network training)</p>\n<p><a href=\"https://leimao.github.io/article/Neural-Networks-Quantization/\" target=\"_blank\">https://leimao.github.io/article/Neural-Networks-Quantization/</a><br>\n<a href=\"https://leimao.github.io/blog/PyTorch-Static-Quantization/\" target=\"_blank\">https://leimao.github.io/blog/PyTorch-Static-Quantization/</a><br>\n<a href=\"https://leimao.github.io/blog/PyTorch-Dynamic-Quantization/\" target=\"_blank\">https://leimao.github.io/blog/PyTorch-Dynamic-Quantization/</a></p>\n<p>differentiable Quantization</p>\n<h2><a href=\"https://github.com/aliyun/alibabacloud-quantization-networks\" target=\"_blank\">https://github.com/aliyun/alibabacloud-quantization-networks</a></h2>\n<p>Actually a embedded implementation has what they called network quantization which both store and compute int8.<br>\nyou can google about it, (e.g. <a href=\"https://developer.nvidia.com/blog/achieving-fp32-accuracy-for-int8-inference-using-quantization-aware-training-with-tensorrt/\" target=\"_blank\">https://developer.nvidia.com/blog/achieving-fp32-accuracy-for-int8-inference-using-quantization-aware-training-with-tensorrt/</a>)</p>\n<p>But for tflite, int8 computation is only for ARM cpu.<br>\n<a href=\"https://github.com/tensorflow/model-optimization/issues/599\" target=\"_blank\">https://github.com/tensorflow/model-optimization/issues/599</a></p>\n<p>,  </p>",
      "rawMarkdown": "\"I see that this method slightly degrades LB(<0.01).\" \n\nthis can be avoided:\n- understand the quantization and measure the numerical error (i.e difference or orginal and quantised values)\n- add quantization noise  during training (e.g. gussian noise)\n\nthis is an easy fix. the quantization is just to save space. values are converted to fp32 during computation.\nyou can also google some work on better quantization error modeling.\n(use search keyword: Quantization Methods for Neural Network training)\n\nhttps://leimao.github.io/article/Neural-Networks-Quantization/\nhttps://leimao.github.io/blog/PyTorch-Static-Quantization/\nhttps://leimao.github.io/blog/PyTorch-Dynamic-Quantization/\n \ndifferentiable Quantization\nhttps://github.com/aliyun/alibabacloud-quantization-networks\n---\n\nActually a embedded implementation has what they called network quantization which both store and compute int8.\nyou can google about it, (e.g. https://developer.nvidia.com/blog/achieving-fp32-accuracy-for-int8-inference-using-quantization-aware-training-with-tensorrt/)\n\nBut for tflite, int8 computation is only for ARM cpu.\nhttps://github.com/tensorflow/model-optimization/issues/599\n\n \n ,",
      "votes": null
    },
    {
      "id": "2180931",
      "postDate": "03/14/2023 07:34:05",
      "content": "<p>It is clear that it is possible to fix it, the question is whether it is necessary to spend time on it. When I saw your 2.75 md model which gives 0.66 LB, and compared it to my 0.67 lb ensembles thirty or forty times larger, I realized that now is not the time to focus on that, it is better to concentrate on building really good singles models and only later will it be seen whether it will be necessary. But I decided to share the result. Good luck</p>",
      "rawMarkdown": "It is clear that it is possible to fix it, the question is whether it is necessary to spend time on it. When I saw your 2.75 md model which gives 0.66 LB, and compared it to my 0.67 lb ensembles thirty or forty times larger, I realized that now is not the time to focus on that, it is better to concentrate on building really good singles models and only later will it be seen whether it will be necessary. But I decided to share the result. Good luck",
      "votes": null
    },
    {
      "id": "2201194",
      "postDate": "03/29/2023 05:58:58",
      "content": "<p>General updates, if you look at the discussions: people are reporting that quantization is slowing things down. From my experience, you need to test each model separately, for different models the indicators are different. You can also add the option keras_model_converter.target_spec.supported_types = [tf.float16] to reduce accuracy and speed drop issues. Search for more details in the discussion on the word Quantization. Thank you all for your comments!</p>",
      "rawMarkdown": "General updates, if you look at the discussions: people are reporting that quantization is slowing things down. From my experience, you need to test each model separately, for different models the indicators are different. You can also add the option keras_model_converter.target_spec.supported_types = [tf.float16] to reduce accuracy and speed drop issues. Search for more details in the discussion on the word Quantization. Thank you all for your comments!",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 2179697,
      "author_name": "megaray",
      "author_url": "",
      "post_date": "03/13/2023 10:34:01",
      "content": "<p>For dynamic range quantization, did you have to use a representative dataset to get the execution time down?</p>",
      "votes": null,
      "replies": [
        {
          "id": 2179702,
          "author_name": "aikhmelnytskyy",
          "author_url": "",
          "post_date": "03/13/2023 10:40:17",
          "content": "<p>I haven't tried it, it needs to be tested. There are several approaches to how to do this in the documentation. I used the simplest method. I think if you dig into the settings there, you can pick up a small drop in lb</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 2179827,
      "author_name": "dschettler8845",
      "author_url": "",
      "post_date": "03/13/2023 12:48:38",
      "content": "<p>Interesting. I’ll add this into my ensemble notebook. Thanks!</p>",
      "votes": null,
      "replies": [
        {
          "id": 2179847,
          "author_name": "aikhmelnytskyy",
          "author_url": "",
          "post_date": "03/13/2023 12:59:57",
          "content": "<p>Note that this slightly affects lb and the overall accuracy of the model. I think the influence there is minimal, but it is due to the fact that: This type of quantization, statically quantizes only the weights from floating point to integer at conversion time, which provides 8-bits of precision. </p>",
          "votes": null,
          "replies": [
            {
              "id": 2179862,
              "author_name": "dschettler8845",
              "author_url": "",
              "post_date": "03/13/2023 13:05:56",
              "content": "<p>I’ll add it in as an option to the NB. I’ll credit you and Hengck (I saw him use it earlier but didn’t know enough to add it). </p>\n<p>I’ll quote you and link this as additional context. I think it’s a valid option, but it should only be used where appropriate due to the performance degradation you described.</p>",
              "votes": null,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 2180565,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "03/14/2023 00:04:27",
      "content": "<p>\"I see that this method slightly degrades LB(&lt;0.01).\" </p>\n<p>this can be avoided:</p>\n<ul>\n<li>understand the quantization and measure the numerical error (i.e difference or orginal and quantised values)</li>\n<li>add quantization noise  during training (e.g. gussian noise)</li>\n</ul>\n<p>this is an easy fix. the quantization is just to save space. values are converted to fp32 during computation.<br>\nyou can also google some work on better quantization error modeling.<br>\n(use search keyword: Quantization Methods for Neural Network training)</p>\n<p><a href=\"https://leimao.github.io/article/Neural-Networks-Quantization/\" target=\"_blank\">https://leimao.github.io/article/Neural-Networks-Quantization/</a><br>\n<a href=\"https://leimao.github.io/blog/PyTorch-Static-Quantization/\" target=\"_blank\">https://leimao.github.io/blog/PyTorch-Static-Quantization/</a><br>\n<a href=\"https://leimao.github.io/blog/PyTorch-Dynamic-Quantization/\" target=\"_blank\">https://leimao.github.io/blog/PyTorch-Dynamic-Quantization/</a></p>\n<p>differentiable Quantization</p>\n<h2><a href=\"https://github.com/aliyun/alibabacloud-quantization-networks\" target=\"_blank\">https://github.com/aliyun/alibabacloud-quantization-networks</a></h2>\n<p>Actually a embedded implementation has what they called network quantization which both store and compute int8.<br>\nyou can google about it, (e.g. <a href=\"https://developer.nvidia.com/blog/achieving-fp32-accuracy-for-int8-inference-using-quantization-aware-training-with-tensorrt/\" target=\"_blank\">https://developer.nvidia.com/blog/achieving-fp32-accuracy-for-int8-inference-using-quantization-aware-training-with-tensorrt/</a>)</p>\n<p>But for tflite, int8 computation is only for ARM cpu.<br>\n<a href=\"https://github.com/tensorflow/model-optimization/issues/599\" target=\"_blank\">https://github.com/tensorflow/model-optimization/issues/599</a></p>\n<p>,  </p>",
      "votes": null,
      "replies": [
        {
          "id": 2180931,
          "author_name": "aikhmelnytskyy",
          "author_url": "",
          "post_date": "03/14/2023 07:34:05",
          "content": "<p>It is clear that it is possible to fix it, the question is whether it is necessary to spend time on it. When I saw your 2.75 md model which gives 0.66 LB, and compared it to my 0.67 lb ensembles thirty or forty times larger, I realized that now is not the time to focus on that, it is better to concentrate on building really good singles models and only later will it be seen whether it will be necessary. But I decided to share the result. Good luck</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 2201194,
      "author_name": "aikhmelnytskyy",
      "author_url": "",
      "post_date": "03/29/2023 05:58:58",
      "content": "<p>General updates, if you look at the discussions: people are reporting that quantization is slowing things down. From my experience, you need to test each model separately, for different models the indicators are different. You can also add the option keras_model_converter.target_spec.supported_types = [tf.float16] to reduce accuracy and speed drop issues. Search for more details in the discussion on the word Quantization. Thank you all for your comments!</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "2179577": "Hi to all! \nI've spent the last week looking for a solution to downsize the model, here are the results:\n1. Reducing the size of the model will not necessarily help us, because there are 60 min. execution time limit.\n2. Remember the limit of 40 mb per model means the limit of 40 mb per unpacked model and not submission. zip archive, even if the archive is smaller, but the unpacked model will be more than mb, you will get a Submission Scoring Error.\n\nNow for the results\n\n1. Pruning (https://www.tensorflow.org/model_optimization/guide/pruning/pruning_with_keras) at first glance reduces the size of the model, but in fact it does not help because the unpacked model is still large (see point two);\n2. Knowledge Distillation (https://keras.io/examples/vision/knowledge_distillation/) significantly degrades CV, plus requires additional moves, which is a waste of time for me.\n3. Dynamic range quantization (https://www.tensorflow.org/lite/performance/post_training_quantization) is the best solution I found. You add one line of code:\nconverter.optimizations = [tf.lite.Optimize.DEFAULT]\nand all your model will be reduced several times (up to 5). But remember that the execution time remains, this method does not help here, I saw a comment that it even increases the execution time, but it all depends on the model. Regarding model accuracy, I see that this method slightly degrades LB(<0.01).\n  You can see an example in my public notebook. Just remove converter.optimizations = [tf.lite.Optimize.DEFAULT]\n  and see the difference in the size of the submission.zip.\n\nGeneral summary:\nUsing Dynamic range quantization can help you successfully use +100mb size model, LB price. But from what I see, you should still try to build many different models with different preprocessing, their simple ensemble will give better results. Good luck to everyone!",
    "2179697": "For dynamic range quantization, did you have to use a representative dataset to get the execution time down?",
    "2179702": "I haven't tried it, it needs to be tested. There are several approaches to how to do this in the documentation. I used the simplest method. I think if you dig into the settings there, you can pick up a small drop in lb",
    "2179827": "Interesting. I’ll add this into my ensemble notebook. Thanks!",
    "2179847": "Note that this slightly affects lb and the overall accuracy of the model. I think the influence there is minimal, but it is due to the fact that: This type of quantization, statically quantizes only the weights from floating point to integer at conversion time, which provides 8-bits of precision.",
    "2179862": "I’ll add it in as an option to the NB. I’ll credit you and Hengck (I saw him use it earlier but didn’t know enough to add it). \n\nI’ll quote you and link this as additional context. I think it’s a valid option, but it should only be used where appropriate due to the performance degradation you described.",
    "2180565": "\"I see that this method slightly degrades LB(<0.01).\" \n\nthis can be avoided:\n- understand the quantization and measure the numerical error (i.e difference or orginal and quantised values)\n- add quantization noise  during training (e.g. gussian noise)\n\nthis is an easy fix. the quantization is just to save space. values are converted to fp32 during computation.\nyou can also google some work on better quantization error modeling.\n(use search keyword: Quantization Methods for Neural Network training)\n\nhttps://leimao.github.io/article/Neural-Networks-Quantization/\nhttps://leimao.github.io/blog/PyTorch-Static-Quantization/\nhttps://leimao.github.io/blog/PyTorch-Dynamic-Quantization/\n \ndifferentiable Quantization\nhttps://github.com/aliyun/alibabacloud-quantization-networks\n---\n\nActually a embedded implementation has what they called network quantization which both store and compute int8.\nyou can google about it, (e.g. https://developer.nvidia.com/blog/achieving-fp32-accuracy-for-int8-inference-using-quantization-aware-training-with-tensorrt/)\n\nBut for tflite, int8 computation is only for ARM cpu.\nhttps://github.com/tensorflow/model-optimization/issues/599\n\n \n ,",
    "2180931": "It is clear that it is possible to fix it, the question is whether it is necessary to spend time on it. When I saw your 2.75 md model which gives 0.66 LB, and compared it to my 0.67 lb ensembles thirty or forty times larger, I realized that now is not the time to focus on that, it is better to concentrate on building really good singles models and only later will it be seen whether it will be necessary. But I decided to share the result. Good luck",
    "2201194": "General updates, if you look at the discussions: people are reporting that quantization is slowing things down. From my experience, you need to test each model separately, for different models the indicators are different. You can also add the option keras_model_converter.target_spec.supported_types = [tf.float16] to reduce accuracy and speed drop issues. Search for more details in the discussion on the word Quantization. Thank you all for your comments!"
  },
  "source": "meta"
}