{
  "id": 392840,
  "title": "[deprecated] how to use tflite tool profiler to check tflite memory?",
  "url": "/competitions/asl-signs/discussion/392840",
  "author_name": "",
  "post_date": "2023-03-07T03:26:53.328559Z",
  "votes": 4,
  "comment_count": 10,
  "views": 0,
  "content": "<h2>** requirement has been changed to \" 40 MB of storage space \" **</h2>\n<h2>(it was previously \"40 MB of memory\")</h2>\n<p>Following the instruction at:</p>\n<p><a href=\"https://github.com/tensorflow/tensorflow/tree/master/tensorflow/lite/tools/benchmark\" target=\"_blank\">https://github.com/tensorflow/tensorflow/tree/master/tensorflow/lite/tools/benchmark</a><br>\n<a href=\"https://www.tensorflow.org/lite/performance/measurement\" target=\"_blank\">https://www.tensorflow.org/lite/performance/measurement</a></p>\n<p>to input worst case xyz: <br>\n\"By default, the tool will use randomized data for model inputs. The following parameters allow users to specify customized input values to the model when running the benchmark tool…\",<br>\nsee <a href=\"https://github.com/tensorflow/tensorflow/blob/master/tensorflow/lite/tools/benchmark/README.md#model-input-parameters\" target=\"_blank\">https://github.com/tensorflow/tensorflow/blob/master/tensorflow/lite/tools/benchmark/README.md#model-input-parameters</a></p>\n<p>if you want to test real values (instead of random values:<br>\nxyz2.tofile('xyz2.binary')</p>\n<pre><code>./benchmark_model --graph='/&lt;path&gt;/transformer-pool-2c-512-80.tflite' \\\n--num_threads=4  \\\n--enable_op_profiling=true \\\n--input_layer='input1' \\\n--input_layer_shape='105,543,3' \\\n--report_peak_memory_footprint=true \n\noptional\n--input_layer_value_files='input1:/&lt;path&gt;/xyz2.binary' \\\n</code></pre>\n<pre><code>INFO: Created TensorFlow Lite XNNPACK delegate for CPU.\nINFO: The input model file size (MB): 2.05214\nINFO: Initialized session in 19.372ms.\nINFO: Running benchmark for at least 1 iterations and at least 0.5 seconds but terminate if exceeding 150 seconds.\nINFO: count=30 first=38626 curr=15592 min=15588 max=38626 avg=16676.5 std=4099\n\nINFO: Running benchmark for at least 50 iterations and at least 1 seconds but terminate if exceeding 150 seconds.\nINFO: count=58 first=15712 curr=46587 min=15585 max=46587 avg=17032.2 std=4617\n\nINFO: Inference timings in us: Init: 19372, First inference: 38626, Warmup (avg): 16676.5, Inference (avg): 17032.2\nINFO: Note: as the benchmark tool itself affects memory footprint, the following is only APPROXIMATE to the actual memory footprint of the model at runtime. Take the information at your discretion.\nINFO: Memory footprint delta from the start of the tool (MB): init=0 overall=23.0625\nINFO: Overall peak memory footprint (MB) via periodic monitoring: 34.3047\nINFO: Memory status at the end of exeution:\nINFO: - VmRSS              : 34 MB\nINFO: + RssAnnon           : 25 MB\nINFO: + RssFile + RssShmem : 9 MB\n....\n</code></pre>\n<p><br>\n</p>\n<p>if you are brave enough, you can setup Android Studio and use memory profiler</p>\n<hr>\n<p>on a side note. it might be asier if kaggle can provide us an evaluation.py script,which include load the tflite model, use a csv and ouput a score. then we can debug at our own</p>",
  "messages": [
    {
      "id": "2171716",
      "postDate": "03/07/2023 03:26:53",
      "content": "<h2>** requirement has been changed to \" 40 MB of storage space \" **</h2>\n<h2>(it was previously \"40 MB of memory\")</h2>\n<p>Following the instruction at:</p>\n<p><a href=\"https://github.com/tensorflow/tensorflow/tree/master/tensorflow/lite/tools/benchmark\" target=\"_blank\">https://github.com/tensorflow/tensorflow/tree/master/tensorflow/lite/tools/benchmark</a><br>\n<a href=\"https://www.tensorflow.org/lite/performance/measurement\" target=\"_blank\">https://www.tensorflow.org/lite/performance/measurement</a></p>\n<p>to input worst case xyz: <br>\n\"By default, the tool will use randomized data for model inputs. The following parameters allow users to specify customized input values to the model when running the benchmark tool…\",<br>\nsee <a href=\"https://github.com/tensorflow/tensorflow/blob/master/tensorflow/lite/tools/benchmark/README.md#model-input-parameters\" target=\"_blank\">https://github.com/tensorflow/tensorflow/blob/master/tensorflow/lite/tools/benchmark/README.md#model-input-parameters</a></p>\n<p>if you want to test real values (instead of random values:<br>\nxyz2.tofile('xyz2.binary')</p>\n<pre><code>./benchmark_model --graph='/&lt;path&gt;/transformer-pool-2c-512-80.tflite' \\\n--num_threads=4  \\\n--enable_op_profiling=true \\\n--input_layer='input1' \\\n--input_layer_shape='105,543,3' \\\n--report_peak_memory_footprint=true \n\noptional\n--input_layer_value_files='input1:/&lt;path&gt;/xyz2.binary' \\\n</code></pre>\n<pre><code>INFO: Created TensorFlow Lite XNNPACK delegate for CPU.\nINFO: The input model file size (MB): 2.05214\nINFO: Initialized session in 19.372ms.\nINFO: Running benchmark for at least 1 iterations and at least 0.5 seconds but terminate if exceeding 150 seconds.\nINFO: count=30 first=38626 curr=15592 min=15588 max=38626 avg=16676.5 std=4099\n\nINFO: Running benchmark for at least 50 iterations and at least 1 seconds but terminate if exceeding 150 seconds.\nINFO: count=58 first=15712 curr=46587 min=15585 max=46587 avg=17032.2 std=4617\n\nINFO: Inference timings in us: Init: 19372, First inference: 38626, Warmup (avg): 16676.5, Inference (avg): 17032.2\nINFO: Note: as the benchmark tool itself affects memory footprint, the following is only APPROXIMATE to the actual memory footprint of the model at runtime. Take the information at your discretion.\nINFO: Memory footprint delta from the start of the tool (MB): init=0 overall=23.0625\nINFO: Overall peak memory footprint (MB) via periodic monitoring: 34.3047\nINFO: Memory status at the end of exeution:\nINFO: - VmRSS              : 34 MB\nINFO: + RssAnnon           : 25 MB\nINFO: + RssFile + RssShmem : 9 MB\n....\n</code></pre>\n<p><br>\n</p>\n<p>if you are brave enough, you can setup Android Studio and use memory profiler</p>\n<hr>\n<p>on a side note. it might be asier if kaggle can provide us an evaluation.py script,which include load the tflite model, use a csv and ouput a score. then we can debug at our own</p>",
      "rawMarkdown": "##  ** requirement has been changed to \" 40 MB of storage space \" **\n(it was previously \"40 MB of memory\")\n--- \n\nFollowing the instruction at:\n\nhttps://github.com/tensorflow/tensorflow/tree/master/tensorflow/lite/tools/benchmark\nhttps://www.tensorflow.org/lite/performance/measurement\n\nto input worst case xyz: \n\"By default, the tool will use randomized data for model inputs. The following parameters allow users to specify customized input values to the model when running the benchmark tool...\",\nsee https://github.com/tensorflow/tensorflow/blob/master/tensorflow/lite/tools/benchmark/README.md#model-input-parameters\n\nif you want to test real values (instead of random values:\nxyz2.tofile('xyz2.binary')\n\n```\n./benchmark_model --graph='/<path>/transformer-pool-2c-512-80.tflite' \\\n--num_threads=4  \\\n--enable_op_profiling=true \\\n--input_layer='input1' \\\n--input_layer_shape='105,543,3' \\\n--report_peak_memory_footprint=true \n\noptional\n--input_layer_value_files='input1:/<path>/xyz2.binary' \\\n\n\n```\n\n\n\n```\nINFO: Created TensorFlow Lite XNNPACK delegate for CPU.\nINFO: The input model file size (MB): 2.05214\nINFO: Initialized session in 19.372ms.\nINFO: Running benchmark for at least 1 iterations and at least 0.5 seconds but terminate if exceeding 150 seconds.\nINFO: count=30 first=38626 curr=15592 min=15588 max=38626 avg=16676.5 std=4099\n\nINFO: Running benchmark for at least 50 iterations and at least 1 seconds but terminate if exceeding 150 seconds.\nINFO: count=58 first=15712 curr=46587 min=15585 max=46587 avg=17032.2 std=4617\n\nINFO: Inference timings in us: Init: 19372, First inference: 38626, Warmup (avg): 16676.5, Inference (avg): 17032.2\nINFO: Note: as the benchmark tool itself affects memory footprint, the following is only APPROXIMATE to the actual memory footprint of the model at runtime. Take the information at your discretion.\nINFO: Memory footprint delta from the start of the tool (MB): init=0 overall=23.0625\nINFO: Overall peak memory footprint (MB) via periodic monitoring: 34.3047\nINFO: Memory status at the end of exeution:\nINFO: - VmRSS              : 34 MB\nINFO: + RssAnnon           : 25 MB\nINFO: + RssFile + RssShmem : 9 MB\n....\n\n```\n\n~~but i got notebbok out of memory for my submission (after running for 20 min), why????~~\n~~https://www.kaggle.com/code/hengck23/pytorch-transformer-solution~~\n\nif you are brave enough, you can setup Android Studio and use memory profiler\n\n---\n\non a side note. it might be asier if kaggle can provide us an evaluation.py script,which include load the tflite model, use a csv and ouput a score. then we can debug at our own",
      "votes": null
    },
    {
      "id": "2171722",
      "postDate": "03/07/2023 03:37:48",
      "content": "<p>some tflite tool:<br>\n<a href=\"https://github.com/eliberis/tflite-tools\" target=\"_blank\">https://github.com/eliberis/tflite-tools</a><br>\n<a href=\"https://www.tensorflow.org/lite/performance/measurement\" target=\"_blank\">https://www.tensorflow.org/lite/performance/measurement</a></p>\n<p><a href=\"https://stackoverflow.com/questions/67542576/what-do-the-tensorflow-lite-benchmark-tools-outputs-mean\" target=\"_blank\">https://stackoverflow.com/questions/67542576/what-do-the-tensorflow-lite-benchmark-tools-outputs-mean</a></p>\n<pre><code>\"Memory usage during initialization time\" - The memory usage difference before/after creating a TFLite interpreter object in C++ and loading the model.\n\n\"Overall memory usage\" - The memory usage difference before creating a TFLite interpreter object in C++ and loading the model and after running the benchmarking tasks.\n</code></pre>\n<pre><code>VmRSS (/proc/meminfo) is broken up into:\n\n VmRSS                       size of memory portions. It contains the three\n                             following parts (VmRSS = RssAnon + RssFile + RssShmem)\n RssAnon                     size of resident anonymous memory\n RssFile                     size of resident file mappings\n RssShmem                    size of resident shmem memory (includes SysV shm,\n                             mapping of tmpfs and shared anonymous mappings)\n</code></pre>",
      "rawMarkdown": "some tflite tool:\nhttps://github.com/eliberis/tflite-tools\nhttps://www.tensorflow.org/lite/performance/measurement\n\n\nhttps://stackoverflow.com/questions/67542576/what-do-the-tensorflow-lite-benchmark-tools-outputs-mean\n```\n\"Memory usage during initialization time\" - The memory usage difference before/after creating a TFLite interpreter object in C++ and loading the model.\n\n\"Overall memory usage\" - The memory usage difference before creating a TFLite interpreter object in C++ and loading the model and after running the benchmarking tasks.\n\n```\n```\nVmRSS (/proc/meminfo) is broken up into:\n\n VmRSS                       size of memory portions. It contains the three\n                             following parts (VmRSS = RssAnon + RssFile + RssShmem)\n RssAnon                     size of resident anonymous memory\n RssFile                     size of resident file mappings\n RssShmem                    size of resident shmem memory (includes SysV shm,\n                             mapping of tmpfs and shared anonymous mappings)\n```",
      "votes": null
    },
    {
      "id": "2172299",
      "postDate": "03/07/2023 12:38:48",
      "content": "<p></p>\n<p></p>\n<p></p>",
      "rawMarkdown": "~~i am suspecting \"Notebook Out of Memory\" is not model inference memory error.~~\n\n~~i recalled that there is once when kaggle evaluation script gave memory error becuase of invalid values, etc in the prediction.~~\n\n~~anyone knows if tflite get_signature_runner() can return invalid values?~~",
      "votes": null
    },
    {
      "id": "2172729",
      "postDate": "03/07/2023 18:25:09",
      "content": "<p>this tool print the allocated tensor:<br>\n<a href=\"https://github.com/PeteBlackerThe3rd/tflite_analyser\" target=\"_blank\">https://github.com/PeteBlackerThe3rd/tflite_analyser</a><br>\nupdate latest op codes from<br>\n<a href=\"https://github.com/zhenhuaw-me/tflite\" target=\"_blank\">https://github.com/zhenhuaw-me/tflite</a></p>\n<pre><code>Done.\nAnalysing graph[0] - b'main'\n\nDynamic tensors requiring memory allocation.\n\nThis model contains 192 operations.\n\n    Tensor Name                                                 (Size Bytes)  [Generating Op]  [Final Use Op]\nShape                                                                      (        24)       [   0]         [  13]\nGatherNd                                                                   (         8)       [   1]         [   5]\nMinimum                                                                    (         8)       [   2]         [   7]\nMinimum_1                                                                  (         8)       [   3]         [  10]\nadd_1                                                                      (         8)       [   4]         [   7]\nadd_2                                                                      (         8)       [   5]         [  10]\nLess_1                                                                     (        -1)       [   6]         [   7]\nSelectV2_1                                                                 (         8)       [   7]         [   8]\nSparseToDense                                                              (        24)       [   8]         [  14]\nLess_2                                                                     (        -1)       [   9]         [  10]\nSelectV2_2                                                                 (         8)       [  10]         [  11]\nSparseToDense_1                                                            (        24)       [  11]         [  13]\n</code></pre>",
      "rawMarkdown": "this tool print the allocated tensor:\nhttps://github.com/PeteBlackerThe3rd/tflite_analyser\nupdate latest op codes from\nhttps://github.com/zhenhuaw-me/tflite\n\n```\nDone.\nAnalysing graph[0] - b'main'\n\nDynamic tensors requiring memory allocation.\n\nThis model contains 192 operations.\n\n    Tensor Name                                                 (Size Bytes)  [Generating Op]  [Final Use Op]\nShape                                                                      (        24)       [   0]         [  13]\nGatherNd                                                                   (         8)       [   1]         [   5]\nMinimum                                                                    (         8)       [   2]         [   7]\nMinimum_1                                                                  (         8)       [   3]         [  10]\nadd_1                                                                      (         8)       [   4]         [   7]\nadd_2                                                                      (         8)       [   5]         [  10]\nLess_1                                                                     (        -1)       [   6]         [   7]\nSelectV2_1                                                                 (         8)       [   7]         [   8]\nSparseToDense                                                              (        24)       [   8]         [  14]\nLess_2                                                                     (        -1)       [   9]         [  10]\nSelectV2_2                                                                 (         8)       [  10]         [  11]\nSparseToDense_1                                                            (        24)       [  11]         [  13]\n\n\n```",
      "votes": null
    },
    {
      "id": "2172756",
      "postDate": "03/07/2023 18:49:07",
      "content": "<p>You can actually just call <code>os.path.getsize()</code> on the unzipped copy of your tflite checkpoint. I went through a similar journey of looking up dedicated tools, but it turns out that because <a href=\"https://www.tensorflow.org/lite/guide#1_generate_a_tensorflow_lite_model\" target=\"_blank\">tflite uses</a> the <a href=\"https://google.github.io/flatbuffers/\" target=\"_blank\">FlatBuffer format</a> the relevant size in memory is essentially the file size on disk. I'll add a note to the evaluation page to clarify this.</p>",
      "rawMarkdown": "You can actually just call `os.path.getsize()` on the unzipped copy of your tflite checkpoint. I went through a similar journey of looking up dedicated tools, but it turns out that because [tflite uses](https://www.tensorflow.org/lite/guide#1_generate_a_tensorflow_lite_model) the [FlatBuffer format](https://google.github.io/flatbuffers/) the relevant size in memory is essentially the file size on disk. I'll add a note to the evaluation page to clarify this.",
      "votes": null
    },
    {
      "id": "2172759",
      "postDate": "03/07/2023 18:53:27",
      "content": "<p>PS - I'll see if we can publish a copy of the evaluation.py script for local testing. I don't expect I'll be able to get to that until later this week, however.</p>",
      "rawMarkdown": "PS - I'll see if we can publish a copy of the evaluation.py script for local testing. I don't expect I'll be able to get to that until later this week, however.",
      "votes": null
    },
    {
      "id": "2172795",
      "postDate": "03/07/2023 19:38:14",
      "content": "<p>\"Your model must also require less than 40 MB in memory\"</p>\n<p><a href=\"https://www.kaggle.com/sohier\" target=\"_blank\">@sohier</a> </p>\n<p>just to clarify: 40 MB refer disk storage space or runtime memory RAM?<br>\n(i always though it is runtime memory RAM)</p>\n<p>Thanks for your comments!</p>",
      "rawMarkdown": "\"Your model must also require less than 40 MB in memory\"\n\n@sohier \n\njust to clarify: 40 MB refer disk storage space or runtime memory RAM?\n(i always though it is runtime memory RAM)\n\nThanks for your comments!",
      "votes": null
    },
    {
      "id": "2172857",
      "postDate": "03/07/2023 21:21:14",
      "content": "<p>For tflite checkpoint files/FlatBuffers the disk storage space and runtime memory RAM are close enough to equivalent for our purposes.</p>",
      "rawMarkdown": "For tflite checkpoint files/FlatBuffers the disk storage space and runtime memory RAM are close enough to equivalent for our purposes.",
      "votes": null
    },
    {
      "id": "2172929",
      "postDate": "03/07/2023 23:53:39",
      "content": "<p>I've gone ahead and updated the evaluation tab language to just talk about size on disk for simplicity.</p>",
      "rawMarkdown": "I've gone ahead and updated the evaluation tab language to just talk about size on disk for simplicity.",
      "votes": null
    },
    {
      "id": "2174694",
      "postDate": "03/09/2023 09:56:32",
      "content": "<p><a href=\"https://blog.tensorflow.org/2020/10/optimizing-tensorflow-lite-runtime.html\" target=\"_blank\">https://blog.tensorflow.org/2020/10/optimizing-tensorflow-lite-runtime.html</a></p>\n<p>understand how runtime memory is allocated in tflite</p>",
      "rawMarkdown": "https://blog.tensorflow.org/2020/10/optimizing-tensorflow-lite-runtime.html\n\nunderstand how runtime memory is allocated in tflite",
      "votes": null
    },
    {
      "id": "2174710",
      "postDate": "03/09/2023 10:13:12",
      "content": "<p>Thanks for the update!</p>",
      "rawMarkdown": "Thanks for the update!",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 2171722,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "03/07/2023 03:37:48",
      "content": "<p>some tflite tool:<br>\n<a href=\"https://github.com/eliberis/tflite-tools\" target=\"_blank\">https://github.com/eliberis/tflite-tools</a><br>\n<a href=\"https://www.tensorflow.org/lite/performance/measurement\" target=\"_blank\">https://www.tensorflow.org/lite/performance/measurement</a></p>\n<p><a href=\"https://stackoverflow.com/questions/67542576/what-do-the-tensorflow-lite-benchmark-tools-outputs-mean\" target=\"_blank\">https://stackoverflow.com/questions/67542576/what-do-the-tensorflow-lite-benchmark-tools-outputs-mean</a></p>\n<pre><code>\"Memory usage during initialization time\" - The memory usage difference before/after creating a TFLite interpreter object in C++ and loading the model.\n\n\"Overall memory usage\" - The memory usage difference before creating a TFLite interpreter object in C++ and loading the model and after running the benchmarking tasks.\n</code></pre>\n<pre><code>VmRSS (/proc/meminfo) is broken up into:\n\n VmRSS                       size of memory portions. It contains the three\n                             following parts (VmRSS = RssAnon + RssFile + RssShmem)\n RssAnon                     size of resident anonymous memory\n RssFile                     size of resident file mappings\n RssShmem                    size of resident shmem memory (includes SysV shm,\n                             mapping of tmpfs and shared anonymous mappings)\n</code></pre>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2172299,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "03/07/2023 12:38:48",
      "content": "<p></p>\n<p></p>\n<p></p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2172729,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "03/07/2023 18:25:09",
      "content": "<p>this tool print the allocated tensor:<br>\n<a href=\"https://github.com/PeteBlackerThe3rd/tflite_analyser\" target=\"_blank\">https://github.com/PeteBlackerThe3rd/tflite_analyser</a><br>\nupdate latest op codes from<br>\n<a href=\"https://github.com/zhenhuaw-me/tflite\" target=\"_blank\">https://github.com/zhenhuaw-me/tflite</a></p>\n<pre><code>Done.\nAnalysing graph[0] - b'main'\n\nDynamic tensors requiring memory allocation.\n\nThis model contains 192 operations.\n\n    Tensor Name                                                 (Size Bytes)  [Generating Op]  [Final Use Op]\nShape                                                                      (        24)       [   0]         [  13]\nGatherNd                                                                   (         8)       [   1]         [   5]\nMinimum                                                                    (         8)       [   2]         [   7]\nMinimum_1                                                                  (         8)       [   3]         [  10]\nadd_1                                                                      (         8)       [   4]         [   7]\nadd_2                                                                      (         8)       [   5]         [  10]\nLess_1                                                                     (        -1)       [   6]         [   7]\nSelectV2_1                                                                 (         8)       [   7]         [   8]\nSparseToDense                                                              (        24)       [   8]         [  14]\nLess_2                                                                     (        -1)       [   9]         [  10]\nSelectV2_2                                                                 (         8)       [  10]         [  11]\nSparseToDense_1                                                            (        24)       [  11]         [  13]\n</code></pre>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2172756,
      "author_name": "sohier",
      "author_url": "",
      "post_date": "03/07/2023 18:49:07",
      "content": "<p>You can actually just call <code>os.path.getsize()</code> on the unzipped copy of your tflite checkpoint. I went through a similar journey of looking up dedicated tools, but it turns out that because <a href=\"https://www.tensorflow.org/lite/guide#1_generate_a_tensorflow_lite_model\" target=\"_blank\">tflite uses</a> the <a href=\"https://google.github.io/flatbuffers/\" target=\"_blank\">FlatBuffer format</a> the relevant size in memory is essentially the file size on disk. I'll add a note to the evaluation page to clarify this.</p>",
      "votes": null,
      "replies": [
        {
          "id": 2172759,
          "author_name": "sohier",
          "author_url": "",
          "post_date": "03/07/2023 18:53:27",
          "content": "<p>PS - I'll see if we can publish a copy of the evaluation.py script for local testing. I don't expect I'll be able to get to that until later this week, however.</p>",
          "votes": null,
          "replies": [
            {
              "id": 2172795,
              "author_name": "hengck23",
              "author_url": "",
              "post_date": "03/07/2023 19:38:14",
              "content": "<p>\"Your model must also require less than 40 MB in memory\"</p>\n<p><a href=\"https://www.kaggle.com/sohier\" target=\"_blank\">@sohier</a> </p>\n<p>just to clarify: 40 MB refer disk storage space or runtime memory RAM?<br>\n(i always though it is runtime memory RAM)</p>\n<p>Thanks for your comments!</p>",
              "votes": null,
              "replies": [
                {
                  "id": 2172857,
                  "author_name": "sohier",
                  "author_url": "",
                  "post_date": "03/07/2023 21:21:14",
                  "content": "<p>For tflite checkpoint files/FlatBuffers the disk storage space and runtime memory RAM are close enough to equivalent for our purposes.</p>",
                  "votes": null,
                  "replies": [
                    {
                      "id": 2172929,
                      "author_name": "sohier",
                      "author_url": "",
                      "post_date": "03/07/2023 23:53:39",
                      "content": "<p>I've gone ahead and updated the evaluation tab language to just talk about size on disk for simplicity.</p>",
                      "votes": null,
                      "replies": [
                        {
                          "id": 2174710,
                          "author_name": "hengck23",
                          "author_url": "",
                          "post_date": "03/09/2023 10:13:12",
                          "content": "<p>Thanks for the update!</p>",
                          "votes": null,
                          "replies": []
                        }
                      ]
                    }
                  ]
                }
              ]
            }
          ]
        }
      ]
    },
    {
      "id": 2174694,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "03/09/2023 09:56:32",
      "content": "<p><a href=\"https://blog.tensorflow.org/2020/10/optimizing-tensorflow-lite-runtime.html\" target=\"_blank\">https://blog.tensorflow.org/2020/10/optimizing-tensorflow-lite-runtime.html</a></p>\n<p>understand how runtime memory is allocated in tflite</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "2171716": "##  ** requirement has been changed to \" 40 MB of storage space \" **\n(it was previously \"40 MB of memory\")\n--- \n\nFollowing the instruction at:\n\nhttps://github.com/tensorflow/tensorflow/tree/master/tensorflow/lite/tools/benchmark\nhttps://www.tensorflow.org/lite/performance/measurement\n\nto input worst case xyz: \n\"By default, the tool will use randomized data for model inputs. The following parameters allow users to specify customized input values to the model when running the benchmark tool...\",\nsee https://github.com/tensorflow/tensorflow/blob/master/tensorflow/lite/tools/benchmark/README.md#model-input-parameters\n\nif you want to test real values (instead of random values:\nxyz2.tofile('xyz2.binary')\n\n```\n./benchmark_model --graph='/<path>/transformer-pool-2c-512-80.tflite' \\\n--num_threads=4  \\\n--enable_op_profiling=true \\\n--input_layer='input1' \\\n--input_layer_shape='105,543,3' \\\n--report_peak_memory_footprint=true \n\noptional\n--input_layer_value_files='input1:/<path>/xyz2.binary' \\\n\n\n```\n\n\n\n```\nINFO: Created TensorFlow Lite XNNPACK delegate for CPU.\nINFO: The input model file size (MB): 2.05214\nINFO: Initialized session in 19.372ms.\nINFO: Running benchmark for at least 1 iterations and at least 0.5 seconds but terminate if exceeding 150 seconds.\nINFO: count=30 first=38626 curr=15592 min=15588 max=38626 avg=16676.5 std=4099\n\nINFO: Running benchmark for at least 50 iterations and at least 1 seconds but terminate if exceeding 150 seconds.\nINFO: count=58 first=15712 curr=46587 min=15585 max=46587 avg=17032.2 std=4617\n\nINFO: Inference timings in us: Init: 19372, First inference: 38626, Warmup (avg): 16676.5, Inference (avg): 17032.2\nINFO: Note: as the benchmark tool itself affects memory footprint, the following is only APPROXIMATE to the actual memory footprint of the model at runtime. Take the information at your discretion.\nINFO: Memory footprint delta from the start of the tool (MB): init=0 overall=23.0625\nINFO: Overall peak memory footprint (MB) via periodic monitoring: 34.3047\nINFO: Memory status at the end of exeution:\nINFO: - VmRSS              : 34 MB\nINFO: + RssAnnon           : 25 MB\nINFO: + RssFile + RssShmem : 9 MB\n....\n\n```\n\n~~but i got notebbok out of memory for my submission (after running for 20 min), why????~~\n~~https://www.kaggle.com/code/hengck23/pytorch-transformer-solution~~\n\nif you are brave enough, you can setup Android Studio and use memory profiler\n\n---\n\non a side note. it might be asier if kaggle can provide us an evaluation.py script,which include load the tflite model, use a csv and ouput a score. then we can debug at our own",
    "2171722": "some tflite tool:\nhttps://github.com/eliberis/tflite-tools\nhttps://www.tensorflow.org/lite/performance/measurement\n\n\nhttps://stackoverflow.com/questions/67542576/what-do-the-tensorflow-lite-benchmark-tools-outputs-mean\n```\n\"Memory usage during initialization time\" - The memory usage difference before/after creating a TFLite interpreter object in C++ and loading the model.\n\n\"Overall memory usage\" - The memory usage difference before creating a TFLite interpreter object in C++ and loading the model and after running the benchmarking tasks.\n\n```\n```\nVmRSS (/proc/meminfo) is broken up into:\n\n VmRSS                       size of memory portions. It contains the three\n                             following parts (VmRSS = RssAnon + RssFile + RssShmem)\n RssAnon                     size of resident anonymous memory\n RssFile                     size of resident file mappings\n RssShmem                    size of resident shmem memory (includes SysV shm,\n                             mapping of tmpfs and shared anonymous mappings)\n```",
    "2172299": "~~i am suspecting \"Notebook Out of Memory\" is not model inference memory error.~~\n\n~~i recalled that there is once when kaggle evaluation script gave memory error becuase of invalid values, etc in the prediction.~~\n\n~~anyone knows if tflite get_signature_runner() can return invalid values?~~",
    "2172729": "this tool print the allocated tensor:\nhttps://github.com/PeteBlackerThe3rd/tflite_analyser\nupdate latest op codes from\nhttps://github.com/zhenhuaw-me/tflite\n\n```\nDone.\nAnalysing graph[0] - b'main'\n\nDynamic tensors requiring memory allocation.\n\nThis model contains 192 operations.\n\n    Tensor Name                                                 (Size Bytes)  [Generating Op]  [Final Use Op]\nShape                                                                      (        24)       [   0]         [  13]\nGatherNd                                                                   (         8)       [   1]         [   5]\nMinimum                                                                    (         8)       [   2]         [   7]\nMinimum_1                                                                  (         8)       [   3]         [  10]\nadd_1                                                                      (         8)       [   4]         [   7]\nadd_2                                                                      (         8)       [   5]         [  10]\nLess_1                                                                     (        -1)       [   6]         [   7]\nSelectV2_1                                                                 (         8)       [   7]         [   8]\nSparseToDense                                                              (        24)       [   8]         [  14]\nLess_2                                                                     (        -1)       [   9]         [  10]\nSelectV2_2                                                                 (         8)       [  10]         [  11]\nSparseToDense_1                                                            (        24)       [  11]         [  13]\n\n\n```",
    "2172756": "You can actually just call `os.path.getsize()` on the unzipped copy of your tflite checkpoint. I went through a similar journey of looking up dedicated tools, but it turns out that because [tflite uses](https://www.tensorflow.org/lite/guide#1_generate_a_tensorflow_lite_model) the [FlatBuffer format](https://google.github.io/flatbuffers/) the relevant size in memory is essentially the file size on disk. I'll add a note to the evaluation page to clarify this.",
    "2172759": "PS - I'll see if we can publish a copy of the evaluation.py script for local testing. I don't expect I'll be able to get to that until later this week, however.",
    "2172795": "\"Your model must also require less than 40 MB in memory\"\n\n@sohier \n\njust to clarify: 40 MB refer disk storage space or runtime memory RAM?\n(i always though it is runtime memory RAM)\n\nThanks for your comments!",
    "2172857": "For tflite checkpoint files/FlatBuffers the disk storage space and runtime memory RAM are close enough to equivalent for our purposes.",
    "2172929": "I've gone ahead and updated the evaluation tab language to just talk about size on disk for simplicity.",
    "2174694": "https://blog.tensorflow.org/2020/10/optimizing-tensorflow-lite-runtime.html\n\nunderstand how runtime memory is allocated in tflite",
    "2174710": "Thanks for the update!"
  },
  "source": "meta"
}