{
  "id": 402258,
  "title": "Inference time error and reduction",
  "url": "/competitions/asl-signs/discussion/402258",
  "author_name": "Leon Sakakibara",
  "post_date": "2023-04-17T16:49:09.187000",
  "votes": 2,
  "comment_count": 3,
  "views": 0,
  "content": "<p>After about an hour of submitting my notebook, it fails with a \"Submission Scoring error\".</p>\n<p>I think it was probably kicked because the inference time exceeded the 100ms limit, but I measured it on my notebook and it is about 80ms.</p>\n<p>Do you guys know anything about this difference in seconds? Also, do you have any tips to shorten the inference time?</p>",
  "messages": [
    {
      "id": 2225304,
      "postDate": "2023-04-18T04:17:40.113Z",
      "content": "<p>You can use post-training model quantization to reduce the inference time of the final TFLite model; for example, by using 16-bit or 8-bit integers to store the weights of the model. See this article: <a href=\"https://www.tensorflow.org/model_optimization/guide/quantization/post_training\" target=\"_blank\">Post-training quantization</a></p>",
      "rawMarkdown": "You can use post-training model quantization to reduce the inference time of the final TFLite model; for example, by using 16-bit or 8-bit integers to store the weights of the model. See this article: [Post-training quantization](https://www.tensorflow.org/model_optimization/guide/quantization/post_training)",
      "votes": 1,
      "replies": [
        {
          "id": 2225655,
          "postDate": "2023-04-18T10:34:14.753Z",
          "content": "<p>That is good information!<br>\nI will give it a try.<br>\nThank you very much.</p>",
          "rawMarkdown": "That is good information!\nI will give it a try.\nThank you very much.",
          "votes": 1
        }
      ]
    },
    {
      "id": 2224782,
      "postDate": "2023-04-17T16:49:09.187Z",
      "content": "<p>After about an hour of submitting my notebook, it fails with a \"Submission Scoring error\".</p>\n<p>I think it was probably kicked because the inference time exceeded the 100ms limit, but I measured it on my notebook and it is about 80ms.</p>\n<p>Do you guys know anything about this difference in seconds? Also, do you have any tips to shorten the inference time?</p>",
      "rawMarkdown": "After about an hour of submitting my notebook, it fails with a \"Submission Scoring error\".\n\nI think it was probably kicked because the inference time exceeded the 100ms limit, but I measured it on my notebook and it is about 80ms.\n\nDo you guys know anything about this difference in seconds? Also, do you have any tips to shorten the inference time?",
      "votes": 2
    },
    {
      "id": 2225244,
      "postDate": "2023-04-18T02:41:55.303Z",
      "rawMarkdown": "",
      "isDeleted": true
    }
  ],
  "comments": [
    {
      "id": 2225304,
      "author_name": "Hassan Abedi",
      "author_url": "",
      "post_date": "2023-04-18T04:17:40.113000",
      "content": "<p>You can use post-training model quantization to reduce the inference time of the final TFLite model; for example, by using 16-bit or 8-bit integers to store the weights of the model. See this article: <a href=\"https://www.tensorflow.org/model_optimization/guide/quantization/post_training\" target=\"_blank\">Post-training quantization</a></p>",
      "votes": 1,
      "replies": [
        {
          "id": 2225655,
          "author_name": "Leon Sakakibara",
          "author_url": "",
          "post_date": "2023-04-18T10:34:14.753000",
          "content": "<p>That is good information!<br>\nI will give it a try.<br>\nThank you very much.</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 2225244,
      "author_name": "",
      "author_url": "",
      "post_date": "2023-04-18T02:41:55.303000",
      "content": "",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "2225304": "You can use post-training model quantization to reduce the inference time of the final TFLite model; for example, by using 16-bit or 8-bit integers to store the weights of the model. See this article: [Post-training quantization](https://www.tensorflow.org/model_optimization/guide/quantization/post_training)",
    "2224782": "After about an hour of submitting my notebook, it fails with a \"Submission Scoring error\".\n\nI think it was probably kicked because the inference time exceeded the 100ms limit, but I measured it on my notebook and it is about 80ms.\n\nDo you guys know anything about this difference in seconds? Also, do you have any tips to shorten the inference time?",
    "2225244": ""
  }
}