{
  "id": 498051,
  "title": "Why the 'mp' parameter does not increase code speed in LightGlueMatcher.",
  "url": "/competitions/image-matching-challenge-2024/discussion/498051",
  "author_name": "",
  "post_date": "2024-04-26T17:57:19.209994Z",
  "votes": 2,
  "comment_count": 2,
  "views": 0,
  "content": "<p>According to github (mp: Enable mixed precision inference. Default: False (off)) which should increase speed and decrease accuracy if set True. However, I made measurements and the time was always the same</p>",
  "messages": [
    {
      "id": "2777569",
      "postDate": "04/26/2024 17:57:19",
      "content": "<p>According to github (mp: Enable mixed precision inference. Default: False (off)) which should increase speed and decrease accuracy if set True. However, I made measurements and the time was always the same</p>",
      "rawMarkdown": "According to github (mp: Enable mixed precision inference. Default: False (off)) which should increase speed and decrease accuracy if set True. However, I made measurements and the time was always the same",
      "votes": null
    },
    {
      "id": "2777775",
      "postDate": "04/26/2024 19:49:13",
      "content": "<p>No idea, but using float16 speeds up for sure.</p>",
      "rawMarkdown": "No idea, but using float16 speeds up for sure.",
      "votes": null
    },
    {
      "id": "2804046",
      "postDate": "05/09/2024 18:54:05",
      "content": "<p>Not sure, but this might be related to this conversation: <a href=\"https://www.kaggle.com/competitions/ai-mathematical-olympiad-prize/discussion/500529#2796083\" target=\"_blank\">https://www.kaggle.com/competitions/ai-mathematical-olympiad-prize/discussion/500529#2796083</a><br>\nBasically, T4 and P100 (generally, all GPUs with CUDA capabilities less than 8.0) does not have true support of certain datatypes (like BFloat16, FP8, etc.), so the computation is actually done in FP32 even if the lower-bit type is used. FP16 is fine though.</p>",
      "rawMarkdown": "Not sure, but this might be related to this conversation: https://www.kaggle.com/competitions/ai-mathematical-olympiad-prize/discussion/500529#2796083\nBasically, T4 and P100 (generally, all GPUs with CUDA capabilities less than 8.0) does not have true support of certain datatypes (like BFloat16, FP8, etc.), so the computation is actually done in FP32 even if the lower-bit type is used. FP16 is fine though.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 2777775,
      "author_name": "oldufo",
      "author_url": "",
      "post_date": "04/26/2024 19:49:13",
      "content": "<p>No idea, but using float16 speeds up for sure.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2804046,
      "author_name": "chankhavu",
      "author_url": "",
      "post_date": "05/09/2024 18:54:05",
      "content": "<p>Not sure, but this might be related to this conversation: <a href=\"https://www.kaggle.com/competitions/ai-mathematical-olympiad-prize/discussion/500529#2796083\" target=\"_blank\">https://www.kaggle.com/competitions/ai-mathematical-olympiad-prize/discussion/500529#2796083</a><br>\nBasically, T4 and P100 (generally, all GPUs with CUDA capabilities less than 8.0) does not have true support of certain datatypes (like BFloat16, FP8, etc.), so the computation is actually done in FP32 even if the lower-bit type is used. FP16 is fine though.</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "2777569": "According to github (mp: Enable mixed precision inference. Default: False (off)) which should increase speed and decrease accuracy if set True. However, I made measurements and the time was always the same",
    "2777775": "No idea, but using float16 speeds up for sure.",
    "2804046": "Not sure, but this might be related to this conversation: https://www.kaggle.com/competitions/ai-mathematical-olympiad-prize/discussion/500529#2796083\nBasically, T4 and P100 (generally, all GPUs with CUDA capabilities less than 8.0) does not have true support of certain datatypes (like BFloat16, FP8, etc.), so the computation is actually done in FP32 even if the lower-bit type is used. FP16 is fine though."
  },
  "source": "meta"
}