{
  "id": 146335,
  "title": "Issues with Tensorflow TPU and mixed precision, XLA",
  "url": "/competitions/jigsaw-multilingual-toxic-comment-classification/discussion/146335",
  "author_name": "",
  "post_date": "2020-04-26T19:36:19.441377500Z",
  "votes": 7,
  "comment_count": 14,
  "views": 0,
  "content": "<p>I was trying to optimize my training time, after using custom training loops I got some boost, but I was hoping to have more improvements with mixed precision and XLA, but what happened was that my training time actually got longer, here were the results:</p>\n\n<ul>\n<li>Normal training <code>(model.fit())</code>: <code>1162s</code> per epoch</li>\n<li>Normal training with mixed precision and XLA<code>(model.fit())</code>: <code>1134s</code> per epoch</li>\n<li>Custom training loop: <code>839</code> per epoch</li>\n<li>Custom training loop with mixed precision: <code>1083s</code> per epoch</li>\n<li>Custom training loop with mixed precision and XLA: <code>1083s</code> per epoch</li>\n</ul>\n\n<p>All were using the same code, and a batch size of <code>128</code>\nAnd this is how I am activating mixed precision and XLA</p>\n\n<p>```</p>\n\n<h1>Mixed precision</h1>\n\n<p>tf.config.optimizer.set_experimental_options({\"auto_mixed_precision\": True})</p>\n\n<h1>XLA</h1>\n\n<p>tf.config.optimizer.set_jit(True)\n```</p>\n\n<p>Anyone had similar issues or know what I might be doing wrong?</p>",
  "messages": [
    {
      "id": "822259",
      "postDate": "04/26/2020 19:36:19",
      "content": "<p>I was trying to optimize my training time, after using custom training loops I got some boost, but I was hoping to have more improvements with mixed precision and XLA, but what happened was that my training time actually got longer, here were the results:</p>\n\n<ul>\n<li>Normal training <code>(model.fit())</code>: <code>1162s</code> per epoch</li>\n<li>Normal training with mixed precision and XLA<code>(model.fit())</code>: <code>1134s</code> per epoch</li>\n<li>Custom training loop: <code>839</code> per epoch</li>\n<li>Custom training loop with mixed precision: <code>1083s</code> per epoch</li>\n<li>Custom training loop with mixed precision and XLA: <code>1083s</code> per epoch</li>\n</ul>\n\n<p>All were using the same code, and a batch size of <code>128</code>\nAnd this is how I am activating mixed precision and XLA</p>\n\n<p>```</p>\n\n<h1>Mixed precision</h1>\n\n<p>tf.config.optimizer.set_experimental_options({\"auto_mixed_precision\": True})</p>\n\n<h1>XLA</h1>\n\n<p>tf.config.optimizer.set_jit(True)\n```</p>\n\n<p>Anyone had similar issues or know what I might be doing wrong?</p>",
      "rawMarkdown": "I was trying to optimize my training time, after using custom training loops I got some boost, but I was hoping to have more improvements with mixed precision and XLA, but what happened was that my training time actually got longer, here were the results:\n\n- Normal training `(model.fit())`: `1162s` per epoch\n- Normal training with mixed precision and XLA`(model.fit())`: `1134s` per epoch\n- Custom training loop: `839` per epoch\n- Custom training loop with mixed precision: `1083s` per epoch\n- Custom training loop with mixed precision and XLA: `1083s` per epoch\n\nAll were using the same code, and a batch size of `128`\nAnd this is how I am activating mixed precision and XLA\n\n```\n# Mixed precision\ntf.config.optimizer.set_experimental_options({\"auto_mixed_precision\": True})\n# XLA\ntf.config.optimizer.set_jit(True)\n```\n\nAnyone had similar issues or know what I might be doing wrong?",
      "votes": null
    },
    {
      "id": "822864",
      "postDate": "04/27/2020 07:50:09",
      "content": "<p>As far as I know huggingface doesn't work properly with mixed precision in TF.</p>",
      "rawMarkdown": "As far as I know huggingface doesn't work properly with mixed precision in TF.",
      "votes": null
    },
    {
      "id": "822925",
      "postDate": "04/27/2020 08:59:11",
      "content": "<p>In theory, you should be able to increase the batch size, which should make then the training faster. but as Psi said, last time I tried with TF I also had an issue with the HuggingFace Transformer repository for xla mixed precision TPU.\nTry to increase the batch size, if you could not, that probably means that mixed precision is not working</p>",
      "rawMarkdown": "In theory, you should be able to increase the batch size, which should make then the training faster. but as Psi said, last time I tried with TF I also had an issue with the HuggingFace Transformer repository for xla mixed precision TPU.\nTry to increase the batch size, if you could not, that probably means that mixed precision is not working",
      "votes": null
    },
    {
      "id": "823097",
      "postDate": "04/27/2020 11:58:37",
      "content": "<p>It seems they have some problems, but it is weird because they have a <a href=\"https://huggingface.co/transformers/benchmarks.html#benchmarking-all-models-for-inference\">benchmark that uses both AMP and XLA</a>, maybe it works properly just for inference?</p>",
      "rawMarkdown": "It seems they have some problems, but it is weird because they have a [benchmark that uses both AMP and XLA](https://huggingface.co/transformers/benchmarks.html#benchmarking-all-models-for-inference), maybe it works properly just for inference?",
      "votes": null
    },
    {
      "id": "823099",
      "postDate": "04/27/2020 12:00:28",
      "content": "<p><a href=\"/ludovick\">@ludovick</a> , I have tried, but it did not work, my max batch size with XLM roBERTa large was 128 with and without mixed precision, that was one of the reasons that made me think that something is wrong.</p>",
      "rawMarkdown": "ludovick , I have tried, but it did not work, my max batch size with XLM roBERTa large was 128 with and without mixed precision, that was one of the reasons that made me think that something is wrong.",
      "votes": null
    },
    {
      "id": "823163",
      "postDate": "04/27/2020 13:11:31",
      "content": "<p><a href=\"/dimitreoliveira\">@dimitreoliveira</a>  Dimitre have you tried\n<code>\npolicy = mixed_precision.Policy('mixed_float16')\nmixed_precision.set_policy(policy)\n</code>\nlike what stated in <a href=\"https://www.tensorflow.org/guide/keras/mixed_precision\">https://www.tensorflow.org/guide/keras/mixed_precision</a> ?</p>\n\n<p>And yes, according to TF2 official presentation if you cannot increase batch_size in fp16, then speed is quite the same. (see the XLA and float16 [before changing batch size] blocks below)</p>\n\n<p>ref : <img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1364892%2Fa6a204559cb46e9aa5b50fe34f14f9e8%2FCD0CAC02-790B-4F00-9C5A-C10C4603DD70.png?generation=1587993342858341&amp;alt=media\" alt=\"\"></p>",
      "rawMarkdown": "dimitreoliveira  Dimitre have you tried\n```\npolicy = mixed_precision.Policy('mixed_float16')\nmixed_precision.set_policy(policy)\n```\nlike what stated in https://www.tensorflow.org/guide/keras/mixed_precision ?\n\nAnd yes, according to TF2 official presentation if you cannot increase batch_size in fp16, then speed is quite the same. (see the XLA and float16 [before changing batch size] blocks below)\n\nref : ![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1364892%2Fa6a204559cb46e9aa5b50fe34f14f9e8%2FCD0CAC02-790B-4F00-9C5A-C10C4603DD70.png?generation=1587993342858341&amp;alt=media)",
      "votes": null
    },
    {
      "id": "823207",
      "postDate": "04/27/2020 13:43:36",
      "content": "<p>Hey <a href=\"/ratthachat\">@ratthachat</a> , I was using \"mixed_bfloat16\", I thought for TPU you were supposed to use bfloat16, isn't it? did you have any success with mixed precision?</p>\n\n<p><code>\npolicy = tf.keras.mixed_precision.experimental.Policy('mixed_bfloat16')\nmixed_precision.set_policy(policy)\n</code></p>",
      "rawMarkdown": "Hey @ratthachat , I was using \"mixed_bfloat16\", I thought for TPU you were supposed to use bfloat16, isn't it? did you have any success with mixed precision?\n\n```\npolicy = tf.keras.mixed_precision.experimental.Policy('mixed_bfloat16')\nmixed_precision.set_policy(policy)\n```",
      "votes": null
    },
    {
      "id": "823285",
      "postDate": "04/27/2020 14:36:22",
      "content": "<p>You are right on bfloat16. I haven't investigated thoroughly on this mix-precision on TPU. I will comeback :)</p>\n\n<p><strong>UPDATED</strong> as psi mentioned, maybe Huggingface really doesn't yet compat with TPU+bfloat16 ... I got this error on Huggingface's model building when I tried to set policy to bfloat16</p>\n\n<p><code>InvalidArgumentError: cannot compute AddV2 as input #1(zero-based) was expected to be a bfloat16 tensor but is a float tensor</code></p>\n\n<p>Nevertheless, today at ICLR2020, Huggingface team said that they are currently working with Google team on TPU, so we can expect more stable version soon :D </p>",
      "rawMarkdown": "You are right on bfloat16. I haven't investigated thoroughly on this mix-precision on TPU. I will comeback :)\n\n**UPDATED** as psi mentioned, maybe Huggingface really doesn't yet compat with TPU+bfloat16 ... I got this error on Huggingface's model building when I tried to set policy to bfloat16\n\n`InvalidArgumentError: cannot compute AddV2 as input #1(zero-based) was expected to be a bfloat16 tensor but is a float tensor`\n\nNevertheless, today at ICLR2020, Huggingface team said that they are currently working with Google team on TPU, so we can expect more stable version soon :D",
      "votes": null
    },
    {
      "id": "823588",
      "postDate": "04/27/2020 18:38:01",
      "content": "<p>Thanks <a href=\"/ratthachat\">@ratthachat</a> , I had this same error when I tried mixed precision the first time, it seems that the optimization does not fully work yet.</p>",
      "rawMarkdown": "Thanks @ratthachat , I had this same error when I tried mixed precision the first time, it seems that the optimization does not fully work yet.",
      "votes": null
    },
    {
      "id": "823726",
      "postDate": "04/27/2020 20:56:18",
      "content": "<p>The mixed precision API is a Keras thing. I'm  not sure it does anything if you are not using Keras model.fit()</p>",
      "rawMarkdown": "The mixed precision API is a Keras thing. I'm  not sure it does anything if you are not using Keras model.fit()",
      "votes": null
    },
    {
      "id": "823783",
      "postDate": "04/27/2020 22:10:47",
      "content": "<p>Hi Martin, </p>\n\n<p>I have experimented 2 cases : </p>\n\n<p>Case1\n<code>policy = mixed_precision.Policy('mixed_bfloat16')\n    mixed_precision.set_policy(policy)</code></p>\n\n<p>the error happened as early as the model building. It happened when Huggingface code tried to construct on embedding layer . I am not expert on this issue but it seems their code forced to use float somehow but TPU required bfloat. So the above reported bfloat error was found.</p>\n\n<p>Case2 Using</p>\n\n<p><code>tf.config.optimizer.set_experimental_options({\"auto_mixed_precision\": True}</code>\nI could run code normally with <code>model.fit</code> but could not double batch size and so no speed gain</p>",
      "rawMarkdown": "Hi Martin, \n\nI have experimented 2 cases : \n\nCase1\n`policy = mixed_precision.Policy('mixed_bfloat16')\n    mixed_precision.set_policy(policy)`\n\nthe error happened as early as the model building. It happened when Huggingface code tried to construct on embedding layer . I am not expert on this issue but it seems their code forced to use float somehow but TPU required bfloat. So the above reported bfloat error was found.\n\nCase2 Using\n\n ` tf.config.optimizer.set_experimental_options({\"auto_mixed_precision\": True}`\nI could run code normally with `model.fit` but could not double batch size and so no speed gain",
      "votes": null
    },
    {
      "id": "823806",
      "postDate": "04/27/2020 22:52:59",
      "content": "<p>Yes <a href=\"/mgornergoogle\">@mgornergoogle</a> , in my experiments using mixed precision and XLA with Keras <code>model.fit()</code> just increased the epoch time with no benefits, it seems that it just gives the compiler some extra workload.</p>",
      "rawMarkdown": "Yes @mgornergoogle , in my experiments using mixed precision and XLA with Keras `model.fit()` just increased the epoch time with no benefits, it seems that it just gives the compiler some extra workload.",
      "votes": null
    },
    {
      "id": "823843",
      "postDate": "04/28/2020 00:22:09",
      "content": "<p>With TPUs, mixed precision bfloat16/float32 is the default mode of operation. You do not need to do anything to enable it. If you do enable it however, you can use it as a memory optimization because some tensors will then be stored in memory in bfloat16 format. Saving on memory can allow you to increase your batch size and that in turn can lead to better TPU utilization and faster training, if the TPU was not fully utilized previously.</p>",
      "rawMarkdown": "With TPUs, mixed precision bfloat16/float32 is the default mode of operation. You do not need to do anything to enable it. If you do enable it however, you can use it as a memory optimization because some tensors will then be stored in memory in bfloat16 format. Saving on memory can allow you to increase your batch size and that in turn can lead to better TPU utilization and faster training, if the TPU was not fully utilized previously.",
      "votes": null
    },
    {
      "id": "854214",
      "postDate": "05/19/2020 21:29:33",
      "content": "<p><a href=\"/mgornergoogle\">@mgornergoogle</a> could you check my post <a href=\"https://www.kaggle.com/c/jigsaw-multilingual-toxic-comment-classification/discussion/152436\">https://www.kaggle.com/c/jigsaw-multilingual-toxic-comment-classification/discussion/152436</a>.</p>\n\n<p>When I didn't enable mixed precision, I get</p>\n\n<p>Compute dtype: float32\nVariable dtype: float32</p>\n\n<p>and if I enable it, I get </p>\n\n<p>Compute dtype: bfloat16\nVariable dtype: float32</p>\n\n<p>and some dtype issue that I had to modify some tf/transformers code.</p>\n\n<p>So <code>mixed precision bfloat16/float32 is the default mode of operation</code> seems not be the case??</p>",
      "rawMarkdown": "mgornergoogle could you check my post [https://www.kaggle.com/c/jigsaw-multilingual-toxic-comment-classification/discussion/152436](https://www.kaggle.com/c/jigsaw-multilingual-toxic-comment-classification/discussion/152436).\n\nWhen I didn't enable mixed precision, I get\n\nCompute dtype: float32\nVariable dtype: float32\n\nand if I enable it, I get \n\nCompute dtype: bfloat16\nVariable dtype: float32\n\nand some dtype issue that I had to modify some tf/transformers code.\n\nSo `mixed precision bfloat16/float32 is the default mode of operation` seems not be the case??",
      "votes": null
    },
    {
      "id": "854249",
      "postDate": "05/19/2020 22:19:13",
      "content": "<p>By default on TPU, matrix multiplications happen on the MXU which is physically a mixed-precision piece of hardware (inputs in float32, converted to bfloat16 on the fly, multiplications in bfloat16 with float32 results, accumulations in float32, results in flaot32.</p>\n\n<p>Enabling mixed precision explicitly on TPU has one additional affect: values will also be stored as bfloat16 in memory. It does not change how the MXU hardware operates.</p>",
      "rawMarkdown": "By default on TPU, matrix multiplications happen on the MXU which is physically a mixed-precision piece of hardware (inputs in float32, converted to bfloat16 on the fly, multiplications in bfloat16 with float32 results, accumulations in float32, results in flaot32.\n\nEnabling mixed precision explicitly on TPU has one additional affect: values will also be stored as bfloat16 in memory. It does not change how the MXU hardware operates.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 822864,
      "author_name": "philippsinger",
      "author_url": "",
      "post_date": "04/27/2020 07:50:09",
      "content": "<p>As far as I know huggingface doesn't work properly with mixed precision in TF.</p>",
      "votes": null,
      "replies": [
        {
          "id": 823097,
          "author_name": "dimitreoliveira",
          "author_url": "",
          "post_date": "04/27/2020 11:58:37",
          "content": "<p>It seems they have some problems, but it is weird because they have a <a href=\"https://huggingface.co/transformers/benchmarks.html#benchmarking-all-models-for-inference\">benchmark that uses both AMP and XLA</a>, maybe it works properly just for inference?</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 822925,
      "author_name": "ludovick",
      "author_url": "",
      "post_date": "04/27/2020 08:59:11",
      "content": "<p>In theory, you should be able to increase the batch size, which should make then the training faster. but as Psi said, last time I tried with TF I also had an issue with the HuggingFace Transformer repository for xla mixed precision TPU.\nTry to increase the batch size, if you could not, that probably means that mixed precision is not working</p>",
      "votes": null,
      "replies": [
        {
          "id": 823099,
          "author_name": "dimitreoliveira",
          "author_url": "",
          "post_date": "04/27/2020 12:00:28",
          "content": "<p><a href=\"/ludovick\">@ludovick</a> , I have tried, but it did not work, my max batch size with XLM roBERTa large was 128 with and without mixed precision, that was one of the reasons that made me think that something is wrong.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 823163,
      "author_name": "ratthachat",
      "author_url": "",
      "post_date": "04/27/2020 13:11:31",
      "content": "<p><a href=\"/dimitreoliveira\">@dimitreoliveira</a>  Dimitre have you tried\n<code>\npolicy = mixed_precision.Policy('mixed_float16')\nmixed_precision.set_policy(policy)\n</code>\nlike what stated in <a href=\"https://www.tensorflow.org/guide/keras/mixed_precision\">https://www.tensorflow.org/guide/keras/mixed_precision</a> ?</p>\n\n<p>And yes, according to TF2 official presentation if you cannot increase batch_size in fp16, then speed is quite the same. (see the XLA and float16 [before changing batch size] blocks below)</p>\n\n<p>ref : <img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1364892%2Fa6a204559cb46e9aa5b50fe34f14f9e8%2FCD0CAC02-790B-4F00-9C5A-C10C4603DD70.png?generation=1587993342858341&amp;alt=media\" alt=\"\"></p>",
      "votes": null,
      "replies": [
        {
          "id": 823207,
          "author_name": "dimitreoliveira",
          "author_url": "",
          "post_date": "04/27/2020 13:43:36",
          "content": "<p>Hey <a href=\"/ratthachat\">@ratthachat</a> , I was using \"mixed_bfloat16\", I thought for TPU you were supposed to use bfloat16, isn't it? did you have any success with mixed precision?</p>\n\n<p><code>\npolicy = tf.keras.mixed_precision.experimental.Policy('mixed_bfloat16')\nmixed_precision.set_policy(policy)\n</code></p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 823285,
          "author_name": "ratthachat",
          "author_url": "",
          "post_date": "04/27/2020 14:36:22",
          "content": "<p>You are right on bfloat16. I haven't investigated thoroughly on this mix-precision on TPU. I will comeback :)</p>\n\n<p><strong>UPDATED</strong> as psi mentioned, maybe Huggingface really doesn't yet compat with TPU+bfloat16 ... I got this error on Huggingface's model building when I tried to set policy to bfloat16</p>\n\n<p><code>InvalidArgumentError: cannot compute AddV2 as input #1(zero-based) was expected to be a bfloat16 tensor but is a float tensor</code></p>\n\n<p>Nevertheless, today at ICLR2020, Huggingface team said that they are currently working with Google team on TPU, so we can expect more stable version soon :D </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 823588,
          "author_name": "dimitreoliveira",
          "author_url": "",
          "post_date": "04/27/2020 18:38:01",
          "content": "<p>Thanks <a href=\"/ratthachat\">@ratthachat</a> , I had this same error when I tried mixed precision the first time, it seems that the optimization does not fully work yet.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 823726,
      "author_name": "mgorner",
      "author_url": "",
      "post_date": "04/27/2020 20:56:18",
      "content": "<p>The mixed precision API is a Keras thing. I'm  not sure it does anything if you are not using Keras model.fit()</p>",
      "votes": null,
      "replies": [
        {
          "id": 823783,
          "author_name": "ratthachat",
          "author_url": "",
          "post_date": "04/27/2020 22:10:47",
          "content": "<p>Hi Martin, </p>\n\n<p>I have experimented 2 cases : </p>\n\n<p>Case1\n<code>policy = mixed_precision.Policy('mixed_bfloat16')\n    mixed_precision.set_policy(policy)</code></p>\n\n<p>the error happened as early as the model building. It happened when Huggingface code tried to construct on embedding layer . I am not expert on this issue but it seems their code forced to use float somehow but TPU required bfloat. So the above reported bfloat error was found.</p>\n\n<p>Case2 Using</p>\n\n<p><code>tf.config.optimizer.set_experimental_options({\"auto_mixed_precision\": True}</code>\nI could run code normally with <code>model.fit</code> but could not double batch size and so no speed gain</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 823806,
          "author_name": "dimitreoliveira",
          "author_url": "",
          "post_date": "04/27/2020 22:52:59",
          "content": "<p>Yes <a href=\"/mgornergoogle\">@mgornergoogle</a> , in my experiments using mixed precision and XLA with Keras <code>model.fit()</code> just increased the epoch time with no benefits, it seems that it just gives the compiler some extra workload.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 823843,
          "author_name": "mgorner",
          "author_url": "",
          "post_date": "04/28/2020 00:22:09",
          "content": "<p>With TPUs, mixed precision bfloat16/float32 is the default mode of operation. You do not need to do anything to enable it. If you do enable it however, you can use it as a memory optimization because some tensors will then be stored in memory in bfloat16 format. Saving on memory can allow you to increase your batch size and that in turn can lead to better TPU utilization and faster training, if the TPU was not fully utilized previously.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 854214,
          "author_name": "yihdarshieh",
          "author_url": "",
          "post_date": "05/19/2020 21:29:33",
          "content": "<p><a href=\"/mgornergoogle\">@mgornergoogle</a> could you check my post <a href=\"https://www.kaggle.com/c/jigsaw-multilingual-toxic-comment-classification/discussion/152436\">https://www.kaggle.com/c/jigsaw-multilingual-toxic-comment-classification/discussion/152436</a>.</p>\n\n<p>When I didn't enable mixed precision, I get</p>\n\n<p>Compute dtype: float32\nVariable dtype: float32</p>\n\n<p>and if I enable it, I get </p>\n\n<p>Compute dtype: bfloat16\nVariable dtype: float32</p>\n\n<p>and some dtype issue that I had to modify some tf/transformers code.</p>\n\n<p>So <code>mixed precision bfloat16/float32 is the default mode of operation</code> seems not be the case??</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 854249,
          "author_name": "mgorner",
          "author_url": "",
          "post_date": "05/19/2020 22:19:13",
          "content": "<p>By default on TPU, matrix multiplications happen on the MXU which is physically a mixed-precision piece of hardware (inputs in float32, converted to bfloat16 on the fly, multiplications in bfloat16 with float32 results, accumulations in float32, results in flaot32.</p>\n\n<p>Enabling mixed precision explicitly on TPU has one additional affect: values will also be stored as bfloat16 in memory. It does not change how the MXU hardware operates.</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "822259": "I was trying to optimize my training time, after using custom training loops I got some boost, but I was hoping to have more improvements with mixed precision and XLA, but what happened was that my training time actually got longer, here were the results:\n\n- Normal training `(model.fit())`: `1162s` per epoch\n- Normal training with mixed precision and XLA`(model.fit())`: `1134s` per epoch\n- Custom training loop: `839` per epoch\n- Custom training loop with mixed precision: `1083s` per epoch\n- Custom training loop with mixed precision and XLA: `1083s` per epoch\n\nAll were using the same code, and a batch size of `128`\nAnd this is how I am activating mixed precision and XLA\n\n```\n# Mixed precision\ntf.config.optimizer.set_experimental_options({\"auto_mixed_precision\": True})\n# XLA\ntf.config.optimizer.set_jit(True)\n```\n\nAnyone had similar issues or know what I might be doing wrong?",
    "822864": "As far as I know huggingface doesn't work properly with mixed precision in TF.",
    "822925": "In theory, you should be able to increase the batch size, which should make then the training faster. but as Psi said, last time I tried with TF I also had an issue with the HuggingFace Transformer repository for xla mixed precision TPU.\nTry to increase the batch size, if you could not, that probably means that mixed precision is not working",
    "823097": "It seems they have some problems, but it is weird because they have a [benchmark that uses both AMP and XLA](https://huggingface.co/transformers/benchmarks.html#benchmarking-all-models-for-inference), maybe it works properly just for inference?",
    "823099": "ludovick , I have tried, but it did not work, my max batch size with XLM roBERTa large was 128 with and without mixed precision, that was one of the reasons that made me think that something is wrong.",
    "823163": "dimitreoliveira  Dimitre have you tried\n```\npolicy = mixed_precision.Policy('mixed_float16')\nmixed_precision.set_policy(policy)\n```\nlike what stated in https://www.tensorflow.org/guide/keras/mixed_precision ?\n\nAnd yes, according to TF2 official presentation if you cannot increase batch_size in fp16, then speed is quite the same. (see the XLA and float16 [before changing batch size] blocks below)\n\nref : ![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1364892%2Fa6a204559cb46e9aa5b50fe34f14f9e8%2FCD0CAC02-790B-4F00-9C5A-C10C4603DD70.png?generation=1587993342858341&amp;alt=media)",
    "823207": "Hey @ratthachat , I was using \"mixed_bfloat16\", I thought for TPU you were supposed to use bfloat16, isn't it? did you have any success with mixed precision?\n\n```\npolicy = tf.keras.mixed_precision.experimental.Policy('mixed_bfloat16')\nmixed_precision.set_policy(policy)\n```",
    "823285": "You are right on bfloat16. I haven't investigated thoroughly on this mix-precision on TPU. I will comeback :)\n\n**UPDATED** as psi mentioned, maybe Huggingface really doesn't yet compat with TPU+bfloat16 ... I got this error on Huggingface's model building when I tried to set policy to bfloat16\n\n`InvalidArgumentError: cannot compute AddV2 as input #1(zero-based) was expected to be a bfloat16 tensor but is a float tensor`\n\nNevertheless, today at ICLR2020, Huggingface team said that they are currently working with Google team on TPU, so we can expect more stable version soon :D",
    "823588": "Thanks @ratthachat , I had this same error when I tried mixed precision the first time, it seems that the optimization does not fully work yet.",
    "823726": "The mixed precision API is a Keras thing. I'm  not sure it does anything if you are not using Keras model.fit()",
    "823783": "Hi Martin, \n\nI have experimented 2 cases : \n\nCase1\n`policy = mixed_precision.Policy('mixed_bfloat16')\n    mixed_precision.set_policy(policy)`\n\nthe error happened as early as the model building. It happened when Huggingface code tried to construct on embedding layer . I am not expert on this issue but it seems their code forced to use float somehow but TPU required bfloat. So the above reported bfloat error was found.\n\nCase2 Using\n\n ` tf.config.optimizer.set_experimental_options({\"auto_mixed_precision\": True}`\nI could run code normally with `model.fit` but could not double batch size and so no speed gain",
    "823806": "Yes @mgornergoogle , in my experiments using mixed precision and XLA with Keras `model.fit()` just increased the epoch time with no benefits, it seems that it just gives the compiler some extra workload.",
    "823843": "With TPUs, mixed precision bfloat16/float32 is the default mode of operation. You do not need to do anything to enable it. If you do enable it however, you can use it as a memory optimization because some tensors will then be stored in memory in bfloat16 format. Saving on memory can allow you to increase your batch size and that in turn can lead to better TPU utilization and faster training, if the TPU was not fully utilized previously.",
    "854214": "mgornergoogle could you check my post [https://www.kaggle.com/c/jigsaw-multilingual-toxic-comment-classification/discussion/152436](https://www.kaggle.com/c/jigsaw-multilingual-toxic-comment-classification/discussion/152436).\n\nWhen I didn't enable mixed precision, I get\n\nCompute dtype: float32\nVariable dtype: float32\n\nand if I enable it, I get \n\nCompute dtype: bfloat16\nVariable dtype: float32\n\nand some dtype issue that I had to modify some tf/transformers code.\n\nSo `mixed precision bfloat16/float32 is the default mode of operation` seems not be the case??",
    "854249": "By default on TPU, matrix multiplications happen on the MXU which is physically a mixed-precision piece of hardware (inputs in float32, converted to bfloat16 on the fly, multiplications in bfloat16 with float32 results, accumulations in float32, results in flaot32.\n\nEnabling mixed precision explicitly on TPU has one additional affect: values will also be stored as bfloat16 in memory. It does not change how the MXU hardware operates."
  },
  "source": "meta"
}