{
  "id": 152436,
  "title": "(Very ugly) Fix to use mixed_precision with Hugging Face",
  "url": "/competitions/jigsaw-multilingual-toxic-comment-classification/discussion/152436",
  "author_name": "",
  "post_date": "2020-05-19T20:24:23.593631500Z",
  "votes": 6,
  "comment_count": 7,
  "views": 0,
  "content": "<p>Hi,</p>\n\n<p>As mentioned in <a href=\"https://www.kaggle.com/c/jigsaw-multilingual-toxic-comment-classification/discussion/146335\">Issues with Tensorflow TPU and mixed precision, XLA</a> by <a href=\"/dimitreoliveira\">@dimitreoliveira</a> and the comments there, there are some issues when trying to use <code>mixed_precision + TPU + Hugging Face models</code> .</p>\n\n<p>I thus went through the errors and tried to modified a few lines in the following files:</p>\n\n<pre><code>- tensorflow_core/python/keras/layers/normalization.py\n- tensorflow_core/python/keras/layers/pooling.py'\n- transformers/modeling_tf_bert.py\n- transformers/modeling_tf_distilbert.py\n</code></pre>\n\n<p>Basically, just did some casts like </p>\n\n<pre><code>     outputs = outputs + offset\n--&amp;gt;  outputs = outputs + math_ops.cast(offset, outputs.dtype)\n</code></pre>\n\n<p>The modifications are done on the fly when running the kernels, just to let people know what modifications are done.</p>\n\n<p>It assumes however the tensorflow version being 2.1 and transformers 2.9.1.</p>\n\n<p>You can find the kernel here</p>\n\n<p><a href=\"https://www.kaggle.com/yihdarshieh/fix-mixed-precision-issue?scriptVersionId=34312736\">https://www.kaggle.com/yihdarshieh/fix-mixed-precision-issue?scriptVersionId=34312736</a></p>\n\n<p>I also run it without using mixed_precision, and the result is </p>\n\n<p><a href=\"https://www.kaggle.com/yihdarshieh/fix-mixed-precision-issue?scriptVersionId=34311172\">https://www.kaggle.com/yihdarshieh/fix-mixed-precision-issue?scriptVersionId=34311172</a></p>\n\n<p>For the outputs (65536 examples, batch size 8, accumulation size 8 --&gt; efffective batch size 512 and 128 updates each epoch):</p>\n\n<pre><code>- without mixed precision\n\nepoch: 1\ntrain loss: 0.29512545466423035\ntrain auc: 0.9460051655769348\ntraining time: 494.237378\nvalid loss: 0.3203233778476715\nvalid auc: 0.9254440069198608\nvalidation time: 45.694682\n---------------------------------\nepoch: 2\ntrain loss: 0.18848922848701477\ntrain auc: 0.97713702917099\ntraining time: 171.611164\nvalid loss: 0.3051695227622986\nvalid auc: 0.9199531674385071\nvalidation time: 7.683248 \n</code></pre>\n\n<p>and </p>\n\n<pre><code>- with mixed precision\n\nepoch: 1\ntrain loss: 0.2803855240345001\ntrain auc: 0.9514274597167969\ntraining time: 542.160114\nvalid loss: 0.267747163772583\nvalid auc: 0.9214180111885071\nvalidation time: 44.528964\n---------------------------------\nepoch: 2\ntrain loss: 0.18564026057720184\ntrain auc: 0.9778314828872681\ntraining time: 143.036698\nvalid loss: 0.29831987619400024\nvalid auc: 0.9271930456161499\nvalidation time: 4.73615\n</code></pre>\n\n<p>So you can see the training time reduces from <code>171 seconds</code> to <code>143 seconds</code> from the 2nd epochs.\nFor the 1st epoch, we got <code>494 vs 542 seconds</code>, but I think it is due to some overhead introduced when building the graph. If you train more epochs or train with more examples, the effect of mixed precision will be more important.</p>\n\n<p>Under mixed precision, I also tried the following extra settings:</p>\n\n<p><a href=\"https://www.kaggle.com/yihdarshieh/fix-mixed-precision-issue?scriptVersionId=34315133\">https://www.kaggle.com/yihdarshieh/fix-mixed-precision-issue?scriptVersionId=34315133</a></p>\n\n<pre><code>- Freeze word embedding layer (256002048 parameters) to avoid sparse gradients conversion overhead.\n\nepoch: 1\ntrain loss: 0.2764839231967926\ntrain auc: 0.9529916644096375\ntraining time: 459.296742\nvalid loss: 0.2673362195491791\nvalid auc: 0.924274742603302\nvalidation time: 46.748431\n---------------------------------\nepoch: 2\ntrain loss: 0.1863064020872116\ntrain auc: 0.9775997996330261\ntraining time: 122.087714\nvalid loss: 0.31233569979667664\nvalid auc: 0.926763653755188\nvalidation time: 4.70213\n</code></pre>\n\n<p>The model performance is almost the same as the previous results, but the training time (from 2nd epoch) reduces to <code>122 seconds</code>. The word embedding layer of <code>tf-xlm-roberta-large</code> has 256002048 parameters while the model itself has <code>559890432</code> parameters.  So now we have a reduction <code>177 seconds --&amp;gt; 122 seconds</code>, which is about 31% less, and still get a comparable result with only 54% number of trainable parameters.</p>\n\n<p>And <a href=\"https://www.kaggle.com/yihdarshieh/fix-mixed-precision-issue?scriptVersionId=34331385\">https://www.kaggle.com/yihdarshieh/fix-mixed-precision-issue?scriptVersionId=34331385</a></p>\n\n<pre><code>- Freeze all pretrained model's weights\n\nepoch: 1\ntrain loss: 0.7130253314971924\ntrain auc: 0.5168130397796631\ntraining time: 106.225788\nvalid loss: 0.7852157354354858\nvalid accuracy: 0.1563895046710968\nvalid auc: 0.4520792067050934\nvalidation time: 44.396948\n---------------------------------\nepoch: 2\ntrain loss: 0.6938205361366272\ntrain auc: 0.5565494298934937\ntraining time: 45.588002\nvalid loss: 0.7376658320426941\nvalid accuracy: 0.1665736585855484\nvalid auc: 0.5545898079872131\nvalidation time: 4.701679\n</code></pre>\n\n<p>which has quite bad results although the timing is only <code>45</code> seconds.</p>\n\n<p>Hope it can help some of you that want to use mixed precision!</p>",
  "messages": [
    {
      "id": "854179",
      "postDate": "05/19/2020 20:24:23",
      "content": "<p>Hi,</p>\n\n<p>As mentioned in <a href=\"https://www.kaggle.com/c/jigsaw-multilingual-toxic-comment-classification/discussion/146335\">Issues with Tensorflow TPU and mixed precision, XLA</a> by <a href=\"/dimitreoliveira\">@dimitreoliveira</a> and the comments there, there are some issues when trying to use <code>mixed_precision + TPU + Hugging Face models</code> .</p>\n\n<p>I thus went through the errors and tried to modified a few lines in the following files:</p>\n\n<pre><code>- tensorflow_core/python/keras/layers/normalization.py\n- tensorflow_core/python/keras/layers/pooling.py'\n- transformers/modeling_tf_bert.py\n- transformers/modeling_tf_distilbert.py\n</code></pre>\n\n<p>Basically, just did some casts like </p>\n\n<pre><code>     outputs = outputs + offset\n--&amp;gt;  outputs = outputs + math_ops.cast(offset, outputs.dtype)\n</code></pre>\n\n<p>The modifications are done on the fly when running the kernels, just to let people know what modifications are done.</p>\n\n<p>It assumes however the tensorflow version being 2.1 and transformers 2.9.1.</p>\n\n<p>You can find the kernel here</p>\n\n<p><a href=\"https://www.kaggle.com/yihdarshieh/fix-mixed-precision-issue?scriptVersionId=34312736\">https://www.kaggle.com/yihdarshieh/fix-mixed-precision-issue?scriptVersionId=34312736</a></p>\n\n<p>I also run it without using mixed_precision, and the result is </p>\n\n<p><a href=\"https://www.kaggle.com/yihdarshieh/fix-mixed-precision-issue?scriptVersionId=34311172\">https://www.kaggle.com/yihdarshieh/fix-mixed-precision-issue?scriptVersionId=34311172</a></p>\n\n<p>For the outputs (65536 examples, batch size 8, accumulation size 8 --&gt; efffective batch size 512 and 128 updates each epoch):</p>\n\n<pre><code>- without mixed precision\n\nepoch: 1\ntrain loss: 0.29512545466423035\ntrain auc: 0.9460051655769348\ntraining time: 494.237378\nvalid loss: 0.3203233778476715\nvalid auc: 0.9254440069198608\nvalidation time: 45.694682\n---------------------------------\nepoch: 2\ntrain loss: 0.18848922848701477\ntrain auc: 0.97713702917099\ntraining time: 171.611164\nvalid loss: 0.3051695227622986\nvalid auc: 0.9199531674385071\nvalidation time: 7.683248 \n</code></pre>\n\n<p>and </p>\n\n<pre><code>- with mixed precision\n\nepoch: 1\ntrain loss: 0.2803855240345001\ntrain auc: 0.9514274597167969\ntraining time: 542.160114\nvalid loss: 0.267747163772583\nvalid auc: 0.9214180111885071\nvalidation time: 44.528964\n---------------------------------\nepoch: 2\ntrain loss: 0.18564026057720184\ntrain auc: 0.9778314828872681\ntraining time: 143.036698\nvalid loss: 0.29831987619400024\nvalid auc: 0.9271930456161499\nvalidation time: 4.73615\n</code></pre>\n\n<p>So you can see the training time reduces from <code>171 seconds</code> to <code>143 seconds</code> from the 2nd epochs.\nFor the 1st epoch, we got <code>494 vs 542 seconds</code>, but I think it is due to some overhead introduced when building the graph. If you train more epochs or train with more examples, the effect of mixed precision will be more important.</p>\n\n<p>Under mixed precision, I also tried the following extra settings:</p>\n\n<p><a href=\"https://www.kaggle.com/yihdarshieh/fix-mixed-precision-issue?scriptVersionId=34315133\">https://www.kaggle.com/yihdarshieh/fix-mixed-precision-issue?scriptVersionId=34315133</a></p>\n\n<pre><code>- Freeze word embedding layer (256002048 parameters) to avoid sparse gradients conversion overhead.\n\nepoch: 1\ntrain loss: 0.2764839231967926\ntrain auc: 0.9529916644096375\ntraining time: 459.296742\nvalid loss: 0.2673362195491791\nvalid auc: 0.924274742603302\nvalidation time: 46.748431\n---------------------------------\nepoch: 2\ntrain loss: 0.1863064020872116\ntrain auc: 0.9775997996330261\ntraining time: 122.087714\nvalid loss: 0.31233569979667664\nvalid auc: 0.926763653755188\nvalidation time: 4.70213\n</code></pre>\n\n<p>The model performance is almost the same as the previous results, but the training time (from 2nd epoch) reduces to <code>122 seconds</code>. The word embedding layer of <code>tf-xlm-roberta-large</code> has 256002048 parameters while the model itself has <code>559890432</code> parameters.  So now we have a reduction <code>177 seconds --&amp;gt; 122 seconds</code>, which is about 31% less, and still get a comparable result with only 54% number of trainable parameters.</p>\n\n<p>And <a href=\"https://www.kaggle.com/yihdarshieh/fix-mixed-precision-issue?scriptVersionId=34331385\">https://www.kaggle.com/yihdarshieh/fix-mixed-precision-issue?scriptVersionId=34331385</a></p>\n\n<pre><code>- Freeze all pretrained model's weights\n\nepoch: 1\ntrain loss: 0.7130253314971924\ntrain auc: 0.5168130397796631\ntraining time: 106.225788\nvalid loss: 0.7852157354354858\nvalid accuracy: 0.1563895046710968\nvalid auc: 0.4520792067050934\nvalidation time: 44.396948\n---------------------------------\nepoch: 2\ntrain loss: 0.6938205361366272\ntrain auc: 0.5565494298934937\ntraining time: 45.588002\nvalid loss: 0.7376658320426941\nvalid accuracy: 0.1665736585855484\nvalid auc: 0.5545898079872131\nvalidation time: 4.701679\n</code></pre>\n\n<p>which has quite bad results although the timing is only <code>45</code> seconds.</p>\n\n<p>Hope it can help some of you that want to use mixed precision!</p>",
      "rawMarkdown": "Hi,\n\nAs mentioned in [Issues with Tensorflow TPU and mixed precision, XLA](https://www.kaggle.com/c/jigsaw-multilingual-toxic-comment-classification/discussion/146335) by @dimitreoliveira and the comments there, there are some issues when trying to use `mixed_precision + TPU + Hugging Face models` .\n\nI thus went through the errors and tried to modified a few lines in the following files:\n\n    - tensorflow_core/python/keras/layers/normalization.py\n    - tensorflow_core/python/keras/layers/pooling.py'\n    - transformers/modeling_tf_bert.py\n    - transformers/modeling_tf_distilbert.py\n\nBasically, just did some casts like \n\n         outputs = outputs + offset\n    --&gt;  outputs = outputs + math_ops.cast(offset, outputs.dtype)\n\nThe modifications are done on the fly when running the kernels, just to let people know what modifications are done.\n\nIt assumes however the tensorflow version being 2.1 and transformers 2.9.1.\n\nYou can find the kernel here\n\n[https://www.kaggle.com/yihdarshieh/fix-mixed-precision-issue?scriptVersionId=34312736](https://www.kaggle.com/yihdarshieh/fix-mixed-precision-issue?scriptVersionId=34312736)\n\nI also run it without using mixed_precision, and the result is \n\n[https://www.kaggle.com/yihdarshieh/fix-mixed-precision-issue?scriptVersionId=34311172](https://www.kaggle.com/yihdarshieh/fix-mixed-precision-issue?scriptVersionId=34311172)\n\nFor the outputs (65536 examples, batch size 8, accumulation size 8 --&gt; efffective batch size 512 and 128 updates each epoch):\n\n    - without mixed precision\n\n    epoch: 1\n    train loss: 0.29512545466423035\n    train auc: 0.9460051655769348\n    training time: 494.237378\n    valid loss: 0.3203233778476715\n    valid auc: 0.9254440069198608\n    validation time: 45.694682\n    ---------------------------------\n    epoch: 2\n    train loss: 0.18848922848701477\n    train auc: 0.97713702917099\n    training time: 171.611164\n    valid loss: 0.3051695227622986\n    valid auc: 0.9199531674385071\n    validation time: 7.683248 \n\nand \n\n    - with mixed precision\n\n    epoch: 1\n    train loss: 0.2803855240345001\n    train auc: 0.9514274597167969\n    training time: 542.160114\n    valid loss: 0.267747163772583\n    valid auc: 0.9214180111885071\n    validation time: 44.528964\n    ---------------------------------\n    epoch: 2\n    train loss: 0.18564026057720184\n    train auc: 0.9778314828872681\n    training time: 143.036698\n    valid loss: 0.29831987619400024\n    valid auc: 0.9271930456161499\n    validation time: 4.73615\n\nSo you can see the training time reduces from `171 seconds` to `143 seconds` from the 2nd epochs.\nFor the 1st epoch, we got `494 vs 542 seconds`, but I think it is due to some overhead introduced when building the graph. If you train more epochs or train with more examples, the effect of mixed precision will be more important.\n\nUnder mixed precision, I also tried the following extra settings:\n\n[https://www.kaggle.com/yihdarshieh/fix-mixed-precision-issue?scriptVersionId=34315133](https://www.kaggle.com/yihdarshieh/fix-mixed-precision-issue?scriptVersionId=34315133)\n    \n    - Freeze word embedding layer (256002048 parameters) to avoid sparse gradients conversion overhead.\n\n    epoch: 1\n    train loss: 0.2764839231967926\n    train auc: 0.9529916644096375\n    training time: 459.296742\n    valid loss: 0.2673362195491791\n    valid auc: 0.924274742603302\n    validation time: 46.748431\n    ---------------------------------\n    epoch: 2\n    train loss: 0.1863064020872116\n    train auc: 0.9775997996330261\n    training time: 122.087714\n    valid loss: 0.31233569979667664\n    valid auc: 0.926763653755188\n    validation time: 4.70213\n\nThe model performance is almost the same as the previous results, but the training time (from 2nd epoch) reduces to `122 seconds`. The word embedding layer of `tf-xlm-roberta-large` has 256002048 parameters while the model itself has `559890432 ` parameters.  So now we have a reduction `177 seconds --&gt; 122 seconds`, which is about 31% less, and still get a comparable result with only 54% number of trainable parameters.\n\n\nAnd [https://www.kaggle.com/yihdarshieh/fix-mixed-precision-issue?scriptVersionId=34331385](https://www.kaggle.com/yihdarshieh/fix-mixed-precision-issue?scriptVersionId=34331385)\n\n    - Freeze all pretrained model's weights\n\n    epoch: 1\n    train loss: 0.7130253314971924\n    train auc: 0.5168130397796631\n    training time: 106.225788\n    valid loss: 0.7852157354354858\n    valid accuracy: 0.1563895046710968\n    valid auc: 0.4520792067050934\n    validation time: 44.396948\n    ---------------------------------\n    epoch: 2\n    train loss: 0.6938205361366272\n    train auc: 0.5565494298934937\n    training time: 45.588002\n    valid loss: 0.7376658320426941\n    valid accuracy: 0.1665736585855484\n    valid auc: 0.5545898079872131\n    validation time: 4.701679\n\nwhich has quite bad results although the timing is only `45` seconds.\n\nHope it can help some of you that want to use mixed precision!",
      "votes": null
    },
    {
      "id": "854192",
      "postDate": "05/19/2020 20:46:10",
      "content": "<p>Thanks for reporting your findings. That's very useful for TF users :)</p>",
      "rawMarkdown": "Thanks for reporting your findings. That's very useful for TF users :)",
      "votes": null
    },
    {
      "id": "854201",
      "postDate": "05/19/2020 20:55:10",
      "content": "<p>Thank you <a href=\"/yihdarshieh\">@yihdarshieh</a> , this is very interesting, I will also try this fix and report results, by the error message I had a feeling that was just a matter os variable cast, as you did but had no idea where to do it, I was hoping that the new HF release would fix this.</p>",
      "rawMarkdown": "Thank you @yihdarshieh , this is very interesting, I will also try this fix and report results, by the error message I had a feeling that was just a matter os variable cast, as you did but had no idea where to do it, I was hoping that the new HF release would fix this.",
      "votes": null
    },
    {
      "id": "854207",
      "postDate": "05/19/2020 21:01:51",
      "content": "<p>I think TF 2.2 might fix something, because when I googled <code>python/keras/layers/normalization</code>, I was firstly redirected to master version, which seems to have a few more castings. Then I realized that I have to check the TF 2.1 version files. For this competition, if people want to use TPU on Kaggle, I think the compatible TF version is 2.1, so the fix in TF 2.2 might not be helpful.</p>",
      "rawMarkdown": "I think TF 2.2 might fix something, because when I googled `python/keras/layers/normalization`, I was firstly redirected to master version, which seems to have a few more castings. Then I realized that I have to check the TF 2.1 version files. For this competition, if people want to use TPU on Kaggle, I think the compatible TF version is 2.1, so the fix in TF 2.2 might not be helpful.",
      "votes": null
    },
    {
      "id": "854259",
      "postDate": "05/19/2020 22:58:05",
      "content": "<p>I suggest trying TF2.2. In TF2.1 there were some problems with Transformers. There is more info <a href=\"https://www.kaggle.com/c/jigsaw-multilingual-toxic-comment-classification/discussion/140022#799781\">here</a></p>",
      "rawMarkdown": "I suggest trying TF2.2. In TF2.1 there were some problems with Transformers. There is more info [here][1]\n\n[1]: https://www.kaggle.com/c/jigsaw-multilingual-toxic-comment-classification/discussion/140022#799781",
      "votes": null
    },
    {
      "id": "854704",
      "postDate": "05/20/2020 08:47:09",
      "content": "<p>Unfortunately, there are still some problems even using TF 2.2, as you can see in this kernel run:</p>\n\n<p><a href=\"https://www.kaggle.com/yihdarshieh/fix-mixed-precision-issue?scriptVersionId=34367172\">https://www.kaggle.com/yihdarshieh/fix-mixed-precision-issue?scriptVersionId=34367172</a></p>",
      "rawMarkdown": "Unfortunately, there are still some problems even using TF 2.2, as you can see in this kernel run:\n\n[https://www.kaggle.com/yihdarshieh/fix-mixed-precision-issue?scriptVersionId=34367172](https://www.kaggle.com/yihdarshieh/fix-mixed-precision-issue?scriptVersionId=34367172)",
      "votes": null
    },
    {
      "id": "854859",
      "postDate": "05/20/2020 11:44:50",
      "content": "<p>Yeah I seem a good Idea <a href=\"/yihdarshieh\">@yihdarshieh</a> , But I meant the Huguingface new release, I think its 2.9.1 or 2.9.2, I saw that they made some improvements with TPUs</p>",
      "rawMarkdown": "Yeah I seem a good Idea @yihdarshieh , But I meant the Huguingface new release, I think its 2.9.1 or 2.9.2, I saw that they made some improvements with TPUs",
      "votes": null
    },
    {
      "id": "854925",
      "postDate": "05/20/2020 12:59:50",
      "content": "<p><a href=\"/dimitreoliveira\">@dimitreoliveira</a> \nThe version in the kernels I published use TF 2.1 + Transformers 2.9.1.\nI also tried TF 2.2 + Transformers 2.9.1, without my modification, and it is not working.</p>",
      "rawMarkdown": "dimitreoliveira \nThe version in the kernels I published use TF 2.1 + Transformers 2.9.1.\nI also tried TF 2.2 + Transformers 2.9.1, without my modification, and it is not working.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 854192,
      "author_name": "seesee",
      "author_url": "",
      "post_date": "05/19/2020 20:46:10",
      "content": "<p>Thanks for reporting your findings. That's very useful for TF users :)</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 854201,
      "author_name": "dimitreoliveira",
      "author_url": "",
      "post_date": "05/19/2020 20:55:10",
      "content": "<p>Thank you <a href=\"/yihdarshieh\">@yihdarshieh</a> , this is very interesting, I will also try this fix and report results, by the error message I had a feeling that was just a matter os variable cast, as you did but had no idea where to do it, I was hoping that the new HF release would fix this.</p>",
      "votes": null,
      "replies": [
        {
          "id": 854207,
          "author_name": "yihdarshieh",
          "author_url": "",
          "post_date": "05/19/2020 21:01:51",
          "content": "<p>I think TF 2.2 might fix something, because when I googled <code>python/keras/layers/normalization</code>, I was firstly redirected to master version, which seems to have a few more castings. Then I realized that I have to check the TF 2.1 version files. For this competition, if people want to use TPU on Kaggle, I think the compatible TF version is 2.1, so the fix in TF 2.2 might not be helpful.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 854859,
          "author_name": "dimitreoliveira",
          "author_url": "",
          "post_date": "05/20/2020 11:44:50",
          "content": "<p>Yeah I seem a good Idea <a href=\"/yihdarshieh\">@yihdarshieh</a> , But I meant the Huguingface new release, I think its 2.9.1 or 2.9.2, I saw that they made some improvements with TPUs</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 854925,
          "author_name": "yihdarshieh",
          "author_url": "",
          "post_date": "05/20/2020 12:59:50",
          "content": "<p><a href=\"/dimitreoliveira\">@dimitreoliveira</a> \nThe version in the kernels I published use TF 2.1 + Transformers 2.9.1.\nI also tried TF 2.2 + Transformers 2.9.1, without my modification, and it is not working.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 854259,
      "author_name": "cdeotte",
      "author_url": "",
      "post_date": "05/19/2020 22:58:05",
      "content": "<p>I suggest trying TF2.2. In TF2.1 there were some problems with Transformers. There is more info <a href=\"https://www.kaggle.com/c/jigsaw-multilingual-toxic-comment-classification/discussion/140022#799781\">here</a></p>",
      "votes": null,
      "replies": [
        {
          "id": 854704,
          "author_name": "yihdarshieh",
          "author_url": "",
          "post_date": "05/20/2020 08:47:09",
          "content": "<p>Unfortunately, there are still some problems even using TF 2.2, as you can see in this kernel run:</p>\n\n<p><a href=\"https://www.kaggle.com/yihdarshieh/fix-mixed-precision-issue?scriptVersionId=34367172\">https://www.kaggle.com/yihdarshieh/fix-mixed-precision-issue?scriptVersionId=34367172</a></p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "854179": "Hi,\n\nAs mentioned in [Issues with Tensorflow TPU and mixed precision, XLA](https://www.kaggle.com/c/jigsaw-multilingual-toxic-comment-classification/discussion/146335) by @dimitreoliveira and the comments there, there are some issues when trying to use `mixed_precision + TPU + Hugging Face models` .\n\nI thus went through the errors and tried to modified a few lines in the following files:\n\n    - tensorflow_core/python/keras/layers/normalization.py\n    - tensorflow_core/python/keras/layers/pooling.py'\n    - transformers/modeling_tf_bert.py\n    - transformers/modeling_tf_distilbert.py\n\nBasically, just did some casts like \n\n         outputs = outputs + offset\n    --&gt;  outputs = outputs + math_ops.cast(offset, outputs.dtype)\n\nThe modifications are done on the fly when running the kernels, just to let people know what modifications are done.\n\nIt assumes however the tensorflow version being 2.1 and transformers 2.9.1.\n\nYou can find the kernel here\n\n[https://www.kaggle.com/yihdarshieh/fix-mixed-precision-issue?scriptVersionId=34312736](https://www.kaggle.com/yihdarshieh/fix-mixed-precision-issue?scriptVersionId=34312736)\n\nI also run it without using mixed_precision, and the result is \n\n[https://www.kaggle.com/yihdarshieh/fix-mixed-precision-issue?scriptVersionId=34311172](https://www.kaggle.com/yihdarshieh/fix-mixed-precision-issue?scriptVersionId=34311172)\n\nFor the outputs (65536 examples, batch size 8, accumulation size 8 --&gt; efffective batch size 512 and 128 updates each epoch):\n\n    - without mixed precision\n\n    epoch: 1\n    train loss: 0.29512545466423035\n    train auc: 0.9460051655769348\n    training time: 494.237378\n    valid loss: 0.3203233778476715\n    valid auc: 0.9254440069198608\n    validation time: 45.694682\n    ---------------------------------\n    epoch: 2\n    train loss: 0.18848922848701477\n    train auc: 0.97713702917099\n    training time: 171.611164\n    valid loss: 0.3051695227622986\n    valid auc: 0.9199531674385071\n    validation time: 7.683248 \n\nand \n\n    - with mixed precision\n\n    epoch: 1\n    train loss: 0.2803855240345001\n    train auc: 0.9514274597167969\n    training time: 542.160114\n    valid loss: 0.267747163772583\n    valid auc: 0.9214180111885071\n    validation time: 44.528964\n    ---------------------------------\n    epoch: 2\n    train loss: 0.18564026057720184\n    train auc: 0.9778314828872681\n    training time: 143.036698\n    valid loss: 0.29831987619400024\n    valid auc: 0.9271930456161499\n    validation time: 4.73615\n\nSo you can see the training time reduces from `171 seconds` to `143 seconds` from the 2nd epochs.\nFor the 1st epoch, we got `494 vs 542 seconds`, but I think it is due to some overhead introduced when building the graph. If you train more epochs or train with more examples, the effect of mixed precision will be more important.\n\nUnder mixed precision, I also tried the following extra settings:\n\n[https://www.kaggle.com/yihdarshieh/fix-mixed-precision-issue?scriptVersionId=34315133](https://www.kaggle.com/yihdarshieh/fix-mixed-precision-issue?scriptVersionId=34315133)\n    \n    - Freeze word embedding layer (256002048 parameters) to avoid sparse gradients conversion overhead.\n\n    epoch: 1\n    train loss: 0.2764839231967926\n    train auc: 0.9529916644096375\n    training time: 459.296742\n    valid loss: 0.2673362195491791\n    valid auc: 0.924274742603302\n    validation time: 46.748431\n    ---------------------------------\n    epoch: 2\n    train loss: 0.1863064020872116\n    train auc: 0.9775997996330261\n    training time: 122.087714\n    valid loss: 0.31233569979667664\n    valid auc: 0.926763653755188\n    validation time: 4.70213\n\nThe model performance is almost the same as the previous results, but the training time (from 2nd epoch) reduces to `122 seconds`. The word embedding layer of `tf-xlm-roberta-large` has 256002048 parameters while the model itself has `559890432 ` parameters.  So now we have a reduction `177 seconds --&gt; 122 seconds`, which is about 31% less, and still get a comparable result with only 54% number of trainable parameters.\n\n\nAnd [https://www.kaggle.com/yihdarshieh/fix-mixed-precision-issue?scriptVersionId=34331385](https://www.kaggle.com/yihdarshieh/fix-mixed-precision-issue?scriptVersionId=34331385)\n\n    - Freeze all pretrained model's weights\n\n    epoch: 1\n    train loss: 0.7130253314971924\n    train auc: 0.5168130397796631\n    training time: 106.225788\n    valid loss: 0.7852157354354858\n    valid accuracy: 0.1563895046710968\n    valid auc: 0.4520792067050934\n    validation time: 44.396948\n    ---------------------------------\n    epoch: 2\n    train loss: 0.6938205361366272\n    train auc: 0.5565494298934937\n    training time: 45.588002\n    valid loss: 0.7376658320426941\n    valid accuracy: 0.1665736585855484\n    valid auc: 0.5545898079872131\n    validation time: 4.701679\n\nwhich has quite bad results although the timing is only `45` seconds.\n\nHope it can help some of you that want to use mixed precision!",
    "854192": "Thanks for reporting your findings. That's very useful for TF users :)",
    "854201": "Thank you @yihdarshieh , this is very interesting, I will also try this fix and report results, by the error message I had a feeling that was just a matter os variable cast, as you did but had no idea where to do it, I was hoping that the new HF release would fix this.",
    "854207": "I think TF 2.2 might fix something, because when I googled `python/keras/layers/normalization`, I was firstly redirected to master version, which seems to have a few more castings. Then I realized that I have to check the TF 2.1 version files. For this competition, if people want to use TPU on Kaggle, I think the compatible TF version is 2.1, so the fix in TF 2.2 might not be helpful.",
    "854259": "I suggest trying TF2.2. In TF2.1 there were some problems with Transformers. There is more info [here][1]\n\n[1]: https://www.kaggle.com/c/jigsaw-multilingual-toxic-comment-classification/discussion/140022#799781",
    "854704": "Unfortunately, there are still some problems even using TF 2.2, as you can see in this kernel run:\n\n[https://www.kaggle.com/yihdarshieh/fix-mixed-precision-issue?scriptVersionId=34367172](https://www.kaggle.com/yihdarshieh/fix-mixed-precision-issue?scriptVersionId=34367172)",
    "854859": "Yeah I seem a good Idea @yihdarshieh , But I meant the Huguingface new release, I think its 2.9.1 or 2.9.2, I saw that they made some improvements with TPUs",
    "854925": "dimitreoliveira \nThe version in the kernels I published use TF 2.1 + Transformers 2.9.1.\nI also tried TF 2.2 + Transformers 2.9.1, without my modification, and it is not working."
  },
  "source": "meta"
}