{
  "id": 119959,
  "title": "TF2 Training kernel - Use Hugging Face's Tensorflow 2 transformer models",
  "url": "/competitions/tensorflow2-question-answering/discussion/119959",
  "author_name": "",
  "post_date": "2019-12-02T19:45:18.395063600Z",
  "votes": 7,
  "comment_count": 9,
  "views": 0,
  "content": "<p>Hi, I just made my first public kernel</p>\n\n<p><a href=\"https://www.kaggle.com/yihdarshieh/use-hugging-face-s-tensorflow-2-transformer-models\">Use Hugging Face's Tensorflow 2 transformer models</a></p>\n\n<p>It provides a training in Tensorflow 2 using Hugging Face's transformers models. I hope it could be helpful to you. There is no inference in this kernel. I will work in a separate kernel ASAP.</p>\n\n<p>I also provide some checkpoints in my public dataset, but don't expect them to perform as well as the Bert Joint baseline. </p>\n\n<p>I committed this kernel with CPU, because I had no more GPU quota this week, and I really want to publish my work since I worked on it for a long time already. I made some condition so it only run for 1 batch when I committed. But the future run will be the whole epochs. Please change it to GPU if you want to run it.</p>\n\n<p>The output checkpoint is named <code>ckpt-3</code>, but it's trained with only 1 batch in the 3rd epoch. So basically, it's just <code>ckpt-2</code>.</p>\n\n<p>I created this new kernel just to make it visible in this competition \"Notebooks\" tab.</p>\n\n<p>For those who want to look the original kernel (GPU), see</p>\n\n<p><a href=\"https://www.kaggle.com/yihdarshieh/use-hugging-face-models\">Use Hugging Face's Tensorflow 2 transformer models for NQ</a></p>\n\n<p>There, you can find the actual ckpt-3 checkpoint in the output.</p>\n\n<p>I really like to have some feedbacks on how to improve this kernel. Thanks.</p>",
  "messages": [
    {
      "id": "686111",
      "postDate": "12/02/2019 19:45:18",
      "content": "<p>Hi, I just made my first public kernel</p>\n\n<p><a href=\"https://www.kaggle.com/yihdarshieh/use-hugging-face-s-tensorflow-2-transformer-models\">Use Hugging Face's Tensorflow 2 transformer models</a></p>\n\n<p>It provides a training in Tensorflow 2 using Hugging Face's transformers models. I hope it could be helpful to you. There is no inference in this kernel. I will work in a separate kernel ASAP.</p>\n\n<p>I also provide some checkpoints in my public dataset, but don't expect them to perform as well as the Bert Joint baseline. </p>\n\n<p>I committed this kernel with CPU, because I had no more GPU quota this week, and I really want to publish my work since I worked on it for a long time already. I made some condition so it only run for 1 batch when I committed. But the future run will be the whole epochs. Please change it to GPU if you want to run it.</p>\n\n<p>The output checkpoint is named <code>ckpt-3</code>, but it's trained with only 1 batch in the 3rd epoch. So basically, it's just <code>ckpt-2</code>.</p>\n\n<p>I created this new kernel just to make it visible in this competition \"Notebooks\" tab.</p>\n\n<p>For those who want to look the original kernel (GPU), see</p>\n\n<p><a href=\"https://www.kaggle.com/yihdarshieh/use-hugging-face-models\">Use Hugging Face's Tensorflow 2 transformer models for NQ</a></p>\n\n<p>There, you can find the actual ckpt-3 checkpoint in the output.</p>\n\n<p>I really like to have some feedbacks on how to improve this kernel. Thanks.</p>",
      "rawMarkdown": "Hi, I just made my first public kernel\n\n[Use Hugging Face's Tensorflow 2 transformer models](https://www.kaggle.com/yihdarshieh/use-hugging-face-s-tensorflow-2-transformer-models)\n\nIt provides a training in Tensorflow 2 using Hugging Face's transformers models. I hope it could be helpful to you. There is no inference in this kernel. I will work in a separate kernel ASAP.\n\nI also provide some checkpoints in my public dataset, but don't expect them to perform as well as the Bert Joint baseline. \n\nI committed this kernel with CPU, because I had no more GPU quota this week, and I really want to publish my work since I worked on it for a long time already. I made some condition so it only run for 1 batch when I committed. But the future run will be the whole epochs. Please change it to GPU if you want to run it.\n\nThe output checkpoint is named `ckpt-3`, but it's trained with only 1 batch in the 3rd epoch. So basically, it's just `ckpt-2`.\n\nI created this new kernel just to make it visible in this competition \"Notebooks\" tab.\n\nFor those who want to look the original kernel (GPU), see\n\n[Use Hugging Face's Tensorflow 2 transformer models for NQ]( https://www.kaggle.com/yihdarshieh/use-hugging-face-models)\n\nThere, you can find the actual ckpt-3 checkpoint in the output.\n\n\n\nI really like to have some feedbacks on how to improve this kernel. Thanks.",
      "votes": null
    },
    {
      "id": "686300",
      "postDate": "12/03/2019 02:10:23",
      "content": "<p>open the notebook,then a **access **button can be found</p>",
      "rawMarkdown": "open the notebook,then a **access **button can be found",
      "votes": null
    },
    {
      "id": "686308",
      "postDate": "12/03/2019 02:25:23",
      "content": "<p><a href=\"/zhaomeng1126\">@zhaomeng1126</a> </p>\n\n<p>It's already public, but when I click the <strong>access</strong> button, I only see </p>\n\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1533864%2F2cac8085599a4c595a2fd62ce51f8c1e%2FCapture.PNG?generation=1575339707742419&amp;alt=media\" alt=\"\"></p>\n\n<p>I created the kernel when I created my dataset, no by \"New notebooks\" in this competition. Is there any difference?</p>",
      "rawMarkdown": "zhaomeng1126 \n\nIt's already public, but when I click the **access** button, I only see \n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1533864%2F2cac8085599a4c595a2fd62ce51f8c1e%2FCapture.PNG?generation=1575339707742419&amp;alt=media)\n\nI created the kernel when I created my dataset, no by \"New notebooks\" in this competition. Is there any difference?",
      "votes": null
    },
    {
      "id": "686335",
      "postDate": "12/03/2019 03:21:15",
      "content": "<p>Your kernel can't be found in this competition,instead ,can be found in your public kernel only</p>",
      "rawMarkdown": "Your kernel can't be found in this competition,instead ,can be found in your public kernel only",
      "votes": null
    },
    {
      "id": "686347",
      "postDate": "12/03/2019 03:44:24",
      "content": "<p>So how to make a kernel that can be found in this competition??</p>",
      "rawMarkdown": "So how to make a kernel that can be found in this competition??",
      "votes": null
    },
    {
      "id": "686352",
      "postDate": "12/03/2019 03:55:24",
      "content": "<p>OK, I need to create a new kernel in this competition and commit it. I have no more GPU quota for this week ...Thanks for your reply</p>",
      "rawMarkdown": "OK, I need to create a new kernel in this competition and commit it. I have no more GPU quota for this week ...Thanks for your reply",
      "votes": null
    },
    {
      "id": "686420",
      "postDate": "12/03/2019 06:11:25",
      "content": "<p>nice！👍 Have you tried mixed precision training with tensorflow 2.0？</p>",
      "rawMarkdown": "nice！👍 Have you tried mixed precision training with tensorflow 2.0？",
      "votes": null
    },
    {
      "id": "686427",
      "postDate": "12/03/2019 06:18:52",
      "content": "<p>I tried a bit, but didn't get it work. I mentioned in the notebook</p>\n\n<p>&gt; Things tried but not working - Solution wanted!¶</p>\n\n<pre><code>I tried to use mixed precision for training, but I couldn't make it work.\n\nUsing tf.keras.mixed_precision.experimental.Policy for mixed precision training causes error when loading pretrained models. (input #1(zero-based) was expected to be a half tensor but is a float tensor [Op:AddV2] name: tf_bert_model_1/bert/embeddings/add/.)\n\nUsing tf.config.optimizer.set_experimental_options doesn't make tranining faster. Probably this is designed only for compiled model??\n</code></pre>",
      "rawMarkdown": "I tried a bit, but didn't get it work. I mentioned in the notebook\n\n&gt; Things tried but not working - Solution wanted!¶\n\n    I tried to use mixed precision for training, but I couldn't make it work.\n\n    Using tf.keras.mixed_precision.experimental.Policy for mixed precision training causes error when loading pretrained models. (input #1(zero-based) was expected to be a half tensor but is a float tensor [Op:AddV2] name: tf_bert_model_1/bert/embeddings/add/.)\n    \n    Using tf.config.optimizer.set_experimental_options doesn't make tranining faster. Probably this is designed only for compiled model??",
      "votes": null
    },
    {
      "id": "686438",
      "postDate": "12/03/2019 06:33:00",
      "content": "<p><a href=\"https://developer.nvidia.com/automatic-mixed-precision\">This page</a> describes a method for tensorflow to use mixed precision. I have tried it on mnist with tf 2.0 but it didn't affect the training process. I am still very confused about this...</p>\n\n<blockquote>\n  <p>TensorFlow\n  Automatic Mixed Precision feature is available both in native TensorFlow and inside the TensorFlow container on NVIDIA NGC container registry:</p>\n  \n  <p>export TF_ENABLE_AUTO_MIXED_PRECISION=1</p>\n  \n  <p>As an alternative, the environment variable can be set inside the TensorFlow Python script:</p>\n  \n  <p>os.environ['TF_ENABLE_AUTO_MIXED_PRECISION'] = '1'</p>\n</blockquote>",
      "rawMarkdown": "[This page](https://developer.nvidia.com/automatic-mixed-precision) describes a method for tensorflow to use mixed precision. I have tried it on mnist with tf 2.0 but it didn't affect the training process. I am still very confused about this...\n&gt; TensorFlow\nAutomatic Mixed Precision feature is available both in native TensorFlow and inside the TensorFlow container on NVIDIA NGC container registry:\n\n&gt; export TF_ENABLE_AUTO_MIXED_PRECISION=1\n\n&gt; As an alternative, the environment variable can be set inside the TensorFlow Python script:\n\n&gt; os.environ['TF_ENABLE_AUTO_MIXED_PRECISION'] = '1'",
      "votes": null
    },
    {
      "id": "719098",
      "postDate": "01/15/2020 06:06:43",
      "content": "<p>how to run this kernel in parallel ？</p>",
      "rawMarkdown": "how to run this kernel in parallel ？",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 686300,
      "author_name": "zhaomeng1126",
      "author_url": "",
      "post_date": "12/03/2019 02:10:23",
      "content": "<p>open the notebook,then a **access **button can be found</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 686308,
      "author_name": "yihdarshieh",
      "author_url": "",
      "post_date": "12/03/2019 02:25:23",
      "content": "<p><a href=\"/zhaomeng1126\">@zhaomeng1126</a> </p>\n\n<p>It's already public, but when I click the <strong>access</strong> button, I only see </p>\n\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1533864%2F2cac8085599a4c595a2fd62ce51f8c1e%2FCapture.PNG?generation=1575339707742419&amp;alt=media\" alt=\"\"></p>\n\n<p>I created the kernel when I created my dataset, no by \"New notebooks\" in this competition. Is there any difference?</p>",
      "votes": null,
      "replies": [
        {
          "id": 686335,
          "author_name": "zhaomeng1126",
          "author_url": "",
          "post_date": "12/03/2019 03:21:15",
          "content": "<p>Your kernel can't be found in this competition,instead ,can be found in your public kernel only</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 686347,
          "author_name": "yihdarshieh",
          "author_url": "",
          "post_date": "12/03/2019 03:44:24",
          "content": "<p>So how to make a kernel that can be found in this competition??</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 686352,
          "author_name": "yihdarshieh",
          "author_url": "",
          "post_date": "12/03/2019 03:55:24",
          "content": "<p>OK, I need to create a new kernel in this competition and commit it. I have no more GPU quota for this week ...Thanks for your reply</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 686420,
      "author_name": "noxuslol",
      "author_url": "",
      "post_date": "12/03/2019 06:11:25",
      "content": "<p>nice！👍 Have you tried mixed precision training with tensorflow 2.0？</p>",
      "votes": null,
      "replies": [
        {
          "id": 686427,
          "author_name": "yihdarshieh",
          "author_url": "",
          "post_date": "12/03/2019 06:18:52",
          "content": "<p>I tried a bit, but didn't get it work. I mentioned in the notebook</p>\n\n<p>&gt; Things tried but not working - Solution wanted!¶</p>\n\n<pre><code>I tried to use mixed precision for training, but I couldn't make it work.\n\nUsing tf.keras.mixed_precision.experimental.Policy for mixed precision training causes error when loading pretrained models. (input #1(zero-based) was expected to be a half tensor but is a float tensor [Op:AddV2] name: tf_bert_model_1/bert/embeddings/add/.)\n\nUsing tf.config.optimizer.set_experimental_options doesn't make tranining faster. Probably this is designed only for compiled model??\n</code></pre>",
          "votes": null,
          "replies": []
        },
        {
          "id": 686438,
          "author_name": "noxuslol",
          "author_url": "",
          "post_date": "12/03/2019 06:33:00",
          "content": "<p><a href=\"https://developer.nvidia.com/automatic-mixed-precision\">This page</a> describes a method for tensorflow to use mixed precision. I have tried it on mnist with tf 2.0 but it didn't affect the training process. I am still very confused about this...</p>\n\n<blockquote>\n  <p>TensorFlow\n  Automatic Mixed Precision feature is available both in native TensorFlow and inside the TensorFlow container on NVIDIA NGC container registry:</p>\n  \n  <p>export TF_ENABLE_AUTO_MIXED_PRECISION=1</p>\n  \n  <p>As an alternative, the environment variable can be set inside the TensorFlow Python script:</p>\n  \n  <p>os.environ['TF_ENABLE_AUTO_MIXED_PRECISION'] = '1'</p>\n</blockquote>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 719098,
      "author_name": "wentixiaogege",
      "author_url": "",
      "post_date": "01/15/2020 06:06:43",
      "content": "<p>how to run this kernel in parallel ？</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "686111": "Hi, I just made my first public kernel\n\n[Use Hugging Face's Tensorflow 2 transformer models](https://www.kaggle.com/yihdarshieh/use-hugging-face-s-tensorflow-2-transformer-models)\n\nIt provides a training in Tensorflow 2 using Hugging Face's transformers models. I hope it could be helpful to you. There is no inference in this kernel. I will work in a separate kernel ASAP.\n\nI also provide some checkpoints in my public dataset, but don't expect them to perform as well as the Bert Joint baseline. \n\nI committed this kernel with CPU, because I had no more GPU quota this week, and I really want to publish my work since I worked on it for a long time already. I made some condition so it only run for 1 batch when I committed. But the future run will be the whole epochs. Please change it to GPU if you want to run it.\n\nThe output checkpoint is named `ckpt-3`, but it's trained with only 1 batch in the 3rd epoch. So basically, it's just `ckpt-2`.\n\nI created this new kernel just to make it visible in this competition \"Notebooks\" tab.\n\nFor those who want to look the original kernel (GPU), see\n\n[Use Hugging Face's Tensorflow 2 transformer models for NQ]( https://www.kaggle.com/yihdarshieh/use-hugging-face-models)\n\nThere, you can find the actual ckpt-3 checkpoint in the output.\n\n\n\nI really like to have some feedbacks on how to improve this kernel. Thanks.",
    "686300": "open the notebook,then a **access **button can be found",
    "686308": "zhaomeng1126 \n\nIt's already public, but when I click the **access** button, I only see \n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1533864%2F2cac8085599a4c595a2fd62ce51f8c1e%2FCapture.PNG?generation=1575339707742419&amp;alt=media)\n\nI created the kernel when I created my dataset, no by \"New notebooks\" in this competition. Is there any difference?",
    "686335": "Your kernel can't be found in this competition,instead ,can be found in your public kernel only",
    "686347": "So how to make a kernel that can be found in this competition??",
    "686352": "OK, I need to create a new kernel in this competition and commit it. I have no more GPU quota for this week ...Thanks for your reply",
    "686420": "nice！👍 Have you tried mixed precision training with tensorflow 2.0？",
    "686427": "I tried a bit, but didn't get it work. I mentioned in the notebook\n\n&gt; Things tried but not working - Solution wanted!¶\n\n    I tried to use mixed precision for training, but I couldn't make it work.\n\n    Using tf.keras.mixed_precision.experimental.Policy for mixed precision training causes error when loading pretrained models. (input #1(zero-based) was expected to be a half tensor but is a float tensor [Op:AddV2] name: tf_bert_model_1/bert/embeddings/add/.)\n    \n    Using tf.config.optimizer.set_experimental_options doesn't make tranining faster. Probably this is designed only for compiled model??",
    "686438": "[This page](https://developer.nvidia.com/automatic-mixed-precision) describes a method for tensorflow to use mixed precision. I have tried it on mnist with tf 2.0 but it didn't affect the training process. I am still very confused about this...\n&gt; TensorFlow\nAutomatic Mixed Precision feature is available both in native TensorFlow and inside the TensorFlow container on NVIDIA NGC container registry:\n\n&gt; export TF_ENABLE_AUTO_MIXED_PRECISION=1\n\n&gt; As an alternative, the environment variable can be set inside the TensorFlow Python script:\n\n&gt; os.environ['TF_ENABLE_AUTO_MIXED_PRECISION'] = '1'",
    "719098": "how to run this kernel in parallel ？"
  },
  "source": "meta"
}