{
  "id": 121572,
  "title": "Using BERT on Cloud TPUs with TensorFlow 2.1 RC",
  "url": "/competitions/tensorflow2-question-answering/discussion/121572",
  "author_name": "Paige Bailey",
  "post_date": "2019-12-14T06:35:02.151000",
  "votes": 21,
  "comment_count": 23,
  "views": 0,
  "content": "<h2>You can now use TensorFlow 2.1 RC on Cloud TPUs via <code>tf-nightly</code>!</h2>\n\n<p>Details can be found in <a href=\"https://github.com/GoogleCloudPlatform/training-data-analyst/blob/master/courses/fast-and-lean-data-science/README-TF2.1.md\">this installation guide</a> from the Kaggle and Google Cloud teams. We are also excited to release an example of <a href=\"https://github.com/tensorflow/models/blob/master/official/nlp/bert/bert_cloud_tpu.md\">BERT fine-tuning with Cloud TPUs</a>, using TensorFlow. </p>\n\n<p><strong>Note</strong>: Given that <code>tf-nightly</code> is a release candidate, not an official release, you may run into rough edges -- so please make sure to <strong>post your questions</strong> in the Kaggle forums, and submit bugs on <a href=\"https://www.github.com/tensorflow/tensorflow\">Github</a>!</p>\n\n<h3>We are also still offering Cloud TPU Quota for entrants in this competition.</h3>\n\n<p>Requests can be made by submitting <a href=\"https://www.kaggle.com/TPU-Tensorflow2-QA\">this form</a> by <strong>December 16, 2019</strong>.</p>\n\n<p>With the upcoming official release of TF 2.1, TPU support with <code>tf.keras</code> and <code>tf.distribute</code> are anticipated to greatly improve developer productivity while taking advantage of the performance boosts that TPUs are already known for. The TensorFlow team is preparing additional tutorial information which will accompany the quota distribution. </p>\n\n<p>If you've been curious about TPUs, now would be a great opportunity to give them a try!</p>\n\n<h3>The fine print, also included on the request form:</h3>\n\n<ul>\n<li>You must complete the pre-requisite setup steps listed on the form in order to receive quota.</li>\n<li>Only one person per team will be issued a TPU quota. This will be verified and multiple requests from the same team will be de-duplicated.</li>\n<li>Be aware that if you are eligible, your request will be fulfilled by December 23, 2019, so you may expect to wait a couple weeks before receiving an email that confirms your quota has been provisioned. Distribution will not be immediate.</li>\n</ul>",
  "messages": [
    {
      "id": 694823,
      "postDate": "2019-12-14T06:35:02.153Z",
      "content": "<h2>You can now use TensorFlow 2.1 RC on Cloud TPUs via <code>tf-nightly</code>!</h2>\n\n<p>Details can be found in <a href=\"https://github.com/GoogleCloudPlatform/training-data-analyst/blob/master/courses/fast-and-lean-data-science/README-TF2.1.md\">this installation guide</a> from the Kaggle and Google Cloud teams. We are also excited to release an example of <a href=\"https://github.com/tensorflow/models/blob/master/official/nlp/bert/bert_cloud_tpu.md\">BERT fine-tuning with Cloud TPUs</a>, using TensorFlow. </p>\n\n<p><strong>Note</strong>: Given that <code>tf-nightly</code> is a release candidate, not an official release, you may run into rough edges -- so please make sure to <strong>post your questions</strong> in the Kaggle forums, and submit bugs on <a href=\"https://www.github.com/tensorflow/tensorflow\">Github</a>!</p>\n\n<h3>We are also still offering Cloud TPU Quota for entrants in this competition.</h3>\n\n<p>Requests can be made by submitting <a href=\"https://www.kaggle.com/TPU-Tensorflow2-QA\">this form</a> by <strong>December 16, 2019</strong>.</p>\n\n<p>With the upcoming official release of TF 2.1, TPU support with <code>tf.keras</code> and <code>tf.distribute</code> are anticipated to greatly improve developer productivity while taking advantage of the performance boosts that TPUs are already known for. The TensorFlow team is preparing additional tutorial information which will accompany the quota distribution. </p>\n\n<p>If you've been curious about TPUs, now would be a great opportunity to give them a try!</p>\n\n<h3>The fine print, also included on the request form:</h3>\n\n<ul>\n<li>You must complete the pre-requisite setup steps listed on the form in order to receive quota.</li>\n<li>Only one person per team will be issued a TPU quota. This will be verified and multiple requests from the same team will be de-duplicated.</li>\n<li>Be aware that if you are eligible, your request will be fulfilled by December 23, 2019, so you may expect to wait a couple weeks before receiving an email that confirms your quota has been provisioned. Distribution will not be immediate.</li>\n</ul>",
      "rawMarkdown": "## You can now use TensorFlow 2.1 RC on Cloud TPUs via `tf-nightly`!\n\nDetails can be found in [this installation guide](https://github.com/GoogleCloudPlatform/training-data-analyst/blob/master/courses/fast-and-lean-data-science/README-TF2.1.md) from the Kaggle and Google Cloud teams. We are also excited to release an example of [BERT fine-tuning with Cloud TPUs](https://github.com/tensorflow/models/blob/master/official/nlp/bert/bert_cloud_tpu.md), using TensorFlow. \n\n**Note**: Given that `tf-nightly` is a release candidate, not an official release, you may run into rough edges -- so please make sure to **post your questions** in the Kaggle forums, and submit bugs on [Github](https://www.github.com/tensorflow/tensorflow)!\n\n### We are also still offering Cloud TPU Quota for entrants in this competition.\n\nRequests can be made by submitting [this form](https://www.kaggle.com/TPU-Tensorflow2-QA) by **December 16, 2019**.\n\nWith the upcoming official release of TF 2.1, TPU support with `tf.keras` and `tf.distribute` are anticipated to greatly improve developer productivity while taking advantage of the performance boosts that TPUs are already known for. The TensorFlow team is preparing additional tutorial information which will accompany the quota distribution. \n\nIf you've been curious about TPUs, now would be a great opportunity to give them a try!\n\n### The fine print, also included on the request form:\n\n* You must complete the pre-requisite setup steps listed on the form in order to receive quota.\n* Only one person per team will be issued a TPU quota. This will be verified and multiple requests from the same team will be de-duplicated.\n* Be aware that if you are eligible, your request will be fulfilled by December 23, 2019, so you may expect to wait a couple weeks before receiving an email that confirms your quota has been provisioned. Distribution will not be immediate.",
      "votes": 21
    },
    {
      "id": 702403,
      "postDate": "2019-12-24T16:39:27.483Z",
      "content": "<p>Will solutions built using TF 2.1 also be considered for the special TF 2.0 prizes? 🙂 </p>",
      "rawMarkdown": "Will solutions built using TF 2.1 also be considered for the special TF 2.0 prizes? 🙂 ",
      "votes": 3,
      "replies": [
        {
          "id": 714775,
          "postDate": "2020-01-09T18:22:37.370Z",
          "content": "<p>Yes.</p>",
          "rawMarkdown": "Yes."
        }
      ]
    },
    {
      "id": 703382,
      "postDate": "2019-12-26T04:36:35.123Z",
      "content": "<p>I have two questions:\n1. if we use tf 2.1 to train a model, can we run inference on kaggle kernel with tf 2.0?\n2. is kaggle kernel gonna be updated to 2.1 as well? </p>\n\n<p>Thank you!</p>",
      "rawMarkdown": "I have two questions:\n1. if we use tf 2.1 to train a model, can we run inference on kaggle kernel with tf 2.0?\n2. is kaggle kernel gonna be updated to 2.1 as well? \n\nThank you!",
      "votes": 1,
      "replies": [
        {
          "id": 703897,
          "postDate": "2019-12-26T19:18:12.603Z",
          "content": "<p>Hi Jiwei,</p>\n\n<p>Good questions.\n1. Yes, this is what you'll have to do until the TF image on Kaggle is updated to 2.1\n2. We will, once TF2.1 has completed its full release.</p>",
          "rawMarkdown": "Hi Jiwei,\n\nGood questions.\n1. Yes, this is what you'll have to do until the TF image on Kaggle is updated to 2.1\n2. We will, once TF2.1 has completed its full release.",
          "votes": 2
        },
        {
          "id": 703994,
          "postDate": "2019-12-26T23:15:06.727Z",
          "content": "<p>Thank you Julia!</p>",
          "rawMarkdown": "Thank you Julia!"
        },
        {
          "id": 704034,
          "postDate": "2019-12-27T00:55:24.100Z",
          "content": "<p>I just figured that upgrading to tf2.1 would require cuda 10.1 and consequently a system wise update of GPU driver. Could you please not upgrade kaggle docker before this competition ends? we already spend two months with tf 2.0 and set up environment as is. With a month left, I would like to focus on modeling rather than configuring environments and learning new APIs. 😂 </p>",
          "rawMarkdown": "I just figured that upgrading to tf2.1 would require cuda 10.1 and consequently a system wise update of GPU driver. Could you please not upgrade kaggle docker before this competition ends? we already spend two months with tf 2.0 and set up environment as is. With a month left, I would like to focus on modeling rather than configuring environments and learning new APIs. 😂 "
        },
        {
          "id": 714766,
          "postDate": "2020-01-09T18:14:43.903Z",
          "content": "<p>Correction: TF2.1 RC0 was already released into notebooks in early December, and the upgrade to the latest release will be made soon. However, our notebooks function such that notebooks created on your original image will be retained with that image. Therefore, if you have a preference for keeping your notebook in TF2.0, you're welcome to do so by not changing your existing notebook to the \"Latest Available\" image.</p>",
          "rawMarkdown": "Correction: TF2.1 RC0 was already released into notebooks in early December, and the upgrade to the latest release will be made soon. However, our notebooks function such that notebooks created on your original image will be retained with that image. Therefore, if you have a preference for keeping your notebook in TF2.0, you're welcome to do so by not changing your existing notebook to the \"Latest Available\" image."
        }
      ]
    },
    {
      "id": 695330,
      "postDate": "2019-12-15T00:42:32.210Z",
      "content": "<blockquote>\n  <p>The fine print, also included on the request form:\n  You must complete the pre-requisite setup steps listed on the form in order to receive quota.</p>\n</blockquote>\n\n<p>I have requested for the TPU quota. Does that mean that now I cannot team up with anyone who had also requested for a TPU quota?</p>",
      "rawMarkdown": "&gt; The fine print, also included on the request form:\nYou must complete the pre-requisite setup steps listed on the form in order to receive quota.\n\nI have requested for the TPU quota. Does that mean that now I cannot team up with anyone who had also requested for a TPU quota?\n",
      "votes": 1,
      "replies": [
        {
          "id": 696014,
          "postDate": "2019-12-16T02:16:08.453Z",
          "content": "<p>You can certainly still team up with other Kagglers who have requested TPU quota; however, only one person from each group will be allocated quota officially.</p>",
          "rawMarkdown": "You can certainly still team up with other Kagglers who have requested TPU quota; however, only one person from each group will be allocated quota officially."
        },
        {
          "id": 696035,
          "postDate": "2019-12-16T03:23:27.767Z",
          "content": "<p>And then what will happen to the TPU quotas if the teaming up happens after credit allocation? What if those credits are also utilised (partially or completely)?</p>",
          "rawMarkdown": "And then what will happen to the TPU quotas if the teaming up happens after credit allocation? What if those credits are also utilised (partially or completely)?",
          "votes": 2
        }
      ]
    },
    {
      "id": 702609,
      "postDate": "2019-12-24T22:55:14.220Z",
      "content": "<p><a href=\"/rohanrao\">@rohanrao</a> good question - yes!</p>",
      "rawMarkdown": "@rohanrao good question - yes!",
      "votes": 2
    },
    {
      "id": 715274,
      "postDate": "2020-01-10T10:41:23.427Z",
      "content": "<p>Does anybody try pytorch on Tpu successfully？I always use pytorch ...</p>",
      "rawMarkdown": "Does anybody try pytorch on Tpu successfully？I always use pytorch ...",
      "replies": [
        {
          "id": 715291,
          "postDate": "2020-01-10T11:03:38.050Z",
          "content": "<p>A few months ago(about October) I have tried to run pytorch on TPU to fine-tune BERT, the training process worked well, but I can't save the trained model, I can't get the solution at that time, then I give up to use pytorch on TPU. I am not sure if it works well now.</p>",
          "rawMarkdown": "A few months ago(about October) I have tried to run pytorch on TPU to fine-tune BERT, the training process worked well, but I can't save the trained model, I can't get the solution at that time, then I give up to use pytorch on TPU. I am not sure if it works well now."
        },
        {
          "id": 715300,
          "postDate": "2020-01-10T11:23:27.060Z",
          "content": "<p>Does metric get worse? I see many bugs in issue</p>",
          "rawMarkdown": "Does metric get worse? I see many bugs in issue"
        },
        {
          "id": 715311,
          "postDate": "2020-01-10T11:37:35.410Z",
          "content": "<p>Only slightly worse, I haven't try to fine-tune more at that time, maybe it can get better.</p>",
          "rawMarkdown": "Only slightly worse, I haven't try to fine-tune more at that time, maybe it can get better."
        },
        {
          "id": 715326,
          "postDate": "2020-01-10T12:00:03.190Z",
          "content": "<p>OK，thank you,but time is not enough to make a try</p>",
          "rawMarkdown": "OK，thank you,but time is not enough to make a try"
        },
        {
          "id": 716400,
          "postDate": "2020-01-11T16:40:44.107Z",
          "content": "<p>PyTorch on TPU is working. Some code additions are needed ( XLA ) </p>",
          "rawMarkdown": "PyTorch on TPU is working. Some code additions are needed ( XLA ) "
        }
      ]
    },
    {
      "id": 698510,
      "postDate": "2019-12-19T10:27:16.513Z",
      "content": "<p>Could someone confirm that these tutorials are working right now? I can't get them to work. Any help is appreciated. Am I doing something wrong?</p>\n\n<p><a href=\"https://github.com/GoogleCloudPlatform/training-data-analyst/issues/786\">https://github.com/GoogleCloudPlatform/training-data-analyst/issues/786</a>\n<a href=\"https://github.com/tensorflow/models/issues/7962\">https://github.com/tensorflow/models/issues/7962</a></p>",
      "rawMarkdown": "Could someone confirm that these tutorials are working right now? I can't get them to work. Any help is appreciated. Am I doing something wrong?\n\nhttps://github.com/GoogleCloudPlatform/training-data-analyst/issues/786\nhttps://github.com/tensorflow/models/issues/7962",
      "replies": [
        {
          "id": 698535,
          "postDate": "2019-12-19T11:04:02.860Z",
          "content": "<p>Hi See, firstly grats on your 1st public LB position. </p>\n\n<p>Our team has successfully set up using the TPU. In order to use the TPU, you will need to create the TPU session with TF-nightly-2.x within the same region, the training data need to be on google storage, then you'll need to pass the TPU name/IP address into experimental_connect_to_cluster. It is much much much faster than GPU training. Finetune the data with the NQ training data takes around 1 hour for 1 epoch.</p>",
          "rawMarkdown": "Hi See, firstly grats on your 1st public LB position. \n\nOur team has successfully set up using the TPU. In order to use the TPU, you will need to create the TPU session with TF-nightly-2.x within the same region, the training data need to be on google storage, then you'll need to pass the TPU name/IP address into experimental_connect_to_cluster. It is much much much faster than GPU training. Finetune the data with the NQ training data takes around 1 hour for 1 epoch.\n\n",
          "votes": 2
        },
        {
          "id": 698539,
          "postDate": "2019-12-19T11:10:42.037Z",
          "content": "<p>Thanks. I too have a VM and TPU. The first tutorial used to work. But since yesterday my code errors:\n```</p>\n\n<h1>Detect hardware, return appropriate distribution strategy</h1>\n\n<p>try:\n    tpu = tf.distribute.cluster_resolver.TPUClusterResolver()  # TPU detection\n    print('Running on TPU ', tpu.cluster_spec().as_dict()['worker'])\nexcept ValueError:\n    tpu = None</p>\n\n<p>if tpu:\n    tf.config.experimental_connect_to_cluster(tpu)\n    tf.tpu.experimental.initialize_tpu_system(tpu)\n    strategy = tf.distribute.experimental.TPUStrategy(tpu)\nelse:\n    strategy = tf.distribute.get_strategy()</p>\n\n<p>print(\"REPLICAS: \", strategy.num_replicas_in_sync)\n```\nDo you use a similar snippet to create the strategy? Does it work now?</p>\n\n<p>This is what I get:\n```\nRunning on TPU  ['192.168.15.2:8470']\nINFO:tensorflow:Initializing the TPU system: rmme</p>\n\n<p>INFO:tensorflow:Initializing the TPU system: rmme</p>\n\n<p>INFO:tensorflow:Clearing out eager caches</p>\n\n<p>INFO:tensorflow:Clearing out eager caches</p>\n\n<hr>\n\n<p>TypeError                                 Traceback (most recent call last)\n```</p>\n\n<p>Edit: I still don't know what caused this error, but it works now. I didn't change anything. Both the VM and TPU were restarted.</p>",
          "rawMarkdown": "Thanks. I too have a VM and TPU. The first tutorial used to work. But since yesterday my code errors:\n```\n# Detect hardware, return appropriate distribution strategy\ntry:\n    tpu = tf.distribute.cluster_resolver.TPUClusterResolver()  # TPU detection\n    print('Running on TPU ', tpu.cluster_spec().as_dict()['worker'])\nexcept ValueError:\n    tpu = None\n\nif tpu:\n    tf.config.experimental_connect_to_cluster(tpu)\n    tf.tpu.experimental.initialize_tpu_system(tpu)\n    strategy = tf.distribute.experimental.TPUStrategy(tpu)\nelse:\n    strategy = tf.distribute.get_strategy()\n\nprint(\"REPLICAS: \", strategy.num_replicas_in_sync)\n```\nDo you use a similar snippet to create the strategy? Does it work now?\n\nThis is what I get:\n```\nRunning on TPU  ['192.168.15.2:8470']\nINFO:tensorflow:Initializing the TPU system: rmme\n\nINFO:tensorflow:Initializing the TPU system: rmme\n\nINFO:tensorflow:Clearing out eager caches\n\nINFO:tensorflow:Clearing out eager caches\n\n---------------------------------------------------------------------------\nTypeError                                 Traceback (most recent call last)\n```\n\nEdit: I still don't know what caused this error, but it works now. I didn't change anything. Both the VM and TPU were restarted."
        },
        {
          "id": 698543,
          "postDate": "2019-12-19T11:19:25.227Z",
          "content": "<p>The variable \"strategy\" we use is exactly the same as below. \n<a href=\"https://github.com/tensorflow/models/blob/master/official/nlp/bert/run_squad.py\">https://github.com/tensorflow/models/blob/master/official/nlp/bert/run_squad.py</a>.</p>\n\n<p>My initial guess is the TPU instance TensorFlow version may not be compatible with your VM TensorFlow version. We use tf-nightly-2.x. It runs flawlessly.</p>",
          "rawMarkdown": "The variable \"strategy\" we use is exactly the same as below. \nhttps://github.com/tensorflow/models/blob/master/official/nlp/bert/run_squad.py.\n\nMy initial guess is the TPU instance TensorFlow version may not be compatible with your VM TensorFlow version. We use tf-nightly-2.x. It runs flawlessly.",
          "votes": 1
        }
      ]
    },
    {
      "id": 696073,
      "postDate": "2019-12-16T04:56:06.263Z",
      "content": "<p>Good news :)</p>",
      "rawMarkdown": "Good news :)"
    },
    {
      "id": 711385,
      "postDate": "2020-01-06T02:38:18.627Z",
      "rawMarkdown": "",
      "isDeleted": true
    }
  ],
  "comments": [
    {
      "id": 702403,
      "author_name": "Vopani",
      "author_url": "",
      "post_date": "2019-12-24T16:39:27.483000",
      "content": "<p>Will solutions built using TF 2.1 also be considered for the special TF 2.0 prizes? 🙂 </p>",
      "votes": 3,
      "replies": [
        {
          "id": 714775,
          "author_name": "Julia Elliott",
          "author_url": "",
          "post_date": "2020-01-09T18:22:37.370000",
          "content": "<p>Yes.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 703382,
      "author_name": "Jiwei Liu",
      "author_url": "",
      "post_date": "2019-12-26T04:36:35.123000",
      "content": "<p>I have two questions:\n1. if we use tf 2.1 to train a model, can we run inference on kaggle kernel with tf 2.0?\n2. is kaggle kernel gonna be updated to 2.1 as well? </p>\n\n<p>Thank you!</p>",
      "votes": 1,
      "replies": [
        {
          "id": 703897,
          "author_name": "Julia Elliott",
          "author_url": "",
          "post_date": "2019-12-26T19:18:12.603000",
          "content": "<p>Hi Jiwei,</p>\n\n<p>Good questions.\n1. Yes, this is what you'll have to do until the TF image on Kaggle is updated to 2.1\n2. We will, once TF2.1 has completed its full release.</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 703994,
          "author_name": "Jiwei Liu",
          "author_url": "",
          "post_date": "2019-12-26T23:15:06.727000",
          "content": "<p>Thank you Julia!</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 704034,
          "author_name": "Jiwei Liu",
          "author_url": "",
          "post_date": "2019-12-27T00:55:24.100000",
          "content": "<p>I just figured that upgrading to tf2.1 would require cuda 10.1 and consequently a system wise update of GPU driver. Could you please not upgrade kaggle docker before this competition ends? we already spend two months with tf 2.0 and set up environment as is. With a month left, I would like to focus on modeling rather than configuring environments and learning new APIs. 😂 </p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 714766,
          "author_name": "Julia Elliott",
          "author_url": "",
          "post_date": "2020-01-09T18:14:43.903000",
          "content": "<p>Correction: TF2.1 RC0 was already released into notebooks in early December, and the upgrade to the latest release will be made soon. However, our notebooks function such that notebooks created on your original image will be retained with that image. Therefore, if you have a preference for keeping your notebook in TF2.0, you're welcome to do so by not changing your existing notebook to the \"Latest Available\" image.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 695330,
      "author_name": "Rohit Agarwal",
      "author_url": "",
      "post_date": "2019-12-15T00:42:32.210000",
      "content": "<blockquote>\n  <p>The fine print, also included on the request form:\n  You must complete the pre-requisite setup steps listed on the form in order to receive quota.</p>\n</blockquote>\n\n<p>I have requested for the TPU quota. Does that mean that now I cannot team up with anyone who had also requested for a TPU quota?</p>",
      "votes": 1,
      "replies": [
        {
          "id": 696014,
          "author_name": "Paige Bailey",
          "author_url": "",
          "post_date": "2019-12-16T02:16:08.453000",
          "content": "<p>You can certainly still team up with other Kagglers who have requested TPU quota; however, only one person from each group will be allocated quota officially.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 696035,
          "author_name": "Rohit Agarwal",
          "author_url": "",
          "post_date": "2019-12-16T03:23:27.767000",
          "content": "<p>And then what will happen to the TPU quotas if the teaming up happens after credit allocation? What if those credits are also utilised (partially or completely)?</p>",
          "votes": 2,
          "replies": []
        }
      ]
    },
    {
      "id": 702609,
      "author_name": "Julia Elliott",
      "author_url": "",
      "post_date": "2019-12-24T22:55:14.220000",
      "content": "<p><a href=\"/rohanrao\">@rohanrao</a> good question - yes!</p>",
      "votes": 2,
      "replies": []
    },
    {
      "id": 715274,
      "author_name": "sakuranew",
      "author_url": "",
      "post_date": "2020-01-10T10:41:23.427000",
      "content": "<p>Does anybody try pytorch on Tpu successfully？I always use pytorch ...</p>",
      "votes": 0,
      "replies": [
        {
          "id": 715291,
          "author_name": "Zhiyu Guo",
          "author_url": "",
          "post_date": "2020-01-10T11:03:38.050000",
          "content": "<p>A few months ago(about October) I have tried to run pytorch on TPU to fine-tune BERT, the training process worked well, but I can't save the trained model, I can't get the solution at that time, then I give up to use pytorch on TPU. I am not sure if it works well now.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 715300,
          "author_name": "sakuranew",
          "author_url": "",
          "post_date": "2020-01-10T11:23:27.060000",
          "content": "<p>Does metric get worse? I see many bugs in issue</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 715311,
          "author_name": "Zhiyu Guo",
          "author_url": "",
          "post_date": "2020-01-10T11:37:35.410000",
          "content": "<p>Only slightly worse, I haven't try to fine-tune more at that time, maybe it can get better.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 715326,
          "author_name": "sakuranew",
          "author_url": "",
          "post_date": "2020-01-10T12:00:03.190000",
          "content": "<p>OK，thank you,but time is not enough to make a try</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 716400,
          "author_name": "Ranko Mosic",
          "author_url": "",
          "post_date": "2020-01-11T16:40:44.107000",
          "content": "<p>PyTorch on TPU is working. Some code additions are needed ( XLA ) </p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 698510,
      "author_name": "See--",
      "author_url": "",
      "post_date": "2019-12-19T10:27:16.513000",
      "content": "<p>Could someone confirm that these tutorials are working right now? I can't get them to work. Any help is appreciated. Am I doing something wrong?</p>\n\n<p><a href=\"https://github.com/GoogleCloudPlatform/training-data-analyst/issues/786\">https://github.com/GoogleCloudPlatform/training-data-analyst/issues/786</a>\n<a href=\"https://github.com/tensorflow/models/issues/7962\">https://github.com/tensorflow/models/issues/7962</a></p>",
      "votes": 0,
      "replies": [
        {
          "id": 698535,
          "author_name": "Shane",
          "author_url": "",
          "post_date": "2019-12-19T11:04:02.860000",
          "content": "<p>Hi See, firstly grats on your 1st public LB position. </p>\n\n<p>Our team has successfully set up using the TPU. In order to use the TPU, you will need to create the TPU session with TF-nightly-2.x within the same region, the training data need to be on google storage, then you'll need to pass the TPU name/IP address into experimental_connect_to_cluster. It is much much much faster than GPU training. Finetune the data with the NQ training data takes around 1 hour for 1 epoch.</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 698539,
          "author_name": "See--",
          "author_url": "",
          "post_date": "2019-12-19T11:10:42.037000",
          "content": "<p>Thanks. I too have a VM and TPU. The first tutorial used to work. But since yesterday my code errors:\n```</p>\n\n<h1>Detect hardware, return appropriate distribution strategy</h1>\n\n<p>try:\n    tpu = tf.distribute.cluster_resolver.TPUClusterResolver()  # TPU detection\n    print('Running on TPU ', tpu.cluster_spec().as_dict()['worker'])\nexcept ValueError:\n    tpu = None</p>\n\n<p>if tpu:\n    tf.config.experimental_connect_to_cluster(tpu)\n    tf.tpu.experimental.initialize_tpu_system(tpu)\n    strategy = tf.distribute.experimental.TPUStrategy(tpu)\nelse:\n    strategy = tf.distribute.get_strategy()</p>\n\n<p>print(\"REPLICAS: \", strategy.num_replicas_in_sync)\n```\nDo you use a similar snippet to create the strategy? Does it work now?</p>\n\n<p>This is what I get:\n```\nRunning on TPU  ['192.168.15.2:8470']\nINFO:tensorflow:Initializing the TPU system: rmme</p>\n\n<p>INFO:tensorflow:Initializing the TPU system: rmme</p>\n\n<p>INFO:tensorflow:Clearing out eager caches</p>\n\n<p>INFO:tensorflow:Clearing out eager caches</p>\n\n<hr>\n\n<p>TypeError                                 Traceback (most recent call last)\n```</p>\n\n<p>Edit: I still don't know what caused this error, but it works now. I didn't change anything. Both the VM and TPU were restarted.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 698543,
          "author_name": "Shane",
          "author_url": "",
          "post_date": "2019-12-19T11:19:25.227000",
          "content": "<p>The variable \"strategy\" we use is exactly the same as below. \n<a href=\"https://github.com/tensorflow/models/blob/master/official/nlp/bert/run_squad.py\">https://github.com/tensorflow/models/blob/master/official/nlp/bert/run_squad.py</a>.</p>\n\n<p>My initial guess is the TPU instance TensorFlow version may not be compatible with your VM TensorFlow version. We use tf-nightly-2.x. It runs flawlessly.</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 696073,
      "author_name": "Chanran Kim",
      "author_url": "",
      "post_date": "2019-12-16T04:56:06.263000",
      "content": "<p>Good news :)</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 711385,
      "author_name": "",
      "author_url": "",
      "post_date": "2020-01-06T02:38:18.627000",
      "content": "",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "694823": "## You can now use TensorFlow 2.1 RC on Cloud TPUs via `tf-nightly`!\n\nDetails can be found in [this installation guide](https://github.com/GoogleCloudPlatform/training-data-analyst/blob/master/courses/fast-and-lean-data-science/README-TF2.1.md) from the Kaggle and Google Cloud teams. We are also excited to release an example of [BERT fine-tuning with Cloud TPUs](https://github.com/tensorflow/models/blob/master/official/nlp/bert/bert_cloud_tpu.md), using TensorFlow. \n\n**Note**: Given that `tf-nightly` is a release candidate, not an official release, you may run into rough edges -- so please make sure to **post your questions** in the Kaggle forums, and submit bugs on [Github](https://www.github.com/tensorflow/tensorflow)!\n\n### We are also still offering Cloud TPU Quota for entrants in this competition.\n\nRequests can be made by submitting [this form](https://www.kaggle.com/TPU-Tensorflow2-QA) by **December 16, 2019**.\n\nWith the upcoming official release of TF 2.1, TPU support with `tf.keras` and `tf.distribute` are anticipated to greatly improve developer productivity while taking advantage of the performance boosts that TPUs are already known for. The TensorFlow team is preparing additional tutorial information which will accompany the quota distribution. \n\nIf you've been curious about TPUs, now would be a great opportunity to give them a try!\n\n### The fine print, also included on the request form:\n\n* You must complete the pre-requisite setup steps listed on the form in order to receive quota.\n* Only one person per team will be issued a TPU quota. This will be verified and multiple requests from the same team will be de-duplicated.\n* Be aware that if you are eligible, your request will be fulfilled by December 23, 2019, so you may expect to wait a couple weeks before receiving an email that confirms your quota has been provisioned. Distribution will not be immediate.",
    "702403": "Will solutions built using TF 2.1 also be considered for the special TF 2.0 prizes? 🙂 ",
    "703382": "I have two questions:\n1. if we use tf 2.1 to train a model, can we run inference on kaggle kernel with tf 2.0?\n2. is kaggle kernel gonna be updated to 2.1 as well? \n\nThank you!",
    "695330": "&gt; The fine print, also included on the request form:\nYou must complete the pre-requisite setup steps listed on the form in order to receive quota.\n\nI have requested for the TPU quota. Does that mean that now I cannot team up with anyone who had also requested for a TPU quota?\n",
    "702609": "@rohanrao good question - yes!",
    "715274": "Does anybody try pytorch on Tpu successfully？I always use pytorch ...",
    "698510": "Could someone confirm that these tutorials are working right now? I can't get them to work. Any help is appreciated. Am I doing something wrong?\n\nhttps://github.com/GoogleCloudPlatform/training-data-analyst/issues/786\nhttps://github.com/tensorflow/models/issues/7962",
    "696073": "Good news :)",
    "711385": ""
  }
}