{
  "id": 130135,
  "title": "Why running on CPU and not TPU ?",
  "url": "/competitions/flower-classification-with-tpus/discussion/130135",
  "author_name": "Catadanna",
  "post_date": "2020-02-12T10:38:44.554000",
  "votes": 2,
  "comment_count": 9,
  "views": 0,
  "content": "<p>Hallo,</p>\n\n<p>I run my code using the Getting started notebook, and I run the cell where the use of TPU is activated. Then, I replaced the initial model with my model. Now, when running the cell where the model is compiled, I see that the CPU is used (at the top left corner of the notebook). </p>\n\n<p>Does someone know  why this happens and what shall I do in order to compile the model on TPU? </p>",
  "messages": [
    {
      "id": 743862,
      "postDate": "2020-02-12T10:38:44.553Z",
      "content": "<p>Hallo,</p>\n\n<p>I run my code using the Getting started notebook, and I run the cell where the use of TPU is activated. Then, I replaced the initial model with my model. Now, when running the cell where the model is compiled, I see that the CPU is used (at the top left corner of the notebook). </p>\n\n<p>Does someone know  why this happens and what shall I do in order to compile the model on TPU? </p>",
      "rawMarkdown": "Hallo,\n\nI run my code using the Getting started notebook, and I run the cell where the use of TPU is activated. Then, I replaced the initial model with my model. Now, when running the cell where the model is compiled, I see that the CPU is used (at the top left corner of the notebook). \n\nDoes someone know  why this happens and what shall I do in order to compile the model on TPU? ",
      "votes": 2
    },
    {
      "id": 1532084,
      "postDate": "2021-10-02T15:51:25.943Z",
      "content": "<p>How to use TPU for machine learning intern projects for a students </p>",
      "rawMarkdown": "How to use TPU for machine learning intern projects for a students "
    },
    {
      "id": 744350,
      "postDate": "2020-02-12T18:58:03.140Z",
      "content": "<p>Hallo, thay is exactly what I did. Model definition is in the <code>with strategy.scope()</code> block but the model compilation, as well as fitting the model,  is outside that loop, right? At least, this is the structure which appears in the getting started notebook. </p>\n\n<p>Thank you for the explanation, Martin, it seems that it is OK! </p>",
      "rawMarkdown": "Hallo, thay is exactly what I did. Model definition is in the `with strategy.scope()` block but the model compilation, as well as fitting the model,  is outside that loop, right? At least, this is the structure which appears in the getting started notebook. \n\nThank you for the explanation, Martin, it seems that it is OK! ",
      "replies": [
        {
          "id": 744384,
          "postDate": "2020-02-12T19:32:43.747Z",
          "content": "<p>I assume you meant \"strategy scope\", not \"while loop\" but yes, that's all correct. That you for confirming that everything is OK now.</p>",
          "rawMarkdown": "I assume you meant \"strategy scope\", not \"while loop\" but yes, that's all correct. That you for confirming that everything is OK now."
        },
        {
          "id": 744386,
          "postDate": "2020-02-12T19:34:42.757Z",
          "content": "<p>Yes <code>strategy.scope()</code>, indeed. I corrected it. In fact it works on CPU until it arrives at that line of code.</p>",
          "rawMarkdown": "Yes `strategy.scope()`, indeed. I corrected it. In fact it works on CPU until it arrives at that line of code."
        }
      ]
    },
    {
      "id": 743981,
      "postDate": "2020-02-12T13:04:30.323Z",
      "content": "<p><a href=\"/catadanna\">@catadanna</a> \nDid you enable TPU training in Tensorflow Keras ?\nLook at <a href=\"https://www.kaggle.com/docs/tpu\">https://www.kaggle.com/docs/tpu</a></p>",
      "rawMarkdown": "@catadanna \nDid you enable TPU training in Tensorflow Keras ?\nLook at [https://www.kaggle.com/docs/tpu](https://www.kaggle.com/docs/tpu)",
      "replies": [
        {
          "id": 744010,
          "postDate": "2020-02-12T13:21:10.413Z",
          "content": "<p>Yes, I took it from the Getting started notebook, as I mentioned before. I run the cell which contains this code  : </p>\n\n<p>```\ntry:\n    tpu = tf.distribute.cluster_resolver.TPUClusterResolver()  # TPU detection. No parameters necessary if TPU_NAME environment variable is set. On Kaggle this is always the case.\n    print('Running on TPU ', tpu.master())\nexcept ValueError:\n    tpu = None</p>\n\n<p>if tpu:\n    tf.config.experimental_connect_to_cluster(tpu)\n    tf.tpu.experimental.initialize_tpu_system(tpu)\n    strategy = tf.distribute.experimental.TPUStrategy(tpu)\nelse:\n    strategy = tf.distribute.get_strategy() # default distribution strategy in Tensorflow. Works on CPU and single GPU.</p>\n\n<p>print(\"REPLICAS: \", strategy.num_replicas_in_sync)\n```</p>\n\n<p>Printed : </p>\n\n<p><code>Running on TPU grpc://10.0.0.2:8470</code></p>",
          "rawMarkdown": "Yes, I took it from the Getting started notebook, as I mentioned before. I run the cell which contains this code  : \n\n\n```\ntry:\n    tpu = tf.distribute.cluster_resolver.TPUClusterResolver()  # TPU detection. No parameters necessary if TPU_NAME environment variable is set. On Kaggle this is always the case.\n    print('Running on TPU ', tpu.master())\nexcept ValueError:\n    tpu = None\n\nif tpu:\n    tf.config.experimental_connect_to_cluster(tpu)\n    tf.tpu.experimental.initialize_tpu_system(tpu)\n    strategy = tf.distribute.experimental.TPUStrategy(tpu)\nelse:\n    strategy = tf.distribute.get_strategy() # default distribution strategy in Tensorflow. Works on CPU and single GPU.\n\nprint(\"REPLICAS: \", strategy.num_replicas_in_sync)\n```\n\nPrinted : \n\n`Running on TPU grpc://10.0.0.2:8470`"
        },
        {
          "id": 744318,
          "postDate": "2020-02-12T18:22:11.093Z",
          "content": "<p>The line that places your model on the TPU is:</p>\n\n<p>```\nwith strategy.scope():\n    # model definition goes here</p>\n\n<p><code>``\nThe</code>strategy<code>is the</code>tf.distribute.experimental.TPUStrategy(tpu)` object you created earlier.\nmodel.compile, model.fit, model.predict and model.evaluate do not need to be in the strategy scope if the model was created in the scope. The model will remember it is on a TPU and use it.</p>\n\n<p>Also, a high CPU usage is normal in the first epoch of training. Your model is being compiled for TPU. Subsequent epochs will be fast and you will see TPU usage in the gauges under TPU MXU (Matrix Multiply Unit):\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F4003597%2F4cb90fd7a5c5f4d1d3069191a65e58b0%2FScreen%20Shot%202020-02-12%20at%2010.18.22.png?generation=1581531697606241&amp;alt=media\" alt=\"\"></p>",
          "rawMarkdown": "The line that places your model on the TPU is:\n\n```\nwith strategy.scope():\n    # model definition goes here\n\n```\nThe `strategy` is the `tf.distribute.experimental.TPUStrategy(tpu)` object you created earlier.\nmodel.compile, model.fit, model.predict and model.evaluate do not need to be in the strategy scope if the model was created in the scope. The model will remember it is on a TPU and use it.\n\nAlso, a high CPU usage is normal in the first epoch of training. Your model is being compiled for TPU. Subsequent epochs will be fast and you will see TPU usage in the gauges under TPU MXU (Matrix Multiply Unit):\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F4003597%2F4cb90fd7a5c5f4d1d3069191a65e58b0%2FScreen%20Shot%202020-02-12%20at%2010.18.22.png?generation=1581531697606241&amp;alt=media)\n\n",
          "votes": 4
        },
        {
          "id": 881060,
          "postDate": "2020-06-10T18:30:03.953Z",
          "content": "<p>Hello <a href=\"/mgornergoogle\">@mgornergoogle</a>, I was wondering if it was possible to train a model greater than 16GB in size?\nI keep running out of memory in the first epoch while training on the TPU while the GPU training goes through fine. </p>",
          "rawMarkdown": "Hello @mgornergoogle, I was wondering if it was possible to train a model greater than 16GB in size?\nI keep running out of memory in the first epoch while training on the TPU while the GPU training goes through fine. "
        }
      ]
    },
    {
      "id": 1314225,
      "postDate": "2021-05-19T03:40:12.500Z",
      "rawMarkdown": "",
      "isDeleted": true
    }
  ],
  "comments": [
    {
      "id": 1532084,
      "author_name": "Sathish S",
      "author_url": "",
      "post_date": "2021-10-02T15:51:25.943000",
      "content": "<p>How to use TPU for machine learning intern projects for a students </p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 744350,
      "author_name": "Catadanna",
      "author_url": "",
      "post_date": "2020-02-12T18:58:03.140000",
      "content": "<p>Hallo, thay is exactly what I did. Model definition is in the <code>with strategy.scope()</code> block but the model compilation, as well as fitting the model,  is outside that loop, right? At least, this is the structure which appears in the getting started notebook. </p>\n\n<p>Thank you for the explanation, Martin, it seems that it is OK! </p>",
      "votes": 0,
      "replies": [
        {
          "id": 744384,
          "author_name": "Martin Görner",
          "author_url": "",
          "post_date": "2020-02-12T19:32:43.747000",
          "content": "<p>I assume you meant \"strategy scope\", not \"while loop\" but yes, that's all correct. That you for confirming that everything is OK now.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 744386,
          "author_name": "Catadanna",
          "author_url": "",
          "post_date": "2020-02-12T19:34:42.757000",
          "content": "<p>Yes <code>strategy.scope()</code>, indeed. I corrected it. In fact it works on CPU until it arrives at that line of code.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 743981,
      "author_name": "Pawel Lubiński",
      "author_url": "",
      "post_date": "2020-02-12T13:04:30.323000",
      "content": "<p><a href=\"/catadanna\">@catadanna</a> \nDid you enable TPU training in Tensorflow Keras ?\nLook at <a href=\"https://www.kaggle.com/docs/tpu\">https://www.kaggle.com/docs/tpu</a></p>",
      "votes": 0,
      "replies": [
        {
          "id": 744010,
          "author_name": "Catadanna",
          "author_url": "",
          "post_date": "2020-02-12T13:21:10.413000",
          "content": "<p>Yes, I took it from the Getting started notebook, as I mentioned before. I run the cell which contains this code  : </p>\n\n<p>```\ntry:\n    tpu = tf.distribute.cluster_resolver.TPUClusterResolver()  # TPU detection. No parameters necessary if TPU_NAME environment variable is set. On Kaggle this is always the case.\n    print('Running on TPU ', tpu.master())\nexcept ValueError:\n    tpu = None</p>\n\n<p>if tpu:\n    tf.config.experimental_connect_to_cluster(tpu)\n    tf.tpu.experimental.initialize_tpu_system(tpu)\n    strategy = tf.distribute.experimental.TPUStrategy(tpu)\nelse:\n    strategy = tf.distribute.get_strategy() # default distribution strategy in Tensorflow. Works on CPU and single GPU.</p>\n\n<p>print(\"REPLICAS: \", strategy.num_replicas_in_sync)\n```</p>\n\n<p>Printed : </p>\n\n<p><code>Running on TPU grpc://10.0.0.2:8470</code></p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 744318,
          "author_name": "Martin Görner",
          "author_url": "",
          "post_date": "2020-02-12T18:22:11.093000",
          "content": "<p>The line that places your model on the TPU is:</p>\n\n<p>```\nwith strategy.scope():\n    # model definition goes here</p>\n\n<p><code>``\nThe</code>strategy<code>is the</code>tf.distribute.experimental.TPUStrategy(tpu)` object you created earlier.\nmodel.compile, model.fit, model.predict and model.evaluate do not need to be in the strategy scope if the model was created in the scope. The model will remember it is on a TPU and use it.</p>\n\n<p>Also, a high CPU usage is normal in the first epoch of training. Your model is being compiled for TPU. Subsequent epochs will be fast and you will see TPU usage in the gauges under TPU MXU (Matrix Multiply Unit):\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F4003597%2F4cb90fd7a5c5f4d1d3069191a65e58b0%2FScreen%20Shot%202020-02-12%20at%2010.18.22.png?generation=1581531697606241&amp;alt=media\" alt=\"\"></p>",
          "votes": 4,
          "replies": []
        },
        {
          "id": 881060,
          "author_name": "Arth Dh",
          "author_url": "",
          "post_date": "2020-06-10T18:30:03.953000",
          "content": "<p>Hello <a href=\"/mgornergoogle\">@mgornergoogle</a>, I was wondering if it was possible to train a model greater than 16GB in size?\nI keep running out of memory in the first epoch while training on the TPU while the GPU training goes through fine. </p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 1314225,
      "author_name": "",
      "author_url": "",
      "post_date": "2021-05-19T03:40:12.500000",
      "content": "",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "743862": "Hallo,\n\nI run my code using the Getting started notebook, and I run the cell where the use of TPU is activated. Then, I replaced the initial model with my model. Now, when running the cell where the model is compiled, I see that the CPU is used (at the top left corner of the notebook). \n\nDoes someone know  why this happens and what shall I do in order to compile the model on TPU? ",
    "1532084": "How to use TPU for machine learning intern projects for a students ",
    "744350": "Hallo, thay is exactly what I did. Model definition is in the `with strategy.scope()` block but the model compilation, as well as fitting the model,  is outside that loop, right? At least, this is the structure which appears in the getting started notebook. \n\nThank you for the explanation, Martin, it seems that it is OK! ",
    "743981": "@catadanna \nDid you enable TPU training in Tensorflow Keras ?\nLook at [https://www.kaggle.com/docs/tpu](https://www.kaggle.com/docs/tpu)",
    "1314225": ""
  }
}