{
  "id": 144182,
  "title": "Can not use tf-nightly with tpu ?",
  "url": "/competitions/jigsaw-multilingual-toxic-comment-classification/discussion/144182",
  "author_name": "",
  "post_date": "2020-04-18T03:10:37.929232600Z",
  "votes": null,
  "comment_count": 3,
  "views": 0,
  "content": "<p>142     if tpu:\n    143         tf.config.experimental_connect_to_cluster(tpu)\n--&gt; 144         tf.tpu.experimental.initialize_tpu_system(tpu)\n    145         strategy = tf.distribute.experimental.TPUStrategy(tpu)\n    146         melt.distributed.set_strategy(strategy)</p>\n\n<p>/opt/conda/lib/python3.6/site-packages/tensorflow/python/tpu/tpu_strategy_util.py in initialize_tpu_system(cluster_resolver)\n    101     context.context()._clear_caches()  # pylint: disable=protected-access\n    102 \n--&gt; 103     serialized_topology = output.numpy()\n    104 \n    105     # TODO(b/134094971): Remove this when lazy tensor copy in multi-device</p>\n\n<p>/opt/conda/lib/python3.6/site-packages/tensorflow/python/framework/ops.py in numpy(self)\n   1106     \"\"\"\n   1107     # TODO(slebedev): Consider avoiding a copy for non-CPU or remote tensors.\n-&gt; 1108     maybe_arr = self._numpy()  # pylint: disable=protected-access\n   1109     return maybe_arr.copy() if isinstance(maybe_arr, np.ndarray) else maybe_arr\n   1110 </p>\n\n<p>/opt/conda/lib/python3.6/site-packages/tensorflow/python/framework/ops.py in _numpy(self)\n   1074       return self._numpy_internal()\n   1075     except core._NotOkStatusException as e:\n-&gt; 1076       six.raise_from(core._status_to_exception(e.code, e.message), None)\n   1077 \n   1078   @property</p>\n\n<p>/opt/conda/lib/python3.6/site-packages/six.py in raise_from(value, from_value)</p>\n\n<p>InvalidArgumentError: NodeDef expected inputs 'string' do not match 0 inputs specified; Op ; attr=T:type; attr=tensor_name:string; attr=send_device:string; attr=send_device_incarnation:int; attr=recv_device:string; attr=client_terminated:bool,default=false; is_stateful=true&gt;; NodeDef: {{node _Send}}</p>",
  "messages": [
    {
      "id": "811534",
      "postDate": "04/18/2020 03:10:37",
      "content": "<p>142     if tpu:\n    143         tf.config.experimental_connect_to_cluster(tpu)\n--&gt; 144         tf.tpu.experimental.initialize_tpu_system(tpu)\n    145         strategy = tf.distribute.experimental.TPUStrategy(tpu)\n    146         melt.distributed.set_strategy(strategy)</p>\n\n<p>/opt/conda/lib/python3.6/site-packages/tensorflow/python/tpu/tpu_strategy_util.py in initialize_tpu_system(cluster_resolver)\n    101     context.context()._clear_caches()  # pylint: disable=protected-access\n    102 \n--&gt; 103     serialized_topology = output.numpy()\n    104 \n    105     # TODO(b/134094971): Remove this when lazy tensor copy in multi-device</p>\n\n<p>/opt/conda/lib/python3.6/site-packages/tensorflow/python/framework/ops.py in numpy(self)\n   1106     \"\"\"\n   1107     # TODO(slebedev): Consider avoiding a copy for non-CPU or remote tensors.\n-&gt; 1108     maybe_arr = self._numpy()  # pylint: disable=protected-access\n   1109     return maybe_arr.copy() if isinstance(maybe_arr, np.ndarray) else maybe_arr\n   1110 </p>\n\n<p>/opt/conda/lib/python3.6/site-packages/tensorflow/python/framework/ops.py in _numpy(self)\n   1074       return self._numpy_internal()\n   1075     except core._NotOkStatusException as e:\n-&gt; 1076       six.raise_from(core._status_to_exception(e.code, e.message), None)\n   1077 \n   1078   @property</p>\n\n<p>/opt/conda/lib/python3.6/site-packages/six.py in raise_from(value, from_value)</p>\n\n<p>InvalidArgumentError: NodeDef expected inputs 'string' do not match 0 inputs specified; Op ; attr=T:type; attr=tensor_name:string; attr=send_device:string; attr=send_device_incarnation:int; attr=recv_device:string; attr=client_terminated:bool,default=false; is_stateful=true&gt;; NodeDef: {{node _Send}}</p>",
      "rawMarkdown": "142     if tpu:\n    143         tf.config.experimental_connect_to_cluster(tpu)\n--&gt; 144         tf.tpu.experimental.initialize_tpu_system(tpu)\n    145         strategy = tf.distribute.experimental.TPUStrategy(tpu)\n    146         melt.distributed.set_strategy(strategy)\n\n/opt/conda/lib/python3.6/site-packages/tensorflow/python/tpu/tpu_strategy_util.py in initialize_tpu_system(cluster_resolver)\n    101     context.context()._clear_caches()  # pylint: disable=protected-access\n    102 \n--&gt; 103     serialized_topology = output.numpy()\n    104 \n    105     # TODO(b/134094971): Remove this when lazy tensor copy in multi-device\n\n/opt/conda/lib/python3.6/site-packages/tensorflow/python/framework/ops.py in numpy(self)\n   1106     \"\"\"\n   1107     # TODO(slebedev): Consider avoiding a copy for non-CPU or remote tensors.\n-&gt; 1108     maybe_arr = self._numpy()  # pylint: disable=protected-access\n   1109     return maybe_arr.copy() if isinstance(maybe_arr, np.ndarray) else maybe_arr\n   1110 \n\n/opt/conda/lib/python3.6/site-packages/tensorflow/python/framework/ops.py in _numpy(self)\n   1074       return self._numpy_internal()\n   1075     except core._NotOkStatusException as e:\n-&gt; 1076       six.raise_from(core._status_to_exception(e.code, e.message), None)\n   1077 \n   1078   @property\n\n/opt/conda/lib/python3.6/site-packages/six.py in raise_from(value, from_value)\n\nInvalidArgumentError: NodeDef expected inputs 'string' do not match 0 inputs specified; Op",
      "votes": null
    },
    {
      "id": "811538",
      "postDate": "04/18/2020 03:22:55",
      "content": "<p>I want to try tf-nigtly, for on my local computer with tf2.2 seems no memory leak but on kaggle kernel with tpu and tf2.1, seems calling model.predict multiple times will cause serious memory leak problem. model.predict_on_batch is a workaround but it is super slow comparing to model.predict.</p>",
      "rawMarkdown": "I want to try tf-nigtly, for on my local computer with tf2.2 seems no memory leak but on kaggle kernel with tpu and tf2.1, seems calling model.predict multiple times will cause serious memory leak problem. model.predict_on_batch is a workaround but it is super slow comparing to model.predict.",
      "votes": null
    },
    {
      "id": "819580",
      "postDate": "04/24/2020 17:25:09",
      "content": "<p>Hi <a href=\"/alicexfeng1987\">@alicexfeng1987</a> - please do not use tf-nightly with TPUs. when you install tf-nightly it will only update half of the software that needs updating, and the TPU side of things <em>doesn't</em> get updated, and you'll ultimately end up with communication problems between your notebook and the TPU. </p>",
      "rawMarkdown": "Hi @alicexfeng1987 - please do not use tf-nightly with TPUs. when you install tf-nightly it will only update half of the software that needs updating, and the TPU side of things _doesn't_ get updated, and you'll ultimately end up with communication problems between your notebook and the TPU.",
      "votes": null
    },
    {
      "id": "819890",
      "postDate": "04/25/2020 00:46:28",
      "content": "<p><a href=\"/jessemostipak\">@jessemostipak</a>  Got it , thanks!</p>",
      "rawMarkdown": "jessemostipak  Got it , thanks!",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 811538,
      "author_name": "alicexfeng1987",
      "author_url": "",
      "post_date": "04/18/2020 03:22:55",
      "content": "<p>I want to try tf-nigtly, for on my local computer with tf2.2 seems no memory leak but on kaggle kernel with tpu and tf2.1, seems calling model.predict multiple times will cause serious memory leak problem. model.predict_on_batch is a workaround but it is super slow comparing to model.predict.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 819580,
      "author_name": "jessemostipak",
      "author_url": "",
      "post_date": "04/24/2020 17:25:09",
      "content": "<p>Hi <a href=\"/alicexfeng1987\">@alicexfeng1987</a> - please do not use tf-nightly with TPUs. when you install tf-nightly it will only update half of the software that needs updating, and the TPU side of things <em>doesn't</em> get updated, and you'll ultimately end up with communication problems between your notebook and the TPU. </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 819890,
      "author_name": "alicexfeng1987",
      "author_url": "",
      "post_date": "04/25/2020 00:46:28",
      "content": "<p><a href=\"/jessemostipak\">@jessemostipak</a>  Got it , thanks!</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "811534": "142     if tpu:\n    143         tf.config.experimental_connect_to_cluster(tpu)\n--&gt; 144         tf.tpu.experimental.initialize_tpu_system(tpu)\n    145         strategy = tf.distribute.experimental.TPUStrategy(tpu)\n    146         melt.distributed.set_strategy(strategy)\n\n/opt/conda/lib/python3.6/site-packages/tensorflow/python/tpu/tpu_strategy_util.py in initialize_tpu_system(cluster_resolver)\n    101     context.context()._clear_caches()  # pylint: disable=protected-access\n    102 \n--&gt; 103     serialized_topology = output.numpy()\n    104 \n    105     # TODO(b/134094971): Remove this when lazy tensor copy in multi-device\n\n/opt/conda/lib/python3.6/site-packages/tensorflow/python/framework/ops.py in numpy(self)\n   1106     \"\"\"\n   1107     # TODO(slebedev): Consider avoiding a copy for non-CPU or remote tensors.\n-&gt; 1108     maybe_arr = self._numpy()  # pylint: disable=protected-access\n   1109     return maybe_arr.copy() if isinstance(maybe_arr, np.ndarray) else maybe_arr\n   1110 \n\n/opt/conda/lib/python3.6/site-packages/tensorflow/python/framework/ops.py in _numpy(self)\n   1074       return self._numpy_internal()\n   1075     except core._NotOkStatusException as e:\n-&gt; 1076       six.raise_from(core._status_to_exception(e.code, e.message), None)\n   1077 \n   1078   @property\n\n/opt/conda/lib/python3.6/site-packages/six.py in raise_from(value, from_value)\n\nInvalidArgumentError: NodeDef expected inputs 'string' do not match 0 inputs specified; Op",
    "811538": "I want to try tf-nigtly, for on my local computer with tf2.2 seems no memory leak but on kaggle kernel with tpu and tf2.1, seems calling model.predict multiple times will cause serious memory leak problem. model.predict_on_batch is a workaround but it is super slow comparing to model.predict.",
    "819580": "Hi @alicexfeng1987 - please do not use tf-nightly with TPUs. when you install tf-nightly it will only update half of the software that needs updating, and the TPU side of things _doesn't_ get updated, and you'll ultimately end up with communication problems between your notebook and the TPU.",
    "819890": "jessemostipak  Got it , thanks!"
  },
  "source": "meta"
}