{
  "id": 229459,
  "title": "TPU Socket closed ERROR",
  "url": "/competitions/bms-molecular-translation/discussion/229459",
  "author_name": "",
  "post_date": "2021-03-30T09:08:40.249361800Z",
  "votes": 6,
  "comment_count": 3,
  "views": 0,
  "content": "<p>Any suggestions about the following error raises while tpu training pipeline,<br>\n🙏🙏🙏🙏🙏</p>\n<p><code>UnavailableError: Socket closed\nAdditional GRPC error information from remote target /job:worker/replica:0/task:0:\n:{\"created\":\"@1617020379.604834479\",\"description\":\"Error received from peer ipv4:10.0.0.2:8470\",\"file\":\"external/com_github_grpc_grpc/src/core/lib/surface/call.cc\",\"file_line\":1056,\"grpc_message\":\"Socket closed\",\"grpc_status\":14}</code></p>\n<p>~thank you</p>",
  "messages": [
    {
      "id": "1256803",
      "postDate": "03/30/2021 09:08:40",
      "content": "<p>Any suggestions about the following error raises while tpu training pipeline,<br>\n🙏🙏🙏🙏🙏</p>\n<p><code>UnavailableError: Socket closed\nAdditional GRPC error information from remote target /job:worker/replica:0/task:0:\n:{\"created\":\"@1617020379.604834479\",\"description\":\"Error received from peer ipv4:10.0.0.2:8470\",\"file\":\"external/com_github_grpc_grpc/src/core/lib/surface/call.cc\",\"file_line\":1056,\"grpc_message\":\"Socket closed\",\"grpc_status\":14}</code></p>\n<p>~thank you</p>",
      "rawMarkdown": "Any suggestions about the following error raises while tpu training pipeline,\n🙏🙏🙏🙏🙏\n\n`UnavailableError: Socket closed\nAdditional GRPC error information from remote target /job:worker/replica:0/task:0:\n:{\"created\":\"@1617020379.604834479\",\"description\":\"Error received from peer ipv4:10.0.0.2:8470\",\"file\":\"external/com_github_grpc_grpc/src/core/lib/surface/call.cc\",\"file_line\":1056,\"grpc_message\":\"Socket closed\",\"grpc_status\":14}`\n\n~thank you",
      "votes": null
    },
    {
      "id": "1256874",
      "postDate": "03/30/2021 10:21:43",
      "content": "<ol>\n<li>Check input dimensions. First (0th) dimension should be multiples of 32</li>\n<li>Are you using too much memory? reduce batch size and check</li>\n<li>Are you using LayerNorm? Try removing some layers and check</li>\n</ol>",
      "rawMarkdown": "1. Check input dimensions. First (0th) dimension should be multiples of 32\n2. Are you using too much memory? reduce batch size and check\n3. Are you using LayerNorm? Try removing some layers and check",
      "votes": null
    },
    {
      "id": "1257068",
      "postDate": "03/30/2021 13:50:31",
      "content": "<p>All req. are fulfilled… but still prob. continues…</p>\n<p>do check relevant notebook </p>\n<p><a href=\"https://www.kaggle.com/akhileshdkapse/vector-to-sequence-modeling-tpu-part-ii?rvi=1\" target=\"_blank\">https://www.kaggle.com/akhileshdkapse/vector-to-sequence-modeling-tpu-part-ii?rvi=1</a></p>",
      "rawMarkdown": "All req. are fulfilled... but still prob. continues...\n\ndo check relevant notebook \n\nhttps://www.kaggle.com/akhileshdkapse/vector-to-sequence-modeling-tpu-part-ii?rvi=1",
      "votes": null
    },
    {
      "id": "1257173",
      "postDate": "03/30/2021 15:24:50",
      "content": "<p>1) As the error occurs on the last epoch step it might be due to the last batch having a different size, you could try setting <code>drop_remainder=True</code>.</p>\n<pre><code>dset = dset.batch(bsize, drop_remainder=True).prefetch(AUTO)\n</code></pre>\n<p>2) Try setting the <code>steps_per_epoch=1</code>  and see whether the error still occurs.</p>\n<p>3) Try removing all callbacks, it is kind of fishy the error happens after the last epoch step.</p>",
      "rawMarkdown": "1) As the error occurs on the last epoch step it might be due to the last batch having a different size, you could try setting `drop_remainder=True`.\n\n```\ndset = dset.batch(bsize, drop_remainder=True).prefetch(AUTO)\n```\n\n2) Try setting the `steps_per_epoch=1`  and see whether the error still occurs.\n\n3) Try removing all callbacks, it is kind of fishy the error happens after the last epoch step.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1256874,
      "author_name": "sirishks",
      "author_url": "",
      "post_date": "03/30/2021 10:21:43",
      "content": "<ol>\n<li>Check input dimensions. First (0th) dimension should be multiples of 32</li>\n<li>Are you using too much memory? reduce batch size and check</li>\n<li>Are you using LayerNorm? Try removing some layers and check</li>\n</ol>",
      "votes": null,
      "replies": [
        {
          "id": 1257068,
          "author_name": "akhileshdkapse",
          "author_url": "",
          "post_date": "03/30/2021 13:50:31",
          "content": "<p>All req. are fulfilled… but still prob. continues…</p>\n<p>do check relevant notebook </p>\n<p><a href=\"https://www.kaggle.com/akhileshdkapse/vector-to-sequence-modeling-tpu-part-ii?rvi=1\" target=\"_blank\">https://www.kaggle.com/akhileshdkapse/vector-to-sequence-modeling-tpu-part-ii?rvi=1</a></p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1257173,
      "author_name": "markwijkhuizen",
      "author_url": "",
      "post_date": "03/30/2021 15:24:50",
      "content": "<p>1) As the error occurs on the last epoch step it might be due to the last batch having a different size, you could try setting <code>drop_remainder=True</code>.</p>\n<pre><code>dset = dset.batch(bsize, drop_remainder=True).prefetch(AUTO)\n</code></pre>\n<p>2) Try setting the <code>steps_per_epoch=1</code>  and see whether the error still occurs.</p>\n<p>3) Try removing all callbacks, it is kind of fishy the error happens after the last epoch step.</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1256803": "Any suggestions about the following error raises while tpu training pipeline,\n🙏🙏🙏🙏🙏\n\n`UnavailableError: Socket closed\nAdditional GRPC error information from remote target /job:worker/replica:0/task:0:\n:{\"created\":\"@1617020379.604834479\",\"description\":\"Error received from peer ipv4:10.0.0.2:8470\",\"file\":\"external/com_github_grpc_grpc/src/core/lib/surface/call.cc\",\"file_line\":1056,\"grpc_message\":\"Socket closed\",\"grpc_status\":14}`\n\n~thank you",
    "1256874": "1. Check input dimensions. First (0th) dimension should be multiples of 32\n2. Are you using too much memory? reduce batch size and check\n3. Are you using LayerNorm? Try removing some layers and check",
    "1257068": "All req. are fulfilled... but still prob. continues...\n\ndo check relevant notebook \n\nhttps://www.kaggle.com/akhileshdkapse/vector-to-sequence-modeling-tpu-part-ii?rvi=1",
    "1257173": "1) As the error occurs on the last epoch step it might be due to the last batch having a different size, you could try setting `drop_remainder=True`.\n\n```\ndset = dset.batch(bsize, drop_remainder=True).prefetch(AUTO)\n```\n\n2) Try setting the `steps_per_epoch=1`  and see whether the error still occurs.\n\n3) Try removing all callbacks, it is kind of fishy the error happens after the last epoch step."
  },
  "source": "meta"
}