{
  "id": 265195,
  "title": "\"TPU Socket closed\" problem",
  "url": "/competitions/rsna-miccai-brain-tumor-radiogenomic-classification/discussion/265195",
  "author_name": "",
  "post_date": "2021-08-15T02:19:22.097095800Z",
  "votes": 3,
  "comment_count": 3,
  "views": 0,
  "content": "<p>While calling <code>model.fit(tf_dataset,...)</code> I'm getting the following error:</p>\n<pre><code>UnavailableError: 2 root error(s) found.\n  (0) Unavailable: Socket closed\n  (1) Unavailable: Unable to find a context_id matching the specified one (6747970432017632323). Perhaps the worker was restarted, or the context was GC'd?\n0 successful operations.\n0 derived errors ignored.\n</code></pre>\n<p>Each TfRecord is only 4 MB size and represents the <strong>tf.float32</strong> matrix of <strong>[500,2000]</strong> size. The case is very hard to debug and playing with batch size, parallel reads, and data type (float64, int32, etc) doesn't help.<br>\nDecreasing the matrix size to [1,2000] didn't help either.</p>\n<p>Your hints are appreciated.</p>",
  "messages": [
    {
      "id": "1472611",
      "postDate": "08/15/2021 02:19:22",
      "content": "<p>While calling <code>model.fit(tf_dataset,...)</code> I'm getting the following error:</p>\n<pre><code>UnavailableError: 2 root error(s) found.\n  (0) Unavailable: Socket closed\n  (1) Unavailable: Unable to find a context_id matching the specified one (6747970432017632323). Perhaps the worker was restarted, or the context was GC'd?\n0 successful operations.\n0 derived errors ignored.\n</code></pre>\n<p>Each TfRecord is only 4 MB size and represents the <strong>tf.float32</strong> matrix of <strong>[500,2000]</strong> size. The case is very hard to debug and playing with batch size, parallel reads, and data type (float64, int32, etc) doesn't help.<br>\nDecreasing the matrix size to [1,2000] didn't help either.</p>\n<p>Your hints are appreciated.</p>",
      "rawMarkdown": "While calling `model.fit(tf_dataset,...)` I'm getting the following error:\n\n```\nUnavailableError: 2 root error(s) found.\n  (0) Unavailable: Socket closed\n  (1) Unavailable: Unable to find a context_id matching the specified one (6747970432017632323). Perhaps the worker was restarted, or the context was GC'd?\n0 successful operations.\n0 derived errors ignored.\n```\n\nEach TfRecord is only 4 MB size and represents the **tf.float32** matrix of **[500,2000]** size. The case is very hard to debug and playing with batch size, parallel reads, and data type (float64, int32, etc) doesn't help.\nDecreasing the matrix size to [1,2000] didn't help either.\n\nYour hints are appreciated.",
      "votes": null
    },
    {
      "id": "1472660",
      "postDate": "08/15/2021 03:44:51",
      "content": "<p>Curious, what happens when you run it with GPU instead of TPU as your notebook backend?</p>\n<p><a href=\"https://github.com/tensorflow/tensorflow/issues/36136#issuecomment-590090660\" target=\"_blank\">https://github.com/tensorflow/tensorflow/issues/36136#issuecomment-590090660</a></p>\n<p>In this Github thread people are saying this is most likely something to do with your host call method. Try testing each of your functions separately to see if there are any issues.</p>\n<p>If possible, you could provide some more specific code so we can help debug your issue?</p>",
      "rawMarkdown": "Curious, what happens when you run it with GPU instead of TPU as your notebook backend?\n\nhttps://github.com/tensorflow/tensorflow/issues/36136#issuecomment-590090660\n\nIn this Github thread people are saying this is most likely something to do with your host call method. Try testing each of your functions separately to see if there are any issues.\n\nIf possible, you could provide some more specific code so we can help debug your issue?",
      "votes": null
    },
    {
      "id": "1473233",
      "postDate": "08/15/2021 12:54:19",
      "content": "<p><a href=\"https://www.kaggle.com/d223chen\" target=\"_blank\">@d223chen</a> with GPU it worked. If you want, I can share the code with you.</p>",
      "rawMarkdown": "d223chen with GPU it worked. If you want, I can share the code with you.",
      "votes": null
    },
    {
      "id": "1473241",
      "postDate": "08/15/2021 12:58:04",
      "content": "<p>Sure, maybe if you can add to the post</p>",
      "rawMarkdown": "Sure, maybe if you can add to the post",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1472660,
      "author_name": "d223chen",
      "author_url": "",
      "post_date": "08/15/2021 03:44:51",
      "content": "<p>Curious, what happens when you run it with GPU instead of TPU as your notebook backend?</p>\n<p><a href=\"https://github.com/tensorflow/tensorflow/issues/36136#issuecomment-590090660\" target=\"_blank\">https://github.com/tensorflow/tensorflow/issues/36136#issuecomment-590090660</a></p>\n<p>In this Github thread people are saying this is most likely something to do with your host call method. Try testing each of your functions separately to see if there are any issues.</p>\n<p>If possible, you could provide some more specific code so we can help debug your issue?</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1473233,
      "author_name": "jhasanov",
      "author_url": "",
      "post_date": "08/15/2021 12:54:19",
      "content": "<p><a href=\"https://www.kaggle.com/d223chen\" target=\"_blank\">@d223chen</a> with GPU it worked. If you want, I can share the code with you.</p>",
      "votes": null,
      "replies": [
        {
          "id": 1473241,
          "author_name": "d223chen",
          "author_url": "",
          "post_date": "08/15/2021 12:58:04",
          "content": "<p>Sure, maybe if you can add to the post</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "1472611": "While calling `model.fit(tf_dataset,...)` I'm getting the following error:\n\n```\nUnavailableError: 2 root error(s) found.\n  (0) Unavailable: Socket closed\n  (1) Unavailable: Unable to find a context_id matching the specified one (6747970432017632323). Perhaps the worker was restarted, or the context was GC'd?\n0 successful operations.\n0 derived errors ignored.\n```\n\nEach TfRecord is only 4 MB size and represents the **tf.float32** matrix of **[500,2000]** size. The case is very hard to debug and playing with batch size, parallel reads, and data type (float64, int32, etc) doesn't help.\nDecreasing the matrix size to [1,2000] didn't help either.\n\nYour hints are appreciated.",
    "1472660": "Curious, what happens when you run it with GPU instead of TPU as your notebook backend?\n\nhttps://github.com/tensorflow/tensorflow/issues/36136#issuecomment-590090660\n\nIn this Github thread people are saying this is most likely something to do with your host call method. Try testing each of your functions separately to see if there are any issues.\n\nIf possible, you could provide some more specific code so we can help debug your issue?",
    "1473233": "d223chen with GPU it worked. If you want, I can share the code with you.",
    "1473241": "Sure, maybe if you can add to the post"
  },
  "source": "meta"
}