{
  "id": 139586,
  "title": "UnavailableError: Socket closed",
  "url": "/competitions/jigsaw-multilingual-toxic-comment-classification/discussion/139586",
  "author_name": "",
  "post_date": "2020-03-29T14:21:50.488302200Z",
  "votes": 14,
  "comment_count": 21,
  "views": 0,
  "content": "<p>What does it mean?</p>",
  "messages": [
    {
      "id": "790370",
      "postDate": "03/29/2020 14:21:50",
      "content": "<p>What does it mean?</p>",
      "rawMarkdown": "What does it mean?",
      "votes": null
    },
    {
      "id": "790482",
      "postDate": "03/29/2020 16:04:33",
      "content": "<p>For me, this usually happens when I run out o memory, a nice way to check is to run the same code with just a few samples.</p>",
      "rawMarkdown": "For me, this usually happens when I run out o memory, a nice way to check is to run the same code with just a few samples.",
      "votes": null
    },
    {
      "id": "799038",
      "postDate": "04/06/2020 05:16:17",
      "content": "<p><a href=\"/ipythonx\">@ipythonx</a> were you able to solve it ? I am getting the same error.</p>",
      "rawMarkdown": "ipythonx were you able to solve it ? I am getting the same error.",
      "votes": null
    },
    {
      "id": "799076",
      "postDate": "04/06/2020 06:23:10",
      "content": "<p>This is what I get lately:</p>\n\n<p><code>5327.8s\n12\n[NbConvertApp] ERROR | Kernel died while waiting for execute reply.\n5329.2s\n13\nTraceback (most recent call last):\n5329.2s\n14\n  File \"/opt/conda/lib/python3.6/site-packages/nbconvert/preprocessors/execute.py\", line 478, in _poll_for_reply\n5329.2s\n15\n    msg = self.kc.shell_channel.get_msg(timeout=timeout)\n5329.2s\n16\n  File \"/opt/conda/lib/python3.6/site-packages/jupyter_client/blocking/channels.py\", line 57, in get_msg\n5329.2s\n17\n    raise Empty\n5329.2s\n18\nqueue.Empty</code></p>",
      "rawMarkdown": "This is what I get lately:\n\n`5327.8s\n12\n[NbConvertApp] ERROR | Kernel died while waiting for execute reply.\n5329.2s\n13\nTraceback (most recent call last):\n5329.2s\n14\n  File \"/opt/conda/lib/python3.6/site-packages/nbconvert/preprocessors/execute.py\", line 478, in _poll_for_reply\n5329.2s\n15\n    msg = self.kc.shell_channel.get_msg(timeout=timeout)\n5329.2s\n16\n  File \"/opt/conda/lib/python3.6/site-packages/jupyter_client/blocking/channels.py\", line 57, in get_msg\n5329.2s\n17\n    raise Empty\n5329.2s\n18\nqueue.Empty`",
      "votes": null
    },
    {
      "id": "799275",
      "postDate": "04/06/2020 10:31:21",
      "content": "<p>I have exactly the same issue, and cannot solve it at all with TPU (code is running fine on GPU).</p>\n\n<p>I can add some details and my attempts to solve this issue . Hopefully, <a href=\"/mgornergoogle\">@mgornergoogle</a> can help us!! : </p>\n\n<p>I got the same \"Socket Closed\" error when I employ <a href=\"/xhlulu\">@xhlulu</a> nice <a href=\"https://www.kaggle.com/xhlulu/jigsaw-tpu-xlm-roberta\">TPU kernel</a>, and try to increase validation data from 8000 to around 60000 (just adding more translated data in es language) </p>\n\n<p>Precisely, the error happen exactly when we call\n<code>model.fit(valid_dataset, ...)</code></p>\n\n<h3>My Attempts</h3>\n\n<p>1) I have found <a href=\"/dimitreoliveira\">@dimitreoliveira</a> <a href=\"https://www.kaggle.com/c/flower-classification-with-tpus/discussion/135443#783021\">thread</a> proposing a method to solve this issue by removing .cache() when creating dataset --&gt; failed\n2) I have tried to decrease BATCH_SIZE --&gt; failed\n3) I have tried changing from <code>model.fit(valid_dataset, ...)</code> to <code>model.fit(x_valid, y_valid ...)</code> --&gt; still failed!!\n4) I have tried skipping<code>model.fit(train_dataset, ...)</code>  and train on valid dataset directly --&gt; still failed!!</p>\n\n<p>I really have no clue since <code>train_dataset</code>is much bigger than this augmented <code>valid_dataset</code> but with <code>train_dataset</code>, the model can still train nicely ...</p>\n\n<p>5) My current dumb solution now is to save model when finish training on train_dataset, and then using GPU session and then re-train on <code>valid_dataset</code>!</p>\n\n<p>I hope somebody can shed some light how to cope with this problem!</p>",
      "rawMarkdown": "I have exactly the same issue, and cannot solve it at all with TPU (code is running fine on GPU).\n\nI can add some details and my attempts to solve this issue . Hopefully, @mgornergoogle can help us!! : \n\nI got the same \"Socket Closed\" error when I employ @xhlulu nice [TPU kernel](https://www.kaggle.com/xhlulu/jigsaw-tpu-xlm-roberta), and try to increase validation data from 8000 to around 60000 (just adding more translated data in es language) \n\nPrecisely, the error happen exactly when we call\n`model.fit(valid_dataset, ...)`\n\n### My Attempts\n1) I have found @dimitreoliveira [thread](https://www.kaggle.com/c/flower-classification-with-tpus/discussion/135443#783021) proposing a method to solve this issue by removing .cache() when creating dataset --&gt; failed\n2) I have tried to decrease BATCH_SIZE --&gt; failed\n3) I have tried changing from `model.fit(valid_dataset, ...)` to `model.fit(x_valid, y_valid ...)` --&gt; still failed!!\n4) I have tried skipping`model.fit(train_dataset, ...)`  and train on valid dataset directly --&gt; still failed!!\n\nI really have no clue since `train_dataset `is much bigger than this augmented `valid_dataset` but with `train_dataset`, the model can still train nicely ...\n\n5) My current dumb solution now is to save model when finish training on train_dataset, and then using GPU session and then re-train on `valid_dataset`!\n\nI hope somebody can shed some light how to cope with this problem!",
      "votes": null
    },
    {
      "id": "799295",
      "postDate": "04/06/2020 10:53:57",
      "content": "<p>I have the same issue and tried the same things as you did. Nothing is killing is the issue!\nI found a discussion corresponding to the bug report.\n<a href=\"https://www.kaggle.com/dimitreoliveira/bug-report-unavailableerror-socket-closed\">https://www.kaggle.com/dimitreoliveira/bug-report-unavailableerror-socket-closed</a></p>",
      "rawMarkdown": "I have the same issue and tried the same things as you did. Nothing is killing is the issue!\nI found a discussion corresponding to the bug report.\nhttps://www.kaggle.com/dimitreoliveira/bug-report-unavailableerror-socket-closed",
      "votes": null
    },
    {
      "id": "799836",
      "postDate": "04/06/2020 20:14:51",
      "content": "<p>I have done many experimentations with this and the flowers competition and got issues related to this multiple times, and its always related to memory usage, so I will post here a checklist that has helped me most of the times, I recommend to try one by time to see what works, if possible run the code in iterative mode to see if the memory usage hits the limit.</p>\n\n<ul>\n<li>Use less data</li>\n<li>Try smaller batches</li>\n<li>Train for fewer epochs</li>\n<li>Remove <code>.cache()</code> from the validation set</li>\n<li>Remove <code>drop_remainder=True</code> from the validation set</li>\n<li>If using custom loop iterate over validation set only once</li>\n</ul>\n\n<p>I have discussed issues like this on <a href=\"https://www.kaggle.com/c/jigsaw-multilingual-toxic-comment-classification/discussion/140007\">these</a> two <a href=\"https://www.kaggle.com/c/flower-classification-with-tpus/discussion/135443#775339\">threads</a></p>",
      "rawMarkdown": "I have done many experimentations with this and the flowers competition and got issues related to this multiple times, and its always related to memory usage, so I will post here a checklist that has helped me most of the times, I recommend to try one by time to see what works, if possible run the code in iterative mode to see if the memory usage hits the limit.\n\n- Use less data\n- Try smaller batches\n- Train for fewer epochs\n- Remove `.cache()` from the validation set\n- Remove `drop_remainder=True` from the validation set\n- If using custom loop iterate over validation set only once\n\nI have discussed issues like this on [these](https://www.kaggle.com/c/jigsaw-multilingual-toxic-comment-classification/discussion/140007) two [threads](https://www.kaggle.com/c/flower-classification-with-tpus/discussion/135443#775339)",
      "votes": null
    },
    {
      "id": "799838",
      "postDate": "04/06/2020 20:15:47",
      "content": "<p>Hey <a href=\"/ratthachat\">@ratthachat</a> it seems to be memory issues, try to see if any of the tips I posted above works for you.</p>",
      "rawMarkdown": "Hey @ratthachat it seems to be memory issues, try to see if any of the tips I posted above works for you.",
      "votes": null
    },
    {
      "id": "799948",
      "postDate": "04/06/2020 23:46:23",
      "content": "<p>I have been able to confirm that many of the TPU stability issues people have been reporting are fixed in TF 2.2. It's not something you can fix on your end. The good news is that the TF 2.2 release is close.</p>",
      "rawMarkdown": "I have been able to confirm that many of the TPU stability issues people have been reporting are fixed in TF 2.2. It's not something you can fix on your end. The good news is that the TF 2.2 release is close.",
      "votes": null
    },
    {
      "id": "811749",
      "postDate": "04/18/2020 07:44:25",
      "content": "<p>Guys, good news. As Martin suggested, I can confirm that my Socket Closed problem is fixed in TF2.2, so since it's \"Release Candidate 1\" is out, we can : </p>\n\n<p><code>!pip install tensorflow==2.2-rc1</code></p>\n\n<p><strong>UPDATED</strong> Martin commented below that I may just got lucky, but anyway it works 50% for me (now rc3 is the latest) . </p>\n\n<p>The other 50% fix is about the float labels as Camaro mentioned below, especially the float64 format.\nSo we can fix it easily by converting labels to either int or float32 .</p>",
      "rawMarkdown": "Guys, good news. As Martin suggested, I can confirm that my Socket Closed problem is fixed in TF2.2, so since it's \"Release Candidate 1\" is out, we can : \n\n`!pip install tensorflow==2.2-rc1`\n\n**UPDATED** Martin commented below that I may just got lucky, but anyway it works 50% for me (now rc3 is the latest) . \n\nThe other 50% fix is about the float labels as Camaro mentioned below, especially the float64 format.\nSo we can fix it easily by converting labels to either int or float32 .",
      "votes": null
    },
    {
      "id": "811964",
      "postDate": "04/18/2020 11:18:22",
      "content": "<p>wow, great. thanks for the info. </p>",
      "rawMarkdown": "wow, great. thanks for the info.",
      "votes": null
    },
    {
      "id": "812063",
      "postDate": "04/18/2020 12:57:31",
      "content": "<p>Great news. Thanks for sharing.</p>",
      "rawMarkdown": "Great news. Thanks for sharing.",
      "votes": null
    },
    {
      "id": "812109",
      "postDate": "04/18/2020 14:01:00",
      "content": "<p>Did anyone try it? does it solve the Unavailable Error?</p>",
      "rawMarkdown": "Did anyone try it? does it solve the Unavailable Error?",
      "votes": null
    },
    {
      "id": "812195",
      "postDate": "04/18/2020 15:21:26",
      "content": "<p>Colab tensorflow is already 2.2.0, but it also has another issue. My code didn't run with tf2.2.0 and I decided to downgrade to 2.1.0😑 It's annoying, but definitely better than \"UnavailableError: Socket closed\", which says nothing indeed!</p>\n\n<p>By the way, when I solved that error before, it was just data type mismatch in label.(Should be int but it was float.)</p>",
      "rawMarkdown": "Colab tensorflow is already 2.2.0, but it also has another issue. My code didn't run with tf2.2.0 and I decided to downgrade to 2.1.0😑 It's annoying, but definitely better than \"UnavailableError: Socket closed\", which says nothing indeed!\n\nBy the way, when I solved that error before, it was just data type mismatch in label.(Should be int but it was float.)",
      "votes": null
    },
    {
      "id": "812199",
      "postDate": "04/18/2020 15:25:34",
      "content": "<p>why we can't train with float label? (i met the same problem, but i want to still keep the float format)</p>",
      "rawMarkdown": "why we can't train with float label? (i met the same problem, but i want to still keep the float format)",
      "votes": null
    },
    {
      "id": "812657",
      "postDate": "04/18/2020 22:54:21",
      "content": "<p>Thanks for the info! <a href=\"/bamps53\">@bamps53</a> </p>",
      "rawMarkdown": "Thanks for the info! @bamps53",
      "votes": null
    },
    {
      "id": "814679",
      "postDate": "04/20/2020 21:29:24",
      "content": "<p>sorry guys, <code>!pip install tensorflow==2.2-rc1</code> has little chance of working in a TPU notebook. It updates the version of TF in the notebook but not on the TPU. Expect communication problems between the two. You might get lucky, but this is still not advised. You have to wait for Kaggle to update the default TF.</p>",
      "rawMarkdown": "sorry guys, `!pip install tensorflow==2.2-rc1` has little chance of working in a TPU notebook. It updates the version of TF in the notebook but not on the TPU. Expect communication problems between the two. You might get lucky, but this is still not advised. You have to wait for Kaggle to update the default TF.",
      "votes": null
    },
    {
      "id": "820052",
      "postDate": "04/25/2020 05:47:53",
      "content": "<p>I was getting the same error. \nAnd when I checked the datatypes I found that my X_label and Y_label are not of same datatype, X_label(float) and Y_label(int). \nCorrecting it solved my error.</p>",
      "rawMarkdown": "I was getting the same error. \nAnd when I checked the datatypes I found that my X_label and Y_label are not of same datatype, X_label(float) and Y_label(int). \nCorrecting it solved my error.",
      "votes": null
    },
    {
      "id": "820423",
      "postDate": "04/25/2020 12:37:56",
      "content": "<p>I think you might have the same issue as i had, try turning your targets to integers. More info in my comment here: <a href=\"https://www.kaggle.com/c/jigsaw-multilingual-toxic-comment-classification/discussion/145948\">https://www.kaggle.com/c/jigsaw-multilingual-toxic-comment-classification/discussion/145948</a>. Hope that's not too late :)</p>",
      "rawMarkdown": "I think you might have the same issue as i had, try turning your targets to integers. More info in my comment here: https://www.kaggle.com/c/jigsaw-multilingual-toxic-comment-classification/discussion/145948. Hope that's not too late :)",
      "votes": null
    },
    {
      "id": "857190",
      "postDate": "05/22/2020 11:40:36",
      "content": "<p>I think if you simply change the output labels to integer values, it might work.\nThe error occurs when the output labels are float64.\nI tried changing the datatype and it works fine.</p>",
      "rawMarkdown": "I think if you simply change the output labels to integer values, it might work.\nThe error occurs when the output labels are float64.\nI tried changing the datatype and it works fine.",
      "votes": null
    },
    {
      "id": "905413",
      "postDate": "06/28/2020 14:33:52",
      "content": "<p><a href=\"/dimitreoliveira\">@dimitreoliveira</a>  did removing <code>.cache()</code> slow down the training since data is not cached?</p>",
      "rawMarkdown": "dimitreoliveira  did removing `.cache()` slow down the training since data is not cached?",
      "votes": null
    },
    {
      "id": "905525",
      "postDate": "06/28/2020 15:59:03",
      "content": "<p>Yes <a href=\"/tonychenxyz\">@tonychenxyz</a> , The training will be slower because the validation data will not be stored in memory, but if you are using smaller models or less data using <code>.cache()</code> may not be a problem.</p>",
      "rawMarkdown": "Yes @tonychenxyz , The training will be slower because the validation data will not be stored in memory, but if you are using smaller models or less data using `.cache()` may not be a problem.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 790482,
      "author_name": "dimitreoliveira",
      "author_url": "",
      "post_date": "03/29/2020 16:04:33",
      "content": "<p>For me, this usually happens when I run out o memory, a nice way to check is to run the same code with just a few samples.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 799038,
      "author_name": "shahules",
      "author_url": "",
      "post_date": "04/06/2020 05:16:17",
      "content": "<p><a href=\"/ipythonx\">@ipythonx</a> were you able to solve it ? I am getting the same error.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 799076,
      "author_name": "philippsinger",
      "author_url": "",
      "post_date": "04/06/2020 06:23:10",
      "content": "<p>This is what I get lately:</p>\n\n<p><code>5327.8s\n12\n[NbConvertApp] ERROR | Kernel died while waiting for execute reply.\n5329.2s\n13\nTraceback (most recent call last):\n5329.2s\n14\n  File \"/opt/conda/lib/python3.6/site-packages/nbconvert/preprocessors/execute.py\", line 478, in _poll_for_reply\n5329.2s\n15\n    msg = self.kc.shell_channel.get_msg(timeout=timeout)\n5329.2s\n16\n  File \"/opt/conda/lib/python3.6/site-packages/jupyter_client/blocking/channels.py\", line 57, in get_msg\n5329.2s\n17\n    raise Empty\n5329.2s\n18\nqueue.Empty</code></p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 799275,
      "author_name": "ratthachat",
      "author_url": "",
      "post_date": "04/06/2020 10:31:21",
      "content": "<p>I have exactly the same issue, and cannot solve it at all with TPU (code is running fine on GPU).</p>\n\n<p>I can add some details and my attempts to solve this issue . Hopefully, <a href=\"/mgornergoogle\">@mgornergoogle</a> can help us!! : </p>\n\n<p>I got the same \"Socket Closed\" error when I employ <a href=\"/xhlulu\">@xhlulu</a> nice <a href=\"https://www.kaggle.com/xhlulu/jigsaw-tpu-xlm-roberta\">TPU kernel</a>, and try to increase validation data from 8000 to around 60000 (just adding more translated data in es language) </p>\n\n<p>Precisely, the error happen exactly when we call\n<code>model.fit(valid_dataset, ...)</code></p>\n\n<h3>My Attempts</h3>\n\n<p>1) I have found <a href=\"/dimitreoliveira\">@dimitreoliveira</a> <a href=\"https://www.kaggle.com/c/flower-classification-with-tpus/discussion/135443#783021\">thread</a> proposing a method to solve this issue by removing .cache() when creating dataset --&gt; failed\n2) I have tried to decrease BATCH_SIZE --&gt; failed\n3) I have tried changing from <code>model.fit(valid_dataset, ...)</code> to <code>model.fit(x_valid, y_valid ...)</code> --&gt; still failed!!\n4) I have tried skipping<code>model.fit(train_dataset, ...)</code>  and train on valid dataset directly --&gt; still failed!!</p>\n\n<p>I really have no clue since <code>train_dataset</code>is much bigger than this augmented <code>valid_dataset</code> but with <code>train_dataset</code>, the model can still train nicely ...</p>\n\n<p>5) My current dumb solution now is to save model when finish training on train_dataset, and then using GPU session and then re-train on <code>valid_dataset</code>!</p>\n\n<p>I hope somebody can shed some light how to cope with this problem!</p>",
      "votes": null,
      "replies": [
        {
          "id": 799295,
          "author_name": "shahules",
          "author_url": "",
          "post_date": "04/06/2020 10:53:57",
          "content": "<p>I have the same issue and tried the same things as you did. Nothing is killing is the issue!\nI found a discussion corresponding to the bug report.\n<a href=\"https://www.kaggle.com/dimitreoliveira/bug-report-unavailableerror-socket-closed\">https://www.kaggle.com/dimitreoliveira/bug-report-unavailableerror-socket-closed</a></p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 799838,
          "author_name": "dimitreoliveira",
          "author_url": "",
          "post_date": "04/06/2020 20:15:47",
          "content": "<p>Hey <a href=\"/ratthachat\">@ratthachat</a> it seems to be memory issues, try to see if any of the tips I posted above works for you.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 820423,
          "author_name": "theonekeyg",
          "author_url": "",
          "post_date": "04/25/2020 12:37:56",
          "content": "<p>I think you might have the same issue as i had, try turning your targets to integers. More info in my comment here: <a href=\"https://www.kaggle.com/c/jigsaw-multilingual-toxic-comment-classification/discussion/145948\">https://www.kaggle.com/c/jigsaw-multilingual-toxic-comment-classification/discussion/145948</a>. Hope that's not too late :)</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 799836,
      "author_name": "dimitreoliveira",
      "author_url": "",
      "post_date": "04/06/2020 20:14:51",
      "content": "<p>I have done many experimentations with this and the flowers competition and got issues related to this multiple times, and its always related to memory usage, so I will post here a checklist that has helped me most of the times, I recommend to try one by time to see what works, if possible run the code in iterative mode to see if the memory usage hits the limit.</p>\n\n<ul>\n<li>Use less data</li>\n<li>Try smaller batches</li>\n<li>Train for fewer epochs</li>\n<li>Remove <code>.cache()</code> from the validation set</li>\n<li>Remove <code>drop_remainder=True</code> from the validation set</li>\n<li>If using custom loop iterate over validation set only once</li>\n</ul>\n\n<p>I have discussed issues like this on <a href=\"https://www.kaggle.com/c/jigsaw-multilingual-toxic-comment-classification/discussion/140007\">these</a> two <a href=\"https://www.kaggle.com/c/flower-classification-with-tpus/discussion/135443#775339\">threads</a></p>",
      "votes": null,
      "replies": [
        {
          "id": 905413,
          "author_name": "tonychenxyz",
          "author_url": "",
          "post_date": "06/28/2020 14:33:52",
          "content": "<p><a href=\"/dimitreoliveira\">@dimitreoliveira</a>  did removing <code>.cache()</code> slow down the training since data is not cached?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 905525,
          "author_name": "dimitreoliveira",
          "author_url": "",
          "post_date": "06/28/2020 15:59:03",
          "content": "<p>Yes <a href=\"/tonychenxyz\">@tonychenxyz</a> , The training will be slower because the validation data will not be stored in memory, but if you are using smaller models or less data using <code>.cache()</code> may not be a problem.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 799948,
      "author_name": "mgorner",
      "author_url": "",
      "post_date": "04/06/2020 23:46:23",
      "content": "<p>I have been able to confirm that many of the TPU stability issues people have been reporting are fixed in TF 2.2. It's not something you can fix on your end. The good news is that the TF 2.2 release is close.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 811749,
      "author_name": "ratthachat",
      "author_url": "",
      "post_date": "04/18/2020 07:44:25",
      "content": "<p>Guys, good news. As Martin suggested, I can confirm that my Socket Closed problem is fixed in TF2.2, so since it's \"Release Candidate 1\" is out, we can : </p>\n\n<p><code>!pip install tensorflow==2.2-rc1</code></p>\n\n<p><strong>UPDATED</strong> Martin commented below that I may just got lucky, but anyway it works 50% for me (now rc3 is the latest) . </p>\n\n<p>The other 50% fix is about the float labels as Camaro mentioned below, especially the float64 format.\nSo we can fix it easily by converting labels to either int or float32 .</p>",
      "votes": null,
      "replies": [
        {
          "id": 811964,
          "author_name": "ipythonx",
          "author_url": "",
          "post_date": "04/18/2020 11:18:22",
          "content": "<p>wow, great. thanks for the info. </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 812063,
          "author_name": "luohongchen1993",
          "author_url": "",
          "post_date": "04/18/2020 12:57:31",
          "content": "<p>Great news. Thanks for sharing.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 812109,
          "author_name": "luohongchen1993",
          "author_url": "",
          "post_date": "04/18/2020 14:01:00",
          "content": "<p>Did anyone try it? does it solve the Unavailable Error?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 812195,
          "author_name": "bamps53",
          "author_url": "",
          "post_date": "04/18/2020 15:21:26",
          "content": "<p>Colab tensorflow is already 2.2.0, but it also has another issue. My code didn't run with tf2.2.0 and I decided to downgrade to 2.1.0😑 It's annoying, but definitely better than \"UnavailableError: Socket closed\", which says nothing indeed!</p>\n\n<p>By the way, when I solved that error before, it was just data type mismatch in label.(Should be int but it was float.)</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 812199,
          "author_name": "zzy990106",
          "author_url": "",
          "post_date": "04/18/2020 15:25:34",
          "content": "<p>why we can't train with float label? (i met the same problem, but i want to still keep the float format)</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 812657,
          "author_name": "luohongchen1993",
          "author_url": "",
          "post_date": "04/18/2020 22:54:21",
          "content": "<p>Thanks for the info! <a href=\"/bamps53\">@bamps53</a> </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 814679,
          "author_name": "mgorner",
          "author_url": "",
          "post_date": "04/20/2020 21:29:24",
          "content": "<p>sorry guys, <code>!pip install tensorflow==2.2-rc1</code> has little chance of working in a TPU notebook. It updates the version of TF in the notebook but not on the TPU. Expect communication problems between the two. You might get lucky, but this is still not advised. You have to wait for Kaggle to update the default TF.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 820052,
      "author_name": "abneetwats24",
      "author_url": "",
      "post_date": "04/25/2020 05:47:53",
      "content": "<p>I was getting the same error. \nAnd when I checked the datatypes I found that my X_label and Y_label are not of same datatype, X_label(float) and Y_label(int). \nCorrecting it solved my error.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 857190,
      "author_name": "bhaveshkotwani",
      "author_url": "",
      "post_date": "05/22/2020 11:40:36",
      "content": "<p>I think if you simply change the output labels to integer values, it might work.\nThe error occurs when the output labels are float64.\nI tried changing the datatype and it works fine.</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "790370": "What does it mean?",
    "790482": "For me, this usually happens when I run out o memory, a nice way to check is to run the same code with just a few samples.",
    "799038": "ipythonx were you able to solve it ? I am getting the same error.",
    "799076": "This is what I get lately:\n\n`5327.8s\n12\n[NbConvertApp] ERROR | Kernel died while waiting for execute reply.\n5329.2s\n13\nTraceback (most recent call last):\n5329.2s\n14\n  File \"/opt/conda/lib/python3.6/site-packages/nbconvert/preprocessors/execute.py\", line 478, in _poll_for_reply\n5329.2s\n15\n    msg = self.kc.shell_channel.get_msg(timeout=timeout)\n5329.2s\n16\n  File \"/opt/conda/lib/python3.6/site-packages/jupyter_client/blocking/channels.py\", line 57, in get_msg\n5329.2s\n17\n    raise Empty\n5329.2s\n18\nqueue.Empty`",
    "799275": "I have exactly the same issue, and cannot solve it at all with TPU (code is running fine on GPU).\n\nI can add some details and my attempts to solve this issue . Hopefully, @mgornergoogle can help us!! : \n\nI got the same \"Socket Closed\" error when I employ @xhlulu nice [TPU kernel](https://www.kaggle.com/xhlulu/jigsaw-tpu-xlm-roberta), and try to increase validation data from 8000 to around 60000 (just adding more translated data in es language) \n\nPrecisely, the error happen exactly when we call\n`model.fit(valid_dataset, ...)`\n\n### My Attempts\n1) I have found @dimitreoliveira [thread](https://www.kaggle.com/c/flower-classification-with-tpus/discussion/135443#783021) proposing a method to solve this issue by removing .cache() when creating dataset --&gt; failed\n2) I have tried to decrease BATCH_SIZE --&gt; failed\n3) I have tried changing from `model.fit(valid_dataset, ...)` to `model.fit(x_valid, y_valid ...)` --&gt; still failed!!\n4) I have tried skipping`model.fit(train_dataset, ...)`  and train on valid dataset directly --&gt; still failed!!\n\nI really have no clue since `train_dataset `is much bigger than this augmented `valid_dataset` but with `train_dataset`, the model can still train nicely ...\n\n5) My current dumb solution now is to save model when finish training on train_dataset, and then using GPU session and then re-train on `valid_dataset`!\n\nI hope somebody can shed some light how to cope with this problem!",
    "799295": "I have the same issue and tried the same things as you did. Nothing is killing is the issue!\nI found a discussion corresponding to the bug report.\nhttps://www.kaggle.com/dimitreoliveira/bug-report-unavailableerror-socket-closed",
    "799836": "I have done many experimentations with this and the flowers competition and got issues related to this multiple times, and its always related to memory usage, so I will post here a checklist that has helped me most of the times, I recommend to try one by time to see what works, if possible run the code in iterative mode to see if the memory usage hits the limit.\n\n- Use less data\n- Try smaller batches\n- Train for fewer epochs\n- Remove `.cache()` from the validation set\n- Remove `drop_remainder=True` from the validation set\n- If using custom loop iterate over validation set only once\n\nI have discussed issues like this on [these](https://www.kaggle.com/c/jigsaw-multilingual-toxic-comment-classification/discussion/140007) two [threads](https://www.kaggle.com/c/flower-classification-with-tpus/discussion/135443#775339)",
    "799838": "Hey @ratthachat it seems to be memory issues, try to see if any of the tips I posted above works for you.",
    "799948": "I have been able to confirm that many of the TPU stability issues people have been reporting are fixed in TF 2.2. It's not something you can fix on your end. The good news is that the TF 2.2 release is close.",
    "811749": "Guys, good news. As Martin suggested, I can confirm that my Socket Closed problem is fixed in TF2.2, so since it's \"Release Candidate 1\" is out, we can : \n\n`!pip install tensorflow==2.2-rc1`\n\n**UPDATED** Martin commented below that I may just got lucky, but anyway it works 50% for me (now rc3 is the latest) . \n\nThe other 50% fix is about the float labels as Camaro mentioned below, especially the float64 format.\nSo we can fix it easily by converting labels to either int or float32 .",
    "811964": "wow, great. thanks for the info.",
    "812063": "Great news. Thanks for sharing.",
    "812109": "Did anyone try it? does it solve the Unavailable Error?",
    "812195": "Colab tensorflow is already 2.2.0, but it also has another issue. My code didn't run with tf2.2.0 and I decided to downgrade to 2.1.0😑 It's annoying, but definitely better than \"UnavailableError: Socket closed\", which says nothing indeed!\n\nBy the way, when I solved that error before, it was just data type mismatch in label.(Should be int but it was float.)",
    "812199": "why we can't train with float label? (i met the same problem, but i want to still keep the float format)",
    "812657": "Thanks for the info! @bamps53",
    "814679": "sorry guys, `!pip install tensorflow==2.2-rc1` has little chance of working in a TPU notebook. It updates the version of TF in the notebook but not on the TPU. Expect communication problems between the two. You might get lucky, but this is still not advised. You have to wait for Kaggle to update the default TF.",
    "820052": "I was getting the same error. \nAnd when I checked the datatypes I found that my X_label and Y_label are not of same datatype, X_label(float) and Y_label(int). \nCorrecting it solved my error.",
    "820423": "I think you might have the same issue as i had, try turning your targets to integers. More info in my comment here: https://www.kaggle.com/c/jigsaw-multilingual-toxic-comment-classification/discussion/145948. Hope that's not too late :)",
    "857190": "I think if you simply change the output labels to integer values, it might work.\nThe error occurs when the output labels are float64.\nI tried changing the datatype and it works fine.",
    "905413": "dimitreoliveira  did removing `.cache()` slow down the training since data is not cached?",
    "905525": "Yes @tonychenxyz , The training will be slower because the validation data will not be stored in memory, but if you are using smaller models or less data using `.cache()` may not be a problem."
  },
  "source": "meta"
}