{
  "id": 199212,
  "title": "TPU \"UnavailableError: Socket closed\"",
  "url": "/competitions/cassava-leaf-disease-classification/discussion/199212",
  "author_name": "DimitreOliveira",
  "post_date": "2020-11-24T21:05:46.943000",
  "votes": 4,
  "comment_count": 0,
  "views": 0,
  "content": "<p>For people that are using TPUs for some time, this error might not be new, but for the others, it might be difficult to debug it, I remember that I spent countless hours debugging similar errors in previous competitions, so here it is a possible solution.</p>\n<p>Most of the weird errors that happen on TPU are caused by memory issues, it was my case, it happened in the middle of the training.</p>\n<pre><code>133/133 - 15s - sparse_categorical_accuracy: 0.8828 - loss: 0.3277 - val_sparse_categorical_accuracy: 0.8593 - val_loss: 0.4045 - lr: 4.4444e-06\nEpoch 25/30\n133/133 - 16s - sparse_categorical_accuracy: 0.8847 - loss: 0.3324 - val_sparse_categorical_accuracy: 0.8614 - val_loss: 0.4051 - lr: 3.7555e-06\nEpoch 26/30\n---------------------------------------------------------------------------\nUnavailableError                          Traceback (most recent call last)\n&lt;ipython-input-14-14103e91c820&gt; in &lt;module&gt;\n     31                         callbacks=[es, LearningRateScheduler(lrfn, verbose=0)],\n     32                         epochs=EPOCHS,\n---&gt; 33                         verbose=2).history\n     34 \n     35     history_list.append(history)\n\n.\n.\n.\n\n/opt/conda/lib/python3.7/site-packages/six.py in raise_from(value, from_value)\n\nUnavailableError: Socket closed\nAdditional GRPC error information:\n{\"created\":\"@1606201969.369747908\",\"description\":\"Error received from peer ipv4:10.0.0.2:8470\",\"file\":\"external/com_github_grpc_grpc/src/core/lib/surface/call.cc\",\"file_line\":1056,\"grpc_message\":\"Socket closed\",\"grpc_status\":14}\n</code></pre>\n<p>The problem was that I used <code>dataset = dataset.cache()</code> on my dataset function, and caching the data while training was overloading the memory, note that using <code>dataset.cache()</code> is not a wrong or bad practice, but if you are using a big model, big images and large batches you are already using a lot of memory, so in this case using cache may cause memory issues.</p>\n<p>The point here is that looking at the error log you probably won't find out that it is related to memory, and in my case, this cost me 27k seconds of TPU quota.</p>",
  "messages": [
    {
      "id": 1089866,
      "postDate": "2020-11-24T21:05:46.943Z",
      "content": "<p>For people that are using TPUs for some time, this error might not be new, but for the others, it might be difficult to debug it, I remember that I spent countless hours debugging similar errors in previous competitions, so here it is a possible solution.</p>\n<p>Most of the weird errors that happen on TPU are caused by memory issues, it was my case, it happened in the middle of the training.</p>\n<pre><code>133/133 - 15s - sparse_categorical_accuracy: 0.8828 - loss: 0.3277 - val_sparse_categorical_accuracy: 0.8593 - val_loss: 0.4045 - lr: 4.4444e-06\nEpoch 25/30\n133/133 - 16s - sparse_categorical_accuracy: 0.8847 - loss: 0.3324 - val_sparse_categorical_accuracy: 0.8614 - val_loss: 0.4051 - lr: 3.7555e-06\nEpoch 26/30\n---------------------------------------------------------------------------\nUnavailableError                          Traceback (most recent call last)\n&lt;ipython-input-14-14103e91c820&gt; in &lt;module&gt;\n     31                         callbacks=[es, LearningRateScheduler(lrfn, verbose=0)],\n     32                         epochs=EPOCHS,\n---&gt; 33                         verbose=2).history\n     34 \n     35     history_list.append(history)\n\n.\n.\n.\n\n/opt/conda/lib/python3.7/site-packages/six.py in raise_from(value, from_value)\n\nUnavailableError: Socket closed\nAdditional GRPC error information:\n{\"created\":\"@1606201969.369747908\",\"description\":\"Error received from peer ipv4:10.0.0.2:8470\",\"file\":\"external/com_github_grpc_grpc/src/core/lib/surface/call.cc\",\"file_line\":1056,\"grpc_message\":\"Socket closed\",\"grpc_status\":14}\n</code></pre>\n<p>The problem was that I used <code>dataset = dataset.cache()</code> on my dataset function, and caching the data while training was overloading the memory, note that using <code>dataset.cache()</code> is not a wrong or bad practice, but if you are using a big model, big images and large batches you are already using a lot of memory, so in this case using cache may cause memory issues.</p>\n<p>The point here is that looking at the error log you probably won't find out that it is related to memory, and in my case, this cost me 27k seconds of TPU quota.</p>",
      "rawMarkdown": "For people that are using TPUs for some time, this error might not be new, but for the others, it might be difficult to debug it, I remember that I spent countless hours debugging similar errors in previous competitions, so here it is a possible solution.\n\nMost of the weird errors that happen on TPU are caused by memory issues, it was my case, it happened in the middle of the training.\n\n```\n133/133 - 15s - sparse_categorical_accuracy: 0.8828 - loss: 0.3277 - val_sparse_categorical_accuracy: 0.8593 - val_loss: 0.4045 - lr: 4.4444e-06\nEpoch 25/30\n133/133 - 16s - sparse_categorical_accuracy: 0.8847 - loss: 0.3324 - val_sparse_categorical_accuracy: 0.8614 - val_loss: 0.4051 - lr: 3.7555e-06\nEpoch 26/30\n---------------------------------------------------------------------------\nUnavailableError                          Traceback (most recent call last)\n<ipython-input-14-14103e91c820> in <module>\n     31                         callbacks=[es, LearningRateScheduler(lrfn, verbose=0)],\n     32                         epochs=EPOCHS,\n---> 33                         verbose=2).history\n     34 \n     35     history_list.append(history)\n\n.\n.\n.\n\n/opt/conda/lib/python3.7/site-packages/six.py in raise_from(value, from_value)\n\nUnavailableError: Socket closed\nAdditional GRPC error information:\n{\"created\":\"@1606201969.369747908\",\"description\":\"Error received from peer ipv4:10.0.0.2:8470\",\"file\":\"external/com_github_grpc_grpc/src/core/lib/surface/call.cc\",\"file_line\":1056,\"grpc_message\":\"Socket closed\",\"grpc_status\":14}\n```\n\nThe problem was that I used `dataset = dataset.cache()` on my dataset function, and caching the data while training was overloading the memory, note that using `dataset.cache()` is not a wrong or bad practice, but if you are using a big model, big images and large batches you are already using a lot of memory, so in this case using cache may cause memory issues.\n\nThe point here is that looking at the error log you probably won't find out that it is related to memory, and in my case, this cost me 27k seconds of TPU quota.",
      "votes": 4
    }
  ],
  "comments": [],
  "raw_markdown_by_id": {
    "1089866": "For people that are using TPUs for some time, this error might not be new, but for the others, it might be difficult to debug it, I remember that I spent countless hours debugging similar errors in previous competitions, so here it is a possible solution.\n\nMost of the weird errors that happen on TPU are caused by memory issues, it was my case, it happened in the middle of the training.\n\n```\n133/133 - 15s - sparse_categorical_accuracy: 0.8828 - loss: 0.3277 - val_sparse_categorical_accuracy: 0.8593 - val_loss: 0.4045 - lr: 4.4444e-06\nEpoch 25/30\n133/133 - 16s - sparse_categorical_accuracy: 0.8847 - loss: 0.3324 - val_sparse_categorical_accuracy: 0.8614 - val_loss: 0.4051 - lr: 3.7555e-06\nEpoch 26/30\n---------------------------------------------------------------------------\nUnavailableError                          Traceback (most recent call last)\n<ipython-input-14-14103e91c820> in <module>\n     31                         callbacks=[es, LearningRateScheduler(lrfn, verbose=0)],\n     32                         epochs=EPOCHS,\n---> 33                         verbose=2).history\n     34 \n     35     history_list.append(history)\n\n.\n.\n.\n\n/opt/conda/lib/python3.7/site-packages/six.py in raise_from(value, from_value)\n\nUnavailableError: Socket closed\nAdditional GRPC error information:\n{\"created\":\"@1606201969.369747908\",\"description\":\"Error received from peer ipv4:10.0.0.2:8470\",\"file\":\"external/com_github_grpc_grpc/src/core/lib/surface/call.cc\",\"file_line\":1056,\"grpc_message\":\"Socket closed\",\"grpc_status\":14}\n```\n\nThe problem was that I used `dataset = dataset.cache()` on my dataset function, and caching the data while training was overloading the memory, note that using `dataset.cache()` is not a wrong or bad practice, but if you are using a big model, big images and large batches you are already using a lot of memory, so in this case using cache may cause memory issues.\n\nThe point here is that looking at the error log you probably won't find out that it is related to memory, and in my case, this cost me 27k seconds of TPU quota."
  }
}