{
  "id": 141212,
  "title": "TPU Bug!",
  "url": "/competitions/jigsaw-multilingual-toxic-comment-classification/discussion/141212",
  "author_name": "",
  "post_date": "2020-04-05T05:50:09.797864500Z",
  "votes": null,
  "comment_count": 1,
  "views": 0,
  "content": "<p>So, the error was this :--</p>\n\n<p>ResourceExhaustedError: {{function_node __inference_distributed_function_153395}} Compilation failure: Ran out of memory in memory space vmem. It should not be possible to run out of vmem - please file a bug against XLA.</p>\n\n<p>Largest program allocations in vmem:</p>\n\n<p>XLA label: register allocator spill slots\n  Allocation type: scoped</p>\n\n<p>XLA label: %fusion.11963 = f32[4,192,1024]{2,0,1:T(4,128)} fusion(f32[4,1,1024]{2,0,1:T(4,128)}, f32[4,1,1024]{2,0,1:T(4,128)}, f32[4,1,1024]{2,0,1:T(4,128)}, f32[4,1,1024]{2,0,1:T(4,128)}, ...(+188)), kind=kLoop, calls=%fused_computation.10858\n  Allocation type: scoped</p>\n\n<p>XLA label: %fusion.11963 = f32[4,192,1024]{2,0,1:T(4,128)} fusion(f32[4,1,1024]{2,0,1:T(4,128)}, f32[4,1,1024]{2,0,1:T(4,128)}, f32[4,1,1024]{2,0,1:T(4,128)}, f32[4,1,1024]{2,0,1:T(4,128)}, ...(+188)), kind=kLoop, calls=%fused_computation.10858\n  Allocation type: scoped</p>\n\n<p>XLA label: %fusion.11963 = f32[4,192,1024]{2,0,1:T(4,128)} fusion(f32[4,1,1024]{2,0,1:T(4,128)}, f32[4,1,1024]{2,0,1:T(4,128)}, f32[4,1,1024]{2,0,1:T(4,128)}, f32[4,1,1024]{2,0,1:T(4,128)}, ...(+188)), kind=kLoop, calls=%fused_computation.10858\n  Allocation type: scoped</p>\n\n<p>XLA label: %fusion.11963 = f32[4,192,1024]{2,0,1:T(4,128)} fusion(f32[4,1,1024]{2,0,1:T(4,128)}, f32[4,1,1024]{2,0,1:T(4,128)}, f32[4,1,1024]{2,0,1:T(4,128)}, f32[4,1,1024]{2,0,1:T(4,128)}, ...(+188)), kind=kLoop, calls=%fused_computation.10858\n  Allocation type: scoped</p>\n\n<pre><code>TPU compilation failed\n [[{{node tpu_compile_succeeded_assert/_5292462253713114861/_5}}]]\n</code></pre>\n\n<p>Hint: If you want to see a list of allocated tensors when OOM happens, add report_tensor_allocations_upon_oom to RunOptions for current allocation info.</p>",
  "messages": [
    {
      "id": "798036",
      "postDate": "04/05/2020 05:50:09",
      "content": "<p>So, the error was this :--</p>\n\n<p>ResourceExhaustedError: {{function_node __inference_distributed_function_153395}} Compilation failure: Ran out of memory in memory space vmem. It should not be possible to run out of vmem - please file a bug against XLA.</p>\n\n<p>Largest program allocations in vmem:</p>\n\n<p>XLA label: register allocator spill slots\n  Allocation type: scoped</p>\n\n<p>XLA label: %fusion.11963 = f32[4,192,1024]{2,0,1:T(4,128)} fusion(f32[4,1,1024]{2,0,1:T(4,128)}, f32[4,1,1024]{2,0,1:T(4,128)}, f32[4,1,1024]{2,0,1:T(4,128)}, f32[4,1,1024]{2,0,1:T(4,128)}, ...(+188)), kind=kLoop, calls=%fused_computation.10858\n  Allocation type: scoped</p>\n\n<p>XLA label: %fusion.11963 = f32[4,192,1024]{2,0,1:T(4,128)} fusion(f32[4,1,1024]{2,0,1:T(4,128)}, f32[4,1,1024]{2,0,1:T(4,128)}, f32[4,1,1024]{2,0,1:T(4,128)}, f32[4,1,1024]{2,0,1:T(4,128)}, ...(+188)), kind=kLoop, calls=%fused_computation.10858\n  Allocation type: scoped</p>\n\n<p>XLA label: %fusion.11963 = f32[4,192,1024]{2,0,1:T(4,128)} fusion(f32[4,1,1024]{2,0,1:T(4,128)}, f32[4,1,1024]{2,0,1:T(4,128)}, f32[4,1,1024]{2,0,1:T(4,128)}, f32[4,1,1024]{2,0,1:T(4,128)}, ...(+188)), kind=kLoop, calls=%fused_computation.10858\n  Allocation type: scoped</p>\n\n<p>XLA label: %fusion.11963 = f32[4,192,1024]{2,0,1:T(4,128)} fusion(f32[4,1,1024]{2,0,1:T(4,128)}, f32[4,1,1024]{2,0,1:T(4,128)}, f32[4,1,1024]{2,0,1:T(4,128)}, f32[4,1,1024]{2,0,1:T(4,128)}, ...(+188)), kind=kLoop, calls=%fused_computation.10858\n  Allocation type: scoped</p>\n\n<pre><code>TPU compilation failed\n [[{{node tpu_compile_succeeded_assert/_5292462253713114861/_5}}]]\n</code></pre>\n\n<p>Hint: If you want to see a list of allocated tensors when OOM happens, add report_tensor_allocations_upon_oom to RunOptions for current allocation info.</p>",
      "rawMarkdown": "So, the error was this :--\n\nResourceExhaustedError: {{function_node __inference_distributed_function_153395}} Compilation failure: Ran out of memory in memory space vmem. It should not be possible to run out of vmem - please file a bug against XLA.\n\nLargest program allocations in vmem:\n\n  XLA label: register allocator spill slots\n  Allocation type: scoped\n\n  XLA label: %fusion.11963 = f32[4,192,1024]{2,0,1:T(4,128)} fusion(f32[4,1,1024]{2,0,1:T(4,128)}, f32[4,1,1024]{2,0,1:T(4,128)}, f32[4,1,1024]{2,0,1:T(4,128)}, f32[4,1,1024]{2,0,1:T(4,128)}, ...(+188)), kind=kLoop, calls=%fused_computation.10858\n  Allocation type: scoped\n\n  XLA label: %fusion.11963 = f32[4,192,1024]{2,0,1:T(4,128)} fusion(f32[4,1,1024]{2,0,1:T(4,128)}, f32[4,1,1024]{2,0,1:T(4,128)}, f32[4,1,1024]{2,0,1:T(4,128)}, f32[4,1,1024]{2,0,1:T(4,128)}, ...(+188)), kind=kLoop, calls=%fused_computation.10858\n  Allocation type: scoped\n\n  XLA label: %fusion.11963 = f32[4,192,1024]{2,0,1:T(4,128)} fusion(f32[4,1,1024]{2,0,1:T(4,128)}, f32[4,1,1024]{2,0,1:T(4,128)}, f32[4,1,1024]{2,0,1:T(4,128)}, f32[4,1,1024]{2,0,1:T(4,128)}, ...(+188)), kind=kLoop, calls=%fused_computation.10858\n  Allocation type: scoped\n\n  XLA label: %fusion.11963 = f32[4,192,1024]{2,0,1:T(4,128)} fusion(f32[4,1,1024]{2,0,1:T(4,128)}, f32[4,1,1024]{2,0,1:T(4,128)}, f32[4,1,1024]{2,0,1:T(4,128)}, f32[4,1,1024]{2,0,1:T(4,128)}, ...(+188)), kind=kLoop, calls=%fused_computation.10858\n  Allocation type: scoped\n\n\tTPU compilation failed\n\t [[{{node tpu_compile_succeeded_assert/_5292462253713114861/_5}}]]\nHint: If you want to see a list of allocated tensors when OOM happens, add report_tensor_allocations_upon_oom to RunOptions for current allocation info.",
      "votes": null
    },
    {
      "id": "799641",
      "postDate": "04/06/2020 15:50:36",
      "content": "<p>If you have a moment, please file an issue here <a href=\"https://github.com/tensorflow/tensorflow/issues\">https://github.com/tensorflow/tensorflow/issues</a>. There seemed to be a couple of reports with a similar issue, e.g.\n<a href=\"https://github.com/tensorflow/tensorflow/issues/35603\">https://github.com/tensorflow/tensorflow/issues/35603</a>\n<a href=\"https://github.com/tensorflow/tensorflow/issues/26893\">https://github.com/tensorflow/tensorflow/issues/26893</a></p>",
      "rawMarkdown": "If you have a moment, please file an issue here [https://github.com/tensorflow/tensorflow/issues](https://github.com/tensorflow/tensorflow/issues). There seemed to be a couple of reports with a similar issue, e.g.\nhttps://github.com/tensorflow/tensorflow/issues/35603\nhttps://github.com/tensorflow/tensorflow/issues/26893",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 799641,
      "author_name": "ifigotin",
      "author_url": "",
      "post_date": "04/06/2020 15:50:36",
      "content": "<p>If you have a moment, please file an issue here <a href=\"https://github.com/tensorflow/tensorflow/issues\">https://github.com/tensorflow/tensorflow/issues</a>. There seemed to be a couple of reports with a similar issue, e.g.\n<a href=\"https://github.com/tensorflow/tensorflow/issues/35603\">https://github.com/tensorflow/tensorflow/issues/35603</a>\n<a href=\"https://github.com/tensorflow/tensorflow/issues/26893\">https://github.com/tensorflow/tensorflow/issues/26893</a></p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "798036": "So, the error was this :--\n\nResourceExhaustedError: {{function_node __inference_distributed_function_153395}} Compilation failure: Ran out of memory in memory space vmem. It should not be possible to run out of vmem - please file a bug against XLA.\n\nLargest program allocations in vmem:\n\n  XLA label: register allocator spill slots\n  Allocation type: scoped\n\n  XLA label: %fusion.11963 = f32[4,192,1024]{2,0,1:T(4,128)} fusion(f32[4,1,1024]{2,0,1:T(4,128)}, f32[4,1,1024]{2,0,1:T(4,128)}, f32[4,1,1024]{2,0,1:T(4,128)}, f32[4,1,1024]{2,0,1:T(4,128)}, ...(+188)), kind=kLoop, calls=%fused_computation.10858\n  Allocation type: scoped\n\n  XLA label: %fusion.11963 = f32[4,192,1024]{2,0,1:T(4,128)} fusion(f32[4,1,1024]{2,0,1:T(4,128)}, f32[4,1,1024]{2,0,1:T(4,128)}, f32[4,1,1024]{2,0,1:T(4,128)}, f32[4,1,1024]{2,0,1:T(4,128)}, ...(+188)), kind=kLoop, calls=%fused_computation.10858\n  Allocation type: scoped\n\n  XLA label: %fusion.11963 = f32[4,192,1024]{2,0,1:T(4,128)} fusion(f32[4,1,1024]{2,0,1:T(4,128)}, f32[4,1,1024]{2,0,1:T(4,128)}, f32[4,1,1024]{2,0,1:T(4,128)}, f32[4,1,1024]{2,0,1:T(4,128)}, ...(+188)), kind=kLoop, calls=%fused_computation.10858\n  Allocation type: scoped\n\n  XLA label: %fusion.11963 = f32[4,192,1024]{2,0,1:T(4,128)} fusion(f32[4,1,1024]{2,0,1:T(4,128)}, f32[4,1,1024]{2,0,1:T(4,128)}, f32[4,1,1024]{2,0,1:T(4,128)}, f32[4,1,1024]{2,0,1:T(4,128)}, ...(+188)), kind=kLoop, calls=%fused_computation.10858\n  Allocation type: scoped\n\n\tTPU compilation failed\n\t [[{{node tpu_compile_succeeded_assert/_5292462253713114861/_5}}]]\nHint: If you want to see a list of allocated tensors when OOM happens, add report_tensor_allocations_upon_oom to RunOptions for current allocation info.",
    "799641": "If you have a moment, please file an issue here [https://github.com/tensorflow/tensorflow/issues](https://github.com/tensorflow/tensorflow/issues). There seemed to be a couple of reports with a similar issue, e.g.\nhttps://github.com/tensorflow/tensorflow/issues/35603\nhttps://github.com/tensorflow/tensorflow/issues/26893"
  },
  "source": "meta"
}