{
  "id": 132341,
  "title": "Error TPU compilation failed",
  "url": "/competitions/flower-classification-with-tpus/discussion/132341",
  "author_name": "",
  "post_date": "2020-02-25T12:49:12.913096Z",
  "votes": 1,
  "comment_count": 5,
  "views": 0,
  "content": "<p>Hi!\nI'm trying to use Tpus in my own dataset. I took the kaggle dataset (nih x rays) and i pass it to TFRecords in my own gcs storage bucket.\nIm avialable to read it and get access to the images and labels from my gcs bucket.\nThe problem comes after i create the model, in fit process. Give me the next error.</p>\n\n<blockquote>\n  <p>{{function_node __inference_distributed_function_20816}} Compilation failure: Asked to propagate a dynamic dimension from hlo dot.397@{}@1 to hlo %all-reduce.402 = f32[50,2]{1,0} all-reduce(f32[50,2]{1,0} %dot.397), replica_groups={{0,1,2,3,4,5,6,7}}, to_apply=%sum.398, metadata={op_type=\"CrossReplicaSum\" op_name=\"RMSprop/CrossReplicaSum_4\"}, which is not implemented.\n  TPU compilation failed\n  [[{{node tpu_compile_succeeded_assert/_10007542197504102739/_7}}]]</p>\n</blockquote>\n\n<p>I'm trying a small net to try to solve it. \nI had the same error in a kaggle notebook and a jupyter notebook from GC AI notebooks .</p>\n\n<p>My Tpu connection works fine.</p>\n\n<p>This is the net example: </p>\n\n<blockquote>\n  <p>with strategy.scope():</p>\n  \n  <p>model = tf.keras.Sequential([\n  tf.keras.layers.Flatten(input_shape=[512, 512, 1]),\n  tf.keras.layers.Dense(500, activation=\"relu\"),\n  tf.keras.layers.Dense(50, activation=\"relu\"),\n  tf.keras.layers.Dense(2, activation='softmax') # classifying into 5 classes\n      ])</p>\n  \n  <p>model.compile(\n  optimizer='adam',\n  loss= 'categorical_crossentropy',\n  metrics=['accuracy']) # % of correct answers</p>\n  \n  <p>model.fit(get_training_dataset(), steps_per_epoch=STEPS_PER_EPOCH, epochs=EPOCHS)</p>\n</blockquote>\n\n<p>Any tips?\nThanks</p>",
  "messages": [
    {
      "id": "756136",
      "postDate": "02/25/2020 12:49:12",
      "content": "<p>Hi!\nI'm trying to use Tpus in my own dataset. I took the kaggle dataset (nih x rays) and i pass it to TFRecords in my own gcs storage bucket.\nIm avialable to read it and get access to the images and labels from my gcs bucket.\nThe problem comes after i create the model, in fit process. Give me the next error.</p>\n\n<blockquote>\n  <p>{{function_node __inference_distributed_function_20816}} Compilation failure: Asked to propagate a dynamic dimension from hlo dot.397@{}@1 to hlo %all-reduce.402 = f32[50,2]{1,0} all-reduce(f32[50,2]{1,0} %dot.397), replica_groups={{0,1,2,3,4,5,6,7}}, to_apply=%sum.398, metadata={op_type=\"CrossReplicaSum\" op_name=\"RMSprop/CrossReplicaSum_4\"}, which is not implemented.\n  TPU compilation failed\n  [[{{node tpu_compile_succeeded_assert/_10007542197504102739/_7}}]]</p>\n</blockquote>\n\n<p>I'm trying a small net to try to solve it. \nI had the same error in a kaggle notebook and a jupyter notebook from GC AI notebooks .</p>\n\n<p>My Tpu connection works fine.</p>\n\n<p>This is the net example: </p>\n\n<blockquote>\n  <p>with strategy.scope():</p>\n  \n  <p>model = tf.keras.Sequential([\n  tf.keras.layers.Flatten(input_shape=[512, 512, 1]),\n  tf.keras.layers.Dense(500, activation=\"relu\"),\n  tf.keras.layers.Dense(50, activation=\"relu\"),\n  tf.keras.layers.Dense(2, activation='softmax') # classifying into 5 classes\n      ])</p>\n  \n  <p>model.compile(\n  optimizer='adam',\n  loss= 'categorical_crossentropy',\n  metrics=['accuracy']) # % of correct answers</p>\n  \n  <p>model.fit(get_training_dataset(), steps_per_epoch=STEPS_PER_EPOCH, epochs=EPOCHS)</p>\n</blockquote>\n\n<p>Any tips?\nThanks</p>",
      "rawMarkdown": "Hi!\nI'm trying to use Tpus in my own dataset. I took the kaggle dataset (nih x rays) and i pass it to TFRecords in my own gcs storage bucket.\nIm avialable to read it and get access to the images and labels from my gcs bucket.\nThe problem comes after i create the model, in fit process. Give me the next error.\n\n\n&gt; {{function_node __inference_distributed_function_20816}} Compilation failure: Asked to propagate a dynamic dimension from hlo dot.397@{}@1 to hlo %all-reduce.402 = f32[50,2]{1,0} all-reduce(f32[50,2]{1,0} %dot.397), replica_groups={{0,1,2,3,4,5,6,7}}, to_apply=%sum.398, metadata={op_type=\"CrossReplicaSum\" op_name=\"RMSprop/CrossReplicaSum_4\"}, which is not implemented.\nTPU compilation failed\n[[{{node tpu_compile_succeeded_assert/_10007542197504102739/_7}}]]\n\nI'm trying a small net to try to solve it. \nI had the same error in a kaggle notebook and a jupyter notebook from GC AI notebooks .\n\nMy Tpu connection works fine.\n\nThis is the net example: \n\n\n&gt;with strategy.scope():\n\n&gt; model = tf.keras.Sequential([\ntf.keras.layers.Flatten(input_shape=[512, 512, 1]),\ntf.keras.layers.Dense(500, activation=\"relu\"),\ntf.keras.layers.Dense(50, activation=\"relu\"),\ntf.keras.layers.Dense(2, activation='softmax') # classifying into 5 classes\n    ])\n\n&gt;model.compile(\noptimizer='adam',\nloss= 'categorical_crossentropy',\nmetrics=['accuracy']) # % of correct answers\n\n&gt;model.fit(get_training_dataset(), steps_per_epoch=STEPS_PER_EPOCH, epochs=EPOCHS)\n\n\nAny tips?\nThanks",
      "votes": null
    },
    {
      "id": "756327",
      "postDate": "02/25/2020 16:03:54",
      "content": "<p>I wonder it can be the channel number, is your own dataset are grey scale ?</p>",
      "rawMarkdown": "I wonder it can be the channel number, is your own dataset are grey scale ?",
      "votes": null
    },
    {
      "id": "756636",
      "postDate": "02/25/2020 22:59:06",
      "content": "<p>How did you set up your dataset ? Can you share the code ?\nI concur with <a href=\"/swordfaith\">@swordfaith</a> below, it might be worth checking into which format your images are decoded.</p>",
      "rawMarkdown": "How did you set up your dataset ? Can you share the code ?\nI concur with @swordfaith below, it might be worth checking into which format your images are decoded.",
      "votes": null
    },
    {
      "id": "757013",
      "postDate": "02/26/2020 10:44:02",
      "content": "<p>I already fix it. Was a problem with the labels. My problem is binary classification, but i would like to use labels [0,1] and [1,0] with a softmax. the reason of that is see the heatmaps of each class after train.\nIf i use binary labels (not one hot) i can train and all its okey.\nI still working to try train with the labels as one hot. </p>\n\n<p>The images were not the problem,a i code and decode as pngs and works fine. </p>",
      "rawMarkdown": "I already fix it. Was a problem with the labels. My problem is binary classification, but i would like to use labels [0,1] and [1,0] with a softmax. the reason of that is see the heatmaps of each class after train.\nIf i use binary labels (not one hot) i can train and all its okey.\nI still working to try train with the labels as one hot. \n\nThe images were not the problem,a i code and decode as pngs and works fine.",
      "votes": null
    },
    {
      "id": "757457",
      "postDate": "02/26/2020 19:29:51",
      "content": "<p>In Keras with one-hot labels, use <code>categorical_crossentropy</code>.\nWith numerical labels, use <code>sparse_categorical_crossentropy</code></p>",
      "rawMarkdown": "In Keras with one-hot labels, use `categorical_crossentropy`.\nWith numerical labels, use `sparse_categorical_crossentropy`",
      "votes": null
    },
    {
      "id": "3338175",
      "postDate": "11/18/2025 23:03:06",
      "content": "<p>Title: 🚨 Urgent: TPU Initialization Fails (OpKernel Missing) Despite Active TPU Runtime</p>\n<p>Hello Kaggle Community,</p>\n<p>I am facing a persistent issue when trying to initialize the TPU in my notebook. I have already confirmed that:</p>\n<p>My account is Verified.\n  The Notebook Accelerator is set to TPU.\n  My TPU quota is currently available.</p>\n<p>However, the standard initialization code consistently throws a <code>NotFoundError</code> because the required OpKernel is missing. I suspect this is an environment configuration issue on the platform itself.</p>\n<p>Has anyone encountered this specific <code>OpKernel not registered</code> error recently while using the TPU runtime and found a workaround?</p>\n<hr>\n<p>Code and Error Details\nCode Used:</p>\n<pre><code> tensorflow  tf\n\n\ntpu = tf.distribute.cluster_resolver.TPUClusterResolver(tpu=)\ntf.tpu.experimental.initialize_tpu_system(tpu)\n\n\nstrategy = tf.distribute.TPUStrategy(tpu)\n()\n</code></pre>\n<p>Traceback Snippet:</p>\n<pre><code>InvalidArgumentError:  OpKernel was registered  support Op  used by {{node ConfigureDistributedTPU}}\n.\nRegistered devices: [CPU]\nRegistered kernels:\n  &lt; registered kernels&gt;\n\nDuring handling of the above exception, another exception occurred:\nNotFoundError: TPUs  found  the cluster. Failed  initialization:  OpKernel was registered  support Op .\n</code></pre>\n<p>Key Observation:\nThe output shows <code>Registered devices: [CPU]</code>, confirming that the environment is not detecting the active TPU accelerator at the TensorFlow software level.</p>\n<p>Any assistance or known workarounds would be greatly appreciated! Thank you.</p>",
      "rawMarkdown": "Title: 🚨 Urgent: TPU Initialization Fails (OpKernel Missing) Despite Active TPU Runtime\n\nHello Kaggle Community,\n\nI am facing a persistent issue when trying to initialize the TPU in my notebook. I have already confirmed that:\n\n  My account is Verified.\n  The Notebook Accelerator is set to TPU.\n  My TPU quota is currently available.\n\nHowever, the standard initialization code consistently throws a `NotFoundError` because the required OpKernel is missing. I suspect this is an environment configuration issue on the platform itself.\n\nHas anyone encountered this specific `OpKernel not registered` error recently while using the TPU runtime and found a workaround?\n\n-----\n\nCode and Error Details\nCode Used:\n```python\nimport tensorflow as tf\n\n# Detect and initialize TPU\ntpu = tf.distribute.cluster_resolver.TPUClusterResolver(tpu='local')\ntf.tpu.experimental.initialize_tpu_system(tpu)\n\n# Create TPU distribution strategy\nstrategy = tf.distribute.TPUStrategy(tpu)\nprint(\"TPU initialized successfully.\")\n```\n\nTraceback Snippet:\n\n```\nInvalidArgumentError: No OpKernel was registered to support Op 'ConfigureDistributedTPU' used by {{node ConfigureDistributedTPU}}\n...\nRegistered devices: [CPU]\nRegistered kernels:\n  <no registered kernels>\n\nDuring handling of the above exception, another exception occurred:\nNotFoundError: TPUs not found in the cluster. Failed in initialization: No OpKernel was registered to support Op 'ConfigureDistributedTPU'...\n```\n\nKey Observation:\nThe output shows `Registered devices: [CPU]`, confirming that the environment is not detecting the active TPU accelerator at the TensorFlow software level.\n\nAny assistance or known workarounds would be greatly appreciated\\! Thank you.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 3338175,
      "author_name": "zahraghgh",
      "author_url": "",
      "post_date": "11/18/2025 23:03:06",
      "content": "<p>Title: 🚨 Urgent: TPU Initialization Fails (OpKernel Missing) Despite Active TPU Runtime</p>\n<p>Hello Kaggle Community,</p>\n<p>I am facing a persistent issue when trying to initialize the TPU in my notebook. I have already confirmed that:</p>\n<p>My account is Verified.\n  The Notebook Accelerator is set to TPU.\n  My TPU quota is currently available.</p>\n<p>However, the standard initialization code consistently throws a <code>NotFoundError</code> because the required OpKernel is missing. I suspect this is an environment configuration issue on the platform itself.</p>\n<p>Has anyone encountered this specific <code>OpKernel not registered</code> error recently while using the TPU runtime and found a workaround?</p>\n<hr>\n<p>Code and Error Details\nCode Used:</p>\n<pre><code> tensorflow  tf\n\n\ntpu = tf.distribute.cluster_resolver.TPUClusterResolver(tpu=)\ntf.tpu.experimental.initialize_tpu_system(tpu)\n\n\nstrategy = tf.distribute.TPUStrategy(tpu)\n()\n</code></pre>\n<p>Traceback Snippet:</p>\n<pre><code>InvalidArgumentError:  OpKernel was registered  support Op  used by {{node ConfigureDistributedTPU}}\n.\nRegistered devices: [CPU]\nRegistered kernels:\n  &lt; registered kernels&gt;\n\nDuring handling of the above exception, another exception occurred:\nNotFoundError: TPUs  found  the cluster. Failed  initialization:  OpKernel was registered  support Op .\n</code></pre>\n<p>Key Observation:\nThe output shows <code>Registered devices: [CPU]</code>, confirming that the environment is not detecting the active TPU accelerator at the TensorFlow software level.</p>\n<p>Any assistance or known workarounds would be greatly appreciated! Thank you.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 756327,
      "author_name": "swordfaith",
      "author_url": "",
      "post_date": "02/25/2020 16:03:54",
      "content": "<p>I wonder it can be the channel number, is your own dataset are grey scale ?</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 756636,
      "author_name": "mgorner",
      "author_url": "",
      "post_date": "02/25/2020 22:59:06",
      "content": "<p>How did you set up your dataset ? Can you share the code ?\nI concur with <a href=\"/swordfaith\">@swordfaith</a> below, it might be worth checking into which format your images are decoded.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 757013,
      "author_name": "ianwilkinson",
      "author_url": "",
      "post_date": "02/26/2020 10:44:02",
      "content": "<p>I already fix it. Was a problem with the labels. My problem is binary classification, but i would like to use labels [0,1] and [1,0] with a softmax. the reason of that is see the heatmaps of each class after train.\nIf i use binary labels (not one hot) i can train and all its okey.\nI still working to try train with the labels as one hot. </p>\n\n<p>The images were not the problem,a i code and decode as pngs and works fine. </p>",
      "votes": null,
      "replies": [
        {
          "id": 757457,
          "author_name": "mgorner",
          "author_url": "",
          "post_date": "02/26/2020 19:29:51",
          "content": "<p>In Keras with one-hot labels, use <code>categorical_crossentropy</code>.\nWith numerical labels, use <code>sparse_categorical_crossentropy</code></p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "756136": "Hi!\nI'm trying to use Tpus in my own dataset. I took the kaggle dataset (nih x rays) and i pass it to TFRecords in my own gcs storage bucket.\nIm avialable to read it and get access to the images and labels from my gcs bucket.\nThe problem comes after i create the model, in fit process. Give me the next error.\n\n\n&gt; {{function_node __inference_distributed_function_20816}} Compilation failure: Asked to propagate a dynamic dimension from hlo dot.397@{}@1 to hlo %all-reduce.402 = f32[50,2]{1,0} all-reduce(f32[50,2]{1,0} %dot.397), replica_groups={{0,1,2,3,4,5,6,7}}, to_apply=%sum.398, metadata={op_type=\"CrossReplicaSum\" op_name=\"RMSprop/CrossReplicaSum_4\"}, which is not implemented.\nTPU compilation failed\n[[{{node tpu_compile_succeeded_assert/_10007542197504102739/_7}}]]\n\nI'm trying a small net to try to solve it. \nI had the same error in a kaggle notebook and a jupyter notebook from GC AI notebooks .\n\nMy Tpu connection works fine.\n\nThis is the net example: \n\n\n&gt;with strategy.scope():\n\n&gt; model = tf.keras.Sequential([\ntf.keras.layers.Flatten(input_shape=[512, 512, 1]),\ntf.keras.layers.Dense(500, activation=\"relu\"),\ntf.keras.layers.Dense(50, activation=\"relu\"),\ntf.keras.layers.Dense(2, activation='softmax') # classifying into 5 classes\n    ])\n\n&gt;model.compile(\noptimizer='adam',\nloss= 'categorical_crossentropy',\nmetrics=['accuracy']) # % of correct answers\n\n&gt;model.fit(get_training_dataset(), steps_per_epoch=STEPS_PER_EPOCH, epochs=EPOCHS)\n\n\nAny tips?\nThanks",
    "756327": "I wonder it can be the channel number, is your own dataset are grey scale ?",
    "756636": "How did you set up your dataset ? Can you share the code ?\nI concur with @swordfaith below, it might be worth checking into which format your images are decoded.",
    "757013": "I already fix it. Was a problem with the labels. My problem is binary classification, but i would like to use labels [0,1] and [1,0] with a softmax. the reason of that is see the heatmaps of each class after train.\nIf i use binary labels (not one hot) i can train and all its okey.\nI still working to try train with the labels as one hot. \n\nThe images were not the problem,a i code and decode as pngs and works fine.",
    "757457": "In Keras with one-hot labels, use `categorical_crossentropy`.\nWith numerical labels, use `sparse_categorical_crossentropy`",
    "3338175": "Title: 🚨 Urgent: TPU Initialization Fails (OpKernel Missing) Despite Active TPU Runtime\n\nHello Kaggle Community,\n\nI am facing a persistent issue when trying to initialize the TPU in my notebook. I have already confirmed that:\n\n  My account is Verified.\n  The Notebook Accelerator is set to TPU.\n  My TPU quota is currently available.\n\nHowever, the standard initialization code consistently throws a `NotFoundError` because the required OpKernel is missing. I suspect this is an environment configuration issue on the platform itself.\n\nHas anyone encountered this specific `OpKernel not registered` error recently while using the TPU runtime and found a workaround?\n\n-----\n\nCode and Error Details\nCode Used:\n```python\nimport tensorflow as tf\n\n# Detect and initialize TPU\ntpu = tf.distribute.cluster_resolver.TPUClusterResolver(tpu='local')\ntf.tpu.experimental.initialize_tpu_system(tpu)\n\n# Create TPU distribution strategy\nstrategy = tf.distribute.TPUStrategy(tpu)\nprint(\"TPU initialized successfully.\")\n```\n\nTraceback Snippet:\n\n```\nInvalidArgumentError: No OpKernel was registered to support Op 'ConfigureDistributedTPU' used by {{node ConfigureDistributedTPU}}\n...\nRegistered devices: [CPU]\nRegistered kernels:\n  <no registered kernels>\n\nDuring handling of the above exception, another exception occurred:\nNotFoundError: TPUs not found in the cluster. Failed in initialization: No OpKernel was registered to support Op 'ConfigureDistributedTPU'...\n```\n\nKey Observation:\nThe output shows `Registered devices: [CPU]`, confirming that the environment is not detecting the active TPU accelerator at the TensorFlow software level.\n\nAny assistance or known workarounds would be greatly appreciated\\! Thank you."
  },
  "source": "meta"
}