{
  "id": 672464,
  "title": "Unable to initialize TPU v5e-8",
  "url": "/competitions/vesuvius-challenge-surface-detection/discussion/672464",
  "author_name": "",
  "post_date": "2026-02-08T12:34:35.648339800Z",
  "votes": 4,
  "comment_count": 13,
  "views": 0,
  "content": "<p>I can't figure out why TPU v5e-8 won't start. <a href=\"https://www.kaggle.com/code/konstantinboyko/train-vesuvius-surface-3d-detection-200-images?scriptVersionId=296537209\" target=\"_blank\">Notebook</a></p>\n<p>In the second option there is a kernel error. <a href=\"https://www.kaggle.com/code/konstantinboyko/train-vesuvius-surface-3d-detection-200-images?scriptVersionId=296541440\" target=\"_blank\">Notebook</a></p>",
  "messages": [
    {
      "id": "3403417",
      "postDate": "02/08/2026 12:34:35",
      "content": "<p>I can't figure out why TPU v5e-8 won't start. <a href=\"https://www.kaggle.com/code/konstantinboyko/train-vesuvius-surface-3d-detection-200-images?scriptVersionId=296537209\" target=\"_blank\">Notebook</a></p>\n<p>In the second option there is a kernel error. <a href=\"https://www.kaggle.com/code/konstantinboyko/train-vesuvius-surface-3d-detection-200-images?scriptVersionId=296541440\" target=\"_blank\">Notebook</a></p>",
      "rawMarkdown": "I can't figure out why TPU v5e-8 won't start. [Notebook](https://www.kaggle.com/code/konstantinboyko/train-vesuvius-surface-3d-detection-200-images?scriptVersionId=296537209)\n\nIn the second option there is a kernel error. [Notebook](https://www.kaggle.com/code/konstantinboyko/train-vesuvius-surface-3d-detection-200-images?scriptVersionId=296541440)",
      "votes": null
    },
    {
      "id": "3403426",
      "postDate": "02/08/2026 12:48:36",
      "content": "<p>It has been broken for many months. To use TPU, you have to set <code>jax</code> backend these days.</p>",
      "rawMarkdown": "It has been broken for many months. To use TPU, you have to set `jax` backend these days.",
      "votes": null
    },
    {
      "id": "3403427",
      "postDate": "02/08/2026 12:49:59",
      "content": "<p>cc. <a href=\"https://www.kaggle.com/herbison\" target=\"_blank\">@herbison</a> \nFor more input on this.</p>",
      "rawMarkdown": "cc. @herbison \nFor more input on this.",
      "votes": null
    },
    {
      "id": "3403435",
      "postDate": "02/08/2026 13:12:59",
      "content": "<p>During the <code>MAP - Charting Student Math Misunderstandings</code> competition it seemed to work.</p>",
      "rawMarkdown": "During the `MAP - Charting Student Math Misunderstandings` competition it seemed to work.",
      "votes": null
    },
    {
      "id": "3403440",
      "postDate": "02/08/2026 13:23:42",
      "content": "<p>In that case, you can try creating a notebook from that competition, or old one which proven to be worked earlier. </p>",
      "rawMarkdown": "In that case, you can try creating a notebook from that competition, or old one which proven to be worked earlier.",
      "votes": null
    },
    {
      "id": "3403473",
      "postDate": "02/08/2026 15:22:10",
      "content": "<p>I think you need to see this <a href=\"https://www.kaggle.com/ipythonx\" target=\"_blank\">@ipythonx</a> <a href=\"https://www.kaggle.com/konstantinboyko\" target=\"_blank\">@konstantinboyko</a> <br>\n<a href=\"https://www.kaggle.com/discussions/product-announcements/670573\" target=\"_blank\">https://www.kaggle.com/discussions/product-announcements/670573</a></p>",
      "rawMarkdown": "I think you need to see this @ipythonx @konstantinboyko <br>\nhttps://www.kaggle.com/discussions/product-announcements/670573",
      "votes": null
    },
    {
      "id": "3403534",
      "postDate": "02/08/2026 17:51:14",
      "content": "<p>Ravi, thanks for this information, but I went through identity authentication a long time ago, when it was first introduced.</p>",
      "rawMarkdown": "Ravi, thanks for this information, but I went through identity authentication a long time ago, when it was first introduced.",
      "votes": null
    },
    {
      "id": "3403562",
      "postDate": "02/08/2026 19:15:18",
      "content": "<p>It looks like the torch+xla option works. <a href=\"https://www.kaggle.com/code/konstantinboyko/map-deepseekmath-7b-it-tpu-train-bf16?scriptVersionId=296596577\" target=\"_blank\">Notebook</a></p>\n<pre><code>WARNING: Logging before InitGoogle() is written to STDERR\nE0000 00:00:1770577160.730043      15 common_lib.cc:621] Could not set metric server port: INVALID_ARGUMENT: Could not find SliceBuilder port 8471 in any of the 0 ports provided in `tpu_process_addresses`=\"local\"\n=== Source Location Trace: ===\nlearning/45eac/tfrc/runtime/common_lib.cc:232\nNum devices: 8\n</code></pre>",
      "rawMarkdown": "It looks like the torch+xla option works. [Notebook](https://www.kaggle.com/code/konstantinboyko/map-deepseekmath-7b-it-tpu-train-bf16?scriptVersionId=296596577)\n\n```\nWARNING: Logging before InitGoogle() is written to STDERR\nE0000 00:00:1770577160.730043      15 common_lib.cc:621] Could not set metric server port: INVALID_ARGUMENT: Could not find SliceBuilder port 8471 in any of the 0 ports provided in `tpu_process_addresses`=\"local\"\n=== Source Location Trace: ===\nlearning/45eac/tfrc/runtime/common_lib.cc:232\nNum devices: 8\n```",
      "votes": null
    },
    {
      "id": "3403574",
      "postDate": "02/08/2026 19:50:31",
      "content": "<p>I think it would be more helpful to follow this <a href=\"https://github.com/Kaggle/docker-python/issues/1367\" target=\"_blank\">ticket</a>.</p>",
      "rawMarkdown": "I think it would be more helpful to follow this [ticket](https://github.com/Kaggle/docker-python/issues/1367).",
      "votes": null
    },
    {
      "id": "3403578",
      "postDate": "02/08/2026 19:54:53",
      "content": "<p>Also, (a bit more work though). Try using colab TPU with tf backend, it works, then check TF version and CUDA version. Then if possible maybe yoo can use that to install here. </p>\n<p>Also, in kaggle see the <a href=\"https://www.kaggle.com/organizations/keras\" target=\"_blank\">keras organizaion</a>, find a noteboo that use tf-backend with Keras 3. You might find something helpful. </p>\n<p>But I would recommened you to use <code>jax</code> backend, why wasting time when itme is limited at this point.</p>",
      "rawMarkdown": "Also, (a bit more work though). Try using colab TPU with tf backend, it works, then check TF version and CUDA version. Then if possible maybe yoo can use that to install here. \n\nAlso, in kaggle see the [keras organizaion](https://www.kaggle.com/organizations/keras), find a noteboo that use tf-backend with Keras 3. You might find something helpful. \n\nBut I would recommened you to use `jax` backend, why wasting time when itme is limited at this point.",
      "votes": null
    },
    {
      "id": "3403588",
      "postDate": "02/08/2026 20:35:31",
      "content": "<p><a href=\"https://www.kaggle.com/konstantinboyko\" target=\"_blank\">@konstantinboyko</a> \nI've found a temporary working soluton. Try this </p>\n<pre><code>!pip install -U -q tensorflow-tpu==2.19.1 -q\n\nimport os\nimport libtpu\n\nos.environ[\"KERAS_BACKEND\"] = \"tensorflow\"\nos.environ[\"PJRT_DEVICE\"] = \"TPU\"\nos.environ[\"NEXT_PLUGGABLE_DEVICE_USE_C_API\"] = \"true\"\nos.environ[\"TF_PLUGGABLE_DEVICE_LIBRARY_PATH\"] = libtpu.get_library_path()\nos.environ[\"TF_XLA_FLAGS\"] = (\n    \"--tf_mlir_enable_mlir_bridge=true \"\n    \"--tf_mlir_enable_convert_control_to_data_outputs_pass=true \"\n    \"--tf_mlir_enable_merge_control_flow_pass=true\"\n)\n</code></pre>\n<pre><code>import numpy as np # linear algebra\nimport pandas as pd # data processing, CSV file I/O (e.g. pd.read_csv)\nimport os\n\nimport keras\nimport tensorflow as tf\ntf.__version__, keras.version()\n# ('2.19.1', '3.13.0')\n</code></pre>\n<pre><code>resolver = tf.distribute.cluster_resolver.TPUClusterResolver(tpu=\"local\")\ntopology = tf.tpu.experimental.initialize_tpu_system(resolver)\ntpu_metadata = resolver.get_tpu_system_metadata()\ndevice_assignment = tf.tpu.experimental.DeviceAssignment.build(\n    topology, num_replicas=tpu_metadata.num_cores\n)\nstrategy = tf.distribute.TPUStrategy(\n    resolver, experimental_device_assignment=device_assignment\n)\nstrategy.num_replicas_in_sync\n8\n</code></pre>\n<pre><code>(x_train, y_train), (x_test, y_test) = keras.datasets.mnist.load_data()\nx_train = x_train.reshape(60000, 784).astype(\"float32\") / 255\nx_test = x_test.reshape(10000, 784).astype(\"float32\") / 255\n\ndef my_model():\n    inputs = keras.Input(shape=(784,))\n    x = layers.Dense(64, activation=\"relu\")(inputs)\n    x = layers.Dense(64, activation=\"relu\")(x)\n    outputs = layers.Dense(10)(x)\n    model = keras.Model(\n        inputs=inputs, outputs=outputs, name=\"mnist_model\"\n    )\n    return model\n\n\nwith strategy.scope():\n    model = my_model()\n    model.compile(\n        optimizer=\"adam\",\n        loss=keras.losses.SparseCategoricalCrossentropy(from_logits=True),\n        metrics=[\"accuracy\"],\n    )\n\nwith strategy.scope():\n    history = model.fit(\n        x_train, y_train, batch_size=64, epochs=2, validation_split=0.2\n        )\n\nwith strategy.scope():\n    test_scores = model.evaluate(x_test, y_test, verbose=2)\nprint(\"Test loss:\", test_scores[0])\nprint(\"Test accuracy:\", test_scores[1])\n</code></pre>\n<pre><code>Epoch 1/2\n750/750 7ms/step - accuracy: 0.8977 - loss: 0.3565 - val_accuracy: 0.9461 - val_loss: 0.1905\nEpoch 2/2\n750/750 6ms/step - accuracy: 0.9521 - loss: 0.1616 - val_accuracy: 0.9576 - val_loss: 0.1400\n313/313 - 2s - 7ms/step - accuracy: 0.9602 - loss: 0.1328\nTest loss: 0.132845938205719\nTest accuracy: 0.9602000117301941\n</code></pre>\n<p>TPU-config <a href=\"https://keras.io/keras_rs/examples/distributed_embedding_tf/\" target=\"_blank\">reference.</a></p>",
      "rawMarkdown": "konstantinboyko \nI've found a temporary working soluton. Try this \n\n```python\n!pip install -U -q tensorflow-tpu==2.19.1 -q\n\nimport os\nimport libtpu\n\nos.environ[\"KERAS_BACKEND\"] = \"tensorflow\"\nos.environ[\"PJRT_DEVICE\"] = \"TPU\"\nos.environ[\"NEXT_PLUGGABLE_DEVICE_USE_C_API\"] = \"true\"\nos.environ[\"TF_PLUGGABLE_DEVICE_LIBRARY_PATH\"] = libtpu.get_library_path()\nos.environ[\"TF_XLA_FLAGS\"] = (\n    \"--tf_mlir_enable_mlir_bridge=true \"\n    \"--tf_mlir_enable_convert_control_to_data_outputs_pass=true \"\n    \"--tf_mlir_enable_merge_control_flow_pass=true\"\n)\n```\n```python\nimport numpy as np # linear algebra\nimport pandas as pd # data processing, CSV file I/O (e.g. pd.read_csv)\nimport os\n\nimport keras\nimport tensorflow as tf\ntf.__version__, keras.version()\n# ('2.19.1', '3.13.0')\n```\n```python\nresolver = tf.distribute.cluster_resolver.TPUClusterResolver(tpu=\"local\")\ntopology = tf.tpu.experimental.initialize_tpu_system(resolver)\ntpu_metadata = resolver.get_tpu_system_metadata()\ndevice_assignment = tf.tpu.experimental.DeviceAssignment.build(\n    topology, num_replicas=tpu_metadata.num_cores\n)\nstrategy = tf.distribute.TPUStrategy(\n    resolver, experimental_device_assignment=device_assignment\n)\nstrategy.num_replicas_in_sync\n8\n```\n```python\n(x_train, y_train), (x_test, y_test) = keras.datasets.mnist.load_data()\nx_train = x_train.reshape(60000, 784).astype(\"float32\") / 255\nx_test = x_test.reshape(10000, 784).astype(\"float32\") / 255\n\ndef my_model():\n    inputs = keras.Input(shape=(784,))\n    x = layers.Dense(64, activation=\"relu\")(inputs)\n    x = layers.Dense(64, activation=\"relu\")(x)\n    outputs = layers.Dense(10)(x)\n    model = keras.Model(\n        inputs=inputs, outputs=outputs, name=\"mnist_model\"\n    )\n    return model\n\n\nwith strategy.scope():\n    model = my_model()\n    model.compile(\n        optimizer=\"adam\",\n        loss=keras.losses.SparseCategoricalCrossentropy(from_logits=True),\n        metrics=[\"accuracy\"],\n    )\n\nwith strategy.scope():\n    history = model.fit(\n        x_train, y_train, batch_size=64, epochs=2, validation_split=0.2\n        )\n\nwith strategy.scope():\n    test_scores = model.evaluate(x_test, y_test, verbose=2)\nprint(\"Test loss:\", test_scores[0])\nprint(\"Test accuracy:\", test_scores[1])\n```\n```bash\n\nEpoch 1/2\n750/750 7ms/step - accuracy: 0.8977 - loss: 0.3565 - val_accuracy: 0.9461 - val_loss: 0.1905\nEpoch 2/2\n750/750 6ms/step - accuracy: 0.9521 - loss: 0.1616 - val_accuracy: 0.9576 - val_loss: 0.1400\n313/313 - 2s - 7ms/step - accuracy: 0.9602 - loss: 0.1328\nTest loss: 0.132845938205719\nTest accuracy: 0.9602000117301941\n```\n\nTPU-config [reference.](https://keras.io/keras_rs/examples/distributed_embedding_tf/)",
      "votes": null
    },
    {
      "id": "3403600",
      "postDate": "02/08/2026 21:58:22",
      "content": "<p>Innat, thank you so much for your help! It works.</p>",
      "rawMarkdown": "Innat, thank you so much for your help! It works.",
      "votes": null
    },
    {
      "id": "3403742",
      "postDate": "02/09/2026 07:26:18",
      "content": "<p>If I take the three metrics required by the competition conditions and calculate them immediately during training, the training converges for me within two to ten epochs. After that, even up to the two-hundredth epoch, things only get worse. And after the 130th epoch, pure overfitting occurs. </p>\n<p>Your notebook converges on the 8th epoch. Manas Choudhary's notebook converges on the 4th epoch, and the combination of these two notebooks converges on the 2nd epoch.</p>\n<p>I saved all the epochs and checked, and the triple metric showed the best epoch correctly.</p>\n<p>The idea for metrics was borrowed from the <code>Metric Failure Demo</code> notebook.</p>",
      "rawMarkdown": "If I take the three metrics required by the competition conditions and calculate them immediately during training, the training converges for me within two to ten epochs. After that, even up to the two-hundredth epoch, things only get worse. And after the 130th epoch, pure overfitting occurs. \n\nYour notebook converges on the 8th epoch. Manas Choudhary's notebook converges on the 4th epoch, and the combination of these two notebooks converges on the 2nd epoch.\n\nI saved all the epochs and checked, and the triple metric showed the best epoch correctly.\n\nThe idea for metrics was borrowed from the `Metric Failure Demo` notebook.",
      "votes": null
    },
    {
      "id": "3403767",
      "postDate": "02/09/2026 08:31:48",
      "content": "<p>Can't say for sure the main cause. I would do</p>\n<ul>\n<li>use 2 or 3 very simple augmentation</li>\n<li>use static learning rate, not scheduler for testing</li>\n<li>use loss function and gradually add more, for example, 1st dice_ce loss, in 2nd experiment, dice_ce + cldice loss, etc. As we know, the loss functions are responsible for optimization. So, either dataloader gives some bad processed sample or, the learing rate scheduler cause the quick converges effect. </li>\n</ul>\n<p>The goal is to identify which may cause the quick converges. </p>",
      "rawMarkdown": "Can't say for sure the main cause. I would do\n\n- use 2 or 3 very simple augmentation\n- use static learning rate, not scheduler for testing\n- use loss function and gradually add more, for example, 1st dice_ce loss, in 2nd experiment, dice_ce + cldice loss, etc. As we know, the loss functions are responsible for optimization. So, either dataloader gives some bad processed sample or, the learing rate scheduler cause the quick converges effect. \n\nThe goal is to identify which may cause the quick converges.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 3403426,
      "author_name": "ipythonx",
      "author_url": "",
      "post_date": "02/08/2026 12:48:36",
      "content": "<p>It has been broken for many months. To use TPU, you have to set <code>jax</code> backend these days.</p>",
      "votes": null,
      "replies": [
        {
          "id": 3403435,
          "author_name": "konstantinboyko",
          "author_url": "",
          "post_date": "02/08/2026 13:12:59",
          "content": "<p>During the <code>MAP - Charting Student Math Misunderstandings</code> competition it seemed to work.</p>",
          "votes": null,
          "replies": [
            {
              "id": 3403440,
              "author_name": "ipythonx",
              "author_url": "",
              "post_date": "02/08/2026 13:23:42",
              "content": "<p>In that case, you can try creating a notebook from that competition, or old one which proven to be worked earlier. </p>",
              "votes": null,
              "replies": [
                {
                  "id": 3403562,
                  "author_name": "konstantinboyko",
                  "author_url": "",
                  "post_date": "02/08/2026 19:15:18",
                  "content": "<p>It looks like the torch+xla option works. <a href=\"https://www.kaggle.com/code/konstantinboyko/map-deepseekmath-7b-it-tpu-train-bf16?scriptVersionId=296596577\" target=\"_blank\">Notebook</a></p>\n<pre><code>WARNING: Logging before InitGoogle() is written to STDERR\nE0000 00:00:1770577160.730043      15 common_lib.cc:621] Could not set metric server port: INVALID_ARGUMENT: Could not find SliceBuilder port 8471 in any of the 0 ports provided in `tpu_process_addresses`=\"local\"\n=== Source Location Trace: ===\nlearning/45eac/tfrc/runtime/common_lib.cc:232\nNum devices: 8\n</code></pre>",
                  "votes": null,
                  "replies": [
                    {
                      "id": 3403574,
                      "author_name": "ipythonx",
                      "author_url": "",
                      "post_date": "02/08/2026 19:50:31",
                      "content": "<p>I think it would be more helpful to follow this <a href=\"https://github.com/Kaggle/docker-python/issues/1367\" target=\"_blank\">ticket</a>.</p>",
                      "votes": null,
                      "replies": [
                        {
                          "id": 3403578,
                          "author_name": "ipythonx",
                          "author_url": "",
                          "post_date": "02/08/2026 19:54:53",
                          "content": "<p>Also, (a bit more work though). Try using colab TPU with tf backend, it works, then check TF version and CUDA version. Then if possible maybe yoo can use that to install here. </p>\n<p>Also, in kaggle see the <a href=\"https://www.kaggle.com/organizations/keras\" target=\"_blank\">keras organizaion</a>, find a noteboo that use tf-backend with Keras 3. You might find something helpful. </p>\n<p>But I would recommened you to use <code>jax</code> backend, why wasting time when itme is limited at this point.</p>",
                          "votes": null,
                          "replies": []
                        }
                      ]
                    }
                  ]
                }
              ]
            }
          ]
        }
      ]
    },
    {
      "id": 3403427,
      "author_name": "ipythonx",
      "author_url": "",
      "post_date": "02/08/2026 12:49:59",
      "content": "<p>cc. <a href=\"https://www.kaggle.com/herbison\" target=\"_blank\">@herbison</a> \nFor more input on this.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 3403473,
      "author_name": "ravi20076",
      "author_url": "",
      "post_date": "02/08/2026 15:22:10",
      "content": "<p>I think you need to see this <a href=\"https://www.kaggle.com/ipythonx\" target=\"_blank\">@ipythonx</a> <a href=\"https://www.kaggle.com/konstantinboyko\" target=\"_blank\">@konstantinboyko</a> <br>\n<a href=\"https://www.kaggle.com/discussions/product-announcements/670573\" target=\"_blank\">https://www.kaggle.com/discussions/product-announcements/670573</a></p>",
      "votes": null,
      "replies": [
        {
          "id": 3403534,
          "author_name": "konstantinboyko",
          "author_url": "",
          "post_date": "02/08/2026 17:51:14",
          "content": "<p>Ravi, thanks for this information, but I went through identity authentication a long time ago, when it was first introduced.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 3403588,
      "author_name": "ipythonx",
      "author_url": "",
      "post_date": "02/08/2026 20:35:31",
      "content": "<p><a href=\"https://www.kaggle.com/konstantinboyko\" target=\"_blank\">@konstantinboyko</a> \nI've found a temporary working soluton. Try this </p>\n<pre><code>!pip install -U -q tensorflow-tpu==2.19.1 -q\n\nimport os\nimport libtpu\n\nos.environ[\"KERAS_BACKEND\"] = \"tensorflow\"\nos.environ[\"PJRT_DEVICE\"] = \"TPU\"\nos.environ[\"NEXT_PLUGGABLE_DEVICE_USE_C_API\"] = \"true\"\nos.environ[\"TF_PLUGGABLE_DEVICE_LIBRARY_PATH\"] = libtpu.get_library_path()\nos.environ[\"TF_XLA_FLAGS\"] = (\n    \"--tf_mlir_enable_mlir_bridge=true \"\n    \"--tf_mlir_enable_convert_control_to_data_outputs_pass=true \"\n    \"--tf_mlir_enable_merge_control_flow_pass=true\"\n)\n</code></pre>\n<pre><code>import numpy as np # linear algebra\nimport pandas as pd # data processing, CSV file I/O (e.g. pd.read_csv)\nimport os\n\nimport keras\nimport tensorflow as tf\ntf.__version__, keras.version()\n# ('2.19.1', '3.13.0')\n</code></pre>\n<pre><code>resolver = tf.distribute.cluster_resolver.TPUClusterResolver(tpu=\"local\")\ntopology = tf.tpu.experimental.initialize_tpu_system(resolver)\ntpu_metadata = resolver.get_tpu_system_metadata()\ndevice_assignment = tf.tpu.experimental.DeviceAssignment.build(\n    topology, num_replicas=tpu_metadata.num_cores\n)\nstrategy = tf.distribute.TPUStrategy(\n    resolver, experimental_device_assignment=device_assignment\n)\nstrategy.num_replicas_in_sync\n8\n</code></pre>\n<pre><code>(x_train, y_train), (x_test, y_test) = keras.datasets.mnist.load_data()\nx_train = x_train.reshape(60000, 784).astype(\"float32\") / 255\nx_test = x_test.reshape(10000, 784).astype(\"float32\") / 255\n\ndef my_model():\n    inputs = keras.Input(shape=(784,))\n    x = layers.Dense(64, activation=\"relu\")(inputs)\n    x = layers.Dense(64, activation=\"relu\")(x)\n    outputs = layers.Dense(10)(x)\n    model = keras.Model(\n        inputs=inputs, outputs=outputs, name=\"mnist_model\"\n    )\n    return model\n\n\nwith strategy.scope():\n    model = my_model()\n    model.compile(\n        optimizer=\"adam\",\n        loss=keras.losses.SparseCategoricalCrossentropy(from_logits=True),\n        metrics=[\"accuracy\"],\n    )\n\nwith strategy.scope():\n    history = model.fit(\n        x_train, y_train, batch_size=64, epochs=2, validation_split=0.2\n        )\n\nwith strategy.scope():\n    test_scores = model.evaluate(x_test, y_test, verbose=2)\nprint(\"Test loss:\", test_scores[0])\nprint(\"Test accuracy:\", test_scores[1])\n</code></pre>\n<pre><code>Epoch 1/2\n750/750 7ms/step - accuracy: 0.8977 - loss: 0.3565 - val_accuracy: 0.9461 - val_loss: 0.1905\nEpoch 2/2\n750/750 6ms/step - accuracy: 0.9521 - loss: 0.1616 - val_accuracy: 0.9576 - val_loss: 0.1400\n313/313 - 2s - 7ms/step - accuracy: 0.9602 - loss: 0.1328\nTest loss: 0.132845938205719\nTest accuracy: 0.9602000117301941\n</code></pre>\n<p>TPU-config <a href=\"https://keras.io/keras_rs/examples/distributed_embedding_tf/\" target=\"_blank\">reference.</a></p>",
      "votes": null,
      "replies": [
        {
          "id": 3403600,
          "author_name": "konstantinboyko",
          "author_url": "",
          "post_date": "02/08/2026 21:58:22",
          "content": "<p>Innat, thank you so much for your help! It works.</p>",
          "votes": null,
          "replies": [
            {
              "id": 3403742,
              "author_name": "konstantinboyko",
              "author_url": "",
              "post_date": "02/09/2026 07:26:18",
              "content": "<p>If I take the three metrics required by the competition conditions and calculate them immediately during training, the training converges for me within two to ten epochs. After that, even up to the two-hundredth epoch, things only get worse. And after the 130th epoch, pure overfitting occurs. </p>\n<p>Your notebook converges on the 8th epoch. Manas Choudhary's notebook converges on the 4th epoch, and the combination of these two notebooks converges on the 2nd epoch.</p>\n<p>I saved all the epochs and checked, and the triple metric showed the best epoch correctly.</p>\n<p>The idea for metrics was borrowed from the <code>Metric Failure Demo</code> notebook.</p>",
              "votes": null,
              "replies": [
                {
                  "id": 3403767,
                  "author_name": "ipythonx",
                  "author_url": "",
                  "post_date": "02/09/2026 08:31:48",
                  "content": "<p>Can't say for sure the main cause. I would do</p>\n<ul>\n<li>use 2 or 3 very simple augmentation</li>\n<li>use static learning rate, not scheduler for testing</li>\n<li>use loss function and gradually add more, for example, 1st dice_ce loss, in 2nd experiment, dice_ce + cldice loss, etc. As we know, the loss functions are responsible for optimization. So, either dataloader gives some bad processed sample or, the learing rate scheduler cause the quick converges effect. </li>\n</ul>\n<p>The goal is to identify which may cause the quick converges. </p>",
                  "votes": null,
                  "replies": []
                }
              ]
            }
          ]
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "3403417": "I can't figure out why TPU v5e-8 won't start. [Notebook](https://www.kaggle.com/code/konstantinboyko/train-vesuvius-surface-3d-detection-200-images?scriptVersionId=296537209)\n\nIn the second option there is a kernel error. [Notebook](https://www.kaggle.com/code/konstantinboyko/train-vesuvius-surface-3d-detection-200-images?scriptVersionId=296541440)",
    "3403426": "It has been broken for many months. To use TPU, you have to set `jax` backend these days.",
    "3403427": "cc. @herbison \nFor more input on this.",
    "3403435": "During the `MAP - Charting Student Math Misunderstandings` competition it seemed to work.",
    "3403440": "In that case, you can try creating a notebook from that competition, or old one which proven to be worked earlier.",
    "3403473": "I think you need to see this @ipythonx @konstantinboyko <br>\nhttps://www.kaggle.com/discussions/product-announcements/670573",
    "3403534": "Ravi, thanks for this information, but I went through identity authentication a long time ago, when it was first introduced.",
    "3403562": "It looks like the torch+xla option works. [Notebook](https://www.kaggle.com/code/konstantinboyko/map-deepseekmath-7b-it-tpu-train-bf16?scriptVersionId=296596577)\n\n```\nWARNING: Logging before InitGoogle() is written to STDERR\nE0000 00:00:1770577160.730043      15 common_lib.cc:621] Could not set metric server port: INVALID_ARGUMENT: Could not find SliceBuilder port 8471 in any of the 0 ports provided in `tpu_process_addresses`=\"local\"\n=== Source Location Trace: ===\nlearning/45eac/tfrc/runtime/common_lib.cc:232\nNum devices: 8\n```",
    "3403574": "I think it would be more helpful to follow this [ticket](https://github.com/Kaggle/docker-python/issues/1367).",
    "3403578": "Also, (a bit more work though). Try using colab TPU with tf backend, it works, then check TF version and CUDA version. Then if possible maybe yoo can use that to install here. \n\nAlso, in kaggle see the [keras organizaion](https://www.kaggle.com/organizations/keras), find a noteboo that use tf-backend with Keras 3. You might find something helpful. \n\nBut I would recommened you to use `jax` backend, why wasting time when itme is limited at this point.",
    "3403588": "konstantinboyko \nI've found a temporary working soluton. Try this \n\n```python\n!pip install -U -q tensorflow-tpu==2.19.1 -q\n\nimport os\nimport libtpu\n\nos.environ[\"KERAS_BACKEND\"] = \"tensorflow\"\nos.environ[\"PJRT_DEVICE\"] = \"TPU\"\nos.environ[\"NEXT_PLUGGABLE_DEVICE_USE_C_API\"] = \"true\"\nos.environ[\"TF_PLUGGABLE_DEVICE_LIBRARY_PATH\"] = libtpu.get_library_path()\nos.environ[\"TF_XLA_FLAGS\"] = (\n    \"--tf_mlir_enable_mlir_bridge=true \"\n    \"--tf_mlir_enable_convert_control_to_data_outputs_pass=true \"\n    \"--tf_mlir_enable_merge_control_flow_pass=true\"\n)\n```\n```python\nimport numpy as np # linear algebra\nimport pandas as pd # data processing, CSV file I/O (e.g. pd.read_csv)\nimport os\n\nimport keras\nimport tensorflow as tf\ntf.__version__, keras.version()\n# ('2.19.1', '3.13.0')\n```\n```python\nresolver = tf.distribute.cluster_resolver.TPUClusterResolver(tpu=\"local\")\ntopology = tf.tpu.experimental.initialize_tpu_system(resolver)\ntpu_metadata = resolver.get_tpu_system_metadata()\ndevice_assignment = tf.tpu.experimental.DeviceAssignment.build(\n    topology, num_replicas=tpu_metadata.num_cores\n)\nstrategy = tf.distribute.TPUStrategy(\n    resolver, experimental_device_assignment=device_assignment\n)\nstrategy.num_replicas_in_sync\n8\n```\n```python\n(x_train, y_train), (x_test, y_test) = keras.datasets.mnist.load_data()\nx_train = x_train.reshape(60000, 784).astype(\"float32\") / 255\nx_test = x_test.reshape(10000, 784).astype(\"float32\") / 255\n\ndef my_model():\n    inputs = keras.Input(shape=(784,))\n    x = layers.Dense(64, activation=\"relu\")(inputs)\n    x = layers.Dense(64, activation=\"relu\")(x)\n    outputs = layers.Dense(10)(x)\n    model = keras.Model(\n        inputs=inputs, outputs=outputs, name=\"mnist_model\"\n    )\n    return model\n\n\nwith strategy.scope():\n    model = my_model()\n    model.compile(\n        optimizer=\"adam\",\n        loss=keras.losses.SparseCategoricalCrossentropy(from_logits=True),\n        metrics=[\"accuracy\"],\n    )\n\nwith strategy.scope():\n    history = model.fit(\n        x_train, y_train, batch_size=64, epochs=2, validation_split=0.2\n        )\n\nwith strategy.scope():\n    test_scores = model.evaluate(x_test, y_test, verbose=2)\nprint(\"Test loss:\", test_scores[0])\nprint(\"Test accuracy:\", test_scores[1])\n```\n```bash\n\nEpoch 1/2\n750/750 7ms/step - accuracy: 0.8977 - loss: 0.3565 - val_accuracy: 0.9461 - val_loss: 0.1905\nEpoch 2/2\n750/750 6ms/step - accuracy: 0.9521 - loss: 0.1616 - val_accuracy: 0.9576 - val_loss: 0.1400\n313/313 - 2s - 7ms/step - accuracy: 0.9602 - loss: 0.1328\nTest loss: 0.132845938205719\nTest accuracy: 0.9602000117301941\n```\n\nTPU-config [reference.](https://keras.io/keras_rs/examples/distributed_embedding_tf/)",
    "3403600": "Innat, thank you so much for your help! It works.",
    "3403742": "If I take the three metrics required by the competition conditions and calculate them immediately during training, the training converges for me within two to ten epochs. After that, even up to the two-hundredth epoch, things only get worse. And after the 130th epoch, pure overfitting occurs. \n\nYour notebook converges on the 8th epoch. Manas Choudhary's notebook converges on the 4th epoch, and the combination of these two notebooks converges on the 2nd epoch.\n\nI saved all the epochs and checked, and the triple metric showed the best epoch correctly.\n\nThe idea for metrics was borrowed from the `Metric Failure Demo` notebook.",
    "3403767": "Can't say for sure the main cause. I would do\n\n- use 2 or 3 very simple augmentation\n- use static learning rate, not scheduler for testing\n- use loss function and gradually add more, for example, 1st dice_ce loss, in 2nd experiment, dice_ce + cldice loss, etc. As we know, the loss functions are responsible for optimization. So, either dataloader gives some bad processed sample or, the learing rate scheduler cause the quick converges effect. \n\nThe goal is to identify which may cause the quick converges."
  },
  "source": "meta"
}